Messy documents in. Complex knowledge graphs out. One command line.
If your pipeline simply compiles data into generic chunks, it won't be able to answer complex questions later.
The guy behind Hyper-Extract just open-sourced a framework to fix exactly this 🔥
Hyper-Extract turns messy documents into typed knowledge targets before you decide how to store them.
Instead of locking you into one architecture, it turns GraphRAG, LightRAG, and KG-Gen into simple engine choices.
What the project includes:
→ 8 different knowledge structures (from simple lists to Spatio-Temporal graphs)
→ 10+ extraction engines ready to go out of the box
→ 80+ zero-code YAML templates for Finance, Legal, Medical, and more
→ Support for local deployment via vLLM (DeepSeek-r1, Qwen, etc.)
100% free and fully open-source under Apache-2.0.
Repo in 🧵 ↓