The first Graph-RAG system that beats GPT-4o on academic benchmarks — running on a single consumer GPU. No cloud required. No data leaves your infrastructure.
Validated on HotpotQA — the standard multi-hop reasoning benchmark. 1,000 questions. No cherry-picking.
| System | EM Score | F1 Score | Infrastructure |
|---|---|---|---|
| Microsoft GraphRAG | 31.70 | 42.74 | GPT-4o (cloud) |
| BGE Dense + GPT-4o | 47.60 | 60.36 | GPT-4o (cloud) |
| Noēsis | 59.50 | 74.74 | 35B model (on-premises) |
| HopRAG | 62.00 | 76.06 | GPT-4o (cloud) |
| StepChain (SOTA) | 66.70 | 79.50 | GPT-4o (cloud) |
Noēsis uses a 35B on-premises model for graph construction. All competitors use GPT-4o for everything. Noēsis retrieves k=10 chunks; baselines use k=20. Source: arXiv:2608.15919
"A 2.3B smartphone-sized model with Noēsis architecture matches GPT-4o dense retrieval. The architecture does ~80% of the work."
— From the paper, Ablation study §5.1We built Noēsis because existing solutions weren't good enough for production.
Each solving a problem no one else has addressed.
Simulates human reading with degrading memory. Forward pass builds context; backward pass reconnects early concepts to late ones. Result: 100% more edges than single-pass extraction on 300+ page documents.
TCP congestion control (1988) transferred to document orchestration. Probes GPU capacity in real-time. 23× faster than sequential. Zero OOM crashes — including 160+ min on 6GB GPU.
Domain-aware selective quantization for Mixture-of-Experts models. Hot layers keep precision, cold layers compress. 6.3× prompt speedup on 12GB. Re-adapts when your domain changes — no cumulative precision loss.
Multiple KBs stay separate. At query time, Mesh discovers emergent connections between them — relationships that exist in no single document. Adaptive threshold, <2ms routing, real-time structural discovery.
Any industry where knowledge is scattered across thousands of documents and multiple teams.
Clinical research, protocols, treatment literature
Contracts, regulations, case precedents
Archives, metadata, rights, production
Literature reviews, cross-disciplinary links
Codebase intelligence, architecture, docs
Air-gapped, sovereign data, compliance
"On a 193-page book: 90% verified precision on causal relationships spanning hundreds of pages."
— From the paper, Source verification studyYour documents never leave your infrastructure. Period.
Noēsis runs entirely on your hardware — a single workstation GPU is enough. No cloud API calls during ingestion or querying. Perfect for healthcare, defence, legal, and any environment where data sovereignty is non-negotiable.
Need cloud power? Plug in any LLM (OpenAI, Anthropic, AWS Bedrock). Same API, same results. Your choice, always.
Minimum hardware. RTX 4080 Laptop class. Full system operational including 35B parameter model.
160+ minutes continuous operation. Zero crashes. The system adapts — it never fails catastrophically.
Start on one machine. Scale to multiple nodes. Same architecture, linear performance gains.
Local GPU, Ollama, OpenAI, Claude, AWS Bedrock. Swap backends without changing anything else.
Universal MCP protocol. Drop-in integrations for the tools your team already uses.
Anthropic's coding agent queries your knowledge graph natively.
Knowledge-aware coding without leaving your editor.
Google's CLI with graph traversal from the terminal.
Universal adapter: works with Continue, Cline, Roo Code, and more.
Full system paper with benchmark results, ablation studies, and architectural details.
Presented "The Sector-Agnostic Shift" alongside James Whitebread (CARE ADHD).
Application No. 102026000023146 — covering all four core algorithms.
Coming: RAI Amsterdam, 11–14 September 2026.
Schedule a demo. See your own documents transformed into an interconnected knowledge graph in minutes.
Or visit us at IBC Amsterdam, 11–14 September 2026