Matching the world's top multi-hop RAG systems, with no GPU, no fine-tuning, just pip install

Reddit r/artificial Tools

Summary

MOTHRAG is a multi-hop RAG system that matches the performance of top GPU-dependent systems (HippoRAG 2, CoRAG, NeocorRAG) using only commodity API calls, with no GPU, no fine-tuning, and deployment via pip install plus API keys.

No content available
Original Article
View Cached Full Text

Cached at: 06/20/26, 02:29 PM

# Matching the world's top multi-hop RAG systems, with no GPU, no fine-tuning, just pip install Source: [https://www.linkedin.com/pulse/matching-worlds-top-multi-hop-rag-systems-gpu-just-pip-geymonat-zbgxe](https://www.linkedin.com/pulse/matching-worlds-top-multi-hop-rag-systems-gpu-just-pip-geymonat-zbgxe) [Julian Geymonat](https://it.linkedin.com/in/julian-geymonat) ### Julian Geymonat Published Jun 19, 2026 The three systems below \(HippoRAG 2, CoRAG, NeocorRAG\) are among the strongest multi\-hop QA frameworks published\. Every one of them depends on a GPU, fine\-tuning, or constrained decoding to get there\. MOTHRAG sits right alongside them on F1, while running entirely on commodity API calls\. No GPU\. No fine\-tuning\. No constrained decoding\. Nonon\-commercial licenses\. System \| Deployment \| HotpotQA \| 2Wiki \| MuSiQue \| AVG HippoRAG 2 \| offline graph \+ GPU \| 75\.5 \| 71\.0 \| 48\.6 \| 65\.0 CoRAG \| trained retrieval \| 75\.1 \| 75\.1 \| 52\.9 \| 67\.7 NeocorRAG \| GPU constrained decode\| 78\.3 \| 76\.1 \| 52\.6 \| 69\.0 MOTHRAG \(ours\) \| commodity APIs only \| 78\.1 \| 76\.3 \| 50\.5 \| 68\.3 Highest average F1 among commercially\-deployable frameworks, within 0\.7 points of the GPU\-bound state of the art, and ahead of it on 2Wiki\. The point isn't beating these systems, it's reaching their tier with none of their infrastructure\. Deployment is a pip install plus API keys: pip install mothrag from mothrag import MothRAG m = MothRAG\.from\_documents\(\["Paris is the capital of France\.", "The Eiffel Tower is in Paris\."\]\) result = m\.query\("In which country is the Eiffel Tower?"\) print\(result\.answer\) print\(result\.confidence\) The pipeline is fully modular\. Readers, embedders and retrieval judges all swap without retraining, installed as optional extras: gemini/openai for API readers and embedders, sentence\-transformers for a local embedding fallback, faiss for vector stores over 100k\-10M chunks, retrieval for classic BM25/graph features, prod for the full stack\. A one\-flag economy tier swaps the retrieval judge and drops cost from ~$0\.032 to ~$0\.018 per query at statistical parity on HotpotQA and 2Wiki\. Every answer is proof\-tree\-structured so you can inspect each reasoning hop, and the per\-query outputs behind every table in the paper are released so you can verify the numbers\. Happy to answer questions about the pipeline or the judge design\. ## More articles by Julian Geymonat ## Explore content categories

Similar Articles

RAG-Stack: Co-Optimizing RAG Serving Performance and Quality

arXiv cs.AI

This paper introduces RAG-Stack, a framework that co-optimizes RAG serving performance and answer quality by efficiently exploring the joint algorithm-system configuration space. It finds Pareto frontiers that cover significantly more quality-performance space than existing configuration-search methods.

ScalableRAG: High-Quality RAG at Zero Ingestion Cost

arXiv cs.AI

This paper introduces ScalableRAG, a retrieval-augmented generation method that achieves high accuracy without any ingestion costs (no vector database or knowledge graph) by using regex-based set creation and aggregative reasoning. It outperforms baselines on multiple datasets and also presents a limited-ingestion variant for further accuracy improvements.