document-qa

Tag

Cards List
#document-qa

@Ryrenz: Want AI to answer based on your own data without building RAG from scratch? These 5 open-source apps turn documents into a Q&A knowledge base. 1. RAGFlow — Advanced layout understanding RAG engine, 83.8k stars. Deep comprehension of complex document layouts, tables, long reports, all parsed accurately with cited answers. A popular choice for enterprise knowledge bases.

X AI KOLs Timeline · 2026-07-01 Cached

Recommends 5 open-source RAG tools (RAGFlow, AnythingLLM, Onyx, Khoj, kotaemon) that turn documents into a Q&A knowledge base with zero code, each with unique features.

0 favorites 0 likes
#document-qa

What are you using to preprocess pdfs before feeding them to a local model?

Reddit r/LocalLLaMA · 2026-06-02

A user seeks recommendations for PDF preprocessing tools to improve input quality for local LLM-based document QA, comparing pymupdf, pdfplumber, docling, and llamaparse for handling messy layouts like tables and multi-column text.

0 favorites 0 likes
#document-qa

Vision-capable LLMs vs. OCR for long-document (including charts, images, tables, etc.) QA

Reddit r/artificial · 2026-05-24

A benchmark comparing vision-capable LLMs (native PDF reading) against OCR-based pipelines on 30 long, image-heavy PDFs finds that OCR with layout extraction still outperforms vision models on chart/table-heavy pages and has a 0% failure rate vs. 7% for native PDF, though the sample size is small and many gaps are within noise.

0 favorites 0 likes
← Back to home

Submit Feedback