@VikParuchuri: If you're blind, most PDFs are a garbled mess - and having one fixed can cost >$100. Our API now makes them usable for …
Summary
An API is announced that makes PDFs usable for blind users by fixing garbled content, drastically reducing costs from over $100 to pennies.
View Cached Full Text
Cached at: 09/09/26, 01:52 PM
If you’re blind, most PDFs are a garbled mess - and having one fixed can cost >$100.
Our API now makes them usable for pennies. https://t.co/K714DnR3wM
Similar Articles
@VikParuchuri: JATS conversion is a huge pain point that slows down science - costs dollars *per page*, and takes weeks. We can do it …
Datalab is releasing a new processor that converts PDFs to JATS XML, reducing cost to cents per page and time to under 5 minutes with 92.6% accuracy in a human-matched benchmark.
@kushalbyatnal: Over 1 billion PDFs are created every day, but your agents still can’t read them reliably. Today we’re releasing Parse …
Extend released Parse 2.0, a state-of-the-art document parsing API that achieves top accuracy on real-world documents, outperforming competitors on the open-source RealDoc-Bench benchmark.
@CoderDaMing: China has open-sourced a peanut-sized OCR that can parse an entire 100-page PDF in one go. It's called 'Unlimited-OCR'. Only 3B parameters. Runs locally. Other OCR tools cut documents page by page, easily losing context. This one reads the entire document at once. → Single 'long-range' parse (32K context window...
China has open-sourced the OCR model Unlimited-OCR with only 3B parameters, which can parse an entire 100-page PDF in one go, supports local execution, achieves 93% accuracy, and is completely free and open-source.
olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models
olmOCR is an open-source toolkit using a fine-tuned vision language model to extract clean text from PDFs while preserving structure, optimized for large-scale batch processing.
@jerryjliu0: A downside with using VLMs to parse PDFs is guaranteeing that the output text is *correct* and output in the correct re…
Jerry Liu discusses challenges with using Vision Language Models for PDF parsing, particularly around ensuring text correctness and maintaining proper reading order while avoiding hallucinations.