Tag
LlamaIndex benchmarked GPT-5.6 on document understanding and found no improvement over GPT-5.5; the model performs well on text and tables but struggles with charts and layout.
Jerry Liu reports updated results for Mistral OCR on ParseBench, showing it outperforms GPT-5.5 and trails only Gemini 3.1 Pro, with strong performance on content faithfulness and semantic formatting.
Jerry Liu's team is presenting ParseBench, a comprehensive document understanding benchmark for VLMs, at CVPR 2026. The benchmark includes 2,000 pages of real-world enterprise documents with evaluation metrics for tables, charts, and visual grounding.