@jerryjliu0: ParseBench is the first benchmark to include VLM chart understanding over enterprise documents. Existing benchmarks (Ch…

X AI KOLs Timeline Papers

Summary

ParseBench introduces the first benchmark evaluating vision-language models on chart comprehension within full enterprise documents, addressing gaps in prior chart-only benchmarks.

ParseBench is the first benchmark to include VLM chart understanding over enterprise documents. Existing benchmarks (ChartQA, ChartXiv) test over charts specifically and not the chart's inclusion in the overall document. Also doesn't contain references to real-world
Original Article
View Cached Full Text

Cached at: 04/22/26, 08:23 AM

ParseBench is the first benchmark to include VLM chart understanding over enterprise documents. Existing benchmarks (ChartQA, ChartXiv) test over charts specifically and not the chart’s inclusion in the overall document. Also doesn’t contain references to real-world

Similar Articles

@jerryjliu0: Our team is at CVPR 2026 if you want to come say hi :)

X AI KOLs Following

Jerry Liu's team is presenting ParseBench, a comprehensive document understanding benchmark for VLMs, at CVPR 2026. The benchmark includes 2,000 pages of real-world enterprise documents with evaluation metrics for tables, charts, and visual grounding.

ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction

Hugging Face Daily Papers

ExtractBench is a new benchmark for schema-guided enterprise document extraction, evaluating value accuracy, record completeness, grounding, and cost across 4,869 pages of enterprise documents. The authors find that commercial VLMs struggle with long documents while coding agents are more accurate but costly, and LlamaExtract AgenticPlus leads on all metrics.

ChartArena: Benchmarking Chart Parsing across Languages, Scenarios, and Formats

Hugging Face Daily Papers

ChartArena is a comprehensive bilingual benchmark for chart parsing that evaluates models across eight chart families and three visual scenarios (digital, printed, hand-drawn), using a human-agent annotation pipeline and format-agnostic evaluation. Evaluations of 26 MLLMs reveal that while proprietary models lead overall, open-source models are catching up, and diagrammatic structures and hand-drawn scenarios remain challenging.