Tag
A comprehensive evaluation of 16 frontier vision language models on document parsing shows that Opus 5.5 offers the best performance relative to its price, especially for tables, while GPT-6 Luna is good for cost-effective parsing. For large-scale use, dedicated tools like LlamaParse are recommended, but Opus 5.5 leads for in-app parsing.
VisTW is a Traditional Chinese vision benchmark for evaluating VLMs on real-world Taiwanese content, and Twinkle Eval has added support for it with accuracy comparisons.
This article evaluates open weights VLMs for egocentric data processing using HFlow, showing that models like Gemma and Qwen achieve high accuracy compared to the Gemini baseline while being cheaper and suitable for self-hosting.
This paper proposes MAVEN, a hierarchical framework for evaluating multimodal content against macro-societal values, featuring a benchmark and optimized evaluators for scalable assessment.
User presents a comprehensive comparison of local text-to-image models using 192 prompts, evaluating capabilities like text rendering, faces, anatomy, and spatial composition, with results and prompts publicly available at imagebench.ai.
An informal experiment using a chessboard reveals that vision language models often fail at spatial reasoning and precise structured output, despite correctly recognizing pieces, highlighting a key gap in VLM evaluation.
A PhD student asks whether submitting vision-language model evaluation work to an EMNLP workshop is worthwhile after rejection from a top imaging venue.