Introducing ASCIITermDraw Bench | Testing the ability of VLMs to Generate and Edit ASCII
Summary
Introduces ASCIITermDraw Bench, a benchmark designed to evaluate vision-language models on ASCII art generation and editing tasks.
Similar Articles
VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and Editing
VCG-Bench is a unified benchmark for evaluating vision-language models on structured diagram generation and editing tasks, introducing a 'Diagram-as-Code' paradigm using symbolic mxGraph XML and a taxonomized dataset of 1,449 diagrams across 6 domains.
Can AI Draw Science? A Benchmark for Evaluating Scientific Figure Generation by Text-to-Image and Multimodal Models
Introduces SciDraw-Bench, a benchmark for evaluating scientific figure generation by text-to-image and multimodal models, with a four-dimensional evaluation protocol. Findings show domain-specific systems outperform general-purpose models, with text fidelity remaining the hardest challenge.
VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?
Introduces VisEditBench, a benchmark of 1,395 human-annotated visualization code-editing tasks for vision-language models, covering feedback-guided repair and reference-guided restyling. Evaluates 20 VLMs and proposes VisEditAgent, a render-grounded editing framework that improves pass rates from 55.75% to 67.99%.
Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture
Terminal-Bench-LILT presents a multilingual coding benchmark with 300 authentic tasks in 10 languages, exposing gaps in AI models' handling of non-English issues like internationalization and cultural conventions.
DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings
Introduces DrawingVQA, the first benchmark for evaluating multimodal large language models on real-world construction drawings, with 33 drawings and 92 QA pairs across three reasoning depths, revealing a gap between model and expert performance.