@contractedai: Can an embodied reasoning model do a brake job? We ran 9 VLMs through an expert mechanic's exam on his own repair, with…

X AI KOLs Following Papers

Summary

An evaluation of 9 vision-language models on an expert mechanic's brake-repair exam shows they know the theory but fail in noticing and adapting during real-world deployment.

Can an embodied reasoning model do a brake job? We ran 9 VLMs through an expert mechanic's exam on his own repair, with @physicreality. In theory they know what to do. Noticing and adapting on their own is where they fall short. This is where deployment would fail. https://t.co/exvDR5XmOV
Original Article
View Cached Full Text

Cached at: 08/13/26, 11:31 PM

Can an embodied reasoning model do a brake job?

We ran 9 VLMs through an expert mechanic’s exam on his own repair, with @physicreality. In theory they know what to do. Noticing and adapting on their own is where they fall short.

This is where deployment would fail. https://t.co/exvDR5XmOV

Similar Articles

Do VLMs Reason Like Engineers? A Benchmark and a Stage-wise Evaluation

arXiv cs.AI

This paper introduces EngVQA, a multimodal benchmark for evaluating engineering reasoning in vision-language models, along with an 8-stage automatic evaluation framework that enables fine-grained analysis of reasoning failures. It reveals substantial limitations in current VLMs' engineering reasoning capabilities.

Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap

arXiv cs.CL

This paper introduces CrossMath, a controlled multimodal reasoning benchmark that reveals a critical limitation in current vision-language models: they perform reasoning primarily in textual space rather than genuine vision-grounded reasoning, with visual input often degrading performance compared to text-only baselines. The authors propose fine-tuning approaches to mitigate this modality gap and improve multimodal reasoning capabilities.

HumanCLAW: Can Vision-Language Models Act Through a Body?

Hugging Face Daily Papers

This paper introduces HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution to assess whether vision-language models can act through a physical body. Testing nine state-of-the-art VLMs on 1,218 episodes across 41 scenes, the best model achieves only a 16.8% success rate, revealing that current VLMs lack embodied self-awareness.