@contractedai: Can an embodied reasoning model do a brake job? We ran 9 VLMs through an expert mechanic's exam on his own repair, with…
Summary
An evaluation of 9 vision-language models on an expert mechanic's brake-repair exam shows they know the theory but fail in noticing and adapting during real-world deployment.
View Cached Full Text
Cached at: 08/13/26, 11:31 PM
Can an embodied reasoning model do a brake job?
We ran 9 VLMs through an expert mechanic’s exam on his own repair, with @physicreality. In theory they know what to do. Noticing and adapting on their own is where they fall short.
This is where deployment would fail. https://t.co/exvDR5XmOV
Similar Articles
Do VLMs Reason Like Engineers? A Benchmark and a Stage-wise Evaluation
This paper introduces EngVQA, a multimodal benchmark for evaluating engineering reasoning in vision-language models, along with an 8-stage automatic evaluation framework that enables fine-grained analysis of reasoning failures. It reveals substantial limitations in current VLMs' engineering reasoning capabilities.
Do Vision-Language Models Truly Perform Vision Reasoning? A Rigorous Study of the Modality Gap
This paper introduces CrossMath, a controlled multimodal reasoning benchmark that reveals a critical limitation in current vision-language models: they perform reasoning primarily in textual space rather than genuine vision-grounded reasoning, with visual input often degrading performance compared to text-only baselines. The authors propose fine-tuning approaches to mitigate this modality gap and improve multimodal reasoning capabilities.
Where Instruction Hierarchy Breaks: Diagnosing and Repairing Failures in Reasoning Language Models
This paper introduces a white-box diagnostic framework that localizes instruction hierarchy failures in reasoning language models into identification, conflict resolution, and response realization stages. It evaluates several models and proposes two training-free self-monitoring mechanisms that reduce non-compliance by 81–99%.
HumanCLAW: Can Vision-Language Models Act Through a Body?
This paper introduces HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution to assess whether vision-language models can act through a physical body. Testing nine state-of-the-art VLMs on 1,218 episodes across 41 scenes, the best model achieves only a 16.8% success rate, revealing that current VLMs lack embodied self-awareness.
@dair_ai: Can an LLM agent actually build a model of an environment it cannot see? This work makes the question gradeable. An age…
A research paper proposes agentic automata learning to evaluate whether LLM agents can infer hidden world models through interaction, finding that performance drops sharply as task complexity increases and that reasoning models outperform non-reasoning ones but still struggle.