Tag
This paper presents a framework that augments Large Language Models with geometric vision parsing and symbolic solving to match state-of-the-art multimodal models on complex geometry problems, using a new benchmark from 2025 Chinese Zhongkao exams for evaluation.
The paper audits frozen LLMs to examine how geometric constraint information is encoded and whether it can be used for generation, influence, and steering, revealing gaps between decodability and actionable outputs.
Introduces JigShape, a jigsaw puzzle benchmark with interlocking pieces to evaluate visual-geometric reasoning in vision-language models. It finds that frontier VLMs largely fail at geometric reasoning and all models collapse on larger puzzle sizes, revealing a 'scaling cliff' for constraint satisfaction.
This paper introduces SR-REAL, a unified framework for spatial vision-language models that combines linguistic deduction and 3D geometric reasoning via reinforcement learning, enabling robust multi-step spatial reasoning across diverse tasks.
This paper introduces P3D-Bench, a benchmark for evaluating multimodal large language models on parametric 3D generation tasks, including text-to-3D, image-to-3D, and assembly-3D, with metrics for geometric precision, semantic alignment, and part-level structure.
Geometric Latent Reasoning (GLR) introduces a geometric path-approximation method for latent reasoning in LLMs, enabling shorter generations while maintaining accuracy across mathematical reasoning benchmarks.