Tag
Introduces MultivationBench, a benchmark for evaluating multimodal large language models' sequential motivation reasoning using story-driven visual narratives based on Maslow's hierarchy and Reiss's desires. Results show all tested models struggle with dynamic motivation inference.
Introduces the Complexity Ceiling Benchmark (CCB) that evaluates LLM reasoning decay as the number of sequential steps increases across three domains. Finds a consistent geometric per-step decay and that all models collapse on transitive social logic within 5 steps, even with strong overall accuracy.
This paper argues that large language models struggle with causal reasoning and long-horizon planning due to a mismatch between sequence prediction and reasoning over latent environment dynamics, and introduces the Latent Dynamics Inference perspective along with the Flux environment to study these limitations.