Tag
Introduces an analytically exact framework for controlled behavioral evaluation of LLMs, using fully crossed factorial experiments and exact token-level probability mass functions to isolate causal biases that aggregate benchmarks obscure.
This paper identifies that failures in visual reasoning often stem from breakdowns in dynamic cross-modal coordination between visual and textual evidence during chain-of-thought generation. It introduces DyCo-RL, a reinforcement learning framework that rewards effective cross-modal coordination, leading to improved reasoning performance.