Tag
This paper introduces FactoSR, a factorized reinforcement learning framework that enhances spatial reasoning in Vision-Language Models by decomposing 4D properties into orthogonal sub-objectives, achieving significant performance boosts on multi-view and video benchmarks.
SpatialClaw is a training-free framework that uses code as an action interface to enable flexible, stateful spatial reasoning in vision-language models, achieving superior performance across diverse 3D/4D spatial reasoning tasks.