Tag
This paper evaluates MiniMax-H3, an omni-modal generative model, by introducing a comprehensive framework to assess its reasoning about the physical world through multimodal inputs. The evaluation reveals that video-based decision reasoning performs best, while audio-based disambiguation reasoning is the weakest.