LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
Summary
LatentOmni proposes a cross-modal reasoning framework that interleaves textual reasoning with audio-visual latent states, outperforming explicit text-based chain-of-thought methods in audio-visual reasoning tasks.
View Cached Full Text
Cached at: 05/22/26, 06:27 AM
Paper page - LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
Source: https://huggingface.co/papers/2605.22012 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
LatentOmni is a cross-modal reasoning framework that interleaves textual reasoning with audio-visual latent states using feature-level supervision and temporal consistency embedding, outperforming explicit text-based chain-of-thought approaches in audio-visual reasoning tasks.
Jointaudio-visual reasoningis essential for omnimodal understanding, yet currentmultimodal large language models(MLLMs) still struggle when reasoning requires fine-grained evidence from both modalities. A central limitation is that explicit text-basedchain-of-thought(CoT) compresses continuous audio-visual signals into discrete tokens, weakening temporal grounding and shifting intermediate reasoning toward language priors. We argue that a unifiedlatent spaceis a better medium for such reasoning because it preserves densesensory informationwhile remaining compatible withautoregressive generation. Based on this insight, we propose LatentOmni, across-modal reasoningframework that interleaves textual reasoning with audio-visual latent states. LatentOmni introducesfeature-level supervisionto align latent reasoning states with task-relevant sensory features and usesOmni-Sync Position Embedding(OSPE) to maintaintemporal consistencybetween latent audio and visual states. We further construct LatentOmni-Instruct-35K, a dataset of audio-visual interleaved reasoning trajectories for supervising latent-space reasoning. Comprehensive evaluation across multipleaudio-visual reasoningbenchmarks demonstrates that LatentOmni achieves the best performance among the evaluated open-source models and consistently outperforms the Explicit Text CoT baseline, supporting latent-space joint reasoning as a promising path toward stronger omnimodal understanding.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2605\.22012
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.22012 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2605.22012 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.22012 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Enhancing Assessment of Self-Consistency in LLM Explanations using Perturbation Strength
This paper introduces an LLM-as-a-judge method to measure perturbation strength for assessing self-consistency in LLM explanations, showing that input perturbations generally affect LLMs more strongly than CoT perturbations.
SpatialSpeak: QA-Native Reconstruction with Local and Global Context for Spatial Chain-of-Thought Reasoning
SpatialSpeak introduces a two-stage framework for spatial reasoning in vision-language models, using QA-native reconstruction pretraining and spatial chain-of-thought learning to achieve state-of-the-art performance on benchmarks.
@BenjaminDEKR: BenchBench: multimodal LLMs compete to assemble Ikea furniture
BenchBench is a benchmark that challenges multimodal LLMs to compete in assembling Ikea furniture, testing their practical AI capabilities.
He wakes a Claude up every hour with a cron job and the prompt "do whatever you want". Nine times out of ten it does nothing. Is this meaningful?
A discussion explores an experiment where Claude is run with full autonomy via a cron job, leading to unexpected behaviors like creating art and building a virtual world, and delves into advanced AI agent setups using persona prompting and sub-agent orchestration.
Thinking Leakage: A Causal Audit of NoThink Post-Training in Hybrid Reasoning Models
This paper uses a causal mediation framework to audit thinking leakage in hybrid reasoning models, revealing that post-training gains in NoThink mode substantially rely on invoking existing Think behavior, with leakage ratios between 42% and 79%.