OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
Summary
This paper introduces OmniVChat, a task for native audio-visual dialogue, and presents a data engine, benchmark, and reinforcement learning method to train and evaluate omni models, demonstrating improvements on synthesized and human-recorded data.
View Cached Full Text
Cached at: 09/21/26, 11:20 AM
Paper page - OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
Source: https://huggingface.co/papers/2609.21465 Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
WedefineOmniVChat(OmniVideoChat)asthetaskofnativeaudio-visualdialoguebetweenauserandanomnimodel.InOmniVChat,omnimodelsdirectlyandsimultaneouslyreceiveaudioandvideofromauserandreturntext.Theuser’squeryisembeddedintheaudioandvideo,withoutaseparatetextquestion,externalcaptioning,orspeechrecognition.Directaudio-visualinputreducesexternallatencyandcomputationwhilepreservingperceptualcues.However,researchonOmniVChatfacestwoconstraints:dataavailabilityandevaluation.Recordingsofpeopleusingtheirowndevicesarescarce.Furthermore,agoodreplyoftenneedstoaccountfortheuser’ssurroundings,facialexpressions,andnearbyobjects,andsuchresponsescanbeexpressedinmanydifferentways,makingkeywordmatchingunreliableforevaluatingreplyquality.Recentprogressinagentsystemsandvideogenerationmakesgenerationforcomprehensionviable,whichmeansusingsynthesizeddialoguesfortrainingandevaluation.Therefore,wepresentOmniVChat-Studio,amulti-agentdataengineforsynthesizingsingle-andmulti-turnaudio-visualdialogues.WeusesynthesizeddialoguestobuildOmniVChat-Bench,anevaluationbenchmarkthatevaluatesomnimodels’basicdialogueabilitiesacrossfiveabilitycategories.WealsopresentOmniVChat-RL,areinforcementlearningrewarddesignthatjointlytargetsreplycorrectness,efficiency,andstyleinOmniVChat.TrainingQwen3-Omni-InstructwithOmniVChat-RLonsynthesizeddialoguesimprovesitsperformanceonbothOmniVChat-Benchandthehuman-recordedOmniVChat-Bench-Human.Thesegainsvalidatetherewarddesignandshowtransfertoreal-worlddialoguesintrainingandevaluation.
View arXiv pageView PDFGitHub10Add to collection
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.21465 in a model README.md to link it from this page.
Datasets citing this paper1
#### Harland/OmniVChat Viewer• Updatedabout 6 hours ago • 2.8k • 68 • 2
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.21465 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence
Ex-Omni-2D is an omni-modal dialogue framework that generates coordinated text, speech, and reference-conditioned video responses via a visual thought plan and a distilled streaming video generator, achieving a practical quality-efficiency trade-off.
OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants
OmniInteract introduces a streaming benchmark for real-time omnimodal LLMs, evaluating online audio-visual processing with temporal grounding and interactive response requirements. Experiments show that current models perform poorly, with the best overall IA-QTF1 score reaching only 0.368.
Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction
This paper introduces Omni-DuplexEval, a benchmark and automatic evaluation framework for real-time duplex interaction in multimodal large language models, assessing continuous response generation and proactive event detection in streaming scenarios.
OmniVideo-100K: A Dataset for Audio-Visual Reasoning through Structured Scripts and Evidence Chains
OmniVideo-100K introduces an automated data engine with entity-anchored scripting and clue-guided QA generation to improve audio-visual reasoning and temporal consistency, achieving significant performance gains across multiple benchmarks.
How does OmniDimension make its AI phone calls sound so natural, fast, and multilingual? Can a solo developer build something similar?
Explores how OmniDimension achieves natural, fast, and multilingual AI phone call synthesis, and discusses the feasibility for solo developers to replicate similar capabilities.