Qwen3.5-Omni Technical Report
Summary
Qwen3.5-Omni is a hundreds-of-billions-parameter multimodal model with advanced audio-visual understanding and generation capabilities, featuring novel Audio-Visual Vibe Coding and achieving SOTA results across 215 benchmarks while matching Gemini-3.1 Pro.
View Cached Full Text
Cached at: 04/20/26, 08:27 AM
Paper page - Qwen3.5-Omni Technical Report
Source: https://huggingface.co/papers/2604.15804
Abstract
Qwen3.5-Omni is a large-scale multimodal model with hundreds of billions of parameters that excels in audio-visual understanding and generation, featuring advanced architectures and novel capabilities like Audio-Visual Vibe Coding.
In this work, we present Qwen3.5-Omni, the latest advancement in the Qwen-Omni model family. Representing a significant evolution over its predecessor, Qwen3.5-Omni scales to hundreds of billions of parameters and supports a 256k context length. By leveraging a massive dataset comprising heterogeneous text-vision pairs and over 100 million hours of audio-visual content, the model demonstrates robust omni-modality capabilities. Qwen3.5-Omni-plus achieves SOTA results across 215 audio and audio-visual understanding (https://huggingface.co/papers?q=audio-visual%20understanding), reasoning, and interaction subtasks and benchmarks, surpassing Gemini-3.1 Pro in key audio tasks and matching it in comprehensive audio-visual understanding (https://huggingface.co/papers?q=audio-visual%20understanding). Architecturally, Qwen3.5-Omni employs a Hybrid Attention Mixture-of-Experts (https://huggingface.co/papers?q=Hybrid%20Attention%20Mixture-of-Experts) (MoE (https://huggingface.co/papers?q=MoE)) framework for both Thinker and Talker, enabling efficient long-sequence inference. The model facilitates sophisticated interaction, supporting over 10 hours of audio understanding and 400 seconds of 720P video (at 1 FPS). To address the inherent instability and unnaturalness in streaming speech synthesis (https://huggingface.co/papers?q=speech%20synthesis), often caused by encoding efficiency discrepancies between text and speech tokenizers, we introduce ARIA (https://huggingface.co/papers?q=ARIA). ARIA (https://huggingface.co/papers?q=ARIA) dynamically aligns text and speech units, significantly enhancing the stability and prosody of conversational speech with minimal latency impact. Furthermore, Qwen3.5-Omni expands linguistic boundaries, supporting multilingual understanding (https://huggingface.co/papers?q=multilingual%20understanding) and speech generation across 10 languages with human-like emotional nuance. Finally, Qwen3.5-Omni exhibits superior audio-visual grounding (https://huggingface.co/papers?q=audio-visual%20grounding) capabilities, generating script-level structured captions with precise temporal synchronization and automated scene segmentation. Remarkably, we observed the emergence of a new capability in omnimodal models: directly performing coding based on audio-visual instructions, which we call Audio-Visual Vibe Coding (https://huggingface.co/papers?q=Audio-Visual%20Vibe%20Coding).
View arXiv page (https://arxiv.org/abs/2604.15804)View PDF (https://arxiv.org/pdf/2604.15804)Add to collection (https://huggingface.co/login?next=%2Fpapers%2F2604.15804)
Get this paper in your agent:
hf papers read 2604.15804
Don’t have the latest CLI?curl -LsSf https://hf.co/cli/install.sh | bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2604.15804 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2604.15804 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2604.15804 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to a collection (https://huggingface.co/new-collection) to link it from this page.
Similar Articles
Qwen3.7 Preview lands on Arena (1 minute read)
Alibaba Qwen announces two major model releases: Qwen3-Omni, the first natively end-to-end omni-modal AI unifying text, image, audio and video, and Qwen3-Next-80B-A3B, an ultra-efficient MoE model with 3B activated parameters per token, achieving SOTA performance and 10x faster inference than Qwen3-32B.
Qwen-Image-2.0 Technical Report
Qwen-Image-2.0 is a new image generation foundation model that unifies high-fidelity synthesis and precise editing using Qwen3-VL and a Multimodal Diffusion Transformer. It excels in text-rich content, multilingual typography, and photorealistic generation.
Qwen/Qwen3.6-35B-A3B-FP8
Alibaba releases Qwen3.6-35B-A3B-FP8, an open-weight quantized variant of Qwen3.6 with 35B parameters and 3B activated via MoE, featuring improved agentic coding capabilities and thinking preservation for iterative development.
Qwen/Qwen3.8-2.4T-A95B-FP8
Qwen releases FP8-quantized weights for Qwen3.8-2.4T-A95B, a 2.4T-parameter MoE model with 95B activated parameters, claiming Qwen-Max-class capability in an open release with strong coding, agentic, and long-context performance.
Qwen/Qwen3.6-35B-A3B
Qwen releases Qwen3.6-35B-A3B, an open-weight Mixture-of-Experts model with 35B total parameters and 3B active parameters, featuring significant improvements in agentic coding and reasoning preservation.