Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

Hugging Face Daily Papers Papers

Summary

Audio-Visual Flamingo (AV-Flamingo) is a fully open audio-visual large language model designed for understanding and reasoning over long and complex videos, outperforming similarly sized open models and competitive with larger models.

We present Audio-Visual Flamingo (AV-Flamingo), a fully open state-of-the-art audio-visual large language model (AV-LLM) for joint understanding and reasoning over audio, images, and long-form videos. Unlike prior AV-LLMs that primarily focus on short clips, AV-Flamingo is designed for understanding and reasoning over long and complex real-world (audio-visual) videos. To support this, we make three key contributions: (i) Audio-Visual-Skills, a large-scale collection of real-world videos with ~7M caption and question-answer training instances designed to emphasize temporal, compositional, and cross-modal audio-visual reasoning; (ii) a novel three-stage curriculum that progressively trains the model from short-range perception to long-horizon multi-event reasoning; and (iii) Temporal Audio-Visual Interleaved Chain-of-Thought, a reasoning framework that explicitly grounds intermediate reasoning steps to timestamps in long audio-visual streams, improving temporal alignment and interpretability. Extensive experiments across 15+ audio-visual, omni-modal, audio, and vision benchmarks show that AV-Flamingo outperforms similarly sized open models by clear margins and remains highly competitive with, and in some cases surpasses, much larger open-weight and closed models, particularly on long and complex real-world audio-visual understanding and reasoning tasks. Beyond benchmark performance, AV-Flamingo exhibits strong real-world utility and transfers well to unseen tasks, highlighting its robustness and generalization ability.
Original Article
View Cached Full Text

Cached at: 07/20/26, 09:40 AM

Paper page - Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

Source: https://huggingface.co/papers/2607.16107 Authors:

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

,

Abstract

WepresentAudio-VisualFlamingo(AV-Flamingo),afullyopenstate-of-the-artaudio-visuallargelanguagemodel(AV-LLM)forjointunderstandingandreasoningoveraudio,images,andlong-formvideos.UnlikepriorAV-LLMsthatprimarilyfocusonshortclips,AV-Flamingoisdesignedforunderstandingandreasoningoverlongandcomplexreal-world(audio-visual)videos.Tosupportthis,wemakethreekeycontributions:(i)Audio-Visual-Skills,alarge-scalecollectionofreal-worldvideoswith~7Mcaptionandquestion-answertraininginstancesdesignedtoemphasizetemporal,compositional,andcross-modalaudio-visualreasoning;(ii)anovelthree-stagecurriculumthatprogressivelytrainsthemodelfromshort-rangeperceptiontolong-horizonmulti-eventreasoning;and(iii)TemporalAudio-VisualInterleavedChain-of-Thought,areasoningframeworkthatexplicitlygroundsintermediatereasoningstepstotimestampsinlongaudio-visualstreams,improvingtemporalalignmentandinterpretability.Extensiveexperimentsacross15+audio-visual,omni-modal,audio,andvisionbenchmarksshowthatAV-Flamingooutperformssimilarlysizedopenmodelsbyclearmarginsandremainshighlycompetitivewith,andinsomecasessurpasses,muchlargeropen-weightandclosedmodels,particularlyonlongandcomplexreal-worldaudio-visualunderstandingandreasoningtasks.Beyondbenchmarkperformance,AV-Flamingoexhibitsstrongreal-worldutilityandtransferswelltounseentasks,highlightingitsrobustnessandgeneralizationability.

View arXiv pageView PDFProject pageAdd to collection

Get this paper in your agent:

hf papers read 2607\.16107

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper1

#### nvidia/nemotron-labs-audio-visual-flamingo-hf Video-Text-to-Text• 9B• Updatedabout 5 hours ago • 19 • 4

Datasets citing this paper1

#### nvidia/av-skills Viewer• Updatedabout 7 hours ago • 23.6k • 10 • 2

Spaces citing this paper1

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Audio-Visual Intelligence in Large Foundation Models

Hugging Face Daily Papers

This survey paper provides a comprehensive review of audio-visual intelligence within large foundation models, establishing a unified taxonomy, synthesizing core methodologies, and outlining key datasets, benchmarks, and open research challenges.

VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding

Papers with Code Trending

VideoChat3 is a fully open, efficient, and generalist video-centric multimodal large language model that introduces Inflated 3D Vision Transformer (I3D-ViT) and Adaptive Frame Resolution for streaming video perception, along with scalable video data synthesis pipelines, achieving superior performance with only 4B parameters.