VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
Summary
VideoChat3 is a fully open, efficient, and generalist video-centric multimodal large language model that introduces Inflated 3D Vision Transformer (I3D-ViT) and Adaptive Frame Resolution for streaming video perception, along with scalable video data synthesis pipelines, achieving superior performance with only 4B parameters.
View Cached Full Text
Cached at: 07/20/26, 09:38 AM
Paper page - VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
Source: https://huggingface.co/papers/2607.14935 Published on Jul 16
#2 Paper of the day Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
Abstract
Recentadvancesinvideounderstandinghavespannedmotion,longvideo,andstreaminginteraction,drivingthisfieldtowardreal-worldapplications.Despitethisprogress,currentopen-sourcemodelsremainlimitedinseveralways.Theyoftenstruggletogeneralizeacrossdiversevideotypes,makingthemeffectiveonlyinspecificdomains.Highcomputationaldemandsfurtherrestricttheirefficiencyandscalability.Moreover,mostmodelsareonlypartiallyopen,withkeycomponentssuchastrainingcode,strategy,ordatasetsunavailable,whichhindersreproducibilityandslowscommunity-drivendevelopment.Toaddresstheseissues,weintroduceVideoChat3,afullyopen,efficient,andgeneralistvideo-centricMLLM.VideoChat3advancesvideounderstandingthroughtwocomplementarydesigns.Forefficiency,weintroduceInflated3DVisionTransformer(I3D-ViT)andAdaptiveFrameResolutionforStreamingVideoPerception,whichenablesefficientspatiotemporalrepresentationandreducesthecostofprocessingvideoinputsduringtrainingandinference.Foreffectiveness,wedevelopascalablevideodatasynthesispipelinethatcuratesthreediverse,high-qualitytrainingdatasets:VideoChat3-Academic2M,VideoChat3-LV116K,andVideoChat3-OL617K,coveringgeneral,long-form,andstreamingvideoscenarios,improvingthemodel’sgeneralizationacrossdomains.Byintegratingthesedesigns,VideoChat3achievesararebalanceofbroadgeneralizationandcomputationalefficiency.Experimentsacrossgeneral,long-form,andstreamingbenchmarksdemonstratethatVideoChat3surpassesprioropen-sourcemodelswithequalorlargerparametercountswithonly4Bparametersandhigherefficiency.
View arXiv pageView PDFProject pageGitHub108Add to collection
Get this paper in your agent:
hf papers read 2607\.14935
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper2
#### MCG-NJU/VideoChat3-4B Video-Text-to-Text• 4B• Updated3 days ago • 342 • 13
#### MCG-NJU/I3D-ViT Image Feature Extraction• 0.4B• Updated3 days ago • 62 • 8
Datasets citing this paper3
#### MCG-NJU/VideoChat3-LV116k Viewer• Updatedabout 21 hours ago • 8.07k • 7.59k • 10 #### MCG-NJU/VideoChat3-Academic2M Viewer• Updatedabout 21 hours ago • 19.2k • 1.98k • 14 #### MCG-NJU/VideoChat3-OL617k Preview• Updated3 days ago • 280 • 8
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.14935 in a Space README.md to link it from this page.
Collections including this paper6
Similar Articles
LoomVideo: Unifying Multimodal Inputs into Video Generation and Editing
LoomVideo introduces a 5B-parameter unified architecture for video generation and editing that reduces computational overhead using novel conditioning mechanisms and multi-modal alignment, achieving competitive performance and faster inference.
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
Mage-VL is an efficient codec-native streaming multimodal foundation model that reduces visual token consumption by over 75% using a custom tokenizer, achieving up to 3.5x inference speedup while matching or outperforming existing models on static and video tasks.
Watch, Remember, Reason: Human-View Video Understanding with MLLMs
A survey presenting a human-view perspective on video understanding with multimodal large language models, organized around watching, remembering, and reasoning abilities, covering challenges, methods, and applications.
InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning
InternVideo3 introduces Multimodal Contextual Reasoning (MCR) and efficient attention mechanisms to enhance long-horizon multimodal tasks, achieving strong results on video understanding benchmarks and demonstrating video agent capabilities.
@_akhaliq: LingBot-Video is out on Hugging Face MoE-based video foundation model built for embodied intelligence 30B params, only …
LingBot-Video, a 30B parameter MoE-based video foundation model for embodied intelligence, has been released on Hugging Face with only 3B active parameters at inference, augmented with 70K hours of embodied data.