facebook/VGGT-Omega
Summary
Meta AI and Oxford VGG released VGGT-Omega, a foundation model for 3D vision, with project page and GitHub repository.
View Cached Full Text
Cached at: 05/19/26, 06:34 PM
facebook/VGGT-Omega · Hugging Face
Source: https://huggingface.co/facebook/VGGT-Omega
Meta AI Research;University of Oxford, VGG
Jianyuan Wang,Minghao Chen,Shangzhan Zhang,Nikita Karaev, Johannes Schönberger,Patrick Labatut,Piotr Bojanowski,David Novotny, Andrea Vedaldi,Christian Rupprecht
https://huggingface.co/facebook/VGGT-Omega#quick-startQuick Start
Please refer to ourGithub Repo
https://huggingface.co/facebook/VGGT-Omega#citationCitation
If you find our repository useful, please consider giving it a star ⭐ and citing our paper in your work:
@inproceedings{wang2026vggtomega,
title={VGGT-{$\Omega$}},
author={Wang, Jianyuan and Chen, Minghao and Zhang, Shangzhan and Karaev, Nikita and Sch{\"o}nberger, Johannes and Labatut, Patrick and Bojanowski, Piotr and Novotny, David and Vedaldi, Andrea and Rupprecht, Christian},
booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year={2026}
}
Similar Articles
Found out the model behind Ox Alpha. It's unreleased z.ai's GLM model
The article reveals that Ox Alpha is based on z.ai's unreleased GLM model, verified through matching token sizes in the tokenizer and image encoder.
Nvidia Cosmos 3
NVIDIA has open-sourced Cosmos 3, a frontier foundation model for physical AI that unifies reasoning, world generation, and action generation within a single Mixture-of-Transformers architecture, releasing model checkpoints, datasets, and training scripts for robotics, autonomous vehicles, and warehouse monitoring.
LiquidAI/LFM2.5-VL-3B · Hugging Face
LiquidAI releases LFM2.5-VL-3B, a 3B multimodal model for on-device deployment with improved OCR, grounding, and efficient inference, available in multiple formats including GGUF, ONNX, and MLX.
open-gigaai/Giga-World-1
Giga-World-1 is a video generation model released by GigaAI Research on Hugging Face, featuring multiple stage checkpoints and LoRA adapters for scene control.
@Modular: It's official: Ox Alpha is @Zai_org's GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series, with 320B…
Z.ai has officially released GLM-5.3-Flash, a natively multimodal AI model with 320B parameters and a 1M-token context window, previously previewed as Ox Alpha, available via Modular Cloud and running on Chinese AI chips under an MIT License.