facebook/VGGT-Omega
摘要
Meta AI 和牛津大学 VGG 发布了 VGGT-Omega,这是一个用于 3D 视觉的基础模型,附有项目页面和 GitHub 仓库。
查看缓存全文
缓存时间: 2026/05/19 18:34
facebook/VGGT-Omega · Hugging Face
来源:https://huggingface.co/facebook/VGGT-Omega 项目页面(http://vggt-omega.github.io/)GitHub 仓库(https://github.com/facebookresearch/vggt-omega)
Meta AI Research(https://ai.facebook.com/research/);牛津大学 VGG(https://www.robots.ox.ac.uk/~vgg/)
Jianyuan Wang(https://jytime.github.io/),Minghao Chen(https://silent-chen.github.io/),Shangzhan Zhang(https://scholar.google.com/citations?user=FUDsZkEAAAAJ&hl=zh-CN),Nikita Karaev(https://nikitakaraevv.github.io/),Johannes Schönberger(https://demuc.de/),Patrick Labatut(https://scholar.google.com/citations?user=IJidh-UAAAAJ&hl=fr),Piotr Bojanowski(https://scholar.google.com/citations?user=lJ_oh2EAAAAJ&hl=en),David Novotny(https://d-novotny.github.io/),Andrea Vedaldi(https://www.robots.ox.ac.uk/~vedaldi/),Christian Rupprecht(https://chrirupp.github.io/)
https://huggingface.co/facebook/VGGT-Omega#quick-start快速开始
请参考我们的GitHub 仓库(https://github.com/facebookresearch/vggt-omega)
https://huggingface.co/facebook/VGGT-Omega#citation引用
如果您觉得我们的仓库有用,欢迎给 ⭐ 并引用我们的论文:
@inproceedings{wang2026vggtomega, title={VGGT-{$\Omega$}}, author={Wang, Jianyuan and Chen, Minghao and Zhang, Shangzhan and Karaev, Nikita and Sch{\"o}nberger, Johannes and Labatut, Patrick and Bojanowski, Piotr and Novotny, David and Vedaldi, Andrea and Rupprecht, Christian}, booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition}, year={2026} }
相似文章
发现了 Ox Alpha 背后的模型。它是未发布的 z.ai GLM 模型。
文章揭示了 Ox Alpha 基于 z.ai 的未发布 GLM 模型,通过分词器和图像编码器中匹配的 token 尺寸得到验证。
Nvidia Cosmos 3
NVIDIA 开源了 Cosmos 3,这是一个物理AI的前沿基础模型,将推理、世界生成和动作生成统一在单一的 Mixture-of-Transformers 架构中,并发布了用于机器人、自动驾驶和仓库监控的模型检查点、数据集和训练脚本。
LiquidAI/LFM2.5-VL-3B · Hugging Face
LiquidAI releases LFM2.5-VL-3B, a 3B multimodal model for on-device deployment with improved OCR, grounding, and efficient inference, available in multiple formats including GGUF, ONNX, and MLX.
open-gigaai/Giga-World-1
Giga-World-1 是 GigaAI Research 在 Hugging Face 上发布的视频生成模型,具备多个阶段检查点和用于场景控制的 LoRA 适配器。
@Modular: 官方确认:Ox Alpha 是 @Zai_org 的 GLM-5.3-Flash,GLM-5 系列中的首个原生多模态模型,拥有 320B 参数…
Z.ai 已正式发布 GLM-5.3-Flash,一款具有 320B 参数和 1M token 上下文窗口的原生多模态 AI 模型,此前以 Ox Alpha 预览,可通过 Modular Cloud 获取,并在 MIT 许可证下运行于中国 AI 芯片。