Meta携Muse Glimmer回归:本地、智能体、多模态且开源

Hugging Face Blog 模型

摘要

Meta发布了Muse Glimmer,这是一个基于Apache 2.0的30B多模态智能体模型,专为本地部署设计,并在Hugging Face库中提供首发支持。

暂无内容
查看原文
查看缓存全文

缓存时间: 2026/08/10 14:01

Meta带着Muse Glimmer回归:本地化、智能体化、多模态且开源

来源:https://huggingface.co/blog/muse-glimmer

返回文章列表 (https://huggingface.co/blog)

  • Benchmarks (https://huggingface.co/blog/muse-glimmer#benchmarks)
  • Architecture (https://huggingface.co/blog/muse-glimmer#architecture)
  • Text Decoder (https://huggingface.co/blog/muse-glimmer#text-decoder)
  • Perception Encoder (https://huggingface.co/blog/muse-glimmer#perception-encoder)
  • Transformers (https://huggingface.co/blog/muse-glimmer#transformers)
  • Text-only Inference (https://huggingface.co/blog/muse-glimmer#text-only-inference)
  • Prompting the model with images and text (https://huggingface.co/blog/muse-glimmer#prompting-the-model-with-images-and-text)
  • Video Inference (https://huggingface.co/blog/muse-glimmer#video-inference)
  • Multimodal tool calling (https://huggingface.co/blog/muse-glimmer#multimodal-tool-calling)
  • Object Detection (https://huggingface.co/blog/muse-glimmer#object-detection)
  • Llama.cpp (https://huggingface.co/blog/muse-glimmer#llamacpp)
  • Speculative Decoding (https://huggingface.co/blog/muse-glimmer#speculative-decoding)
  • Speculative Decoding with transformers (https://huggingface.co/blog/muse-glimmer#speculative-decoding-with-transformers)
  • Speculative Decoding with llama.cpp (https://huggingface.co/blog/muse-glimmer#speculative-decoding-with-llamacpp)
  • Inference Endpoints (https://huggingface.co/blog/muse-glimmer#inference-endpoints)
  • Support for Muse Glimmer vLLM with transformers backend (https://huggingface.co/blog/muse-glimmer#support-for-muse-glimmer-vllm-with-transformers-backend)
  • Fine-tuning with TRL (https://huggingface.co/blog/muse-glimmer#fine-tuning-with-trl)
  • Demos (https://huggingface.co/blog/muse-glimmer#demos)
  • Connect OpenClaw to Muse Glimmer (https://huggingface.co/blog/muse-glimmer#connect-openclaw-to-muse-glimmer)
  • Hey Muse Glimmer, quantize yourself (https://huggingface.co/blog/muse-glimmer#hey-muse-glimmer-quantize-yourself)
  • Hey Muse Glimmer, deploy yourself (https://huggingface.co/blog/muse-glimmer#hey-muse-glimmer-deploy-yourself)
  • Hey Muse Glimmer, optimize yourself (https://huggingface.co/blog/muse-glimmer#hey-muse-glimmer-optimize-yourself)
  • Hey Muse Glimmer, research the Hub (https://huggingface.co/blog/muse-glimmer#hey-muse-glimmer-research-the-hub)
  • Wrapping Up (https://huggingface.co/blog/muse-glimmer#wrapping-up)

开源大语言模型元老们的好消息!今日发布的 Muse Glimmer 是 Meta 的新多模态模型,专为本地智能体(agentic)应用场景设计。它从 Muse 蒸馏至 300 亿参数,并以 Apache 2.0 许可证发布,非常适合本地部署以保护隐私、降低成本,或者纯粹拿来折腾。它面向注重隐私的应用场景,如编码、文档分析、个人助理、Claw 或 Hermes 类设置。为了庆祝,我们与 Meta 合作在 transformersllama.cppvLLM、Inference Endpoints 以及其他库中提供了首日支持。我们构建了一些很酷的东西,并在本博客中分享我们的发现。请查看下面的演示以获取灵感 (https://huggingface.co/blog/muse-glimmer#demos)。

你可以在 Hugging Face Hub 上找到 Muse Glimmer (https://huggingface.co/meta-models/Muse-Glimmer-30B)。

https://huggingface.co/blog/muse-glimmer#benchmarks Benchmarks

基准测试结果

分数按已发布的报告给出。粗体表示比较模型中的最佳结果;↓ 表示越低越好。

类别基准Muse Glimmer-30B 高推理Gemma4-31B 思考模式Qwen3.6-27B 思考模式
通用智能体MCP Atlas75.554.262.5
通用智能体DeepSearch QA74.661.771.1
通用智能体τ3-Banking23.515.116.7
通用智能体WildClawBench47.637.643.2
通用智能体GDPval-AA95381141
通用智能体GAIA243.336.440.0
通用智能体SkillsBench(带技能)44.332.446.6
通用智能体OSWorld-Verified65.958.575.6
智能体编码SWE-Bench Pro51.236.950.2
智能体编码SWE-Bench Verified76.066.677.2
智能体编码TerminalBench 2.151.743.460.7
智能体编码SciCode43.643.439.8
多模态Charxiv Reasoning78.877.778.4
多模态ScreenSpot Pro75.475.976.1
多模态OmniDocBench v1.575.872.577.8
多模态MMMU Pro747375
安全CI Memories 违规(↓):26.4 覆盖率:64.8违规(↓):12.1 覆盖率:53.0违规(↓):53.4 覆盖率:66.9
安全Siren AgentDojo 攻击成功率(↓):28.4 效用:94.2攻击成功率(↓):25.6 效用:90.8攻击成功率(↓):40.3 效用:92.7
通用能力与推理IFBench77.076.070.8
通用能力与推理AIME 202694.789.294.1
通用能力与推理GPQA Diamond83.585.784.2
通用能力与推理Humanity’s Last Exam(文本+无工具)22.023.623.1
通用能力与推理AA-LCR80.068.373.3
通用能力与推理Beam 128K65.158.263.0

https://huggingface.co/blog/muse-glimmer#architecture Architecture

Muse Glimmer 是一个稠密的 300 亿参数模型,由以下部分组成:

  • 用于视觉的 20 亿 ViT 风格编码器(感知编码器)
  • 280 亿参数文本解码器

除了主 VLM 之外,还有一个基于 DFlash 实现的投机解码草稿器。该模块的使用是可选的,它可以提供更快的生成速度,但会额外占用一些内存。我们发现这个草稿器特别适合结构化内容生成,比如编码。

https://huggingface.co/blog/muse-glimmer#text-decoder Text Decoder

该语言模型使用以下架构组件:

  • 混合注意力: 交替使用三个滑动窗口层(2048 个 token),采用旋转位置嵌入,随后是使用全注意力和 NoPE(无位置嵌入)的第四层。因此模式为(SWA, SWA, SWA, Full),重复 13 次,共 52 层。这使模型能够通过 RoPE 保留相对顺序和距离信息,并通过 NoPE 全局保留信息。
  • 门控分组查询注意力: 每个键值头由 16 个查询头共享,这将 KV 缓存内存减少了 16 倍,并使生成更快、更便宜。
  • Q-K 归一化与额外的查询缩放: 在计算注意力之前,Muse Glimmer 对每个查询头和键头应用 RMS 归一化以保持注意力 logits 稳定。之后,查询会乘以一个缩放因子,以在归一化后设置目标 logit 缩放。额外的查询缩放行为类似于 softmax 层面的逆温度。

https://huggingface.co/blog/muse-glimmer#perception-encoder Perception Encoder

Muse Glimmer 使用一个图像编码器来处理图像和视频。与其他 VLM 中使用的相对较小的视觉编码器不同,这是一个体积较大的 20 亿参数 ViT 类模型,设计基于 Perception Encoder 架构。Perception Encoder 此前由 Meta 作为各种下游空间和多模态任务的主干引入 (https://huggingface.co/papers/2504.13181)。该编码器将图像切成 2 帧 x 3 通道 x 14 x 14 的块,并通过线性层进行投影。然后,从学习到的位置表中插值的绝对位置嵌入会被添加到这些嵌入中。接着这些嵌入被送入由 50 层和 GELU MLP 组成的视觉塔。与语言模型类似,注意力模式由三个窗口注意力层和一个全注意力层组成。在注意力层内部,对查询和键应用 2D RoPE。在 transformer 之后,像素洗牌(pixel shuffle)将相邻空间 token 的 2x2 组合并,使图像 token 数量减少 4 倍,同时不会丢失通道信息。合并后的特征随后被投影到文本解码器的共享嵌入空间。

视频逐帧通过相同的编码器,其中每一帧被转换为块(形状为 [batch, temporal groups, grid height, grid width, 2 frames, 3 channels, 14, 14])。处理器目标为每秒 2 帧,并将片段上限设为 96 帧,在整个视频中均匀采样。处理器会创建带时间戳的视频占位符,将文本与帧交错,例如“Time: 0.0s <|video|> x N”,在最终投影层之前会替换为最终的视频嵌入。

https://huggingface.co/blog/muse-glimmer#transformers Transformers

升级 transformers 到最新版本以使用 Muse Glimmer。

pip install --upgrade transformers accelerate

Muse Glimmer 在 transformers 中拥有首日支持,无论是主模型还是投机解码草稿器。你可以使用 AutoModelForMultimodalLMAutoProcessor 类来加载模型和处理器。

from transformers import AutoProcessor, AutoModelForMultimodalLM

MODEL_ID = "meta-models/Muse-Glimmer-30B"

# Load model
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
    MODEL_ID,
    dtype="auto",
    device_map="auto"
)

相同的代码片段可以在 NVIDIA(CUDA)、AMD(ROCm)和 Intel(XPU)GPU 上不加修改地运行,device_map="auto" 会将模型放置在可用的加速器上。

https://huggingface.co/blog/muse-glimmer#text-only-inference Text-only Inference

加载模型后,你可以通过如下方式执行纯文本推理。

from transformers import AutoProcessor, AutoModelForMultimodalLM

MODEL_ID = "meta-models/Muse-Glimmer-30B"

# Load model
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
    MODEL_ID,
    dtype="auto",
    device_map="auto"
)

# Prompt
messages = [
    {"role": "user", "content": "Write a short joke about saving RAM."},
]

# Process input
inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
    add_generation_prompt=True,
    reasoning_strength="low"
).to(model.device)

input_len = inputs["input_ids"].shape[-1]

# Generate output
outputs = model.generate(**inputs)
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)
print(response)

https://huggingface.co/blog/muse-glimmer#prompting-the-model-with-images-and-text Prompting the model with images and text

我们需要 torchvision 才能使用图像和文本。

pip install torchvision

Muse Glimmer 接受图像作为输入,如下所示:

from transformers import AutoProcessor, AutoModelForMultimodalLM

MODEL_ID = "meta-models/Muse-Glimmer-30B"

# Load model
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
    MODEL_ID,
    dtype="auto",
    device_map="auto"
)

# Images + Text
messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "image": "https://huggingface.co/datasets/merve/vl-test-suite/resolve/main/SF.png"},
            {"type": "text", "text": "What is shown in this image?"}
        ]
    }
]

inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
    add_generation_prompt=True,
    reasoning_strength="low"
).to(model.device)

input_len = inputs["input_ids"].shape[-1]

# Generate output
outputs = model.generate(**inputs)
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)
print(response)

https://huggingface.co/blog/muse-glimmer#video-inference Video Inference

要处理视频,我们建议在环境中安装 torchcodec

pip install torchcodec

Muse Glimmer 可以回答关于视频不带音频的复杂问题。你可以像下面这样进行视频推理,这是一个来自 VideoMME2 的示例,它是目前最受欢迎的视频问答基准。

from transformers import AutoProcessor, AutoModelForMultimodalLM

MODEL_ID = "meta-models/Muse-Glimmer-30B"

# Load model
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
    MODEL_ID,
    dtype="auto",
    device_map="auto"
)

# Videos + Text
messages = [
    {
        "role": "user",
        "content": [
            {"type": "video", "video": "https://huggingface.co/datasets/merve/vl-test-suite/resolve/main/IMG_8137.mp4"},
            {"type": "text", "text": "Describe what happens in this video."},
        ],
    },
]

inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
    add_generation_prompt=True,
    reasoning_strength="low",
    processor_kwargs={"num_frames": 96},
).to(model.device)

input_len = inputs["input_ids"].shape[-1]

outputs = model.generate(**inputs)
response = processor.decode(
    outputs[0, input_len:],
    skip_special_tokens=False,
)
print(response)

https://huggingface.co/blog/muse-glimmer#multimodal-tool-calling Multimodal tool calling

Muse Glimmer 可以进行多模态工具调用,下面是如何操作。在以下示例中,我们要求模型根据图像中的城市调用天气工具。

import json
import re
from transformers import AutoProcessor, AutoModelForMultimodalLM

MODEL_ID = "meta-models/Muse-Glimmer-30B"

# Load model
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
    MODEL_ID,
    dtype="auto",
    device_map="auto"
)

tools = [
    {
        "type": "function",
        "function": {
            "name": "weather.get",
            "description": "Get the current weather for a city.",
            "parameters": {
                "type": "object",
                "properties": {
                    "city": {"type": "string"},
                },
                "required": ["city"],
            },
        },
    }
]

messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "image": "https://huggingface.co/datasets/merve/vl-test-suite/resolve/main/SF.png"},
            {"type": "text", "text": "I'm going to the city in this picture. What clothes should I wear?"},
        ],
    },
]

inputs = processor.apply_chat_template(
    messages,
    tools=tools,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
    add_generation_prompt=True,
    reasoning_strength="low"
).to(model.device)

input_len = inputs["input_ids"].shape[-1]

outputs = model.generate(**inputs)
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)
print(response)

https://huggingface.co/blog/muse-glimmer#object-detection Object Detection

你可以使用 Muse Glimmer 在图像中进行开放式对象检测,如下所示。

from transformers import AutoProcessor, AutoModelForMultimodalLM

MODEL_ID = "meta-models/Muse-Glimmer-30B"

# Load model
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
    MODEL_ID,
    dtype="auto",
    device_map="auto"
)

messages = [{
    "role": "user",
    "content": [
        {"type": "image", "image": "https://huggingface.co/datasets/merve/vl-test-suite/resolve/main/SF.png"},
        {
            "type": "text",
            "text": (
                "Detect the bridge. Return only the detection in the model's "
                "native object-detection format, with no explanation."
            ),
        },
    ],
}]

inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
    add_generation_prompt=True,
    reasoning_strength="low",
).to(model.device)

input_len = inputs["input_ids"].shape[-1]

outputs = model.generate(**inputs, max_new_tokens=128)
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)
detections = json.loads(response.removesuffix("<|eot|>"))
print(detections)

这里有一个执行对象检测的端到端脚本:GitHub Gist (https://gist.github.com/ariG23498/4f3587eb7753c0ff77c269e2c1efe2c0)

https://huggingface.co/blog/muse-glimmer#llamacpp Llama.cpp

Muse Glimmer 拥有首日 llama.cpp 支持。Meta 已在此仓库中发布了校准量化版本 (https://huggingface.co/meta-models/Muse-Glimmer-30B-GGUF),Unsloth 也在发布优化的量化版本。同时也支持 DFlash 投机解码。

你可以使用预构建的 llama 二进制文件来启动 llama 服务器或 CLI。要安装 llama.cpp,请运行:

curl -LsSf https://llama.app/install.sh | sh

然后你可以按如下方式启动服务器:

llama serve -hf meta-models/Muse-Glimmer-30B-GGUF

服务器启动后,你可以访问 localhost:8080 与内置 WebUI 聊天。你还可以

相似文章

Meta 开源 Muse Glimmer 30B Agent 模型,扎克伯格推动个人超级智能愿景

Reddit r/ArtificialInteligence

Meta 开源了 Muse Glimmer,这是一个采用 Apache 2.0 协议的 30B 参数多模态 Agent 模型,针对本地工具使用和编码进行了优化,通过 4-bit 量化压缩至 20GB 以下,可在消费级 GPU 上运行。扎克伯格还承诺开源 Muse Spark 1.2 的权重,并设立 10 亿美元社区基金用于数据中心所在地区。