Meta携Muse Glimmer回归:本地、智能体、多模态且开源

Hugging Face Blog 模型

摘要

Meta发布了Muse Glimmer,这是一个基于Apache 2.0的30B多模态智能体模型,专为本地部署设计,并在Hugging Face库中提供首发支持。

暂无内容
查看原文
查看缓存全文

缓存时间: 2026/08/10 14:01

Meta带着Muse Glimmer回归:本地化、智能体化、多模态且开源

来源:https://huggingface.co/blog/muse-glimmer

返回文章列表 (https://huggingface.co/blog)

  • Benchmarks (https://huggingface.co/blog/muse-glimmer#benchmarks)
  • Architecture (https://huggingface.co/blog/muse-glimmer#architecture)
  • Text Decoder (https://huggingface.co/blog/muse-glimmer#text-decoder)
  • Perception Encoder (https://huggingface.co/blog/muse-glimmer#perception-encoder)
  • Transformers (https://huggingface.co/blog/muse-glimmer#transformers)
  • Text-only Inference (https://huggingface.co/blog/muse-glimmer#text-only-inference)
  • Prompting the model with images and text (https://huggingface.co/blog/muse-glimmer#prompting-the-model-with-images-and-text)
  • Video Inference (https://huggingface.co/blog/muse-glimmer#video-inference)
  • Multimodal tool calling (https://huggingface.co/blog/muse-glimmer#multimodal-tool-calling)
  • Object Detection (https://huggingface.co/blog/muse-glimmer#object-detection)
  • Llama.cpp (https://huggingface.co/blog/muse-glimmer#llamacpp)
  • Speculative Decoding (https://huggingface.co/blog/muse-glimmer#speculative-decoding)
  • Speculative Decoding with transformers (https://huggingface.co/blog/muse-glimmer#speculative-decoding-with-transformers)
  • Speculative Decoding with llama.cpp (https://huggingface.co/blog/muse-glimmer#speculative-decoding-with-llamacpp)
  • Inference Endpoints (https://huggingface.co/blog/muse-glimmer#inference-endpoints)
  • Support for Muse Glimmer vLLM with transformers backend (https://huggingface.co/blog/muse-glimmer#support-for-muse-glimmer-vllm-with-transformers-backend)
  • Fine-tuning with TRL (https://huggingface.co/blog/muse-glimmer#fine-tuning-with-trl)
  • Demos (https://huggingface.co/blog/muse-glimmer#demos)
  • Connect OpenClaw to Muse Glimmer (https://huggingface.co/blog/muse-glimmer#connect-openclaw-to-muse-glimmer)
  • Hey Muse Glimmer, quantize yourself (https://huggingface.co/blog/muse-glimmer#hey-muse-glimmer-quantize-yourself)
  • Hey Muse Glimmer, deploy yourself (https://huggingface.co/blog/muse-glimmer#hey-muse-glimmer-deploy-yourself)
  • Hey Muse Glimmer, optimize yourself (https://huggingface.co/blog/muse-glimmer#hey-muse-glimmer-optimize-yourself)
  • Hey Muse Glimmer, research the Hub (https://huggingface.co/blog/muse-glimmer#hey-muse-glimmer-research-the-hub)
  • Wrapping Up (https://huggingface.co/blog/muse-glimmer#wrapping-up)

开源大语言模型元老们的好消息!今日发布的 Muse Glimmer 是 Meta 的新多模态模型,专为本地智能体(agentic)应用场景设计。它从 Muse 蒸馏至 300 亿参数,并以 Apache 2.0 许可证发布,非常适合本地部署以保护隐私、降低成本,或者纯粹拿来折腾。它面向注重隐私的应用场景,如编码、文档分析、个人助理、Claw 或 Hermes 类设置。为了庆祝,我们与 Meta 合作在 transformers、llama.cpp、vLLM、Inference Endpoints 以及其他库中提供了首日支持。我们构建了一些很酷的东西,并在本博客中分享我们的发现。请查看下面的演示以获取灵感 (https://huggingface.co/blog/muse-glimmer#demos)。

你可以在 Hugging Face Hub 上找到 Muse Glimmer (https://huggingface.co/meta-models/Muse-Glimmer-30B)。

https://huggingface.co/blog/muse-glimmer#benchmarks Benchmarks

基准测试结果

分数按已发布的报告给出。粗体表示比较模型中的最佳结果;↓ 表示越低越好。

类别基准Muse Glimmer-30B 高推理Gemma4-31B 思考模式Qwen3.6-27B 思考模式
通用智能体MCP Atlas75.554.262.5
通用智能体DeepSearch QA74.661.771.1
通用智能体τ3-Banking23.515.116.7
通用智能体WildClawBench47.637.643.2
通用智能体GDPval-AA95381141
通用智能体GAIA243.336.440.0
通用智能体SkillsBench(带技能)44.332.446.6
通用智能体OSWorld-Verified65.958.575.6
智能体编码SWE-Bench Pro51.236.950.2
智能体编码SWE-Bench Verified76.066.677.2
智能体编码TerminalBench 2.151.743.460.7
智能体编码SciCode43.643.439.8
多模态Charxiv Reasoning78.877.778.4
多模态ScreenSpot Pro75.475.976.1
多模态OmniDocBench v1.575.872.577.8
多模态MMMU Pro747375
安全CI Memories 违规(↓):26.4 覆盖率:64.8违规(↓):12.1 覆盖率:53.0违规(↓):53.4 覆盖率:66.9
安全Siren AgentDojo 攻击成功率(↓):28.4 效用:94.2攻击成功率(↓):25.6 效用:90.8攻击成功率(↓):40.3 效用:92.7
通用能力与推理IFBench77.076.070.8
通用能力与推理AIME 202694.789.294.1
通用能力与推理GPQA Diamond83.585.784.2
通用能力与推理Humanity’s Last Exam(文本+无工具)22.023.623.1
通用能力与推理AA-LCR80.068.373.3
通用能力与推理Beam 128K65.158.263.0

https://huggingface.co/blog/muse-glimmer#architecture Architecture

Muse Glimmer 是一个稠密的 300 亿参数模型,由以下部分组成:

  • 用于视觉的 20 亿 ViT 风格编码器(感知编码器)
  • 280 亿参数文本解码器

除了主 VLM 之外,还有一个基于 DFlash 实现的投机解码草稿器。该模块的使用是可选的,它可以提供更快的生成速度,但会额外占用一些内存。我们发现这个草稿器特别适合结构化内容生成,比如编码。

https://huggingface.co/blog/muse-glimmer#text-decoder Text Decoder

该语言模型使用以下架构组件:

  • 混合注意力: 交替使用三个滑动窗口层(2048 个 token),采用旋转位置嵌入,随后是使用全注意力和 NoPE(无位置嵌入)的第四层。因此模式为(SWA, SWA, SWA, Full),重复 13 次,共 52 层。这使模型能够通过 RoPE 保留相对顺序和距离信息,并通过 NoPE 全局保留信息。
  • 门控分组查询注意力: 每个键值头由 16 个查询头共享,这将 KV 缓存内存减少了 16 倍,并使生成更快、更便宜。
  • Q-K 归一化与额外的查询缩放: 在计算注意力之前,Muse Glimmer 对每个查询头和键头应用 RMS 归一化以保持注意力 logits 稳定。之后,查询会乘以一个缩放因子,以在归一化后设置目标 logit 缩放。额外的查询缩放行为类似于 softmax 层面的逆温度。

https://huggingface.co/blog/muse-glimmer#perception-encoder Perception Encoder

Muse Glimmer 使用一个图像编码器来处理图像和视频。与其他 VLM 中使用的相对较小的视觉编码器不同,这是一个体积较大的 20 亿参数 ViT 类模型,设计基于 Perception Encoder 架构。Perception Encoder 此前由 Meta 作为各种下游空间和多模态任务的主干引入 (https://huggingface.co/papers/2504.13181)。该编码器将图像切成 2 帧 x 3 通道 x 14 x 14 的块,并通过线性层进行投影。然后,从学习到的位置表中插值的绝对位置嵌入会被添加到这些嵌入中。接着这些嵌入被送入由 50 层和 GELU MLP 组成的视觉塔。与语言模型类似,注意力模式由三个窗口注意力层和一个全注意力层组成。在注意力层内部,对查询和键应用 2D RoPE。在 transformer 之后,像素洗牌(pixel shuffle)将相邻空间 token 的 2x2 组合并,使图像 token 数量减少 4 倍,同时不会丢失通道信息。合并后的特征随后被投影到文本解码器的共享嵌入空间。

视频逐帧通过相同的编码器,其中每一帧被转换为块(形状为 [batch, temporal groups, grid height, grid width, 2 frames, 3 channels, 14, 14])。处理器目标为每秒 2 帧,并将片段上限设为 96 帧,在整个视频中均匀采样。处理器会创建带时间戳的视频占位符,将文本与帧交错,例如“Time: 0.0s <|video|> x N”,在最终投影层之前会替换为最终的视频嵌入。

https://huggingface.co/blog/muse-glimmer#transformers Transformers

升级 transformers 到最新版本以使用 Muse Glimmer。

pip install --upgrade transformers accelerate

Muse Glimmer 在 transformers 中拥有首日支持,无论是主模型还是投机解码草稿器。你可以使用 AutoModelForMultimodalLM 和 AutoProcessor 类来加载模型和处理器。

from transformers import AutoProcessor, AutoModelForMultimodalLM

MODEL_ID = "meta-models/Muse-Glimmer-30B"

# Load model
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
    MODEL_ID,
    dtype="auto",
    device_map="auto"
)

相同的代码片段可以在 NVIDIA(CUDA)、AMD(ROCm)和 Intel(XPU)GPU 上不加修改地运行,device_map="auto" 会将模型放置在可用的加速器上。

https://huggingface.co/blog/muse-glimmer#text-only-inference Text-only Inference

加载模型后,你可以通过如下方式执行纯文本推理。

from transformers import AutoProcessor, AutoModelForMultimodalLM

MODEL_ID = "meta-models/Muse-Glimmer-30B"

# Load model
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
    MODEL_ID,
    dtype="auto",
    device_map="auto"
)

# Prompt
messages = [
    {"role": "user", "content": "Write a short joke about saving RAM."},
]

# Process input
inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
    add_generation_prompt=True,
    reasoning_strength="low"
).to(model.device)

input_len = inputs["input_ids"].shape[-1]

# Generate output
outputs = model.generate(**inputs)
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)
print(response)

https://huggingface.co/blog/muse-glimmer#prompting-the-model-with-images-and-text Prompting the model with images and text

我们需要 torchvision 才能使用图像和文本。

pip install torchvision

Muse Glimmer 接受图像作为输入,如下所示:

from transformers import AutoProcessor, AutoModelForMultimodalLM

MODEL_ID = "meta-models/Muse-Glimmer-30B"

# Load model
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
    MODEL_ID,
    dtype="auto",
    device_map="auto"
)

# Images + Text
messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "image": "https://huggingface.co/datasets/merve/vl-test-suite/resolve/main/SF.png"},
            {"type": "text", "text": "What is shown in this image?"}
        ]
    }
]

inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
    add_generation_prompt=True,
    reasoning_strength="low"
).to(model.device)

input_len = inputs["input_ids"].shape[-1]

# Generate output
outputs = model.generate(**inputs)
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)
print(response)

https://huggingface.co/blog/muse-glimmer#video-inference Video Inference

要处理视频,我们建议在环境中安装 torchcodec。

pip install torchcodec

Muse Glimmer 可以回答关于视频不带音频的复杂问题。你可以像下面这样进行视频推理,这是一个来自 VideoMME2 的示例,它是目前最受欢迎的视频问答基准。

from transformers import AutoProcessor, AutoModelForMultimodalLM

MODEL_ID = "meta-models/Muse-Glimmer-30B"

# Load model
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
    MODEL_ID,
    dtype="auto",
    device_map="auto"
)

# Videos + Text
messages = [
    {
        "role": "user",
        "content": [
            {"type": "video", "video": "https://huggingface.co/datasets/merve/vl-test-suite/resolve/main/IMG_8137.mp4"},
            {"type": "text", "text": "Describe what happens in this video."},
        ],
    },
]

inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
    add_generation_prompt=True,
    reasoning_strength="low",
    processor_kwargs={"num_frames": 96},
).to(model.device)

input_len = inputs["input_ids"].shape[-1]

outputs = model.generate(**inputs)
response = processor.decode(
    outputs[0, input_len:],
    skip_special_tokens=False,
)
print(response)

https://huggingface.co/blog/muse-glimmer#multimodal-tool-calling Multimodal tool calling

Muse Glimmer 可以进行多模态工具调用,下面是如何操作。在以下示例中,我们要求模型根据图像中的城市调用天气工具。

import json
import re
from transformers import AutoProcessor, AutoModelForMultimodalLM

MODEL_ID = "meta-models/Muse-Glimmer-30B"

# Load model
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
    MODEL_ID,
    dtype="auto",
    device_map="auto"
)

tools = [
    {
        "type": "function",
        "function": {
            "name": "weather.get",
            "description": "Get the current weather for a city.",
            "parameters": {
                "type": "object",
                "properties": {
                    "city": {"type": "string"},
                },
                "required": ["city"],
            },
        },
    }
]

messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "image": "https://huggingface.co/datasets/merve/vl-test-suite/resolve/main/SF.png"},
            {"type": "text", "text": "I'm going to the city in this picture. What clothes should I wear?"},
        ],
    },
]

inputs = processor.apply_chat_template(
    messages,
    tools=tools,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
    add_generation_prompt=True,
    reasoning_strength="low"
).to(model.device)

input_len = inputs["input_ids"].shape[-1]

outputs = model.generate(**inputs)
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)
print(response)

https://huggingface.co/blog/muse-glimmer#object-detection Object Detection

你可以使用 Muse Glimmer 在图像中进行开放式对象检测,如下所示。

from transformers import AutoProcessor, AutoModelForMultimodalLM

MODEL_ID = "meta-models/Muse-Glimmer-30B"

# Load model
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
    MODEL_ID,
    dtype="auto",
    device_map="auto"
)

messages = [{
    "role": "user",
    "content": [
        {"type": "image", "image": "https://huggingface.co/datasets/merve/vl-test-suite/resolve/main/SF.png"},
        {
            "type": "text",
            "text": (
                "Detect the bridge. Return only the detection in the model's "
                "native object-detection format, with no explanation."
            ),
        },
    ],
}]

inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
    add_generation_prompt=True,
    reasoning_strength="low",
).to(model.device)

input_len = inputs["input_ids"].shape[-1]

outputs = model.generate(**inputs, max_new_tokens=128)
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)
detections = json.loads(response.removesuffix("<|eot|>"))
print(detections)

这里有一个执行对象检测的端到端脚本:GitHub Gist (https://gist.github.com/ariG23498/4f3587eb7753c0ff77c269e2c1efe2c0)

https://huggingface.co/blog/muse-glimmer#llamacpp Llama.cpp

Muse Glimmer 拥有首日 llama.cpp 支持。Meta 已在此仓库中发布了校准量化版本 (https://huggingface.co/meta-models/Muse-Glimmer-30B-GGUF),Unsloth 也在发布优化的量化版本。同时也支持 DFlash 投机解码。

你可以使用预构建的 llama 二进制文件来启动 llama 服务器或 CLI。要安装 llama.cpp,请运行:

curl -LsSf https://llama.app/install.sh | sh

然后你可以按如下方式启动服务器:

llama serve -hf meta-models/Muse-Glimmer-30B-GGUF

服务器启动后,你可以访问 localhost:8080 与内置 WebUI 聊天。你还可以

相似文章

Meta 开源 Muse Glimmer 30B Agent 模型,扎克伯格推动个人超级智能愿景

Reddit r/ArtificialInteligence

Meta 开源了 Muse Glimmer,这是一个采用 Apache 2.0 协议的 30B 参数多模态 Agent 模型,针对本地工具使用和编码进行了优化,通过 4-bit 量化压缩至 20GB 以下,可在消费级 GPU 上运行。扎克伯格还承诺开源 Muse Spark 1.2 的权重,并设立 10 亿美元社区基金用于数据中心所在地区。

推出 Muse Glimmer

Simon Willison's Blog

Meta 推出 Muse Glimmer,一款基于 Apache 2.0 协议的全新 30B 开放权重模型,针对智能体任务完成、可靠工具使用和多步推理进行了优化。Simon Willison 使用 LM Studio 和 llm-coding-agent 在本地对其进行了测试。