@UnslothAI:Gemma 4 12B 现在可以通过 Dynamic GGUFs 在仅 8GB 内存上本地运行。Google 的新模型 Gemma 4 12B Unified 支持图像…

X AI KOLs Timeline 模型

摘要

Gemma 4 12B,Google 的多模态开放模型,支持图像、音频和 256K 上下文,现在可以通过 Unsloth 的 Dynamic GGUFs 在仅 8GB 内存上本地运行,并通过 Unsloth Studio 实现本地训练和推理。

Gemma 4 12B 现在可以通过 Dynamic GGUFs 在仅 8GB 内存上本地运行。Google 的新模型 Gemma 4 12B Unified 支持图像、音频和 256K 上下文。您可以通过 Unsloth Studio 运行和训练该模型。GGUF: https://huggingface.co/unsloth/gemma-4-12b-it-GGUF… 指南: https://unsloth.ai/docs/models/gemma-4…
查看原文
查看缓存全文

缓存时间: 2026/06/03 19:53

Gemma 4 12B 现在可以通过 Dynamic GGUFs 在仅 8GB RAM 的本地设备上运行。Google 的新模型 Gemma 4 12B Unified 支持图像、音频和 256K 上下文。你可以通过 Unsloth Studio 运行和训练该模型。GGUF:https://huggingface.co/unsloth/gemma-4-12b-it-GGUF… 指南:https://unsloth.ai/docs/models/gemma-4… —

unsloth/gemma-4-12b-it-GGUF · Hugging Face

来源:https://huggingface.co/unsloth/gemma-4-12b-it-GGUF

https://huggingface.co/unsloth/gemma-4-12b-it-GGUF#read-our-how-to-run-gemma-4-12b-guide

阅读我们的《如何运行 Gemma 4 12B 指南》!(https://docs.unsloth.ai/models/gemma-4)

gemma 4 在 unsloth studio 中 —

Hugging Face (https://huggingface.co/collections/google/gemma-4)|GitHub (https://github.com/google-gemma)|Launch Blog (https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12B/)|Documentation (https://ai.google.dev/gemma/docs/core)

License:Apache 2.0 (https://ai.google.dev/gemma/docs/gemma_4_license)|Authors:Google DeepMind (https://deepmind.google/models/gemma/)

此模型卡片适用于 Gemma 4 12B Unified 模型,该模型是 Gemma 4 系列开放模型的一部分。它采用与 Gemma 4 E2B 和 E4B(文本、音频、图像和视频输入)相同的多模态功能,将原生音频和视觉理解直接引入本地环境,无需单独的编码器。这种统一的多模态方法使模型无需编码器,部署规模非常适合消费类设备和简化的本地执行。

Gemma 是 Google DeepMind 构建的开放模型系列。Gemma 4 模型是多模态的,处理文本和图像输入(E2B、E4B 和 12B 支持音频)并生成文本输出。此版本包括预训练和指令调优两种变体的开放权重模型。Gemma 4 拥有高达 256K 令牌的上下文窗口,并支持超过 140 种语言的多语言能力。Gemma 4 兼具 Dense 和 Mixture-of-Experts (MoE) 架构,非常适合文本生成、编码和推理等任务。模型提供五种不同尺寸:E2B、E4B、12B、26B A4B 和 31B。其多样化的尺寸使其可部署于从高端手机到笔记本电脑和服务器的各种环境,让尖端人工智能的获取更加民主化。

Gemma 4 引入了关键的能力和架构进步:

  • 推理——系列中的所有模型都设计为高能力的推理器,具有可配置的思考模式。
  • 扩展的多模态能力——处理文本、具有可变纵横比和分辨率支持的图像(所有模型)、视频以及音频(E2B、E4B 和 12B 模型原生支持)。
  • 多样且高效的架构——提供不同尺寸的 Dense 和 Mixture-of-Experts (MoE) 变体,以实现可扩展部署。
  • 针对设备端优化——较小模型专为在笔记本电脑和移动设备上高效本地执行而设计。
  • 增加的上下文窗口——小模型拥有 128K 上下文窗口,中模型支持 256K。
  • 增强的编码与智能体能力——在编码基准测试中取得显著改进,同时原生支持函数调用,为高性能自主智能体提供支持。
  • 原生系统提示支持——Gemma 4 引入了对 system 角色的原生支持,实现更结构化、更可控的对话。

https://huggingface.co/unsloth/gemma-4-12b-it-GGUF#models-overview

模型概览

Gemma 4 模型旨在每种尺寸下提供前沿性能,针对从移动和边缘设备(E2B、E4B)到消费级 GPU 和工作站(12B、26B A4B、31B)的部署场景。它们非常适合于推理、智能体工作流、编码和多模态理解。模型采用混合注意力机制,将局部滑动窗口注意力与全全局注意力交错,确保最终层始终为全局层。这种混合设计提供了轻量级模型的处理速度和低内存占用,同时不牺牲复杂长上下文任务所需的深度感知能力。为了优化长上下文的内存,全局层采用统一的键和值,并应用比例 RoPE (p-RoPE)。

https://huggingface.co/unsloth/gemma-4-12b-it-GGUF#dense-models

Dense 模型

属性E2BE4B12B Unified31B Dense
总参数量23 亿有效(含嵌入层 51 亿)45 亿有效(含嵌入层 80 亿)119.5 亿307 亿
层数35424860
滑动窗口512 tokens512 tokens1024 tokens1024 tokens
上下文长度128K tokens128K tokens256K tokens256K tokens
词汇量262K262K262K262K
支持模态文本、图像、音频文本、图像、音频文本、图像、音频文本、图像
视觉编码器参数量~1.5 亿~1.5 亿-~5.5 亿
音频编码器参数量~3 亿~3 亿-无音频

E2B 和 E4B 中的 “E” 代表 “有效” 参数。较小模型采用每层嵌入(PLE)技术,以最大化设备端部署的参数效率。PLE 不是向模型添加更多层或参数,而是为每个解码器层提供其自身的小型嵌入表,用于每个令牌。这些嵌入表虽然很大,但仅用于快速查找,这就是有效参数量远小于总量的原因。

Gemma 4 12B Unified 中的 “Unified” 指的是其无编码器架构。其他 Gemma 4 模型使用专用编码器处理多模态数据,然后传递给 LLM。Gemma 4 12B 完全消除了这些编码器,通过轻量级线性层将原始图像块和音频波形直接投影到 LLM 的嵌入空间。这种统一方法意味着所有模态直接流入单个仅解码器 Transformer,减少了多模态延迟,并允许整个模型一次完成微调。

https://huggingface.co/unsloth/gemma-4-12b-it-GGUF#mixture-of-experts-moe-model

Mixture-of-Experts (MoE) 模型

属性26B A4B MoE
总参数量252 亿
激活参数量38 亿
层数30
滑动窗口1024 tokens
上下文长度256K tokens
词汇量262K
专家数量8 活跃 / 128 总计及 1 共享
支持模态文本、图像
视觉编码器参数量~5.5 亿

26B A4B 中的 “A” 代表 “激活参数”,与模型包含的总参数量相对。通过在推理过程中仅激活 40 亿参数的子集,MoE 模型的运行速度远快于其 252 亿总参数所暗示的速度。与 Dense 31B 模型相比,这使得它成为快速推理的绝佳选择,因为它的运行速度几乎与 40 亿参数模型一样快。

https://huggingface.co/unsloth/gemma-4-12b-it-GGUF#benchmark-results

基准测试结果

这些模型针对大量不同的数据集和指标进行了评估,以涵盖文本生成的不同方面。表中标注的评估结果针对指令调优模型。

Gemma 4 31BGemma 4 26B A4BGemma 4 12B UnifiedGemma 4 E4BGemma 4 E2BGemma 3 27B (无思考)
MMLU Pro85.2%82.6%77.2%69.4%60.0%67.6%
AIME 2026 no tools89.2%88.3%77.5%42.5%37.5%20.8%
LiveCodeBench v680.0%77.1%72.0%52.0%44.0%29.1%
Codeforces ELO215017181659940633110
GPQA Diamond84.3%82.3%78.8%58.6%43.4%42.4%
Tau2 (平均 3 次)76.9%68.2%69.0%42.2%24.5%16.2%
HLE no tools19.5%8.7%5.2%---
HLE with search26.5%17.2%----
BigBench Extra Hard74.4%64.8%53.0%33.1%21.9%19.3%
MMMLU88.4%86.3%83.4%76.6%67.4%70.7%
Vision
MMMU Pro76.9%73.8%69.1%52.6%44.2%49.7%
OmniDocBench 1.5 (平均编辑距离,越低越好)0.1310.1490.1640.1810.2900.365
MATH-Vision85.6%82.4%79.7%59.5%52.4%46.0%
MedXPertQA MM61.3%58.1%48.7%28.7%23.5%-
Audio
CoVoST--38.5*35.5433.47-
FLEURS (越低越好)--0.069*0.080.09-
Long Context
MRCR v2 8 needle 128k (平均)66.4%44.1%43.4%25.4%19.1%13.5%

*排除中文。

https://huggingface.co/unsloth/gemma-4-12b-it-GGUF#core-capabilities

核心能力

Gemma 4 模型处理文本、视觉和音频方面的广泛任务。关键能力包括:

  • 思考——内置推理模式,让模型在回答前逐步思考。
  • 长上下文——上下文窗口最高达 128K tokens (E2B/E4B) 和 256K tokens (12B, 26B A4B/31B)。
  • 图像理解——目标检测、文档/PDF 解析、屏幕和 UI 理解、图表理解、OCR(包括多语言)、手写识别和指示。图像可按可变纵横比和分辨率处理。
  • 视频理解——通过处理帧序列分析视频。
  • 交错多模态输入——在单个提示中自由混合文本和图像,顺序不限。
  • 函数调用——原生支持结构化工具使用,实现智能体工作流。
  • 编码——代码生成、补全和修正。
  • 多语言——开箱即用支持 35 种以上语言,预训练数据覆盖 140 种以上语言。
  • 音频(仅 E2B、E4B 和 12B)——自动语音识别(ASR)和跨语言语音到文本翻译。

https://huggingface.co/unsloth/gemma-4-12b-it-GGUF#getting-started

快速上手

您可以使用最新版本的 Transformers 来使用所有 Gemma 4 模型。要开始使用,请在您的环境中安装必要的依赖项:

pip install -U transformers torch accelerate

安装完成后,您可以使用以下代码加载模型:

from transformers import AutoProcessor, AutoModelForMultimodalLM

MODEL_ID = "google/gemma-4-12B-it"

# Load model
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
    MODEL_ID,
    dtype="auto",
    device_map="auto"
)

加载模型后,您可以开始生成输出:

# Prompt
messages = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Write a short joke about saving RAM."},
]

# Process input
inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
    add_generation_prompt=True,
    enable_thinking=False
).to(model.device)

input_len = inputs["input_ids"].shape[-1]

# Generate output
outputs = model.generate(**inputs, max_new_tokens=1024)
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)

# Parse output
processor.parse_response(response)

要启用推理功能,请设置 enable_thinking=True,parse_response 函数将负责解析思考输出。

下面,您还可以找到处理音频(仅 E2B、E4B、12B)、图像和视频以及文本的代码片段:

处理音频的代码

确保安装以下包: pip install -U transformers torch torchvision librosa accelerate

然后您可以使用以下代码加载模型:

from transformers import AutoProcessor, AutoModelForMultimodalLM

MODEL_ID = "google/gemma-4-12B-it"

# Load model
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
    MODEL_ID,
    dtype="auto",
    device_map="auto"
)

加载模型后,您可以通过在提示中直接引用音频 URL 来开始生成输出:

# Prompt - add audio after text
messages = [
    {
        "role": "user",
        "content": [
            {"type": "text", "text": "Transcribe the following speech segment in its original language. Follow these specific instructions for formatting the answer:\n* Only output the transcription, with no newlines.\n* When transcribing numbers, write the digits, i.e. write 1.7 and not one point seven, and write 3 instead of three."},
            {"type": "audio", "audio": "https://raw.githubusercontent.com/google-gemma/cookbook/refs/heads/main/apps/sample-data/journal1.wav"},
        ]
    }
]

# Process input
inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
    add_generation_prompt=True,
).to(model.device)

input_len = inputs["input_ids"].shape[-1]

# Generate output
outputs = model.generate(**inputs, max_new_tokens=512)
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)

# Parse output
processor.parse_response(response)

处理图像的代码

确保安装以下包: pip install -U transformers torch torchvision accelerate

然后您可以使用以下代码加载模型:

from transformers import AutoProcessor, AutoModelForMultimodalLM

MODEL_ID = "google/gemma-4-12B-it"

# Load model
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
    MODEL_ID,
    dtype="auto",
    device_map="auto"
)

加载模型后,您可以通过在提示中直接引用图像 URL 来开始生成输出:

# Prompt - add image before text
messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "url": "https://raw.githubusercontent.com/google-gemma/cookbook/refs/heads/main/apps/sample-data/GoldenGate.png"},
            {"type": "text", "text": "What is shown in this image?"}
        ]
    }
]

# Process input
inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
    add_generation_prompt=True,
).to(model.device)

input_len = inputs["input_ids"].shape[-1]

# Generate output
outputs = model.generate(**inputs, max_new_tokens=512)
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)

# Parse output
processor.parse_response(response)

处理视频的代码

确保安装以下包: pip install -U transformers torch torchvision librosa accelerate

然后您可以使用以下代码加载模型:

from transformers import AutoProcessor, AutoModelForMultimodalLM

MODEL_ID = "google/gemma-4-12B-it"

# Load model
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForMultimodalLM.from_pretrained(
    MODEL_ID,
    dtype="auto",
    device_map="auto"
)

加载模型后,您可以通过在提示中直接引用视频 URL 来开始生成输出:

# Prompt - add video before text
messages = [
    {
        'role': 'user',
        'content': [
            {"type": "video", "video": "https://github.com/bebechien/gemma/raw/refs/heads/main/videos/ForBiggerBlazes.mp4"},
            {'type': 'text', 'text': 'Describe this video.'}
        ]
    }
]

# Process input
inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
    add_generation_prompt=True,
).to(model.device)

input_len = inputs["input_ids"].shape[-1]

# Generate output
outputs = model.generate(**inputs, max_new_tokens=512)
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)

# Parse output
processor.parse_response(response)

https://huggingface.co/unsloth/gemma-4-12b-it-GGUF#best-practices

最佳实践

为获得最佳性能,请使用以下配置和最佳实践:

https://huggingface.co/unsloth/gemma-4-12b-it-GGUF#1-sampling-parameters

  1. 采样参数

在所有用例中使用以下标准化采样配置:

  • temperature=1.0
  • top_p=0.95
  • top_k=64

https://huggingface.co/unsloth/gemma-4-12b-it-GGUF#2-thinking-mode-configuration

  1. 思考模式配置

与 Gemma 3 相比,这些模型使用标准的 system、assistant 和 user 角色。要正确管理思考过程,请使用以下控制令牌:

  • 触发思考: 通过在系统提示开头包含 <|think|> 令牌来启用思考。要禁用思考,请移除该令牌。
  • 标准生成: 启用思考后,模型将输出其内部推理,然后使用以下结构输出最终答案: <|channel>thought\n ** [内部推理] ** ``
  • 禁用思考行为: 对于除 E2B 和 E4B 变体之外的所有模型,如果禁用思考,模型仍将生成标签,但思考块为空。

相似文章

unsloth/gemma-4-26B-A4B-it-GGUF

Hugging Face Models Trending

# unsloth/gemma-4-26B-A4B-it-GGUF · Hugging Face 来源:[https://huggingface.co/unsloth/gemma-4-26B-A4B-it-GGUF](https://huggingface.co/unsloth/gemma-4-26B-A4B-it-GGUF) ## [https://huggingface.co/unsloth/gemma-4-26B-A4B-it-GGUF#read-our-how-to-run-gemma-4-guide](https://huggingface.co/unsloth/gemma-4-26B-A4B-it-GGUF#read-our-how-to-run-gemma-4-guide)阅读我们的[如何运行 Gemma 4 指南](https://docs.unsloth.ai/models/gemma-4)! *请参阅[Unsloth Dynamic 2.0 GGUFs](https://unsloth.ai/docs/basics/unslot

unsloth/gemma-4-12B-it-qat-GGUF

Hugging Face Models Trending

Unsloth 发布了Google DeepMind的Gemma 4模型的GGUF量化版本,通过量化感知训练(QAT)优化,在保持质量的同时降低内存需求,支持多种格式和大小,适用于不同的部署场景。

google/gemma-4-26B-A4B-it

Hugging Face Models Trending

Google DeepMind 发布 Gemma 4,一系列开放权重的多模态模型,参数量从2.3B到31B,支持文本、图像、视频和音频输入。模型具有256K上下文窗口,MoE和密集架构,增强的推理能力,并针对从移动设备到服务器的部署进行优化。