@seclink: 官方页面地址: https://huggingface.co/moonshotai/Kimi-K3… 这是 Moonshot AI(月之暗面)发布的 2.8 万亿参数(MoE 架构,激活约 104B)开源权重模型,支持原生多模态(文本+图…

X AI KOLs Timeline 模型

摘要

Moonshot AI releases Kimi K3, a 2.8 trillion parameter open-weight MoE model with native multimodal capabilities and a 1 million token context window, claiming it as the world's first open 3T-class model.

官方页面地址: https://huggingface.co/moonshotai/Kimi-K3… 这是 Moonshot AI(月之暗面)发布的 2.8 万亿参数(MoE 架构,激活约 104B)开源权重模型,支持原生多模态(文本+图像)和 100 万 token 上下文。 权重以 Safetensors 等格式提供,采用 Kimi K3 License。 Kimi K3 的开放权重已于 2026 年 7 月 27 日正式发布,主要可在 Hugging Face 查看和下载。
查看原文
查看缓存全文

缓存时间: 2026/07/28 08:24

官方页面地址: https://huggingface.co/moonshotai/Kimi-K3…

这是 Moonshot AI(月之暗面)发布的 2.8 万亿参数(MoE 架构,激活约 104B)开源权重模型,支持原生多模态(文本+图像)和 100 万 token 上下文。

权重以 Safetensors 等格式提供,采用 Kimi K3 License。

Kimi K3 的开放权重已于 2026 年 7 月 27 日正式发布,主要可在 Hugging Face 查看和下载。


moonshotai/Kimi-K3 · Hugging Face

Source: https://huggingface.co/moonshotai/Kimi-K3 Kimi K3


ChatHomepage

Hugging FaceTwitter FollowDiscordModelScope

License

📰Tech Blog|📄Full Report

https://huggingface.co/moonshotai/Kimi-K3#1-model-introduction1. Model Introduction

Kimi K3 is an open-weight, native multimodal agentic model and our most capable model to date. It is a 2.8T-parameter model built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), with native vision capabilities and a 1-million-token context window. It is the world’s first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning.

https://huggingface.co/moonshotai/Kimi-K3#key-featuresKey Features

  • New Architecture: Kimi K3 is built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), and scales up MoE sparsity with a Stable LatentMoE framework that activates 16 out of 896 experts — yielding an approximate 2.5× improvement in overall scaling efficiency over Kimi K2.
  • Long-Horizon Coding: Operating with minimal human oversight, Kimi K3 sustains long engineering sessions, navigates massive repositories, and orchestrates terminal tools — from GPU kernel optimization and compiler development to vision-in-the-loop game dev, CAD, and even chip design.
  • Agentic Knowledge Work: Kimi K3 advances end-to-end knowledge work, producing deep research with interactive visualizations, widgets and dashboards, and motion design and video editing, powered by its native multimodal architecture.
  • Native Multimodality & Long Context: Kimi K3 understands text, images, and video within the same model, and supports a 1-million-token context window.
  • Open Frontier Weights: We release the full Kimi K3 model weights under the Kimi K3 License, making frontier intelligence openly available for research, deployment, and further innovation.

https://huggingface.co/moonshotai/Kimi-K3#2-model-summary2. Model Summary

ArchitectureMixture-of-Experts (MoE)Total Parameters2.8TActivated Parameters104BNumber of Layers93Number of Dense Layers1Attention-Layer Composition69 KDA + 24 Gated MLAAttention Hidden Dimension7168Number of Attention Heads96Latent MoE Dimension3584MoE Hidden Dimension(per Expert)3072Number of Experts896Selected Experts per Token16Number of Shared Experts2Vocabulary Size160KContext Length1048576Attention MechanismKDA & Gated MLAActivation FunctionSiTU-GLUVision EncoderMoonViT-V2Parameters of Vision Encoder401MQuantizationMXFP4 weights / MXFP8 activations (quantization-aware training)ModalityText, Image

https://huggingface.co/moonshotai/Kimi-K3#3-evaluation-results3. Evaluation Results

BenchmarkKimi K3 (max)Claude Fable 5 (max, w/ fallback)GPT-5.6 Sol (max)Claude Opus 4.8 (max)GPT-5.5 (xhigh)GLM-5.2 (max)Reasoning & KnowledgeGPQA Diamond93.592.694.191.093.591.2CritPt23.428.632.320.927.120.9AA-LCR74.770.073.767.774.371.3HLE-Full43.5 / 56.053.3 / 63.044.5 / 58.049.8 / 57.941.4 / 52.2—CodingDeepSWE67.570.073.059.067.046.2ProgramBench77.876.877.671.970.863.7Terminal-Bench 2.188.388.088.884.683.482.7FrontierSWE81.286.671.366.764.967.3SWE-Marathon42.035.039.040.014.013.0PostTrainBench36.641.434.634.128.434.3MLS-Bench-Lite48.349.946.242.835.540.4SciCode58.760.256.153.556.150.5Kimi Code Bench 2.072.976.964.871.769.064.2AgenticBrowseComp91.288.090.484.384.4—DeepSearchQA (F1)95.094.2—93.1——ResearchRubrics76.2—73.873.564.071.1GDPval-AA v2 (Elo)168617471736159314911510Toolathlon-Verified76.577.974.976.273.559.9MCPMark-Verified94.587.492.976.492.9—MCP-Atlas84.284.783.683.682.882.6AutomationBench30.829.129.727.222.712.9JobBench54.357.445.448.438.343.4AA-Briefcase (Elo)154815831495135411581260Agents’ Last Exam28.325.7†29.627.026.620.4APEX-Agents41.043.339.939.438.535.6OfficeQA Pro63.369.963.263.960.941.4SpreadsheetBench 234.834.732.431.629.128.1OSWorld-Verified84.885.083.083.479.0—OSWorld 2.058.366.162.655.749.5—SaaS-Bench60.1—61.456.143.8—τ³-Banking33.426.833.027.631.326.8Harvey Lab-AA94.693.687.291.186.391.0CorpFin v271.671.864.466.768.466.1Finance Agent v254.456.353.853.951.849.7Legal Research Bench44.249.548.143.840.431.3VisionWorldVQA ForceAnswer51.056.741.839.138.5—OmniDocBench91.189.885.887.989.4—PerceptionBench58.557.259.747.255.8—Video-MME (w. sub)90.0—89.586.089.3—MMVU82.1—81.279.281.7—BabyVision w/ python85.790.588.981.283.6—MMMU-Pro81.6 / 83.481.2 / 86.583.0 / 84.678.9 / 82.781.2 / 83.2—CharXiv (RQ)84.8 / 91.388.9 / 93.584.6 / 89.180.5 / 89.984.1 / 89.0—MathVision94.3 / 97.894.8 / 98.695.8 / 97.886.7 / 97.192.2 / 96.8—ZeroBench (pass@5)23.0 / 41.023.0 / 46.017.0 / 35.017.0 / 34.022.0 / 41.0— FootnotesAll Kimi K3 results are obtained with reasoning effort set to ‘max’ and temperature = 1.0. For single-step tasks, such as GPQA Diamond, HLE-Full, and vision benchmarks without tools, we set top-p = 0.95; for agentic tasks, we set top-p = 1.0. For HLE-Full, MMMU-Pro, CharXiv (RQ), MathVision, and ZeroBench, each cell reports the scores without and with tool augmentation (general tools for HLE-Full, Python for the vision benchmarks), in that order.

  1. Reasoning & knowledge benchmarks- **CritPt and AA-LCR.**Scores are cited fromArtificial Analysisas of July 23, 2026.
  2. Coding benchmarks- **DeepSWE.**Kimi K3 is evaluated with the Kimi Code harness. The GLM-5.2 score is taken from theGLM-5.2 release blog; all remaining scores are from the officialDeepSWE leaderboard, under which Kimi K3 attains 67.3 with the mini-SWE-agent harness. We report the DeepSWE v1.1 tasks. - **Terminal-Bench 2.1.**Kimi K3 is evaluated with the Kimi Code harness. For all other models, we report the best score across harnesses: GLM-5.2 with Claude Code (GLM-5.2 release blog); Claude Opus 4.8 and Claude Fable 5 with Terminus 2 (Artificial Analysis); GPT-5.5 and GPT-5.6 Sol with Codex (OpenAI). - **ProgramBench.**Kimi K3 is evaluated with the Kimi Code harness. The GLM-5.2 score is from theGLM-5.2 release blog; all other scores are fromVals AI. - **SWE-Marathon.**Kimi K3, Claude Opus 4.8, and Claude Fable 5 are evaluated with the Claude Code harness; GPT-5.6 Sol is evaluated with the Codex harness. The GLM-5.2 score is from theGLM-5.2 release blog. Our evaluation is based on an H20-calibrated branch of theofficial tasksas of July 9, 2026, prior to the final v1.1 release: the Docker images, performance gates, and reference oracles for the GPU tasks have been recalibrated for H20, while the correctness and anti-cheat validators remain unchanged. Additionally, Claude Fable 5 hit fallbacks on 35% of the tasks in our evaluation, which may have negatively impacted its measured performance. - **FrontierSWE.**Kimi K3 is evaluated with the Kimi Code harness and GPT-5.6 Sol with the Codex harness; all other results are fromFrontierSWE. Dominance scores are recomputed from the raw scores using the official evaluation script and are current as of July 16, 2026. - **PostTrainBench.**Scores for GLM-5.2, GPT-5.5, and Claude Opus 4.8 are adopted from the officialPostTrainBenchresults. Kimi K3, Claude Fable 5, and GPT-5.6 Sol are evaluated with the official Harbor implementation at maximum reasoning effort, averaged over three runs on H20 GPUs (instead of H100 in the official setting) — Kimi K3 and Claude Fable 5 with the Claude Code harness, and GPT-5.6 Sol with the Codex harness. - **MLS-Bench-Lite.**Kimi K3 is evaluated with the Kimi Code harness; GLM-5.2 and the Claude models with the Claude Code harness; GPT-5.5 and GPT-5.6 Sol with the Codex harness. - **SciCode.**Scores are cited fromArtificial Analysisas of July 23, 2026. - **Kimi Code Bench 2.0 (in-house).**Kimi K3 is evaluated with the Kimi Code harness (it attains 73.7 with the Claude Code harness); GLM-5.2, Claude Opus 4.8, and Claude Fable 5 with the Claude Code harness; GPT-5.5 and GPT-5.6 Sol with the Codex harness. All models are evaluated at maximum reasoning effort, except GPT-5.5, which uses the “xhigh” setting. As the benchmark includes cybersecurity and safety-related tasks, we also disclose the fraction of refused or fallback tasks: Claude Fable 5 hit 13 fallbacks and 1 refusal out of 80 tasks; 10 refusals out of 80 tasks entered GPT-5.6 Sol’s cyber guard; GPT-5.5 had 3 refusals out of 80 tasks.
  3. Agentic benchmarks- **OfficeQA Pro.**Each test case provides the agent with the entire PDF corpus, with all PDFs rendered as images and no machine-readable text available. - **OfficeQA Pro and SpreadsheetBench 2.**Kimi K3, GLM-5.2, Claude Opus 4.8, and Claude Fable 5 are evaluated with the Claude Code harness; GPT-5.5 and GPT-5.6 Sol are evaluated with the Codex harness. - **MCP-Atlas.**All models are evaluated on the 500-task public subset with a 100-turn limit, using Gemini 3.1 Pro as the judge. - **AutomationBench.**All models are evaluated on the 600-task public subset, following the official GitHub setup in all other respects. - **BrowseComp.**We adopt a context-compaction strategy triggered at 300K tokens. When evaluated with the full 1M-token context window and no context management, Kimi K3 achieves a score of 90.4. The results of Claude Fable 5, Claude Opus 4.8, GPT-5.6 Sol, and GPT-5.5 are cited fromAnthropicandOpenAI. - **GDPval-AA v2, AA-Briefcase, τ³-Banking, Harvey Lab-AA, and APEX-Agents.**Scores are cited fromArtificial Analysisand theAPEX-Agents leaderboardas of July 23, 2026. For Harvey Lab-AA, we report the criterion pass rate. - **CorpFin v2, Finance Agent v2, and Legal Research Bench.**Scores are cited fromVals AI. - **Agents’ Last Exam.**Scores are cited from theofficial leaderboardas of July 23, 2026; we report the leaderboard’s primary pass-rate metric. On the leaderboard, each model is paired with a specific harness: Kimi K3 with Kimi Code; GPT-5.6 Sol and GPT-5.5 with Codex; Claude Fable 5, Claude Opus 4.8, and GLM-5.2 with Claude Code.†The Claude Fable 5 entry runs at xhigh effort with 40% of tasks annotated as downgraded.
  4. Multimodal benchmarks- Except for ZeroBench, which follows the official setting and is run five times, all multimodal scores are averaged over three runs. MMMU-Pro is evaluated following the official protocol, preserving the original input order and prepending images to the text input. - PerceptionBenchis an in-house benchmark that focuses on atomic visual perception capabilities.

https://huggingface.co/moonshotai/Kimi-K3#4-native-mxfp4-quantization4. Native MXFP4 Quantization

Kimi K3 applies quantization-aware training from the SFT stage onward, using MXFP4 weights with MXFP8 activations for broad hardware compatibility.

https://huggingface.co/moonshotai/Kimi-K3#5-deployment5. Deployment

You can access Kimi K3’s API onhttps://platform.kimi.aiby selectingkimi\-k3, and we provide OpenAI/Anthropic-compatible API for you. Currently, Kimi K3 is recommended to run on the following inference engines:


https://huggingface.co/moonshotai/Kimi-K3#6-model-usage6. Model Usage

Kimi K3 always has thinking enabled, and will returnreasoning\_content. Thinking effort is configured with the top-levelreasoning\_effortrequest field, which supports"low","high", and"max"(default"max").

Kimi K3 was trained in the preserved thinking history mode. For multi-turn conversations and tool calls, Kimi K3 requires the complete assistant message returned by the API to be passed back tomessagesas-is — includingreasoning\_contentandtool\_calls, not justcontent:

import openai

def chat_with_preserved_thinking(client: openai.OpenAI, model_name: str):
    messages = [
        {
            "role": "user",
            "content": "Tell me three random numbers."
        },
        {
            "role": "assistant",
            "reasoning_content": "I'll start by listing five numbers: 473, 921, 235, 215, 222, and I'll tell you the first three.",
            "content": "473, 921, 235"
        },
        {
            "role": "user",
            "content": "What are the other two numbers you have in mind?"
        }
    ]

    response = client.chat.completions.create(
        model=model_name,
        messages=messages,
        stream=False,
        max_tokens=4096,
        reasoning_effort="max",
    )
    # the assistant should mention 215 and 222 that appear in the prior reasoning content
    print(f"response: {response.choices[0].message.reasoning}")
    return response.choices[0].message.content

For full guides and examples (vision input, structured output, partial mode, tool choice, dynamic tool loading, context caching), see theKimi K3 QuickstartandThinking Effort.

https://huggingface.co/moonshotai/Kimi-K3#coding-agent-frameworkCoding Agent Framework

Kimi K3 works best withKimi Code CLIas its agent framework. We warmly invite you to give it a try — run Kimi Code in your terminal and select Kimi K3 using the/modelcommand. We hope you enjoy building with Kimi K3, and we would love to hear your feedback!


https://huggingface.co/moonshotai/Kimi-K3#7-license7. License

Both the code repository and the model weights are released under theKimi K3 License.


https://huggingface.co/moonshotai/Kimi-K3#8-contact-us8. Contact Us

If you have any questions, please reach out at[email protected].

Downloads last month2,850

Model tree formoonshotai/Kimi-K3https://huggingface.co/docs/hub/model-cards#specifying-a-base-model

Spaces usingmoonshotai/Kimi-K35

Collection includingmoonshotai/Kimi-K3

Evaluation resultshttps://huggingface.co/docs/hub/eval-results

相似文章

Kimi-K3 技术报告 [pdf]

Hacker News Top

MoonshotAI发布了Kimi-K3,一个拥有2.8万亿参数的开源权重多模态智能体模型,具备100万token的上下文窗口,基于全新的Kimi Delta Attention和Attention Residuals架构,实现了显著的扩展性改进。

Kimi-K3 在 HuggingFace 上发布

Reddit r/artificial

Moonshot AI 发布了 Kimi-K3,这是一个拥有 2.8 万亿参数的混合专家模型,支持 100 万 tokens 的上下文窗口。该模型已在 HuggingFace 上以具有商业限制的宽松许可证提供。

发布 Kimi K3 模型权重和技术报告(2分钟阅读)

TLDR AI

Kimi Moonshot 发布了 Kimi K3,这是一个 2.8 万亿参数的多模态模型,拥有 100 万 token 的上下文窗口,以及 Kimi Delta Attention 和 Attention Residuals 等架构创新。公司声称其效率显著提升,在内部基准测试中超越了 Claude Opus 4.8 和 GPT-5.5。