DeepSeek V4.1 Flash 正在令人惊讶地接近 GPT-5.6 Sol 的领域,同时价格低得离谱

Reddit r/singularity 模型

摘要

DeepSeek 发布了 V4.1 Flash,一个 552B MoE 模型,具有高效的活动参数,性能接近 GPT-5.6 Sol,而成本却低得多。

DeepSeek 刚刚发布了 V4.1 Flash,一个 552B MoE 模型,输入仅有 8B 活动参数,输出为 16B。 - X 帖子链接:https://x.com/deepseek\_ai/status/2097930608790167907 - Hugging Face 模型页面链接:https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash - 论文链接:https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek\_V41\_Tech\_Report.pdf
查看原文

相似文章

DeepSeek v4 Flash 能力显著提升

Reddit r/LocalLLaMA

DeepSeek V4 Flash 在预览更新中展现出显著的基准测试提升,在智能体编码任务上与 GPT-5.6 Terra 互有胜负。