[新模型!] Supra2-100M Base 和 Instruct——快去看看吧!

Reddit r/LocalLLaMA 模型

摘要

SupraLabs 发布了 Supra2-100M Base 和 Instruct 模型,这是一个新的小型语言模型系列,包含社区驱动的改进、基准测试和一个 GGUF 版本。

嘿,各位!在我们之前的模型(如 Supra-50M-Instruct 和 -Reasoning)上收到了大量好评,以及许多社区的点赞、关注和投票后,我们看到许多社区请求要求推出新模型。我们的工作激励了很多人——现在我们要推出全新的 Supra2 系列。我们的首个版本是:Supra2-100M Base 和 Instruct。以下是一些示例:Prompt: "What is google?" Answer: Google is a web-based platform that allows users to search and find information on the web. It's a social media platform that uses algorithms to make recommendations based on various factors such as location, time of day, and interests. Google has a number of features that make it easy to find information, including: - Searching for keywords and phrases related to various topics, like books, movies, or music - Analyzing website traffic and traffic patterns - Creating a custom search interface - Suggesting alternative ways to find the information - Providing recommendations for books, articles, and other content - Allowing users to customize the search results One of the main advantages of Google is that it's easy to use, as users can search for the information they need, and then filter the results based on their interests. This makes it accessible to a wider audience. 更多内容请参阅 HF 上的 README:HF 模型链接:Base:https://huggingface.co/SupraLabs/Supra2-100M Instruct:https://huggingface.co/SupraLabs/Supra2-100M-Instruct 以下是与更小甚至更大模型的基准对比:https://preview.redd.it/gbkjvw07j6hh1.png?width=919&format=png&auto=webp&s=09bdb05952e28915726d1ca9c5ec7971f5a14f28 祝你们使用愉快!🤗🔥 Instruct 模型的 GGUF 版本已经在 HF 仓库中了!接下来是什么?--> Supra2-Nano、-Small、-Medium ... 以及 ... IMG!敬请期待接下来的发布!!如果你愿意,可以通过点赞和关注支持我们!😺
查看原文

相似文章

[新发布] Supra-50M 正式推出!

Reddit r/LocalLLaMA

SupraLabs 发布了 Supra-50M,一个紧凑的 5000 万参数因果语言模型,包含基础版和指令版,基于 fineweb-edu 的 200 亿个 token 训练,在多项关键基准测试中达到了可与 GPT-2 和 SmolLM 等更大模型竞争的水平。

[新模型] Supra-Title-0.3B 刚刚发布!

Reddit r/LocalLLaMA

Supra Labs 发布了 Supra Title,这是一个参数为 350M 的专用模型,用于生成聊天对话标题。该模型基于 LFM2.5 构建,以 GGUF 格式运行在任何硬件上,且无需系统提示。

[发布] SupraBrain-50M-v0.1

Reddit r/LocalLLaMA

SupraLabs 发布了 SupraBrain-50M,这是一种混合语言模型,结合了 Gated DeltaNet、滑动窗口注意力(Sliding-Window Attention)和惊喜门控(Surprise-Gated)机制,尽管训练使用的 token 数量少得多,但性能与 Supra-Base-50M 几乎持平。