small-language-model

Tag

Cards List
#small-language-model

I trained a 75M parameter LLM from scratch on 18B tokens and it beats a model almost double its size

Reddit r/LocalLLaMA ↗ · 2026-06-02

Trained a 75M parameter LLM called KeyLM from scratch on 18B tokens, achieving competitive instruction-following scores against larger models while using fewer parameters and less data.

0 favorites 0 likes
#small-language-model

OpenBMB releases MiniCPM5-1B LLM. Currently one of the most powerful LLMs for its size. ( 17.9 on the Artificial Analysis Intelligence Index)

Reddit r/singularity ↗ · 2026-05-27 Cached

OpenBMB releases MiniCPM5-1B, a leading 1B open weights LLM that achieves the highest Artificial Analysis Intelligence Index score (17.9) in its size class, surpassing larger models like Qwen3.5 2B while using fewer parameters.

0 favorites 0 likes
#small-language-model

@AdinaYakup: MiniCPM5-1B is an impressive release in the 1B class! @OpenBMB https://huggingface.co/collections/openbmb/minicpm5… 1B …

X AI KOLs Following ↗ · 2026-05-25 Cached

MiniCPM5-1B is a new 1B parameter AI model from OpenBMB featuring hybrid reasoning with Think/No-Think modes, 128K context, and Apache 2.0 license, running on various hardware.

0 favorites 0 likes
#small-language-model

@ModelScope2022: MiniCPM5-1B is now fully open source, including weights, training data, and deployment code. 1B params, #1 on Artificia…

X AI KOLs Following ↗ · 2026-05-25 Cached

MiniCPM5-1B is fully open-sourced with weights, training data, and deployment code; it achieves top scores among sub-2B models and runs on edge devices.

0 favorites 0 likes
#small-language-model

[NEW] Supra-50M Released!

Reddit r/LocalLLaMA ↗ · 2026-05-22

SupraLabs released Supra-50M, a compact 50M-parameter causal language model with base and instruct versions, trained on 20B tokens from fineweb-edu, achieving competitive benchmarks against larger models like GPT-2 and SmolLM.

0 favorites 0 likes
#small-language-model

@Sapient_Int: Introducing HRM-Text. An ultra-lean 1B-parameter reasoning language model designed to deliver strong general performanc…

X AI KOLs Timeline ↗ · 2026-05-18 Cached

Sapient Intelligence introduces HRM-Text, a 1B-parameter reasoning language model trained on only 40B tokens with a budget of $1,000, achieving competitive performance while drastically reducing data and compute requirements.

0 favorites 0 likes
#small-language-model

@_vmlops: MICROSOFT'S FARA-7B CAN USE YOUR COMPUTER FOR YOU 7b params...clicks, scrolls, fills forms, books tickets all on its ow…

X AI KOLs Timeline ↗ · 2026-05-18 Cached

Microsoft released Fara-7B, a 7-billion parameter small language model that can autonomously control a computer to perform tasks like clicking, scrolling, and filling forms, running on-device and beating larger models like OpenAI's computer-use agent on benchmarks.

0 favorites 0 likes
#small-language-model

@DJLougen: To whoever trained this @Microsoft , god bless you and your soul this is impressive with browserOS

X AI KOLs Timeline ↗ · 2026-05-14 Cached

Microsoft released Fara-7B, a 7 billion parameter agentic small language model for computer use, achieving state-of-the-art performance among models of its size and competitive with larger systems.

0 favorites 0 likes
#small-language-model

Needle: We Distilled Gemini Tool Calling Into a 26M Model

Reddit r/LocalLLaMA ↗ · 2026-05-12

Cactus-Compute released Needle, a 26M parameter open-source model distilled from Gemini for efficient on-device function calling using a novel Simple Attention Network architecture without MLPs.

0 favorites 0 likes
#small-language-model

@ash_csx: We’re dropping two open source SLMs this week. 1. One of them matches SOTA accuracy at up to 93x smaller. 2. The other …

X AI KOLs Following ↗ · 2026-05-11 Cached

Two new open-source small language models are being released: one matches state-of-the-art accuracy at up to 93x smaller size, and the other outperforms a recent OpenAI model. The first model drops tomorrow.

0 favorites 0 likes
#small-language-model

@AdinaYakup: MiniCPM V4.6 a 1B MLLM that actually runs on your phone, just released by @OpenBMB 1B - Apache2.0 Runs on iOS, Android,…

X AI KOLs Following ↗ · 2026-05-11 Cached

OpenBMB has released MiniCPM V4.6, a 1B-parameter multimodal large language model optimized for mobile devices under the Apache 2.0 license. It features mixed visual token compression and claims approximately 1.5x faster throughput than Qwen3.5 0.8B while running natively on iOS, Android, and HarmonyOS.

0 favorites 0 likes
#small-language-model

DR-Venus: Towards Frontier Edge-Scale Deep Research Agents with Only 10K Open Data

Hugging Face Daily Papers ↗ · 2026-04-21 Cached

DR-Venus-4B is a 4B-parameter deep-research agent trained on only 10K open samples via agentic SFT+RL with turn-level rewards, outrunning prior sub-9B agents and rivaling 30B models on research benchmarks while staying deployable on edge devices.

0 favorites 0 likes
#small-language-model

@zhijianliu_: DFlash for Qwen3.6-35B-A3B just dropped The community was running the day-1 preview before we even finished training. N…

X AI KOLs Following ↗ · 2026-04-20

Z-lab releases DFlash for Qwen3.6-35B-A3B, a model fine-tuning/compression technique, with training complete and weights now available on GitHub and HuggingFace.

0 favorites 0 likes
#small-language-model

Cactus-Compute/needle

Hugging Face Models Trending ↗ · 2026-03-16 Cached

Cactus-Compute releases Needle, a 26M parameter distilled model from Gemini 3.1, using a pure attention architecture optimized for on-device inference and local fine-tuning.

0 favorites 0 likes
#small-language-model

microsoft/Fara-7B

Hugging Face Models Trending ↗ · 2025-10-30 Cached

Microsoft released Fara-7B, an efficient 7 billion parameter agentic small language model (SLM) for computer use tasks, achieving state-of-the-art performance within its size class and competitive with larger systems.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback