small-language-model

Tag

Cards List
#small-language-model

Scaling former VibeThinker-1.5B to 3B — now it reaches frontier math & coding performance

Reddit r/LocalLLaMA · 2026-06-16

The VibeThinker-3B model achieves state-of-the-art math and coding reasoning performance, scoring 94.3 on AIME'26 and 96.1% on unseen LeetCode problems, demonstrating that small models can reach frontier-level reasoning in verifiable domains.

0 favorites 0 likes
#small-language-model

Cleo: trying to fit full analyst behavior in a 2B model [P]

Reddit r/MachineLearning · 2026-06-15

Cleo is a finetuned version of Qwen3.5-2B-Base designed for text-to-SQL tasks, using a unified harness for training and inference that supports live execution evidence and safety checks. All code, model, and datasets are open-source.

0 favorites 0 likes
#small-language-model

@GitTrend0x: A pure local desktop automation powerhouse, and most importantly, saves money! https://github.com/microsoft/fara This is Fara-7B, an efficient Computer Use Agent small model from Microsoft! In a word, it surpasses traditional large model CUA: only 7B parameters...

X AI KOLs Timeline · 2026-06-15 Cached

Microsoft launches Fara-7B, an efficient Computer Use Agent with only 7B parameters, surpassing larger models on web tasks, supporting pure local deployment, and achieving low-cost desktop automation.

0 favorites 0 likes
#small-language-model

@nini_incrypto_: Microsoft's recent practical release lets a 7B model take over your mouse and keyboard! FARA abandons pointless chat and focuses purely on local desktop automation. Its core advantages boil down to two words: obedient and cost-effective. 1. Pure desktop execution: opens web pages, fills forms, and automatically runs all repetitive mechanical workflows. 2. ...

X AI KOLs Timeline · 2026-06-14 Cached

Microsoft has released Fara-7B, a small 7B-parameter language model focused on pure local desktop automation. It can directly take over your mouse and keyboard to execute repetitive workflows, with low cost and no need for internet connectivity.

0 favorites 0 likes
#small-language-model

[NEW MODEL] Supra-Title-0.3B Just released!

Reddit r/LocalLLaMA · 2026-06-12

Supra Labs released Supra Title, a 350M parameter model specialized for generating chat conversation titles. Built on LFM2.5, it runs on any hardware in GGUF format and requires no system prompt.

0 favorites 0 likes
#small-language-model

WeiboAI/VibeThinker-3B

Hugging Face Models Trending · 2026-06-12 Cached

VibeThinker-3B is a 3B-parameter model that achieves frontier-level reasoning performance on math, coding, and STEM benchmarks by optimizing the Spectrum-to-Signal Principle (SSP) post-training pipeline, reaching performance comparable to much larger models.

0 favorites 0 likes
#small-language-model

IAPO: Input Attribution-Aware Policy Optimization for Tool Use in Small Multimodal Agents

arXiv cs.LG · 2026-06-11 Cached

This paper introduces IAPO, a reinforcement learning algorithm that improves tool-calling capabilities in multimodal small language models by aligning input attribution with a stronger teacher. Experiments on Qwen2.5-VL-3B show an average 3% improvement in visual question answering accuracy across six test sets.

0 favorites 0 likes
#small-language-model

@harshbhatt7585: https://x.com/harshbhatt7585/status/2063593933314113587

X AI KOLs Timeline · 2026-06-07 Cached

The author shares learnings from training a 160M parameter LLM from scratch, experimenting with architectures like multi-token prediction and hierarchical reasoning models. They emphasize the importance of fast iteration, simplifying ideas, and understanding why architectures work.

0 favorites 0 likes
#small-language-model

Microsoft Aion 1.0 Instruct and Aion 1.0 Plan models!

Reddit r/LocalLLaMA · 2026-06-03

Microsoft announced two new on-device AI models at Build 2026: Aion 1.0 Instruct, an open-weights small language model, and Aion 1.0 Plan, a 14B parameter reasoning and tool-calling model for local agentic workflows.

0 favorites 0 likes
#small-language-model

I trained a 75M parameter LLM from scratch on 18B tokens and it beats a model almost double its size

Reddit r/LocalLLaMA · 2026-06-02

Trained a 75M parameter LLM called KeyLM from scratch on 18B tokens, achieving competitive instruction-following scores against larger models while using fewer parameters and less data.

0 favorites 0 likes
#small-language-model

OpenBMB releases MiniCPM5-1B LLM. Currently one of the most powerful LLMs for its size. ( 17.9 on the Artificial Analysis Intelligence Index)

Reddit r/singularity · 2026-05-27 Cached

OpenBMB releases MiniCPM5-1B, a leading 1B open weights LLM that achieves the highest Artificial Analysis Intelligence Index score (17.9) in its size class, surpassing larger models like Qwen3.5 2B while using fewer parameters.

0 favorites 0 likes
#small-language-model

@AdinaYakup: MiniCPM5-1B is an impressive release in the 1B class! @OpenBMB https://huggingface.co/collections/openbmb/minicpm5… 1B …

X AI KOLs Following · 2026-05-25 Cached

MiniCPM5-1B is a new 1B parameter AI model from OpenBMB featuring hybrid reasoning with Think/No-Think modes, 128K context, and Apache 2.0 license, running on various hardware.

0 favorites 0 likes
#small-language-model

@ModelScope2022: MiniCPM5-1B is now fully open source, including weights, training data, and deployment code. 1B params, #1 on Artificia…

X AI KOLs Following · 2026-05-25 Cached

MiniCPM5-1B is fully open-sourced with weights, training data, and deployment code; it achieves top scores among sub-2B models and runs on edge devices.

0 favorites 0 likes
#small-language-model

[NEW] Supra-50M Released!

Reddit r/LocalLLaMA · 2026-05-22

SupraLabs released Supra-50M, a compact 50M-parameter causal language model with base and instruct versions, trained on 20B tokens from fineweb-edu, achieving competitive benchmarks against larger models like GPT-2 and SmolLM.

0 favorites 0 likes
#small-language-model

@Sapient_Int: Introducing HRM-Text. An ultra-lean 1B-parameter reasoning language model designed to deliver strong general performanc…

X AI KOLs Timeline · 2026-05-18 Cached

Sapient Intelligence introduces HRM-Text, a 1B-parameter reasoning language model trained on only 40B tokens with a budget of $1,000, achieving competitive performance while drastically reducing data and compute requirements.

0 favorites 0 likes
#small-language-model

@_vmlops: MICROSOFT'S FARA-7B CAN USE YOUR COMPUTER FOR YOU 7b params...clicks, scrolls, fills forms, books tickets all on its ow…

X AI KOLs Timeline · 2026-05-18 Cached

Microsoft released Fara-7B, a 7-billion parameter small language model that can autonomously control a computer to perform tasks like clicking, scrolling, and filling forms, running on-device and beating larger models like OpenAI's computer-use agent on benchmarks.

0 favorites 0 likes
#small-language-model

@DJLougen: To whoever trained this @Microsoft , god bless you and your soul this is impressive with browserOS

X AI KOLs Timeline · 2026-05-14 Cached

Microsoft released Fara-7B, a 7 billion parameter agentic small language model for computer use, achieving state-of-the-art performance among models of its size and competitive with larger systems.

0 favorites 0 likes
#small-language-model

Needle: We Distilled Gemini Tool Calling Into a 26M Model

Reddit r/LocalLLaMA · 2026-05-12

Cactus-Compute released Needle, a 26M parameter open-source model distilled from Gemini for efficient on-device function calling using a novel Simple Attention Network architecture without MLPs.

0 favorites 0 likes
#small-language-model

@ash_csx: We’re dropping two open source SLMs this week. 1. One of them matches SOTA accuracy at up to 93x smaller. 2. The other …

X AI KOLs Following · 2026-05-11 Cached

Two new open-source small language models are being released: one matches state-of-the-art accuracy at up to 93x smaller size, and the other outperforms a recent OpenAI model. The first model drops tomorrow.

0 favorites 0 likes
#small-language-model

@AdinaYakup: MiniCPM V4.6 a 1B MLLM that actually runs on your phone, just released by @OpenBMB 1B - Apache2.0 Runs on iOS, Android,…

X AI KOLs Following · 2026-05-11 Cached

OpenBMB has released MiniCPM V4.6, a 1B-parameter multimodal large language model optimized for mobile devices under the Apache 2.0 license. It features mixed visual token compression and claims approximately 1.5x faster throughput than Qwen3.5 0.8B while running natively on iOS, Android, and HarmonyOS.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback