small-language-model

Tag

Cards List
#small-language-model

LFM2.5-2.6B: Deploy Agents Everywhere (8 minute read)

TLDR AI · 4d ago Cached

Liquid AI releases LFM2.5-2.6B, a compact agentic model designed to run entirely on-device, enabling free inference, low latency, and privacy. The post details its training pipeline including SFT, teacher specialization, distillation, and agentic RL.

0 favorites 0 likes
#small-language-model

Deploy local agents everywhere with LFM2.5-2.6B

Hugging Face Blog · 5d ago Cached

Liquid AI releases LFM2.5-2.6B, a compact agentic model designed for on-device deployment, supporting tool calling and multi-step workflows with efficient inference on CPUs and GPUs.

0 favorites 0 likes
#small-language-model

[NEW MODELS!] Supra2-100M Base and Instruct - go check them out!

Reddit r/LocalLLaMA · 5d ago

SupraLabs releases Supra2-100M Base and Instruct models, a new small language model family with community-driven improvements, benchmarks, and a GGUF version.

0 favorites 0 likes
#small-language-model

Models for minimalist RAG: B1ade 335M Embedding and 1B Parameter Small Language Models

arXiv cs.CL · 2026-07-31 Cached

Presents B1ade, a minimalist RAG architecture with a 335M zero-training embedding model and a 1B SLM trained via GRPO on 723M tokens, showing emergent attribution behavior and competitive QA performance without large-scale pretraining.

0 favorites 0 likes
#small-language-model

Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Mode

arXiv cs.CL · 2026-07-27 Cached

Nanbeige4.2-3B is a compact 3B parameter general agentic model pretrained from scratch with a Looped Transformer, achieving strong agentic and reasoning performance, outperforming larger models on diverse benchmarks. The model and code are open-sourced.

0 favorites 0 likes
#small-language-model

Agentic AI at Two Different Scales: Nanbeige4.2-3B and Laguna S2.1 (9 minute read)

TLDR AI · 2026-07-27 Cached

Compares two new AI models for agentic workloads: the compact Nanbeige4.2-3B with a looped transformer architecture and the large Mixture-of-Experts Laguna S2.1, both released on Hugging Face.

0 favorites 0 likes
#small-language-model

OpenLanguageModel: Readable and Composable Small-Language-Model Pretraining for Education and Research

arXiv cs.CL · 2026-07-21 Cached

OpenLanguageModel (OLM) is an open-source PyTorch library for building and pretraining small language models with readable, composable architecture code, bridging education and research.

0 favorites 0 likes
#small-language-model

openbmb released MiniCPM5-2B, not yet available at huggingface

Reddit r/LocalLLaMA · 2026-07-20

OpenBMB released MiniCPM5-2B, a 2 billion parameter language model, currently not yet available on Hugging Face.

0 favorites 0 likes
#small-language-model

A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism

arXiv cs.AI · 2026-07-15 Cached

This paper investigates whether GRPO post-training improves a small (4B-8B) language and vision-language model web agent. It finds a controlled null result: no configuration yields credible gains on mastered tasks, and moderate-to-high learning rates cause degradation or collapse, revealing a double dissociation between degrade and collapse regimes.

0 favorites 0 likes
#small-language-model

Index SLM Technical Report

arXiv cs.CL · 2026-07-14 Cached

Bilibili releases Index-1.9B, a series of open small language models pre-trained on 2.8 trillion tokens, achieving competitive performance on benchmarks. The four models include base, pure (no instruction data), chat, and a character model with retrieval-augmented generation for role-playing.

0 favorites 0 likes
#small-language-model

I developed a 270 million parameter language model entirely from scratch as an independent research project

Reddit r/LocalLLaMA · 2026-07-05 Cached

A 270M parameter language model trained from scratch on English Wikipedia and instruction-tuned for conversational AI, developed as an independent research project.

1 favorites 1 likes
#small-language-model

The Wiola Architecture for Efficient Small Language Models

arXiv cs.AI · 2026-07-03 Cached

Wiola is a novel Small Language Model architecture introducing five independently designed components—SRPE, GCLA, ATM, DSFF, and WiolaRMSNorm—aimed at improving efficiency and coherence, released in sizes from 120M to 1.5B parameters and integrated with HuggingFace Transformers.

0 favorites 0 likes
#small-language-model

@akshay_pachaar: If you use LLM-as-judge, this one is for you. (bookmark it) Most teams validate their agent's outputs by calling a fron…

X AI KOLs Following · 2026-06-30 Cached

Details an approach to train a small LLM judge for evaluating agent outputs, replacing costly frontier models, with a Claude Code plugin for deployment.

0 favorites 0 likes
#small-language-model

@liquidai: Introducing LFM2.5-230M: our smallest model yet, built to run fast anywhere (CPUs, NPUs, and GPUs) to enable agentic ta…

X AI KOLs Timeline · 2026-06-25 Cached

Liquid AI releases LFM2.5-230M, a small 230M parameter model optimized for fast inference on CPUs, NPUs, and GPUs, targeting agentic tasks on devices like phones and robots.

0 favorites 0 likes
#small-language-model

LiquidAI/LFM2.5-230M

Hugging Face Models Trending · 2026-06-24 Cached

Liquid AI released LFM2.5-230M, a compact 230M-parameter hybrid model optimized for on-device deployment with fast edge inference speeds (213 tok/s on Galaxy S25 Ultra) and built for agentic tasks via reinforcement learning.

0 favorites 0 likes
#small-language-model

[NEW MODEL] SupraLabs just released supra-title-FFT-preview, 115K samples, almost 10x our first chat title dataset

Reddit r/LocalLLaMA · 2026-06-20

SupraLabs released supra-title-FFT-preview, a full fine-tuned 0.4B parameter model for chat title generation, trained on 115K samples — nearly 10x larger than their previous dataset.

0 favorites 0 likes
#small-language-model

@xdotli: my friend @xeophon thinks coding is solved here's validation that a 3b model is trained with focus on algo efficiency a…

X AI KOLs Timeline · 2026-06-20 Cached

Nanbeige 4.1, a 3B model, outperforms Qwen3-30b-A3b and Qwen 3.5 4b in coding tasks with focus on algorithmic efficiency, achieving long horizon tasks with 600+ tool calls.

0 favorites 0 likes
#small-language-model

@cjzafir: A 3B parameter SLM: VibeThinker (fine-tuned on Qwen 2.5) matches Claude Opus 4.5 performance. Same performance as: > De…

X AI KOLs Timeline · 2026-06-17 Cached

VibeThinker, a 3B parameter model fine-tuned on Qwen 2.5, achieves performance comparable to Claude Opus 4.5 and much larger models like DeepSeek v3 through innovative post-training that includes multi-path thinking and staged training on math, coding, and science.

0 favorites 0 likes
#small-language-model

Microsoft Tests Phi Silica for Windows AI on Nvidia GPUs (6 minute read)

TLDR AI · 2026-06-17 Cached

Microsoft is testing Phi Silica support on Nvidia GPUs, allowing developers to run the small language model locally on Windows devices with RTX 30-series or newer GPUs, though it lacks NPU-only features like prompt compression.

0 favorites 0 likes
#small-language-model

Why Weibo's tiny VibeThinker-3B has the AI world arguing over benchmarks again (15 minute read)

TLDR AI · 2026-06-17 Cached

Weibo's VibeThinker-3B, a 3B parameter model, claims to match or exceed the reasoning performance of much larger models like DeepSeek V3.2 and Gemini 3 Pro on math and coding benchmarks, sparking debate over benchmark reliability and the necessity of scaling.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback