tiny-models

Tag

Cards List
#tiny-models

Why are tiny models (<50M parameters) or swarms of specialised micro-models so rarely deployed in production?

Reddit r/LocalLLaMA · 6d ago

The author questions why tiny AI models with fewer than 50M parameters or swarms of specialized micro-models are rarely deployed in production, speculating on reasons like tooling biases or the convenience of generalist models.

0 favorites 0 likes
#tiny-models

@tobi: Training tiny models for special purpose use cases works so incredibly well if you have a great self improving recursiv…

X AI KOLs Timeline · 2026-09-01 Cached

A tweet from @tobi discusses how training tiny models for specialized use cases with a self-improving flywheel is highly effective, noting that Shopify ML team's finetuned 0.8b model outperforms GPT 5.6-sol xhigh in a specific task.

0 favorites 0 likes
#tiny-models

@svpino: Tiny, specialized, and open models are the future! The TwiL-LM family of models is now available on HuggingFace for Tra…

X AI KOLs Timeline · 2026-08-10 Cached

TwiL-LM, a family of tiny specialized open models, is released on HuggingFace. The 3B version outperforms OpenAI's 120B gpt-oss on formal reasoning benchmarks and runs efficiently on consumer hardware.

0 favorites 0 likes
#tiny-models

Running a 28.9M parameter LLM on an $8 microcontroller

Hacker News Top · 2026-07-25 Cached

A developer demonstrates running a 28.9 million parameter language model on an $8 ESP32-S3 microcontroller using Google's Per-Layer Embeddings to store most parameters in flash, achieving around 9.5 tokens per second on-device text generation.

0 favorites 0 likes
#tiny-models

[BIG DATASET RELEASE] - SupraLabs/reasoning-corpus-4K-5M-v1 - Train your tiny SLMs to think!

Reddit r/LocalLLaMA · 2026-07-24

SupraLabs releases reasoning-corpus-4K-5M-v1, a 5M-sample reasoning dataset for training small language models (SLMs), featuring chain-of-thought traces and ChatML format, hosted on Hugging Face.

0 favorites 0 likes
#tiny-models

@HarshalsinghCN: introducing tinyrouter i reverse engineered the routing architecture behind Skana AI's Fugu and built replication for o…

X AI KOLs Timeline · 2026-07-04 Cached

TinyRouter is a tiny 10K-parameter LLM router that learns to route each question to the best specialist model from a pool of open-source LLMs, using evolutionary training. It achieves performance matching or exceeding individual models on MMLU and math benchmarks.

0 favorites 0 likes
#tiny-models

Tiny Scale Is All I Can Spare To Play With Transformer

Reddit r/LocalLLaMA · 2026-06-11

A student introduces Silia, a novel transformer architecture that combines attention and FFN into a unified operation to save parameters at scales ≤10M, achieving comparable performance to GPT-2 with fewer parameters despite limited compute resources.

0 favorites 0 likes
#tiny-models

@neural_avb: Next video is on training tiny (<1B) models for preference tuning. Plus how to generate preference datasets with local …

X AI KOLs Timeline · 2026-05-26 Cached

Announces an upcoming video on training tiny models for preference tuning, covering reward models, RLHF, DPO, ORPO with Unsloth and TRL.

0 favorites 0 likes
#tiny-models

Flavors of Moonshine: Tiny Specialized ASR Models for Edge Devices

Papers with Code Trending · 2025-09-02 Cached

This paper presents Flavors of Moonshine, a suite of tiny specialized ASR models for edge devices. The authors show that monolingual models trained on a balanced mix of human-labeled, pseudo-labeled, and synthetic data outperform larger multilingual models like Whisper, achieving state-of-the-art error rates for small models and enabling on-device ASR for underrepresented languages.

0 favorites 0 likes
← Back to home

Submit Feedback