A.L.F.R.E.D. - 2B models with template can match 35B Models with 4x less speed
Summary
A.L.F.R.E.D. proposes a system that distills knowledge from large models into small ones and routes simple tasks to the small models, achieving performance of 35B models with 2B models while reducing inference cost by 4x.
Similar Articles
@jinyuhou0: On popular benchmarks, our 30B model matches systems 20-30x its size (gpt-5.4-xhigh, DeepSeek-V3.2, Kimi-K2.5), while u…
A new 30B model matches systems 20-30x its size on popular benchmarks while using up to 95% fewer reasoning tokens than comparable agentic LLMs, achieved through a learned configurator that decides when and how to reason. Model and code are openly available.
@sheriyuo: A 35B-parameter MoE agentic model with only 3B active that claims to match or surpass 100B-class models through post-tr…
A 35B-parameter MoE model with only 3B active parameters matches or surpasses 100B-class models using post-training RL, achieving significant efficiency gains.
@liquidai: Introducing LFM2.5-230M: our smallest model yet, built to run fast anywhere (CPUs, NPUs, and GPUs) to enable agentic ta…
Liquid AI releases LFM2.5-230M, a small 230M parameter model optimized for fast inference on CPUs, NPUs, and GPUs, targeting agentic tasks on devices like phones and robots.
A 4b model is now beating 30b ones at web research and the reason is not size
A 4 billion parameter open model from the Apodex family outperforms 30 billion parameter models on web research benchmarks, attributed to careful training data and self-verification techniques rather than raw scale, suggesting a more democratic trajectory for AI capability.
@VukRosic99: Li Auto's Mach-Mind-4-Flash: a 35B MoE with only 3B active params that matches 100B-class models through post-training …
Li Auto's Mach-Mind-4-Flash is a 35B MoE model with only 3B active parameters that, through post-training optimization using a unified RL/on-policy distillation loss, achieves performance matching or surpassing 100B-parameter models across multiple benchmarks.