Tag
This article provides an update on the Engram model, detailing its 2.6b parameter architecture with a large Engram table and initial training progress at 100m tokens, showing improved completions with context-aware data offloading.
The paper details ufakzeka-1, a 151M-parameter Turkish language model built from scratch with a total cost of about $286, describing the training pipeline, evaluation methods, and key findings on small-model training limitations.
AgentCloak is a free browser extension that uses a small AI model to locally detect and replace personal details in AI prompts, ensuring privacy across services like ChatGPT and Claude.
Cua has open-sourced CUA-S1-FORMS, a tiny 2.8MB AI model specialized for form-filling tasks, achieving 99.7% accuracy and enabling local deployment.
Cactus Needle 3 is a small, sliceable foundation model for automation tasks that runs on-device, achieving performance comparable to larger models on function calling and structured extraction.
The author provides an update on building a small 2B parameter AI model with an Engram component, trained on 15m tokens from Wikipedia to achieve surprising coherence, with plans for an Apache 2.0 open-source release.
The article describes 'jevlike', a reverse-engineered AI model for selecting among text options in one pass, with demos on games like Doom and chess, hosted on GitHub.
Needle 3 is a compact AI foundation model optimized for edge devices like mobiles and wearables, offering tool calling, structured extraction, and text embedding in a single 8-29 MB file.
SHADOW 50M is a compact 44M-parameter AI model with ternary weights and fixed 512-bit codes, trained from scratch on 45B tokens, enabling offline inference at high speed on consumer hardware.
Microsoft Research presents FrogNano, a 4B coding agent trained exclusively via reinforcement learning with online task synthesis, achieving competitive performance without distillation from larger models.
Tencent has open-sourced a 1.5B parameter AI model called AuK that can replace multiple audio tools, handling tasks like TTS, voice cloning, and denoising via natural language instructions.
A tiny 27.5KB AI model called gpu-lexer from Vercel Labs runs locally in the browser using WebGPU to syntax-highlight code in any programming language without prior knowledge.
The article asks users which small AI models are most useful and mentions three competing models in the same size class.
OpenBMB releases MiniCPM5-2B, a 2B-parameter dense Transformer model achieving state-of-the-art performance in its size class for on-device deployment, along with open-source high-quality training datasets.
sanoTTS is a family of compact TTS models, with the smallest being 294k parameters, optimized for microcontrollers and outperforming larger models in benchmarks.
This guide details a fine-tuning recipe using Group Relative Policy Optimization (GRPO) with the TRL library to enhance the LFM2.5-350M model's structured output compliance, improving IFStruct benchmark performance from 22.6% to 29.7%.
A multilingual 3.7B parameter reasoning MoE model has been pretrained from scratch on a consumer-grade GPU over several months, with support for 13 languages and available on Hugging Face.
Needle 2 is an open, 45M-parameter AI model for tool calling and structured extraction, optimized to run in browsers at 14MB with guaranteed JSON output via constrained sampling.
The article describes the development of a 48M parameter model specialized for tool calling in AI agents, which uses grammar to ensure valid JSON outputs and is open-source for customization on specific API catalogs.
The tweet praises the LFM2.5-2.6B model for its strong performance in on-device evaluations, highlighting its generalization capabilities beyond agentic tasks in partnership with Artificial Analysis for testing on mobile devices.