Tag
A 35B-parameter MoE model with only 3B active parameters matches or surpasses 100B-class models using post-training RL, achieving significant efficiency gains.
LuxTTS is a lightweight voice cloning TTS model, supporting 48kHz high-fidelity output, achieving 150x real-time speed on a single GPU, requiring only 1GB VRAM for local operation, with performance comparable to models ten times its size.
A user praises the Gemma 4 e2b model for its speed and output quality on low-end hardware, comparing it favorably to ChatGPT 3.5 and 4, and asks for recommendations on other small models that work well on older computers.
InnerZoom proposes a single-forward framework for cross-layer evidence bridging in GUI grounding, achieving state-of-the-art performance on multiple benchmarks while reducing latency by up to 31.8%.
Cohere Labs releases North Mini Code, a 30B parameter (3B active) open-source coding model under Apache 2.0, optimized for code generation and agentic tasks, capable of running locally with 20GB RAM via 4-bit quantization.
Lite Any Stereo V2 presents an efficient stereo matching approach achieving state-of-the-art accuracy with significantly reduced latency through optimized architecture and training strategies, including a 2D-only cost aggregation framework and a three-stage training strategy.
Unsloth enables free fine-tuning of a 31B parameter multimodal model on Kaggle using 4-bit quantization, requiring only 22-24GB VRAM for local runs.
Introduces SEVRA, a selective verification controller for budget-aware reasoning that decides when to accept a model's initial answer versus spending extra compute on verification, improving accuracy and reducing unnecessary tokens on benchmarks like MATH500 and GSM8K.
VibeThinker, a 3B parameter model fine-tuned on Qwen 2.5, achieves performance comparable to Claude Opus 4.5 and much larger models like DeepSeek v3 through innovative post-training that includes multi-path thinking and staged training on math, coding, and science.
Baidu Wenxin releases PP-OCRv6, offering three model tiers: Tiny, Small, and Medium, supporting over 50 languages. The Tiny version is only 1.5MB and can run locally in a browser, with the fastest single-image inference at 97ms, proving that small specialized models can outperform large models on OCR tasks.
Cohere released North Mini Code, its first open-source coding model under Apache 2.0, designed to be small, cost-effective, locally deployable, and focused on agentic performance.
A team of 8 released a 30B-A3B coding model under Apache 2.0 that matches Claude Haiku 4.5 performance and beats NVIDIA's 120B-A12B Nemotron 3 Super on the Artificial Analysis Coding Index.
Cohere released its first open-source coding model, North Mini Code Small, designed for efficient agentic performance and community input.
The author demonstrates that the Gemma-4-26B-A4B model runs efficiently on a CPU-only system using Koboldcpp, achieving 7 tokens per second on an old desktop, suggesting that powerful GPUs may not be necessary for local LLM inference.
We just launched Gemma 4 12B, a mid-sized multimodal model with native audio inputs, requiring only 16GB memory and released under Apache 2.0.
Presents WeCon, a weight-conditioned neural solver for multi-objective combinatorial optimization problems that achieves comparable hypervolume to the state-of-the-art while reducing inference time by 40%.
OpenBMB releases MiniCPM-V 4.6, a 1.3B-parameter multimodal LLM with 262k context and significantly reduced visual encoding FLOPs, achieving strong benchmark performance and broad inference framework support.
Proposes delta-Mem, a lightweight online memory mechanism that uses a compact state matrix updated by delta-rule learning to improve long-context performance of frozen LLMs without full fine-tuning or context extension.
SANA World Model is a new AI model that uses hybrid linear attention for efficiency and speed.
SANA-WM is a 2.6B-parameter open-source world model that generates high-fidelity 720p minute-scale videos with precise camera control, achieving industrial-level quality while significantly reducing computational requirements.