Tag
A new open-source tool called claude-code-local allows running a 122B parameter model locally on a MacBook, achieving 65 tokens per second with full Claude Code support, beating cloud Opus in speed.
Google's 2009 study demonstrates that slower search response times reduce user engagement and satisfaction.
A benchmark shows Diffusion Gemma is 4x faster than Gemma4 but makes 6x more factual mistakes, especially on obscure topics, trading factual accuracy for smooth text generation.
Steeve Morin reports that after 5 days of work, his implementation is now within 10% of llama.cpp's speed, achieving 64 tok/s vs 70 tok/s, with more work to do.
A Twitter thread discussing whether a database filesystem abstraction (PostgresFS) or a skill-based approach with local Bash is better for agent workflows. The skill approach wins on composability and speed.
Mixedbread's reranker achieves GPT 5.5-level performance on OBLIQ-bench while being 27x faster, according to early results.
Jerry Liu announces LiteParse v2, a Rust-based PDF parser that is claimed to be the fastest and most accurate open-source, model-free PDF parser available.
Introducing LFM2.5 8b A1b, a new AI model with performance on par with Nemotron 3 Nano but at higher speed. Support is being added to SmallCode for non-standard tool calls.
A blog post arguing that speed in software development leads to better learning and decision-making, offering practical advice like avoiding delay and sharing work early.
Google released Gemini 3.5 Flash, a hybrid speed model that rivals Opus 4.7 and GPT-5.5 in speed and cost while performing well on agentic and coding benchmarks.
Promoting Atlas Inference, an open-source inference serving tool that achieved 200+ tok/s on a Qwen3.6-35B-A3B benchmark.
Mozilla announces Project Nova, a redesign of Firefox focusing on privacy, speed, and a cleaner, warmer design, with updates to tabs, settings, and compact mode.
User shares an optimized recipe for running Qwen 3.5 122B Int4 on a single DGX Spark with vLLM, achieving over 40 tokens per second. They invite others to try and further optimize it.
Cerebras is now running Kimi K2.6, a trillion-parameter model, in enterprise trials at ~1,000 tokens/s, the fastest frontier model performance ever measured by Artificial Analysis.
A tweet highlighting Google's Gemini 3.5 Flash as a fast, capable, and affordable AI model release, emphasizing its impressive benchmarks and price/performance ratio.
Julien C explains how to run llama.cpp with Multi-token prediction (MTP) for ~2x generation speed, using either the Dense 27B or MoE 35B model, with instructions for installation and configuration.
Unsloth Qwen3.6 27B Q6_K achieves over 100 tokens per second with MTP on RTX 5090, up from 45-50 t/s without MTP.
UnslothAI founder Daniel Han released the experimental MTP GGUF version of Qwen3.6, achieving 140 tokens/s for the 27B model and 220 tokens/s for the 35B-A3B version on consumer GPUs — a 1.4x speedup with zero accuracy loss.
TanStack Devtools migrated to OxcProject parser and magic-string, achieving a 3.56× speedup with per-file transform dropping from 1.65 ms to 0.46 ms.
OpenAI previewed the Ultrafast mode for GPT-5.6 Sol, offering up to 14x faster speed while maintaining intelligence of the same quality as the standard mode, enabling near-real-time experiences in scenarios such as monitoring, data processing, search, and coding.