speed

Tag

Cards List
#speed

@HowToPrompt__: You can now run Claude Code 100% locally on a MacBook for $0/month. It's called claude-code-local. It runs a 122B param…

X AI KOLs Timeline ↗ · 2026-06-16 Cached

A new open-source tool called claude-code-local allows running a 122B parameter model locally on a MacBook, achieving 65 tokens per second with full Claude Code support, beating cloud Opus in speed.

0 favorites 0 likes
#speed

Speed Matters for Google Web Search [2009]

Lobsters Hottest ↗ · 2026-06-15 Cached

Google's 2009 study demonstrates that slower search response times reduce user engagement and satisfaction.

0 favorites 0 likes
#speed

Diffusion Gemma is 4x faster, but makes 6x more mistakes!

Reddit r/LocalLLaMA ↗ · 2026-06-12

A benchmark shows Diffusion Gemma is 4x faster than Gemma4 but makes 6x more factual mistakes, especially on obscure topics, trading factual accuracy for smooth text generation.

0 favorites 0 likes
#speed

@steeve: aaaaaand we're faster (i know i know)

X AI KOLs Following ↗ · 2026-06-08 Cached

Steeve Morin reports that after 5 days of work, his implementation is now within 10% of llama.cpp's speed, achieving 64 tok/s vs 70 tok/s, with more work to do.

0 favorites 0 likes
#speed

@aparnadhinak: https://x.com/aparnadhinak/status/2062233330196926720

X AI KOLs Timeline ↗ · 2026-06-03 Cached

A Twitter thread discussing whether a database filesystem abstraction (PostgresFS) or a skill-based approach with local Bash is better for agent workflows. The skill approach wins on composability and speed.

0 favorites 0 likes
#speed

@RuiTheBaker: GPT 5.5-level ranking but 27x faster?! @mixedbreadai

X AI KOLs Following ↗ · 2026-06-02 Cached

Mixedbread's reranker achieves GPT 5.5-level performance on OBLIQ-bench while being 27x faster, according to early results.

0 favorites 0 likes
#speed

@jerryjliu0: Parse PDFs at lightspeed (this video is at 1x) Absolute cinema

X AI KOLs Following ↗ · 2026-05-29 Cached

Jerry Liu announces LiteParse v2, a Rust-based PDF parser that is claimed to be the fastest and most accurate open-source, model-free PDF parser available.

0 favorites 0 likes
#speed

New LFM2.5 8b A1b model!!

Reddit r/LocalLLaMA ↗ · 2026-05-29

Introducing LFM2.5 8b A1b, a new AI model with performance on par with Nemotron 3 Nano but at higher speed. Support is being added to SmallCode for non-standard tool calls.

0 favorites 0 likes
#speed

Fast is better than slow

Lobsters Hottest ↗ · 2026-05-27 Cached

A blog post arguing that speed in software development leads to better learning and decision-making, offering practical advice like avoiding delay and sharing work early.

0 favorites 0 likes
#speed

Gemini 3.5 Flash Looks Good For How Fast It Is (8 minute read)

TLDR AI ↗ · 2026-05-26 Cached

Google released Gemini 3.5 Flash, a hybrid speed model that rivals Opus 4.7 and GPT-5.5 in speed and cost while performing well on agentic and coding benchmarks.

0 favorites 0 likes
#speed

@no_stp_on_snek: In progress

X AI KOLs Following ↗ · 2026-05-23 Cached

Promoting Atlas Inference, an open-source inference serving tool that achieved 200+ tok/s on a Qwen3.6-35B-A3B benchmark.

0 favorites 0 likes
#speed

Designing Firefox for the future

Lobsters Hottest ↗ · 2026-05-22 Cached

Mozilla announces Project Nova, a redesign of Firefox focusing on privacy, speed, and a cleaner, warmer design, with updates to tabs, settings, and compact mode.

0 favorites 0 likes
#speed

40+tok/s - optimized recipe for Qwen 3.5 122B Int4 on a single DGX Spark with vLLM

Reddit r/LocalLLaMA ↗ · 2026-05-20

User shares an optimized recipe for running Qwen 3.5 122B Int4 on a single DGX Spark with vLLM, achieving over 40 tokens per second. They invite others to try and further optimize it.

0 favorites 0 likes
#speed

@YRSM_Simon: This is big news! Kimi 2.6 is a generative-level model. In this age of overflowing LLM capabilities, speed will become the deciding factor in competition. Is the chip sector about to see another 'sector rotation'? 😅

X AI KOLs Following ↗ · 2026-05-20 Cached

Cerebras is now running Kimi K2.6, a trillion-parameter model, in enterprise trials at ~1,000 tokens/s, the fastest frontier model performance ever measured by Artificial Analysis.

0 favorites 0 likes
#speed

@VraserX: Gemini 3.5 Flash might be Google’s most dangerous release yet. The benchmarks are impressive, but the real story is spe…

X AI KOLs Following ↗ · 2026-05-19 Cached

A tweet highlighting Google's Gemini 3.5 Flash as a fast, capable, and affordable AI model release, emphasizing its impressive benchmarks and price/performance ratio.

0 favorites 0 likes
#speed

@julien_c: I've seen some confusion online on how to run llama.cpp with MTP (Multi-token prediction) in the simplest way possible.…

X AI KOLs Following ↗ · 2026-05-19 Cached

Julien C explains how to run llama.cpp with Multi-token prediction (MTP) for ~2x generation speed, using either the Dense 27B or MoE 35B model, with instructions for installation and configuration.

0 favorites 0 likes
#speed

@populartourist: Unsloth Qwen3.6 27B Q6_K doing over 100 t/s with MTP on RTX 5090. That's coming up from 45-50 t/s without MTP. That's i…

X AI KOLs Timeline ↗ · 2026-05-16 Cached

Unsloth Qwen3.6 27B Q6_K achieves over 100 tokens per second with MTP on RTX 5090, up from 45-50 t/s without MTP.

0 favorites 0 likes
#speed

@berryxia: Damn, even my eyes can't keep up with this speed! Daniel Han, founder of UnslothAI, YC S24, previously at NVIDIA doing ML, just released the experimental MTP GGUF of Qwen3.6. The 27B model hits 140 tokens/s on a single GPU. 35B-A...

X AI KOLs Timeline ↗ · 2026-05-14

UnslothAI founder Daniel Han released the experimental MTP GGUF version of Qwen3.6, achieving 140 tokens/s for the 27B model and 220 tokens/s for the 35B-A3B version on consumer GPUs — a 1.4x speedup with zero accuracy loss.

0 favorites 0 likes
#speed

@tan_stack: TanStack Devtools just migrated to @OxcProject parser + magic-string! The results: Per-file transform: 1.65 ms → 0.46 m…

X AI KOLs Following ↗ · 2026-05-13 Cached

TanStack Devtools migrated to OxcProject parser and magic-string, achieving a 3.56× speedup with per-file transform dropping from 1.65 ms to 0.46 ms.

0 favorites 0 likes
#speed

Previewing Ultrafast mode: GPT‑5.6 Sol at up to 14X the speed

YouTube AI Channels ↗ · 2026-08-14 Cached

OpenAI previewed the Ultrafast mode for GPT-5.6 Sol, offering up to 14x faster speed while maintaining intelligence of the same quality as the standard mode, enabling near-real-time experiences in scenarios such as monitoring, data processing, search, and coding.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback