high-speed-inference

Tag

Cards List
#high-speed-inference

@StefanoErmon: Today we're excited to announce Mercury 2.5 It’s the most capable diffusion LLM on the market. It is a 40% jump in inte…

X AI KOLs Timeline ↗ · 2026-09-08 Cached

Mercury 2.5 is announced as the most capable diffusion language model, with a 40% increase in intelligence over Mercury 2, operating at over 1,100 tokens/sec on NVIDIA GPUs, and optimized for production with low latency and cost.

0 favorites 0 likes
#high-speed-inference

Nifer is insane. 700t/s with Qwen 3.6 35B (no thinking). Purpose build for RTX5090. Full 250k context too.

Reddit r/LocalLLaMA ↗ · 2026-07-27

Nifer is a tool that achieves 700 tokens per second inference on Qwen 3.6 35B without thinking, specifically optimized for the RTX 5090, with support for full 250k context.

0 favorites 0 likes
#high-speed-inference

@ClementDelangue: Kog open-sourced on @huggingface the 2B model that they used to show a model running at 3,000+ tokens per second. Very …

X AI KOLs Timeline ↗ · 2026-06-24 Cached

Kog has open-sourced the Laneformer 2B model, a 2.3B parameter instruction-tuned coding model designed for high-speed decoding, achieving over 3,000 tokens per second by prioritizing latency from the architecture stage.

0 favorites 0 likes
#high-speed-inference

@zephyr_z9: This is super big I think this is the first useful speculative decoding method deployed on a big quasi frontier model M…

X AI KOLs Following ↗ · 2026-06-08 Cached

Xiaomi MiMo releases MiMo-V2.5-Pro-UltraSpeed, achieving over 1,000 tokens per second on a 1 trillion parameter model using speculative decoding, the first practical deployment of such speed at scale.

0 favorites 0 likes
← Back to home

Submit Feedback