inference-speed

Tag

Cards List
#inference-speed

Trained my first small language model

Reddit r/LocalLLaMA ↗ · 5h ago

The author trained a small language model to replace Gemini Flash for a summarization task, achieving 97% accuracy with 0.06s latency, suitable for deployment in an internal app.

0 favorites 0 likes
#inference-speed

I added OpenVINO support to Laya: 40 ms per question on CPU, 3.4x faster than PyTorch

Reddit r/LocalLLaMA ↗ · 23h ago

Added OpenVINO support to Laya, achieving 40 ms per question on CPU, which is 3.4 times faster than PyTorch.

0 favorites 0 likes
#inference-speed

@mattshumer_: I've been testing GPT-6 Sol for a bit now. It's solid, but I still prefer Astra/Fable 5.1 (and now, likely Opus 5.5) fo…

X AI KOLs Following ↗ · 3d ago Cached

Matt Shumer tests GPT-6 Sol and shares his preference for Astra/Fable 5.1 and Opus 5.5 models, while referencing OpenAI's announcement of faster and more affordable GPT-6 Sol and Luna models.

0 favorites 0 likes
#inference-speed

@FinanceYF5: 1/ Claude has optimized the inference performance of over 30 open-source biomolecular models, boosting their average ru…

X AI KOLs Timeline ↗ · 5d ago Cached

Claude has optimized the inference performance of over 30 open-source biomolecular models, boosting their average running speed by 4 times with custom GPU software and open-sourced code.

0 favorites 0 likes
#inference-speed

I tested Qwen3.8 27B IQ3_XXS (10.18GiB) vs Bonsai Ternary PQ2 (6.42GiB)

Reddit r/LocalLLaMA ↗ · 6d ago

The author compared the performance of Qwen3.8 27B IQ3_XXS and Bonsai Ternary PQ2 on limited VRAM, finding that Qwen is faster and uses fewer tokens, while Bonsai has a smaller file size but longer generation times.

0 favorites 0 likes
#inference-speed

@LangChain: Hot topic livestream: Learn about Jev A buzzy new model Jev, by @typesafeai, reports up to 200x faster inference and 40…

X AI KOLs Timeline ↗ · 2026-09-18 Cached

LangChain is hosting a livestream to discuss TypeSafe AI's new model Jev, which promises up to 200x faster inference and 400x lower cost for classification tasks, aiming to improve agent loops.

0 favorites 0 likes
#inference-speed

Bonsai 1.7B can solve simple physics problems on a 12 watt intel N97 at about 9.1 t/s

Reddit r/LocalLLaMA ↗ · 2026-09-18

The Bonsai 1.7B AI model can solve simple physics problems efficiently on a low-power Intel N97 processor, achieving inference speeds of about 9.1 tokens per second.

0 favorites 0 likes
#inference-speed

@usutaku_channel: To verify the swift judgment of "Jev," which is currently the hottest topic, I tried having it classify emails. The one…

X AI KOLs Timeline ↗ · 2026-09-18

This article tests the 'Jev' AI model for email classification, comparing it with other fast models from AI companies, and reports that 'Jev' performed best.

0 favorites 0 likes
#inference-speed

@sydneyrunkle: https://x.com/sydneyrunkle/status/2100754364545761643

X AI KOLs Timeline ↗ · 2026-09-18 Cached

TypeSafe AI releases Jev, a System One model for fast, structured decisions in agent loops, offering up to 200x faster inference and 400x lower cost for classification tasks compared to traditional LLMs.

0 favorites 0 likes
#inference-speed

600tok/s single request on qwen3.6 35ba3b with Ninfer on an RTX Pro 6000. Anybody remember that Comcast ad "stupid fast"?

Reddit r/LocalLLaMA ↗ · 2026-09-17

Achieving 600 tokens per second on the Qwen3.6 model using Ninfer on an RTX Pro 6000, noted as useful for brute-force tasks despite not being the most advanced model.

0 favorites 0 likes
#inference-speed

DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression

Hacker News Top ↗ · 2026-09-17 Cached

This article analyzes the DeepSeek-V4.1 Flash model, detailing its technical report on KV cache compression and architectural optimizations that enable efficient long-context processing and high-speed inference.

0 favorites 0 likes
#inference-speed

Breaking the 1.58-bit Barrier for Ternary LLMs

arXiv cs.AI ↗ · 2026-09-16 Cached

This paper introduces BITCOS, a distribution-adaptive layout for storing ternary LLM weights more efficiently, achieving up to 1.28× speedup in matrix-vector multiplication and 1.27× in inference throughput on GPUs.

0 favorites 0 likes
#inference-speed

@gabriel1: we forgot how insane the concept of "doing parallel work in the background" and i'm so excited for models that are astr…

X AI KOLs Timeline ↗ · 2026-09-15 Cached

A user expresses excitement for future AI models that can perform parallel tasks in the background and respond within 10 seconds, highlighting how faster models could transform user experience.

0 favorites 0 likes
#inference-speed

@dongxi_nlp: The speed of LLMs involves two types of experiences: how long it takes to start responding, and whether the response is…

X AI KOLs Timeline ↗ · 2026-09-15 Cached

The article explains two key metrics for LLM response speed: Time to First Token (TTFT) related to prefill, and Inter-Token Latency (ITL) related to decode, affecting responsiveness and smoothness.

0 favorites 0 likes
#inference-speed

Would you consider 5t/s usable for a local model?

Reddit r/LocalLLaMA ↗ · 2026-09-09

User discusses the usability of running the Qwen3.8 27b model locally at 5 tokens per second, comparing performance on different hardware setups and noting that lower speed can still be acceptable if system resources are managed.

0 favorites 0 likes
#inference-speed

GLM 5.3 Flash Q4 @ 60tps / 550tps on M3 Ultra

Reddit r/LocalLLaMA ↗ · 2026-09-09

Optimizations for GLM 5.3 Flash on Apple M3 Ultra achieve up to 550 t/s prefill and 38 t/s inference speed through kernel fusion and efficient memory use, without quality loss.

0 favorites 0 likes
#inference-speed

@MiaAI_lab: GLM 5.3 Flash on 2x DGX Sparks Look at the tok/s at the top :) This is prose. Releasing soon

X AI KOLs Timeline ↗ · 2026-09-08 Cached

GLM 5.3 Flash is an upcoming AI model that showcases high token-per-second performance on 2x DGX Spark systems, with an imminent release.

0 favorites 0 likes
#inference-speed

@LotusDecoder: DeepSeek-V4.1-Flash-0910 decode 400 token/s 😋 Would deploying this to my home DGX spark also achieve this speed?

X AI KOLs Timeline ↗ · 2026-09-08

A user on X/Twitter asks if deploying the DeepSeek-V4.1-Flash-0910 model on a home DGX Spark could achieve a decode speed of 400 tokens per second.

0 favorites 0 likes
#inference-speed

Qwen3.8-Flash-Next on 2x3090 + DDR4: 17 → 25-29 t/s decode with the expert cache PR

Reddit r/LocalLLaMA ↗ · 2026-09-03

User benchmarks and details a GPU-resident expert cache PR in llama.cpp that boosts decode speed for Qwen3.8-Flash-Next on a dual RTX 3090 system from 17 to 25-29 tokens per second.

0 favorites 0 likes
#inference-speed

Confirmed bolting Q8 NGram into IQ4 Qwen no speed degradation

Reddit r/LocalLLaMA ↗ · 2026-09-02

A user replaced the lower-precision N-gram layer in a Qwen model with Q8 quantization and found no significant speed degradation during inference, with output quality still being tested.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback