llm-performance

Tag

Cards List
#llm-performance

@DavidFSWD: THE MOST Amzing thing in Local AI this week is this optimized Qwen 3.8 model: ukisai/Swift-1.5-Qwen3.8-27B-GSQ-RCO-GGUF…

X AI KOLs Timeline ↗ · 6d ago Cached

This tweet highlights an optimized version of the Qwen 3.8 model that can run efficiently on local hardware with under 16GB VRAM, achieving 30-80 tokens per second on older GPU setups.

0 favorites 0 likes
#llm-performance

@wafer_ai: Wafer beat Cerebras on latency for @ycombinator's AI Office Hours. GLM-5.2 on Wafer averaged 379 ms versus 674 ms for G…

X AI KOLs Following ↗ · 2026-09-25 Cached

Wafer outperformed Cerebras on latency for YC's AI Office Hours, with GLM-5.2 achieving 379 ms versus 674 ms for Gemma 4, resulting in 44% lower latency and increased user engagement.

0 favorites 0 likes
#llm-performance

@no_stp_on_snek: Great video @digitalix . Apples new little toaster is a beast! https://youtu.be/c_58D7ixOQI Now which kidney do I need …

X AI KOLs Following ↗ · 2026-09-23 Cached

Apple's M5 Ultra chip shows up to 1.5x faster token generation and 4x faster prompt processing for local LLMs compared to the M3 Ultra, with improved memory bandwidth and storage speeds, but at a high cost.

0 favorites 0 likes
#llm-performance

Upgraded my local setup with 2 rtx pros and it's amazing.

Reddit r/LocalLLaMA ↗ · 2026-09-16

A software developer shares their experience upgrading a local setup with two RTX Pro GPUs, troubleshooting power issues, and achieving high performance for running LLMs like Qwen and Deepseek. They discuss configuration details and seek advice on optimizations and model suggestions.

0 favorites 0 likes
#llm-performance

@dongxi_nlp: The speed of LLMs involves two types of experiences: how long it takes to start responding, and whether the response is…

X AI KOLs Timeline ↗ · 2026-09-15 Cached

The article explains two key metrics for LLM response speed: Time to First Token (TTFT) related to prefill, and Inter-Token Latency (ITL) related to decode, affecting responsiveness and smoothness.

0 favorites 0 likes
#llm-performance

Nex-N2.5-mini-MLX-4bit on Apple M5 Max — 133.6 tok/s — llm-bench.io

Reddit r/LocalLLaMA ↗ · 2026-09-12 Cached

Benchmark results for the Nex-N2.5-mini-MLX-4bit model on Apple M5 Max hardware, achieving 133.6 tokens per second generation speed and quality scores up to 85.80 in research tasks.

0 favorites 0 likes
#llm-performance

3.8-27B has ruined 3.5/3.6-35B’s for me. It’s just *absurdly* superior.

Reddit r/LocalLLaMA ↗ · 2026-09-12

The author praises the superior performance and efficiency of the 3.8-27B AI model over others like 3.5/3.6-35B, highlighting its attention to detail and lower token usage in replicated projects.

0 favorites 0 likes
#llm-performance

Ran GPT-6 Astra on an expert level medical knowledge benchmark. Scored 87.4% beating every other LLM

Reddit r/ArtificialInteligence ↗ · 2026-09-10

GPT-6 Astra scored 87.4% on the MedXpertQA medical benchmark, surpassing all other LLMs and marking a significant improvement from GPT-4o's 42.8% score two years ago.

0 favorites 0 likes
#llm-performance

MoE offloaded - advice on difference between Intel vs AMD CPU instruction sets

Reddit r/LocalLLaMA ↗ · 2026-09-08

The article discusses performance differences between Intel and AMD CPUs for LLM inference, comparing instruction sets and memory bandwidth impacts on offloading speeds.

0 favorites 0 likes
#llm-performance

CyrillicQA: The Influence of Phonetically Encoded Secret Language on LLM Performance

arXiv cs.CL ↗ · 2026-08-25 Cached

The paper investigates whether large language models can phonetically decode encoded languages, such as German written in Cyrillic characters, to assess their abstraction capabilities and performance beyond standard Latin-based training data.

0 favorites 0 likes
#llm-performance

Ornith-1.5-35B-A3B Q4 running 60tk/s on 4070Ti.

Reddit r/LocalLLaMA ↗ · 2026-08-20

A user reports achieving 60 tokens per second with the Ornith-1.5-35B-A3B MOE model on an NVIDIA RTX 4070 Ti, demonstrating efficient local inference optimization.

0 favorites 0 likes
#llm-performance

I'm (mostly) picking models on speed now, not intelligence

Lobsters Hottest ↗ · 2026-08-02 Cached

The author argues that frontier LLMs have reached a 'good enough' intelligence threshold, so they now prioritize speed over raw intelligence when choosing models, citing fast open-weights models like GLM5.2 and DeepSeek V4 Flash as daily drivers.

0 favorites 0 likes
#llm-performance

How fast is 10 tokens per second really?

Simon Willison's Blog ↗ · 2026-05-20 Cached

Simon Willison explores the practical meaning of 10 tokens per second speed for large language models, offering context on how fast that feels and its implications for usability.

0 favorites 0 likes
#llm-performance

Position: Let's Develop Data Probes to Fundamentally Understand How Data Affects LLM Performance

arXiv cs.AI ↗ · 2026-05-20 Cached

This position paper advocates for developing 'data probes'—synthetic sequences from random processes—to systematically study how data characteristics affect LLM performance, aiming to move beyond empirical heuristics.

0 favorites 0 likes
#llm-performance

@ClementDelangue: Local open-weight AI on a laptop has been improving more than twice as fast as Moore's Law! Between May 2024 and May 20…

X AI KOLs Following ↗ · 2026-05-11

Hugging Face CEO Clement Delangue claims local open-weight AI performance on laptops is improving 4.7x faster than Moore's Law, citing progress from Llama 3 70B to DeepSeek V4 Flash on unchanged hardware.

0 favorites 0 likes
#llm-performance

@davis7: @0xSero helped me setup local models properly and I uh, had no idea these things had gotten this good Are they frontier…

X AI KOLs Following ↗ · 2026-05-09

The author highlights the impressive capabilities of the open-source Qwen 3.6-27B model running locally on an RTX 5090, noting its strong performance on programming tasks and comparing it favorably to commercial models, despite the complexity of local deployment.

0 favorites 0 likes
#llm-performance

@GigaAI: Introducing hallucination correction. We have reduced hallucination by 70%. Giga's hallucination rate is at ~1%. Better…

X AI KOLs Timeline ↗ · 2026-05-07 Cached

GigaAI announces a new hallucination correction feature that reduces the model's hallucination rate to approximately 1%, claiming superior reliability compared to frontier models.

0 favorites 0 likes
← Back to home

Submit Feedback