ai-inference

Tag

Cards List
#ai-inference

Qwen 3.8 2.4T at 288k tokens/s on Nvidia GB300 NVL72

Reddit r/LocalLLaMA · 3h ago

NVIDIA showcases the high-throughput performance of serving the Qwen3-8B 2.4T parameter model on GB300 NVL72 hardware, achieving over 4k tokens per second per GPU.

0 favorites 0 likes
#ai-inference

5090: Windows or Linux for Qwen3.8.27b

Reddit r/LocalLLaMA · 16h ago

User seeks advice on the best operating system (Windows or Linux) and inference server to run the Qwen3.8.27b model on a dedicated AI rig with RTX 5090 and 96GB RAM for optimal performance.

0 favorites 0 likes
#ai-inference

A 397-Billion AI Just Ran on an iPhone

Reddit r/ArtificialInteligence · 23h ago

A 397-billion-parameter AI model has been successfully run on an iPhone, demonstrating on-device AI capabilities with a mixture-of-experts design, though facing challenges in speed, storage, and heat.

0 favorites 0 likes
#ai-inference

LLMRouter open-sources 16+ router library with xRouteBench

Reddit r/ArtificialInteligence · yesterday

A new paper on arXiv introduces an open-source library called LLMRouter with over 16 router implementations and a benchmark xRouteBench, demonstrating that learned routers can outperform fixed-model baselines by 14.6%.

0 favorites 0 likes
#ai-inference

@akshay_pachaar: GPU architecture, clearly explained. The usual assumption is that a faster GPU means more compute, so a chip rated for …

X AI KOLs Timeline · yesterday Cached

The article clarifies that GPU performance in AI inference is limited by memory bandwidth rather than compute power, using the NVIDIA H100 as an example to explain GPU architecture and its effect on token generation rates.

0 favorites 0 likes
#ai-inference

Comparing confidential inference APIs

Reddit r/singularity · 2d ago

The article compares confidential inference APIs from Privatemode, Tinfoil, NEAR AI, and Chutes, highlighting their security features like end-to-end encryption and trusted execution environments, along with tradeoffs in model selection and verification maturity.

0 favorites 0 likes
#ai-inference

NInfer day0 support for Qwen3.8 27b: ~200 tok/s generation, with tons of engine improvments

Reddit r/LocalLLaMA · 2d ago

NInfer adds day-0 support for the Qwen3.8-27B model, achieving around 200 tokens per second on a single RTX 5090 with speculative decoding, and includes engine improvements like concurrent requests and kernel optimizations.

0 favorites 0 likes
#ai-inference

Kog is going deeper to squeeze more inference out of GPUs

TechCrunch AI · 2d ago Cached

French startup Kog is using software optimization to dramatically speed up LLM inference on standard datacenter GPUs, claiming up to 30x faster throughput. After a demo of 3,000 tokens per second with its open-sourced Laneformer 2B model, the company is now focusing on larger models and targeting software engineering workflows.

0 favorites 0 likes
#ai-inference

Cascadia Launches Distributed AI Inference for Intel Hardware

Reddit r/artificial · 3d ago

Cascadia has launched a distributed AI inference system designed for Intel hardware, enabling scalable and efficient inference workloads.

0 favorites 0 likes
#ai-inference

I built a weird, low-power llama.cpp server using an Intel N100 + RTX 5060Ti

Reddit r/LocalLLaMA · 5d ago

A hobbyist describes building a low-power llama.cpp server using an Intel N100 motherboard and a refurbished RTX 5060 Ti, sharing performance numbers, power consumption, and model choices.

0 favorites 0 likes
#ai-inference

Two Bets on Standing Still, and a Dark Horse (17 minute read)

TLDR AI · 6d ago Cached

An analysis of AI inference hardware, comparing Taalas and Groq's approaches to etching model weights into silicon, and noting recent investments by Nvidia and AMD.

0 favorites 0 likes
#ai-inference

@robertnishihara: This is *the* vllm event of the year. Inference is one of the most rapidly evolving parts of infra. This is the best ev…

X AI KOLs Following · 2026-08-05 Cached

The vLLM Conference is coming up in San Francisco, Aug 24–26, hosted by Inferact at Ray Summit, featuring speakers from major AI infrastructure companies.

0 favorites 0 likes
#ai-inference

SK hynix, In Collaboration With SanDisk, Unveils The New High Bandwidth Flash (HBF) Standard, Helping To Resolve AI Inference Bottlenecks, Targeting Up To 3TB/s Bandwidth

Reddit r/LocalLLaMA · 2026-08-04 Cached

SK hynix and SanDisk unveiled the High Bandwidth Flash (HBF) standard to bridge the performance gap between HBM memory and SSDs, targeting up to 3TB/s bandwidth and 512GB capacity to speed AI inference.

0 favorites 0 likes
#ai-inference

The inference cost for Astra to solve 10 long-open math problems was roughly $2,000. Lean proofs are on GitHub.

Reddit r/singularity · 2026-08-03

Astra, an unreleased AI system, produced machine-checkable Lean 4 proofs for 10 long-open math problems at roughly $2,000 inference cost, sparking debate about the true cost and significance of AI-discovered mathematics.

0 favorites 0 likes
#ai-inference

70-class VRAM stagnation

Reddit r/LocalLLaMA · 2026-08-03

The author observes that Nvidia's desktop 70-class GPUs have stayed at 12GB VRAM across two generations, and suggests Nvidia may be intentionally limiting memory to preserve demand for higher-margin AI-focused hardware.

0 favorites 0 likes
#ai-inference

Are you ready for Le Chaton FAT or still wasting money on GPUs?

Reddit r/LocalLLaMA · 2026-08-02

The author shares their storage server build optimized for local AI inference, anticipating a rumored 26T-a3b model called "Le Chaton FAT" and using high-capacity NVMe drives with ZFS for model storage.

0 favorites 0 likes
#ai-inference

@ycombinator: At our latest YC Paper Club, researchers and builders presented on multi-GPU kernels, intelligence per watt, heterogene…

X AI KOLs Timeline · 2026-07-29 Cached

Y Combinator hosted a Paper Club where researchers presented innovations in multi-GPU kernel optimization, including ParallelKittens, a CUDA framework that simplifies development of overlapped multi-GPU kernels and achieves significant speedups across workloads.

0 favorites 0 likes
#ai-inference

Viable ways to run K3 locally

Reddit r/LocalLLaMA · 2026-07-27

Discussion of viable low-cost hardware configurations to run the Kimi K3 AI model locally.

0 favorites 0 likes
#ai-inference

PSA: DO NOT use Intel consumer platforms for multi-GPU setups

Reddit r/LocalLLaMA · 2026-07-25

Testing reveals that Intel consumer platforms like Z890 with Arrow Lake CPUs have hardware/firmware limitations that prevent proper PCIe Peer-to-Peer (P2P) communication between multiple GPUs, making them unsuitable for multi-GPU AI workloads despite adequate lane counts.

0 favorites 0 likes
#ai-inference

AMD and Cerebras Launch AI Inference Solution (10 minute read)

TLDR AI · 2026-07-24 Cached

AMD and Cerebras announced a joint AI inference solution combining AMD Helios rackscale solutions with Cerebras Wafer-Scale Engine, aiming for ultra-low latency and high throughput. The disaggregated inference workflow is expected to deliver up to 5x higher tokens per second per watt.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback