ai-inference

Tag

Cards List
#ai-inference

@ivanalog_com: Debugged it for a bit, and now the speed of running Qwen Flash on Google A100 (3000+ prefill, 90+ tps) is already faste…

X AI KOLs Timeline ↗ · 2026-09-25 Cached

A GitHub repository collabosm provides an optimized setup for running the Qwen3.8-Flash-Next model on a Google Colab A100, achieving inference speeds faster than commercial APIs with detailed performance metrics and instructions.

0 favorites 0 likes
#ai-inference

@TheAhmadOsman: Mac Studio M5 Ultra vs 2x DGX Spark for DeepSeek V4 Flash - DGX Sparks: 2.41x faster on prefill (compute-bound) - Mac S…

X AI KOLs Timeline ↗ · 2026-09-24 Cached

The article compares the inference performance of Apple's Mac Studio M5 Ultra with two NVIDIA DGX Spark units when running the DeepSeek V4 Flash model, showing DGX Sparks are faster in prefill while Mac Studio is slightly faster in generation.

0 favorites 0 likes
#ai-inference

MacBook Pro M5 Max LSE LLM running an AMD Radeon AI PRO R9700 over Thunderbolt 5 in a Razer enclosure

Reddit r/LocalLLaMA ↗ · 2026-09-24

This article demonstrates a setup where an AMD Radeon AI Pro R9700 GPU is used on a MacBook Pro via Thunderbolt 5 with a custom driver and LemonSeed-Engine for running LLMs like Qwen3.6-3.8.

0 favorites 0 likes
#ai-inference

My foray into local ai. Two BC-250 ex mining apus running Qwen3.6-35B-A3B Q4_K_M at 60 tok/s with 64k context

Reddit r/LocalLLaMA ↗ · 2026-09-24

A user shares their local AI setup using two BC-250 ex-mining APUs to run the Qwen3.6-35B-A3B model with llama.cpp, achieving 60 tok/s and 64k context for under $300.

0 favorites 0 likes
#ai-inference

@charles_irl: After this, people kept asking us how @modal is able to serve agent inference so well. So we wrote it all down. New blo…

X AI KOLs Timeline ↗ · 2026-09-23 Cached

Modal details how they optimized inference performance for trillion-parameter coding agents, achieving significant improvements in throughput and interactivity for their service.

0 favorites 0 likes
#ai-inference

@MSFTResearch: Robots are getting smarter, but how can their hardware match that growth? New Microsoft Research findings show that mov…

X AI KOLs Following ↗ · 2026-09-23 Cached

Microsoft Research findings demonstrate that offloading AI inference from robots to edge or cloud systems enhances task success rates, efficiency, and battery life for physical AI applications.

0 favorites 0 likes
#ai-inference

I built a dual R9700 rig. Looking for advice.

Reddit r/LocalLLaMA ↗ · 2026-09-23

A user has built a dual R9700 rig and is seeking community advice on various AI inference, fine-tuning, and system optimization topics.

0 favorites 0 likes
#ai-inference

@PyTorch: How do you keep @vllm_project moving at the speed of light without excluding users who run diverse models on diverse ha…

X AI KOLs Following ↗ · 2026-09-22 Cached

Introduces hardware-agnostic layers in vLLM to maintain high performance while ensuring portability across diverse hardware, as announced in a PyTorch Foundation blog post.

0 favorites 0 likes
#ai-inference

@taroleo: When calling Jev from the US West Coast, a single request compiling 6 questions takes about 130 ms, or 20-25 ms per jud…

X AI KOLs Timeline ↗ · 2026-09-21

The tweet highlights the low latency and scalability of calling Jev from the US West Coast, with 130 ms per request for 6 questions and constant latency under high parallelism, indicating good design and future potential with specialized models.

0 favorites 0 likes
#ai-inference

@rohanpaul_ai: "American companies such as Modal, Fireworks, and Baseten will be able to serve Kimi K3, at one-tenth the cost of their…

X AI KOLs Timeline ↗ · 2026-09-21 Cached

American companies can serve Kimi K3 at one-tenth the cost due to access to advanced Nvidia and AMD chips, with irony as R&D shifts to China but hardware optimization could further reduce costs.

0 favorites 0 likes
#ai-inference

@TheAhmadOsman: The future of inference isn’t necessarily in any of the current hardware providers btw No disrespect to the incumbents,…

X AI KOLs Timeline ↗ · 2026-09-21 Cached

A tweet suggests that future AI inference hardware may not come from current providers like NVIDIA, highlighting acquisitions of startups such as Groq because GPUs are not optimally designed for inference.

0 favorites 0 likes
#ai-inference

The Inference Gap (56 minute read)

TLDR AI ↗ · 2026-09-21 Cached

The article investigates how changes in the inference regime for Anthropic's Fable 5 model led to performance degradation, emphasizing that inference quality is crucial for realizing frontier AI capabilities.

0 favorites 0 likes
#ai-inference

Success running Qwen 3.8 27B EXL3 on RTX 3060 + 5060 Ti

Reddit r/LocalLLaMA ↗ · 2026-09-20

A user successfully runs the Qwen 3.8 27B AI model on a mixed setup of RTX 3060 and 5060 Ti GPUs using tensor parallelism with exllamav3, achieving around 50 tokens per second with MTP enabled.

0 favorites 0 likes
#ai-inference

@rohanpaul_ai: Baseten CEO Tuhin Srivastava (@tuhinone) just dropped 2 numbers. shows how fast AI inference is exploding: token volume…

X AI KOLs Following ↗ · 2026-09-20 Cached

Baseten CEO Tuhin Srivastava reveals that token volume on Baseten has grown 40x year-over-year, with revenue increasing about 10x in the last 12 months, highlighting the explosion in AI inference.

0 favorites 0 likes
#ai-inference

Laya (OS Jev) on Mac M4 CoreML Offline (45 decisions per second)

Hacker News Top ↗ · 2026-09-20

The Laya model runs offline on Apple's M4 chip using CoreML, achieving 45 decisions per second in inference.

0 favorites 0 likes
#ai-inference

@LinQ444: jev 和laya的对比 https://github.com/mizorewww/laya-mlx…

X AI KOLs Timeline ↗ · 2026-09-20 Cached

Laya-MLX is an open-weight tool for running typed decision AI models locally on Apple Silicon with low latency, providing native inference without cloud APIs. It includes benchmarks showing fast performance on devices like M3 Max.

0 favorites 0 likes
#ai-inference

focus-llama: a llama.cpp fork implementing Declarative Attention (arXiv:2609.02737)

Reddit r/LocalLLaMA ↗ · 2026-09-20

This article describes a fork of llama.cpp called focus-llama that implements Declarative Attention from a recent paper, allowing models to declare needed context chunks during inference to optimize KV cache usage and reduce decode time.

0 favorites 0 likes
#ai-inference

@QuixiAI: By 2031, open/on-prem models handle a majority of economically routine inference in privacy-sensitive and cost-sensitiv…

X AI KOLs Following ↗ · 2026-09-20 Cached

By 2031, open and on-prem models are predicted to dominate routine inference in privacy-sensitive and cost-sensitive organizations, sidelining cloud-based AI companies like OpenAI and AnthropicAI.

0 favorites 0 likes
#ai-inference

@no_stp_on_snek: Fyi already been testing DFlash2 on Atlas since Aug 30 on mainline. :) https://github.com/Avarok-Cybersecurity/atlas-re…

X AI KOLs Timeline ↗ · 2026-09-19 Cached

The article announces testing of DFlash2 on Atlas and details atlasctl, a command-line tool for deploying and running LLM inference models on NVIDIA DGX Spark systems.

0 favorites 0 likes
#ai-inference

Qwen3.8-Flash-Next at 1M context on Strix Halo: 38 tok/s decode, 18 min prefill (halogen 0.12.0)

Reddit r/LocalLLaMA ↗ · 2026-09-19

Halogen version 0.12.0 fixes performance degradation at high context depths, showing improved decode and prefill speeds for Qwen3.8-Flash-Next at 1 million tokens of context on AMD Ryzen AI Max+ hardware.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback