ai-inference

Tag

Cards List
#ai-inference

@ManusAI: Introducing Manus Flex. Bring your own API key to Manus. Use a supported provider with the Manus agent harness, tools, …

X AI KOLs Timeline ↗ · yesterday Cached

Manus introduces Flex, a new module that allows users to bring their own API key to power Manus's agent infrastructure, with initial partners including OpenRouter, Fireworks, and Modal.

0 favorites 0 likes
#ai-inference

Inference Engines will become a series of one-offs

Reddit r/LocalLLaMA ↗ · yesterday

The article argues that specialized, one-off inference engines will outperform general ones like llama.cpp due to optimization for specific models and hardware, predicting their rise as the norm in AI inference.

0 favorites 0 likes
#ai-inference

RAM Offloading with vLLM - tcclaviger appreciation post

Reddit r/LocalLLaMA ↗ · yesterday

Appreciation post for tcclaviger adding expert RAM offloading support to vLLM, enabling running large AI models like DeepSeek-V4-Flash-Vision-Exp on local setups with multiple GPUs.

0 favorites 0 likes
#ai-inference

Source: Inference provider Modal Labs closing in on $750M round at $15.75B valuation

TechCrunch AI ↗ · 2d ago Cached

Modal Labs, an AI inference infrastructure provider, is closing in on a $750 million funding round at a $15.75 billion valuation, more than tripling its valuation from four months ago amid soaring demand for AI inference services.

0 favorites 0 likes
#ai-inference

95+ TPS through 100K generated for qwen3.8 27b, 262K ctx, on a single 3090

Reddit r/LocalLLaMA ↗ · 2d ago

LlamAmpere v0.4 is released with Ampere-specific improvements, achieving over 95 tokens per second and supporting 262K context for the Qwen3.8-27B model on a single NVIDIA 3090 GPU.

0 favorites 0 likes
#ai-inference

I’m calling this the Monstrosity. 5 ex mining BC-250 boards Qwen3-Coder-Next Q4 at 40 tok/s

Reddit r/LocalLLaMA ↗ · 2d ago

The user built a low-cost setup using five ex-mining BC-250 boards to run the Qwen3-Coder-Next AI model, achieving around 40 tokens per second at 30k context with plans to expand.

0 favorites 0 likes
#ai-inference

@TheAhmadOsman: Seriously how could people tolerate these slow ass cloud models anymore Painfully slow in comparison to anything I self…

X AI KOLs Timeline ↗ · 3d ago Cached

A tweet criticizing the slow performance of cloud AI models compared to self-hosted solutions.

0 favorites 0 likes
#ai-inference

Let's talk about trading compute (16 minute read)

TLDR AI ↗ · 3d ago Cached

The article explores the emerging market for compute derivatives and their potential to transform how neoclouds manage GPU rental risks and pricing in the AI inference cloud industry.

0 favorites 0 likes
#ai-inference

2x Tesla p100s, q6_k quant, Qwen 3.8 27B ~60tps V3.0

Reddit r/LocalLLaMA ↗ · 3d ago

A user shares kernel optimizations for Tesla P100 GPUs to improve performance when running the Qwen 3.8 model with llama.cpp, achieving significant speedups in inference.

0 favorites 0 likes
#ai-inference

@PyTorch: Curious about how to make enterprise agentic inference production-ready with @PyTorch and @vllm_project & to learn more…

X AI KOLs Following ↗ · 3d ago Cached

The article promotes a talk at the PyTorch Conference North America focused on making enterprise agentic inference production-ready using PyTorch and vLLM, covering ecosystem updates and registration details.

0 favorites 0 likes
#ai-inference

Splash fork optimised for M5 Max: ~1.5× faster (1.25× single request)

Reddit r/LocalLLaMA ↗ · 3d ago Cached

Splish is an unofficial fork of Splash that optimizes Metal kernels for Apple M5 Max chips, delivering up to 1.5× faster AI inference speeds for models like Qwen3.8-27B without compromising quality.

0 favorites 0 likes
#ai-inference

Improved and fixed template for GPT-OSS (again). Includes preserve_thinking and fix for Unsloth-induced bug

Reddit r/LocalLLaMA ↗ · 4d ago

The article details a bug fix for the GPT-OSS template from Unsloth, where chat history rendering incorrectly drops model answers during multi-turn inference, causing model degradation. The author shares an updated template that preserves thinking to improve performance.

0 favorites 0 likes
#ai-inference

Getting stupidly good results on my 4x3060ti setup.

Reddit r/LocalLLaMA ↗ · 4d ago

A user optimized a 4x3060ti GPU rig for AI inference using tensor parallelism with Exl3 and vllm, achieving up to 120 tokens per second with large context windows.

0 favorites 0 likes
#ai-inference

2400cc Inference Racer: Dual RTX 3090 motors, NVLink turbo, naked 7840U ThinkPad ECU, VW Golf radiator

Reddit r/LocalLLaMA ↗ · 4d ago

A DIY inference machine built with dual RTX 3090 GPUs, a ThinkPad motherboard, and a VW Golf radiator, capable of running AI models like Qwen3.8-27B via vLLM.

0 favorites 0 likes
#ai-inference

@Oluwaphilemon1: Qwen3.8-Flash-Next running at 15 tok/s on a 12GB RTX 5070 apparently wasn’t acceptable. So instead of buying a bigger G…

X AI KOLs Timeline ↗ · 4d ago Cached

A user built an open-source inference engine called Strata to optimize running Qwen3.8-Flash-Next on modest hardware, achieving up to 65.1 tok/s from an initial 15 tok/s.

0 favorites 0 likes
#ai-inference

@PyTorch: Elastic Expert Parallelism in @vllm_project lets you add or remove GPUs from an active Mixture-of-Experts deployment du…

X AI KOLs Following ↗ · 4d ago Cached

The article promotes a presentation on Elastic Expert Parallelism in vLLM at the PyTorch Conference North America 2026, discussing how to dynamically add or remove GPUs in Mixture-of-Experts deployments with minimal downtime.

0 favorites 0 likes
#ai-inference

Got 85.6tok/s MTP with Qwen3.8:27b. Single RTX 5090

Reddit r/LocalLLaMA ↗ · 4d ago

Achieved 85.6 tokens per second using the Qwen3.8:27b model on a single RTX 5090 GPU.

0 favorites 0 likes
#ai-inference

QWEN3.6-27B-MLX-8bit (29.5GIG) Is Excellent

Reddit r/openclaw ↗ · 5d ago

A user shares their experience switching to the QWEN3.6-27B-MLX-8bit model for local AI tasks, finding it performs comparably to larger models while saving significant RAM on a Mac Studio, improving workflow stability.

0 favorites 0 likes
#ai-inference

@ivanalog_com: Debugged it for a bit, and now the speed of running Qwen Flash on Google A100 (3000+ prefill, 90+ tps) is already faste…

X AI KOLs Timeline ↗ · 6d ago Cached

A GitHub repository collabosm provides an optimized setup for running the Qwen3.8-Flash-Next model on a Google Colab A100, achieving inference speeds faster than commercial APIs with detailed performance metrics and instructions.

0 favorites 0 likes
#ai-inference

@TheAhmadOsman: Mac Studio M5 Ultra vs 2x DGX Spark for DeepSeek V4 Flash - DGX Sparks: 2.41x faster on prefill (compute-bound) - Mac S…

X AI KOLs Timeline ↗ · 6d ago Cached

The article compares the inference performance of Apple's Mac Studio M5 Ultra with two NVIDIA DGX Spark units when running the DeepSeek V4 Flash model, showing DGX Sparks are faster in prefill while Mac Studio is slightly faster in generation.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback