inference-engine

Tag

Cards List
#inference-engine

@charles_irl: What do you get when you cross a query planner with an inference engine? Over a billion tokens a minute! Introducing Qu…

X AI KOLs Timeline ↗ · yesterday Cached

Quail is an open source AI-SQL engine that integrates query planning with LLM inference to achieve over 1 billion input tokens per minute on an H100 GPU.

0 favorites 0 likes
#inference-engine

DeepSeek-V4-Flash-0731 at ~40–50 tok/s on 2× Radeon AI PRO R9700 with the affinity engine (prebuilt quant + fixes)

Reddit r/LocalLLaMA ↗ · 3d ago

The article reports on running DeepSeek-V4-Flash-0731 on two Radeon AI PRO R9700 GPUs with the affinity inference engine, achieving 40-50 tok/s decode speed via a prebuilt quantization and stability fixes.

0 favorites 0 likes
#inference-engine

Gewell - Gemma4 inference engine

Reddit r/LocalLLaMA ↗ · 5d ago

Gewell announces an inference engine designed for Gemma4 AI models to streamline deployment and execution.

0 favorites 0 likes
#inference-engine

@YichiZ03: zero-to-sglang just hit 1,000 GitHub stars! Thanks for all your support! Our new chapters: Chapter 2.1: mini-sglang — W…

X AI KOLs Timeline ↗ · 6d ago Cached

The zero-to-sglang course has hit 1,000 GitHub stars, with new chapters released covering mini-sglang and the journey of a request including code walkthroughs.

0 favorites 0 likes
#inference-engine

Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro

Reddit r/LocalLLaMA ↗ · 2026-09-19

Inco Splash is an open-source inference engine optimized for Apple silicon, offering significant speed improvements for running AI models like Qwen3.8-27B on M-series MacBooks.

0 favorites 0 likes
#inference-engine

@jianchen1799: Local models can now handle agentic workloads. Local inference engines need to catch up. Today we’re releasing Splash: …

X AI KOLs Timeline ↗ · 2026-09-19 Cached

Inco AI releases Splash, an open-source inference engine optimized for Apple silicon, claiming up to 3× faster decode speeds for local model serving, enabling agentic workloads on devices like M5 Max MacBook Pro.

0 favorites 0 likes
#inference-engine

@vllm_project: Thank you to everyone who showed up, spoke, and stayed throughout the vLLM Conference. Lines out the door for talks say…

X AI KOLs Following ↗ · 2026-09-19 Cached

The vLLM Conference concluded successfully with high community attendance, and session recordings are available in the thread.

0 favorites 0 likes
#inference-engine

Intel releases OpenVINO 2026.4

Reddit r/LocalLLaMA ↗ · 2026-09-17

Intel releases OpenVINO 2026.4 with expanded AI model support, performance enhancements like multi-token prediction, and new features for profiling and inference across CPUs, GPUs, and NPUs.

0 favorites 0 likes
#inference-engine

The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction

arXiv cs.AI ↗ · 2026-09-17 Cached

Edge0 is a streaming MoE inference engine that uses trained routing prediction to serve 35B Mixture-of-Experts models from SSD, achieving near-fp16 performance on consumer hardware with low memory usage.

0 favorites 0 likes
#inference-engine

harshatheg/Qwen-2.5-1B-RLCD

Hugging Face Models Trending ↗ · 2026-09-16 Cached

A high-throughput inference engine for structured information extraction on Apple Silicon using MLX, offering parallel constrained decoding with 5.6x to 7.0x latency reductions and 100% schema validity.

0 favorites 0 likes
#inference-engine

@Huahuazo: The inference engine race has been competitive up to today, with many still seeing 'getting it to run' as the finish line. SGLang takes a different path. LMSYS's open-source high-performance service framework, with its standout feature being RadixAttention—where KV Cache for common prefixes can be reused, benefiting multi-turn dialogues, Agents, structured output, and more...

X AI KOLs Timeline ↗ · 2026-09-12 Cached

SGLang is an open-source high-performance inference service framework from LMSYS, leveraging RadixAttention technology to achieve KV Cache reuse, supporting various hardware and models, and widely used in production environments.

0 favorites 0 likes
#inference-engine

Faster than Light in Air: 8-22 tg/s Qwen3.8-Flash-Next (Q4/Q4ish) on a 32GB M4 MacBook Air

Reddit r/LocalLLaMA ↗ · 2026-09-10

Cherenkov is a new inference engine for Apple Silicon that enables efficient memory-constrained inference of large AI models like Qwen3.8-Flash-Next using predictive expert streaming and mixed-precision execution.

0 favorites 0 likes
#inference-engine

Inside the megakernel serving engine for North Mini Code (22 minute read)

TLDR AI ↗ · 2026-09-09 Cached

Cohere introduces a megakernel serving engine for North Mini Code that achieves 1.25× to 1.41× faster inference than vLLM on H100 GPUs by optimizing memory bandwidth for autoregressive decoding.

0 favorites 0 likes
#inference-engine

@no_stp_on_snek: Development time on inference engines I’m working on

X AI KOLs Timeline ↗ · 2026-09-08 Cached

The author shares insights about the development time for inference engines they are working on.

0 favorites 0 likes
#inference-engine

Our doc-QA agent runs four small models as four separate services - has anyone actually consolidated this?

Reddit r/AI_Agents ↗ · 2026-09-08

The author describes consolidating four small AI models from separate services into a single server using Superlinked's inference engine to reduce operational overhead, while discussing trade-offs like GPU sharing and blast radius concerns.

0 favorites 0 likes
#inference-engine

We open-sourced Paddock, our Rust/C++ inference engine with its own CUDA kernels (MIT/Apache-2.0)

Reddit r/LocalLLaMA ↗ · 2026-09-04

Paddock is an open-source Rust/C++ inference engine with custom CUDA kernels, demonstrating competitive performance against vLLM and SGLang in benchmarks while supporting OpenAI/Anthropic style APIs and GGUF/safetensors formats.

0 favorites 0 likes
#inference-engine

@LinusEkenstam: Quite a big deal. 1.35x faster inference than MLX-LM 1.23x faster at prefill having the hybrid option to pick from insi…

X AI KOLs Timeline ↗ · 2026-09-02 Cached

Perplexity has open-sourced Lily, a local inference engine optimized for Qwen3.6-35B-A3B on Apple Silicon, achieving 1.35x faster inference than MLX-LM.

0 favorites 0 likes
#inference-engine

VoxGen, an AMD-optimized TTS inference engine for VoxCPM 2 models

Reddit r/LocalLLaMA ↗ · 2026-09-02

VoxGen is a new lightweight native inference engine for VoxCPM2 models, optimized for AMD cards using Rust and Vulkan compute to enhance performance and remove Python/PyTorch dependencies.

0 favorites 0 likes
#inference-engine

Question: Why is prefill unbelievably faster in vLLM than other inference engines?

Reddit r/LocalLLaMA ↗ · 2026-09-01

The user shares benchmark results showing vLLM's significantly faster prefill performance compared to llama.cpp and other engines, and questions the technical reasons behind this speed difference.

0 favorites 0 likes
#inference-engine

(NInfer Fork) I wanted to have a 1M context Qwen-3.8 27B, tp2, dual 5090s

Reddit r/LocalLLaMA ↗ · 2026-08-29

A developer forked NInfer, a C++20/CUDA inference engine, to add tensor-parallelism and YaRN rope scaling, enabling Qwen3.8-27B to run with a 1M token context on dual 5090 GPUs and outperforming vLLM in specific decode scenarios.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback