inference-engine

Tag

Cards List
#inference-engine

Spent two weeks on a kernel that benchmarked 29x faster. End to end it's maybe 6-10%, and it's not even wired in yet.

Reddit r/LocalLLaMA ↗ · 2026-07-24

The author optimized a matmul kernel for BitNet's ternary models on CPU, achieving 29x speedup in isolation, but found that the model is memory-bound, resulting in only 6-10% end-to-end gain. The inference engine is available as open-source.

0 favorites 0 likes
#inference-engine

Built a from-scratch BitNet inference engine in pure C — 1.8× faster than bitnet.cpp on Xeon (36 tok/s), zero dependencies [BitNet & Bonsai CPU testers wanted]

Reddit r/LocalLLaMA ↗ · 2026-07-22

Project Zero is a from-scratch C99 LLM inference engine that runs BitNet and Qwen Bonsai-27B on CPU with zero dependencies, achieving 1.8× speedup over bitnet.cpp on Xeon. The project seeks community benchmarks for both models.

0 favorites 0 likes
#inference-engine

@QuixiAI: QuixiAI/embeddinggemma.c - fast as fuck cross-platform embedding inference engine. Written in c, runs anywhere. Inferen…

X AI KOLs Following ↗ · 2026-07-20 Cached

QuixiAI releases embeddinggemma.c, a fast cross-platform embedding inference engine written in C, supporting multiple backends (CPU, Metal, CUDA, ROCm, SYCL) and Matryoshka embeddings with a standard HTTP API.

0 favorites 0 likes
#inference-engine

@no_stp_on_snek: can the inference engine itself change model behavior? ran two quant and speculative-decode stacks of the same base mod…

X AI KOLs Following ↗ · 2026-07-14 Cached

A developer compares two inference stacks (production build vs SignalNine's q27) on the same Qwen model and finds they produce different honesty under pressure, with one fabricating progress and the other refusing appropriately, suggesting inference engines can affect model behavior beyond speed and quality metrics.

0 favorites 0 likes
#inference-engine

moondream3.1-9B-A2B

Reddit r/LocalLLaMA ↗ · 2026-07-12 Cached

Moondream 3.1 is a vision language model with mixture-of-experts architecture (9B total parameters, 2B active), delivering state-of-the-art visual reasoning, detection, pointing, and captioning, deployable locally via the Photon inference engine or through the Moondream Cloud API.

0 favorites 0 likes
#inference-engine

@jino_rohit: over the last 6-8 months, ive been trying to move towards the ml systems and ai infra space. these are some of my favor…

X AI KOLs Timeline ↗ · 2026-07-11 Cached

The author shares their work over 6-8 months in ML systems and AI infrastructure, including a lightweight Python LLM inference engine (tachyon) that achieves 600+ tokens/s on consumer hardware with continuous batching and prefix caching, alongside blog posts on CUDA/CUTE DSL and collective communication, and contributions to SGLang and vLLM.

0 favorites 0 likes
#inference-engine

@JiaZhihao: Excited to share Lithos’ serving stack for Kimi K2.7 Code, a 1T-parameter frontier coding model. On a single 8×B200 nod…

X AI KOLs Timeline ↗ · 2026-07-10 Cached

Lithos announces its inference engine serving Kimi K2.7 Code, achieving over 1,000 tokens/sec per user on a single 8×B200 node at native precision, 3.4–5.7× faster than major providers.

0 favorites 0 likes
#inference-engine

@yibie: Recommend this project—a single person wrote an inference engine in pure C, making the 744-billion-parameter GLM-5.2 run on a consumer machine with 25GB RAM. No GPU, no BLAS, no Python runtime—about 1300 lines of C. The core insight is simple: MoE…

X AI KOLs Timeline ↗ · 2026-07-10 Cached

Colibri is an inference engine written in pure C, approximately 1300 lines of code, zero dependencies. It can run the 744-billion-parameter GLM-5.2 MoE model on a consumer machine with 25GB RAM, achieved by streaming loaded routing experts and efficient caching, no GPU or Python runtime needed.

0 favorites 0 likes
#inference-engine

TensorSharp : Open Source Local LLM Inference Engine

Reddit r/ArtificialInteligence ↗ · 2026-07-04 Cached

TensorSharp is an open-source .NET library for running LLM inference locally, supporting GGUF models and offering a CLI, web chatbot, and OpenAI-compatible APIs with multiple backend options (CUDA, Metal, CPU).

0 favorites 0 likes
#inference-engine

@mgoin_: Today I deleted PagedAttention from vLLM

X AI KOLs Timeline ↗ · 2026-07-02 Cached

Michael Goin announced removing PagedAttention from vLLM, a significant change to the open-source LLM inference engine.

0 favorites 0 likes
#inference-engine

Popping the GPU Bubble

Hacker News Top ↗ · 2026-06-30 Cached

Moondream's Photon inference engine eliminates GPU bubbles through pipelined decoding, achieving near-realtime VLM inference with up to 35% higher decode throughput.

0 favorites 0 likes
#inference-engine

@SuJinYan123: Just 6 hours after DeepSeek open-sourced the Qwen DSpark weights, OpenInfer already has DSpark support running on RTX 5…

X AI KOLs Timeline ↗ · 2026-06-28 Cached

OpenInfer, a pure Rust+CUDA LLM inference engine, quickly added support for DeepSeek's DSpark speculative decoding technique on RTX 5090, achieving nearly 500 tok/s per user and scaling to ~2.4K aggregate tok/s, outperforming DFlash on non-random workloads.

0 favorites 0 likes
#inference-engine

A barebones CPU-only inference engine for Qwen 3, written from scratch in pure C

Reddit r/LocalLLaMA ↗ · 2026-06-28

A minimal CPU-only inference engine for Qwen 3 models implemented from scratch in pure C.

0 favorites 0 likes
#inference-engine

@yoheinakajima: who wants to help eyal poke holes in this approach to run LLM inference... in browser?

X AI KOLs Following ↗ · 2026-06-25 Cached

Eyal Toledano built an LLM inference engine using pure WebGPU/WGSL, running on-device in browser and Node without API keys, and is seeking peer review.

0 favorites 0 likes
#inference-engine

@songhan_mit: We develop an agent-native approach to accelerate genAI, continuing the success of KDA (Kernel Design Agent) at a highe…

X AI KOLs Following ↗ · 2026-06-25 Cached

Enze Xie announces Sol Video Inference Engine, an agent-native, training-free full-stack accelerator for video diffusion that auto-tunes cache, sparse attention, token pruning, quantization, and kernel fusion, achieving >2× end-to-end speedup on large models like 64B Cosmos3-Super and 22B LTX-2.3.

0 favorites 0 likes
#inference-engine

@Mayhem4Markets: https://x.com/Mayhem4Markets/status/2069090022117019928

X AI KOLs Following ↗ · 2026-06-22 Cached

A detailed technical comparison of two dominant LLM serving frameworks, SGLang and vLLM, covering architectural differences in KV cache management (RadixAttention vs PagedAttention), throughput, latency, and deployment considerations for self-hosted environments.

0 favorites 0 likes
#inference-engine

@QuixiAI: https://x.com/QuixiAI/status/2068776183102067086

X AI KOLs Following ↗ · 2026-06-21 Cached

DwarfStar is a self-contained native inference engine optimized for DeepSeek V4 Flash and PRO models, supporting Metal, CUDA, and ROCm backends, with a focus on high-end personal machines and Mac Studios.

0 favorites 0 likes
#inference-engine

@h100envy: Ying Sheng co-wrote SGLang, the inference engine now serving Grok at xAI on a hundred thousand GPUs. She also built Fle…

X AI KOLs Timeline ↗ · 2026-06-19 Cached

Ying Sheng co-wrote SGLang, the inference engine now serving Grok at xAI on a hundred thousand GPUs, achieving 5x cost cuts over DeepSeek's API; she also built FlexGen and helped build Chatbot Arena.

0 favorites 0 likes
#inference-engine

Fearless Concurrency on the GPU: Safe GPU inference in Rust, competitive with vLLM/SGLang [R]

Reddit r/MachineLearning ↗ · 2026-06-18

cuTile Rust introduces a tile-based programming model that leverages Rust's ownership to guarantee memory safety and data-race freedom for GPU kernels, and the Grout inference engine built on it achieves competitive throughput with vLLM/SGLang for Qwen3 models.

0 favorites 0 likes
#inference-engine

@ProfBuehlerMIT: For science, AI sovereignty and physics-grounded reasoning are non-negotiable. But how can we teach a small LLM like Ge…

X AI KOLs Timeline ↗ · 2026-06-18 Cached

mistral.rs now natively supports Agent Skills, enabling locally-run small LLMs to perform complex agentic workflows for scientific tasks, with full control over models, data, and execution.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback