inference-server

Tag

Cards List
#inference-server

GitHub - coder543/minnow: Fast LLaDA2.2 inference server

Reddit r/LocalLLaMA ↗ · 2026-09-09 Cached

Minnow is a high-performance Rust inference server for LLaDA2.2 models, offering optimized inference with quantization, GPU acceleration, and significant speed improvements over standard Transformers implementations.

0 favorites 0 likes
#inference-server

Perplexity open-sourced their Mac inference server for Qwen 3.6

Reddit r/LocalLLaMA ↗ · 2026-09-02

Perplexity has open-sourced a Mac inference server optimized for the Qwen 3.6 model to achieve best performance on Apple Silicon.

0 favorites 0 likes
#inference-server

@akshay_pachaar: run agent harnesses 100% private & offline. (no token costs, no API keys, 100% open-source) your agent runs locally. th…

X AI KOLs Following ↗ · 2026-09-02 Cached

Magnitude is an open-source inference server that runs AI models locally on your hardware, integrating with coding agents to ensure privacy and no data leakage, with setup in one command.

0 favorites 0 likes
#inference-server

@ao_qu18465: https://x.com/ao_qu18465/status/2094867930081337730

X AI KOLs Timeline ↗ · 2026-09-01 Cached

Reef is an open-source infrastructure that enables AI agents to continuously improve themselves by learning from inference experience, focusing on evolving both models and harnesses.

0 favorites 0 likes
#inference-server

@akshay_pachaar: https://x.com/akshay_pachaar/status/2094765529231929361

X AI KOLs Following ↗ · 2026-09-01 Cached

This article is a practitioner's guide to running local AI models for agent work, highlighting hardware trade-offs and introducing Magnitude, an open-source inference server that simplifies configuration for optimized performance.

0 favorites 0 likes
#inference-server

Show HN: Reame – a CPU inference server that gets faster as it runs

Hacker News Top ↗ · 2026-07-11 Cached

Reame is an LLM inference server built on llama.cpp that optimizes for CPU hardware by caching prompt prefixes and generated n-grams, becoming faster with repeated use. It is designed for cheap hardware like shared vCPUs and free tiers, targeting repetitive AI workloads such as document extraction and batch pipelines.

0 favorites 0 likes
#inference-server

@MiaAI_lab: If you're looking for a simple start/stop script to run your Qwen3.6 27B/35B, check this out. It's optimized for speed …

X AI KOLs Timeline ↗ · 2026-06-28 Cached

MiaAI-Lab provides a simple Bash start/stop script for running Qwen3.6 27B/35B GGUF models via llama-server, optimized for speed and coding performance.

0 favorites 0 likes
#inference-server

@JaydevTonde: https://x.com/JaydevTonde/status/2068361821002846418

X AI KOLs Timeline ↗ · 2026-06-20 Cached

A detailed tutorial on implementing CUDA Graphs in an LLM inference server Tokn, covering FastAPI server setup, engine initialization, and CUDA Graph capture for optimized decode phases.

0 favorites 0 likes
#inference-server

If you had $150K for building a production-class local inference server to serve 300 people, what would you buy?

Reddit r/LocalLLaMA ↗ · 2026-05-29

A user seeks advice on purchasing a failover inference server under $150K to serve 300 people, discussing options like used H100s, RTX Pro 6000, and DGX Station for running 122b AWQ models with vLLM.

0 favorites 0 likes
#inference-server

@gyro_ai: Running large models locally for your own tools involves a mountain of Python dependencies and endless backend configuration — the environment alone scares off many. In reality, most people just want a local interface that works instantly. Shimmy is a Rust-based local inference service, compiled into a single binary, offering an interface identical to OpenAI's…

X AI KOLs Timeline ↗ · 2026-05-24 Cached

Shimmy is a lightweight single-binary local inference server that provides a drop-in OpenAI-compatible API for running GGUF models, supporting hot-swapping models and requiring no Python dependencies.

0 favorites 0 likes
#inference-server

magnitudedev/magnitude

GitHub Trending (daily) ↗ · 2026-09-03 Cached

Magnitude is an open-source inference server that runs AI models locally on your hardware, ensuring privacy and offline use with integration into various agent tools like Pi and Claude Code.

0 favorites 0 likes
#inference-server

modular/modular

GitHub Trending (daily) ↗ · 2026-08-20 Cached

The Modular Platform offers open-source tools for AI development and deployment, featuring the MAX Framework and Mojo Language.

0 favorites 0 likes
← Back to home

Submit Feedback