Tag
Minnow is a high-performance Rust inference server for LLaDA2.2 models, offering optimized inference with quantization, GPU acceleration, and significant speed improvements over standard Transformers implementations.
Perplexity has open-sourced a Mac inference server optimized for the Qwen 3.6 model to achieve best performance on Apple Silicon.
Magnitude is an open-source inference server that runs AI models locally on your hardware, integrating with coding agents to ensure privacy and no data leakage, with setup in one command.
Reef is an open-source infrastructure that enables AI agents to continuously improve themselves by learning from inference experience, focusing on evolving both models and harnesses.
This article is a practitioner's guide to running local AI models for agent work, highlighting hardware trade-offs and introducing Magnitude, an open-source inference server that simplifies configuration for optimized performance.
Reame is an LLM inference server built on llama.cpp that optimizes for CPU hardware by caching prompt prefixes and generated n-grams, becoming faster with repeated use. It is designed for cheap hardware like shared vCPUs and free tiers, targeting repetitive AI workloads such as document extraction and batch pipelines.
MiaAI-Lab provides a simple Bash start/stop script for running Qwen3.6 27B/35B GGUF models via llama-server, optimized for speed and coding performance.
A detailed tutorial on implementing CUDA Graphs in an LLM inference server Tokn, covering FastAPI server setup, engine initialization, and CUDA Graph capture for optimized decode phases.
A user seeks advice on purchasing a failover inference server under $150K to serve 300 people, discussing options like used H100s, RTX Pro 6000, and DGX Station for running 122b AWQ models with vLLM.
Shimmy is a lightweight single-binary local inference server that provides a drop-in OpenAI-compatible API for running GGUF models, supporting hot-swapping models and requiring no Python dependencies.
Magnitude is an open-source inference server that runs AI models locally on your hardware, ensuring privacy and offline use with integration into various agent tools like Pi and Claude Code.
The Modular Platform offers open-source tools for AI development and deployment, featuring the MAX Framework and Mojo Language.