@samhogan: introducing fast inference (https://fast.inference.net) fast inference is an LLM API for devs who want to go faster acc…
Summary
Fast Inference is an LLM API service that offers fast and affordable access to top open-source and closed-source models for developers, with integrations for coding agents and tools like Claude Code and Codex.
View Cached Full Text
Cached at: 08/27/26, 09:41 PM
introducing fast inference (https://t.co/zYhuq93ukF)
fast inference is an LLM API for devs who want to go faster
access the fastest Kimi K3, GLM, DeepSeek, etc for $9/mo
integrations with Claude Code, Codex, OpenCode & more
speeds compounds. waiting for Claude does not https://t.co/5R4xRiiG0u
Fast Inference | Fast hosted models, one API
Source: https://fast.inference.net/ Fast open-source models
Fast Inferencefor AI agents
Absurdly fast LLM inference for AI agents. Compatible with Claude, Codex, Cursor, Pi, Hermes, and more. Up to 10x faster than normal.
Go faster
More tokens, no limits.
Up to 10x more tokens than your Claude or Codex subscription. Never let a subscription limit break your focus and slow you down.
View plans Go smarter
head to headkimi ≈ opus
- SWE72 74
- GPQA81 83
- tok/s184 38
- $/1M in4.5 5
Smart, fast, & affordable.
Fast inference on the top open-source models. Just as smart as Claude and Codex at a fraction of the cost. And did we mention much faster?
View integrations One API
model catalogopen + closed
open
- kimi-k3-fast
- glm-5.2-fast
- deepseek-v4-fast
closed
- opus-5
- gpt-5.6
- fable-5
All the best models.
One click to switch between the best open-source models and closed-source models like Claude and Codex. Find the right model for you.
Coding agents
Fast inference with any agent.
Route Claude Code, Codex, Grok, OpenCode, Pi, Hermes, OpenClaw, Fx, and Cursor through the Inference.net gateway. One shared key. Native surfaces. Safe reversal.
Models
Fast frontier models.
Kimi, GLM, DeepSeek, Fable, and GPT — routed through Fast Inference. Open the full catalog for every model we serve.
Pricing
Simple, predictable pricing
$9/ mo
$10 of Gateway credit, renewed each month.
Pay as you go after the credit. Same rates as going direct.
- Works with Claude Code and Codex
- Automatic provider failover
- Personal usage meter
- Cancel any time
Get started $49/ mo
$60 of Gateway credit, renewed each month.
Pay as you go after the credit. Same rates as going direct.
- Everything in Operator
- More monthly Gateway allowance
- Hosted and provider model access
- Cancel any time
Private by default
Your data is Your data.
Prompts and outputs are never used to train models. Data retention is kept to the minimum needed to operate the service.
01No training
Your data never enters a training set.
02Scoped access
Every key belongs to one project.
03Clear usage
See exactly where tokens and spend go.
04Easy revocation
Disable a credential in one click.
FAQ
Frequently askedquestions.
Quick answers to help you get started
faq.sh8questions
Fast Inference is an OpenAI-compatible API that routes to every model and provider. One key, with no per-request markup.
Similar Articles
@asterailabs: Introducing Aster Inference -- The world's fastest inference API created by AI research agents We serve the world's fas…
Aster Labs launches Aster Inference, claiming the world's fastest inference API using AI research agents, with benchmark speeds for models like OpenAI's gpt-oss-120b and GLM 5.2.
Show HN: Free Inference Engineer and Model Training Roadmap
InferQuest is a free web application offering open roadmaps and verifiable milestones for learning inference engineering and LLM training, with paths focused on making models fast and cheap in production or efficient on minimal hardware.
Hetzner is working on LLM Inference
Hetzner has launched an experimental LLM inference API service, offering an OpenAI-compatible endpoint with the Qwen3.6-35B-A3B-FP8 model. The service is free during the experiment period, has no SLA, and is intended to gather user feedback.
@0xSero: Here's everything you need to know about inference and hosting LLMs. Have you ever seen: - vllm - sglang - llama.cpp - …
An overview of popular open-source inference engines including vLLM, SGLang, llama.cpp, and ExLlamaV3 for hosting and running large language models.
InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents
InferenceBench is a benchmark that evaluates AI agents on optimizing LLM inference speed using an H100 GPU across multiple bottleneck scenarios. Results show agents improve over naive baselines but frequently converge on single frameworks and underperform simple hyperparameter searches, indicating a need for better exploration strategies.