@samhogan: introducing fast inference (https://fast.inference.net) fast inference is an LLM API for devs who want to go faster acc…

X AI KOLs Timeline Tools

Summary

Fast Inference is an LLM API service that offers fast and affordable access to top open-source and closed-source models for developers, with integrations for coding agents and tools like Claude Code and Codex.

introducing fast inference (https://t.co/zYhuq93ukF) fast inference is an LLM API for devs who want to go faster access the fastest Kimi K3, GLM, DeepSeek, etc for $9/mo integrations with Claude Code, Codex, OpenCode & more speeds compounds. waiting for Claude does not https://t.co/5R4xRiiG0u
Original Article
View Cached Full Text

Cached at: 08/27/26, 09:41 PM

introducing fast inference (https://t.co/zYhuq93ukF)

fast inference is an LLM API for devs who want to go faster

access the fastest Kimi K3, GLM, DeepSeek, etc for $9/mo

integrations with Claude Code, Codex, OpenCode & more

speeds compounds. waiting for Claude does not https://t.co/5R4xRiiG0u


Fast Inference | Fast hosted models, one API

Source: https://fast.inference.net/ Fast open-source models

Fast Inferencefor AI agents

Absurdly fast LLM inference for AI agents. Compatible with Claude, Codex, Cursor, Pi, Hermes, and more. Up to 10x faster than normal.

Go faster

More tokens, no limits.

Up to 10x more tokens than your Claude or Codex subscription. Never let a subscription limit break your focus and slow you down.

View plans Go smarter

head to headkimi ≈ opus

  • SWE72 74
  • GPQA81 83
  • tok/s184 38
  • $/1M in4.5 5

Smart, fast, & affordable.

Fast inference on the top open-source models. Just as smart as Claude and Codex at a fraction of the cost. And did we mention much faster?

View integrations One API

model catalogopen + closed

open

  • kimi-k3-fast
  • glm-5.2-fast
  • deepseek-v4-fast

closed

  • opus-5
  • gpt-5.6
  • fable-5

All the best models.

One click to switch between the best open-source models and closed-source models like Claude and Codex. Find the right model for you.

Explore models

Coding agents

Fast inference with any agent.

Route Claude Code, Codex, Grok, OpenCode, Pi, Hermes, OpenClaw, Fx, and Cursor through the Inference.net gateway. One shared key. Native surfaces. Safe reversal.

Models

Fast frontier models.

Kimi, GLM, DeepSeek, Fable, and GPT — routed through Fast Inference. Open the full catalog for every model we serve.

View all models

Pricing

Simple, predictable pricing

$9/ mo

$10 of Gateway credit, renewed each month.

Pay as you go after the credit. Same rates as going direct.

  • Works with Claude Code and Codex
  • Automatic provider failover
  • Personal usage meter
  • Cancel any time

Get started $49/ mo

$60 of Gateway credit, renewed each month.

Pay as you go after the credit. Same rates as going direct.

  • Everything in Operator
  • More monthly Gateway allowance
  • Hosted and provider model access
  • Cancel any time

Get started

Private by default

Your data is Your data.

Prompts and outputs are never used to train models. Data retention is kept to the minimum needed to operate the service.

01No training

Your data never enters a training set.

02Scoped access

Every key belongs to one project.

03Clear usage

See exactly where tokens and spend go.

04Easy revocation

Disable a credential in one click.

FAQ

Frequently askedquestions.

Quick answers to help you get started

faq.sh8questions

Fast Inference is an OpenAI-compatible API that routes to every model and provider. One key, with no per-request markup.

Similar Articles

Show HN: Free Inference Engineer and Model Training Roadmap

Hacker News Top

InferQuest is a free web application offering open roadmaps and verifiable milestones for learning inference engineering and LLM training, with paths focused on making models fast and cheap in production or efficient on minimal hardware.

Hetzner is working on LLM Inference

Hacker News Top

Hetzner has launched an experimental LLM inference API service, offering an OpenAI-compatible endpoint with the Qwen3.6-35B-A3B-FP8 model. The service is free during the experiment period, has no SLA, and is intended to gather user feedback.

InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents

arXiv cs.AI

InferenceBench is a benchmark that evaluates AI agents on optimizing LLM inference speed using an H100 GPU across multiple bottleneck scenarios. Results show agents improve over naive baselines but frequently converge on single frameworks and underperform simple hyperparameter searches, indicating a need for better exploration strategies.