@asterailabs: Introducing Aster Inference -- The world's fastest inference API created by AI research agents We serve the world's fas…
Summary
Aster Labs launches Aster Inference, claiming the world's fastest inference API using AI research agents, with benchmark speeds for models like OpenAI's gpt-oss-120b and GLM 5.2.
View Cached Full Text
Cached at: 07/16/26, 02:19 PM
Introducing Aster Inference – The world’s fastest inference API created by AI research agents
We serve the world’s fastest inference on GPU:
- OpenAI’s gpt-oss-120b @ 644 tps
- http://Z.ai’s GLM 5.2 @ 281 tps
At Aster, we’re automating open-ended research, and we use inference optimization as a task to benchmark our agents against. We’re creating a product out of the inference discoveries made from our system.
As our agents discover more, we plan to further improve our inference product and ship new, SOTA AI products.
Similar Articles
@samhogan: introducing fast inference (https://fast.inference.net) fast inference is an LLM API for devs who want to go faster acc…
Fast Inference is an LLM API service that offers fast and affordable access to top open-source and closed-source models for developers, with integrations for coding agents and tools like Claude Code and Codex.
Jalapeño’s first results show industry-leading speed and efficiency in AI inference
OpenAI's Jalapeño custom inference chip shows industry-leading speed and efficiency, delivering higher throughput and lower latency across various AI models.
@k1rallik: NVIDIA IS LITERALLY GIVING AWAY FREE AI INFERENCE I literally set it up in 5 minutes and couldn't believe it was free D…
NVIDIA offers free AI inference via DGX Cloud with OpenAI-compatible API for popular models like DeepSeek, MiniMax, Kimi, GLM, and Llama, claimable in 5 minutes.
Real-time LLM Inference on Standard GPUs: 3k tokens/s per request
Kog AI launches a tech preview of the Kog Inference Engine, achieving 3,000 tokens/s per request on standard datacenter GPUs by co-designing model architecture, runtime, and low-level GPU code, targeting latency-critical AI agent workflows.
@akshay_pachaar: Massive breakthrough here! Self-hosting LLMs just got ~75% cheaper: Most agent pipelines now run 4-5 small models under…
Superlinked releases SIE, an open-source inference engine that serves 85+ models behind one API with on-demand loading and LRU eviction, cutting self-hosting GPU costs by ~75% for agent pipelines.