Tag
Hugging Face announces Baseten as a supported Inference Provider, enabling serverless access to popular open-weight LLMs like Kimi K3, DeepSeek V4 Flash, and GLM-5.2 directly from model pages and SDKs.
Baseten details how it built the fastest API for GLM-5.2, achieving over double the launch-day performance and introducing a latency-optimized Fast version for coding and agents, with further improvements planned.
Baseten announces the world's fastest API for the GLM-5.2 open model, achieving over 280 tokens per second via NVFP4 quantization, disaggregated inference, and other optimizations.
Baseten processes over 1 billion inference calls per day and has raised $1.5B to scale its infrastructure, highlighting inference as a massive market.
Apoorv Agrawal from Altimeter Capital explains why they are doubling down on their investment in Baseten, arguing that inference will become the largest market and that post-trained open source models offer the best combination of capability, cost, and control.
Baseten announces a $1.5B Series F funding round led by multiple investors, citing 20x revenue growth and 40x inference volume growth as evidence of the market's shift toward inference as the key AI layer.
Baseten, a $13 billion AI startup, provides software and computing capacity to companies using lower-cost AI models as alternatives to OpenAI and Anthropic.
A tweet from Philip Kiely highlights cost savings by switching from closed-source AI models to open-source alternatives, using Baseten's ROI calculator tool.
Cursor AI launches a new interview series with developers, starting with a conversation with the Baseten team about their use of coding agents, current workflows, and future predictions.