Tag
NVIDIA is offering free access to four AI models, including DeepSeek V4.1 Flash, GLM 5.3, GLM 5.3 Flash, and Kimi K3, through their platform without requiring a credit card.
Nunchux AI has launched Modelverse, a multimodal generative AI inference service providing fast, affordable access to over 30 image, video, and avatar models through a single API.
A blog post argues that AI inference providers are undermining session portability by returning opaque, provider-bound state (encrypted reasoning tokens, hidden subagent messages, unexportable context) rather than user-owned transcripts, and proposes practical tests for session ownership.
Kimi K3, a 2.8-trillion-parameter open-source model from Moonshot AI with 1M token context and native vision, is now available on the Telnyx Inference API for tasks like coding, reasoning, and multimodality.
BentoML is an open-source Python framework that simplifies packaging, serving, and deploying AI models as REST APIs with containerization and multi-model orchestration.
Aster Labs launches Aster Inference, claiming the world's fastest inference API using AI research agents, with benchmark speeds for models like OpenAI's gpt-oss-120b and GLM 5.2.
The article argues that open-weight AI models are catching up to closed ones not via distillation but due to the modularisation of the AI stack—stable interfaces (Transformer architecture, OpenAI-compatible APIs, agentic harnesses) allow innovations to diffuse rapidly across the ecosystem, shrinking the capability gap while keeping a massive price advantage, potentially leading to a commoditisation of frontier AI.
Baseten announces the world's fastest API for the GLM-5.2 open model, achieving over 280 tokens per second via NVFP4 quantization, disaggregated inference, and other optimizations.
Kimi K2.6 is released as an open-weight model with strong agentic capabilities, accessible via FireworksAI’s fast inference APIs.