Tag
Qwen3.8-2.4T-A95B by Alibaba Qwen and Alibaba Cloud is now available on Modal, served with a custom DFlash speculator trained on tool-call-heavy data and a full 1M context window.
DeepSeek-V4-Flash is a 284B-parameter MoE model with 13B active parameters per token, featuring a hybrid compressed attention mechanism that reduces KV cache needs for 1M-token contexts. It can be served with SGLang on Modal for fast decoding on a single B300.
Kimi K3, an open model that leads agentic coding benchmarks with native vision and a 1M-token context window, is now available to run in Codex on Modal via its Shared Endpoint.
Thinking Machines released Inkling-Small, a 276B-parameter mixture-of-experts model with 12B active parameters, 1M context, and native image/audio understanding, now available on Modal with NVIDIA B300 support.
Modal's CTO Akshat Bubna clarifies that a security incident involving a rogue agent was caused by a customer's unauthenticated endpoint, not a compromise of Modal's platform isolation.
The author describes the liberating feeling of using an open AI model (Kimi K3) on their own inference endpoint, contrasting it with the experience of using proprietary services.
Modal is a serverless cloud platform designed for AI workloads, supporting inference, training, and sandboxes. It enables instant scaling from zero to thousands of GPUs with pure Python code, significantly reducing latency and accelerating time to market.
Modal announced its first annual conference, Runtime, to be held on October 1st at The Midway in San Francisco.
Serverless GPUs may have higher hourly rates but can be more cost-effective overall depending on workload peak-to-average demand. The article on Modal's blog illustrates this with a widget.
Charles wrote an article explaining what Modal is while on a flight to ICML in Seoul.
In a Twitter thread, the founder of Modal explains that the platform is best understood as a computer, drawing parallels between traditional computer architecture and Modal's serverless cloud infrastructure.
Matt Pocock praises prototyping tool Wayfinder for its AI-powered modal that enables editing course text from anywhere.
Modal announces its integration with Claude Science, providing elastic compute infrastructure for life sciences researchers, with up to $100K in committed resources.
Modal announces integration with Claude Science, providing elastic compute for life sciences researchers with up to $100K in committed resources.
Yong Quan highlights that better speculative decoders can unlock near-linear throughput gains in LLM inference, as presented at a Modal workshop by Charles.
Modal announces Auto Endpoints for effortless inference, praised by developer Anthony Corletti as a top-level abstraction over compute, storage, and networking.
Modal introduces Auto Endpoints, a self-serve service for optimized, production-grade LLM inference with full code ownership, transparent metrics, and autoscaling, built on their serverless GPU infrastructure.
Modal announces managed private LLM endpoints available to everyone, with easy deployment via UI or CLI and full code access for customers.
Modal announces Auto Endpoints, a service enabling optimized open-source AI inference with a single click, aiming to counter the trend of proprietary models and services.
Modal announces Auto Endpoints, a new feature for owning and deploying AI inference.