Tag
GLM 5.2 demonstrates fast performance on Modal's cloud platform.
Modal discusses the importance of managing the entire lifecycle of sandbox systems beyond initial boot, highlighting tools like .wait_until_ready().
Modal explains the complexities of building performant sandbox systems beyond initial container boot and shares tools for lifecycle management.
A new page in the LLM Engineer's Almanac provides a block quant visualizer to help engineers understand quantization formats for owning their LLM inference.
On Friday, we released six new state-of-the-art drafters for accelerated inference, along with a blog post on speculative decoding and a roofline model tool to estimate speedups.
Discusses how sandbox startup latency and scaling in RL training infrastructure can significantly impact training performance, referencing a detailed analysis by SemiAnalysis on matching trainer and generator throughput.
DFlash, a block-diffusion drafter with KV injection, is now running at frontier scale, achieving up to 4.3x greater throughput over baseline, integrated with Modal and SGLang for Qwen 397B.
A developer built a multimodal semantic search over 68k artworks from the National Gallery of Art using Qwen3-VL-Embedding, FAISS, Modal, and Cloudflare R2. The system achieves warm response times of ~1.3s and cold starts of ~44s, supporting both text-to-image and image-to-image queries.
A Modal tutorial demonstrating how to scale protein binder design using ESMFold2 and ESMC models, with code for iterative optimization and autoscaling infrastructure.
A tweet highlights that frontier reinforcement learning is now an infrastructure problem, noting the use of the open-source slime library in Modal's RL stack and upstream contributions.
A user expresses excitement about working on reinforcement learning at Modal, referencing Modal's announcement of an open-source library and lessons learned for scaling RL training.
Modal announces an open-source library for reinforcement learning on its platform, addressing infrastructure challenges in post-training RL with scalable deployment.
Modal announces day 0 support for Step 3.7 Flash, a 198B parameter MoE model with 256K context and native image/video understanding.
Modal announces day 0 support for the Step 3.7 Flash AI model, a 198B parameter MoE with 11B active parameters, 256K context, three reasoning levels, and native image and video understanding.
A technical guide on deploying a Hermes Agent using Fly, Modal, OpenRouter, and Cloudflare, with a detailed walkthrough from architecture to deployment.
OpenAI is co-hosting an autoresearch hackathon this Saturday with Raindrop AI and Modal, focusing on building self-improving systems like agents and models.
AI infrastructure startups Modal, Cerebras, Exa, and TurboPuffer have shown outstanding performance in the past week.
A user demonstrates running Microsoft Word inside a Modal sandbox on the day of Modal's Series C funding announcement.
The CEO of AI infrastructure platform Modal argues that while CUDA currently enjoys a strong moat, it will erode over 2-3 years as software improves to allow running CUDA code on alternative accelerators like TPUs.
Modal announces that AppliedCompute is using its platform to train custom agent workforces for companies like DoorDash, Mercor, and Cognition, highlighting the shift from frontier models to specialized models.