Tag
Modal is a serverless cloud platform designed for AI workloads, supporting inference, training, and sandboxes. It enables instant scaling from zero to thousands of GPUs with pure Python code, significantly reducing latency and accelerating time to market.
A user reports near-linear performance scaling when adding a second RTX 3090 for inference with a Qwen model, achieving roughly 1.8x decode TPS improvement without NVLink.
A Modal tutorial demonstrating how to scale protein binder design using ESMFold2 and ESMC models, with code for iterative optimization and autoscaling infrastructure.
OpenAI shares their deep learning infrastructure approach and open-sources kubernetes-ec2-autoscaler, a batch-optimized scaling manager for Kubernetes, emphasizing how infrastructure quality multiplies research progress.