Tag
GPUHedge is an open-source tool that uses speculative execution to hedge between serverless GPU providers, reducing cold start p95 latency from 117s to 30s.
Serverless GPUs may have higher hourly rates but can be more cost-effective overall depending on workload peak-to-average demand. The article on Modal's blog illustrates this with a widget.
Modal has announced that replicas of vLLM and SGLang servers now start up 3-10x faster, leveraging improvements in GPU health management and CUDA context checkpointing.