GPUHedge: Hedging serverless GPU providers improves cold start p95 latency from 117s to 30s [P]
Summary
GPUHedge is an open-source tool that uses speculative execution to hedge between serverless GPU providers, reducing cold start p95 latency from 117s to 30s.
Similar Articles
How to achieve truly serverless GPUs (20 minute read)
Modal explains the four key ingredients they developed to spin up serverless GPU inference replicas in seconds instead of minutes, enabling efficient GPU allocation for variable AI workloads.
Launch HN: Expanse (YC P26) – Unlock Wasted GPU Capacity
Expanse is a startup that improves GPU/HPC cluster utilization by predicting job resource needs and providing optimizations, addressing the common problem of over-requesting resources that leads to 30-40% effective utilization.
Reduce GVisor Cold Starts with GPU Snapshotting
Cerebrium reduces GPU cold starts for AI workloads by checkpointing CPU and GPU memory, restoring fully initialized containers in seconds, cutting startup time by over 80%.
@svpino: Serveless, but for models, is here! I stopped renting servers 12 years ago. I mostly moved everything to serverless fun…
Runpod announces FlashBoot, a serverless approach for AI models that moves idle models to cheaper storage and pages them back to GPU, achieving cold starts under 200ms and cutting costs by 90% compared to traditional clouds.
GPU access is still broken in 2026 — and someone's trying to fix it with a compute futures market
Inferra is building a derivatives exchange for GPU compute, offering perpetual futures for chips like H100 and B200 to enable price discovery and cost hedging, aiming to fix the opaque GPU market.