Why Does the Cloud Stop Computing?: Lessons from Hundreds of Service Outages (2016)
Summary
Analyzes hundreds of cloud service outages to derive lessons on improving reliability and understanding why cloud computing fails.
Similar Articles
Omnipresent availability risks in cloud software
The article discusses omnipresent availability risks in cloud software, covering common failure areas like saturation, networking, and security, with examples from major incidents.
Lessons Learned from Building Cloud Agents (12 minute read)
Cursor shares key lessons from building cloud agents, emphasizing that providing a full development environment is critical for agent output quality, and that long-running agents require durable execution and enterprise-like infrastructure.
21 years and counting of 'eight fallacies of distributed computing' (2025)
A retrospective on the eight fallacies of distributed computing, originally formulated by Sun Microsystems engineers, examining their continued relevance 21 years later for network operators and developers.
GitHub, autoscaling, and the component substitution fallacy
The blog post analyzes a recent GitHub outage caused by a misconfigured autoscaling policy, discussing the challenges of autoscaling in cloud infrastructure and referencing the component substitution fallacy to highlight systemic issues.
The Cloud Is Becoming a Geopolitical Risk
The article discusses how cloud computing infrastructure is increasingly becoming a point of geopolitical tension, posing risks to global stability and security.