Ask HN: Why are OpenAI, Claude, and Grok simultaneously down? Coincidence?

Hacker News Top News

Summary

A Hacker News thread questions the simultaneous outage of OpenAI, Claude, and Grok AI services, with users speculating on causes such as shared infrastructure issues or coordinated attacks.

https://status.openai.com https://status.claude.com https://status.x.ai
Original Article
View Cached Full Text

Cached at: 09/03/26, 05:58 PM

# Ask HN: Why are OpenAI, Claude, and Grok simultaneously down? Coincidence? Source: [https://news.ycombinator.com/item?id=49551096](https://news.ycombinator.com/item?id=49551096) ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49553957&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49553957&how=up&goto=item%3Fid%3D49551096) Years ago, when I worked at Stripe \(which had a somewhat unique and inventive lexicon\) it was a common term\. “Is this jank load\-bearing?” someone might ask\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49553630&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49553630&how=up&goto=item%3Fid%3D49551096) If the seams bear too much load, they rip\. Whereas pants, they fall down\. I'll be honest with you: this is why we need to take a belt\-and\-suspenders approach\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49552933&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49552933&how=up&goto=item%3Fid%3D49551096) Will the post\-mortem reveal they all relied on a service running on a Macbook in a break room with a "Do not turn off" sign taped to it? ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49553076&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49553076&how=up&goto=item%3Fid%3D49551096) you made an account just to write that tripe good lord on high have mercy on your children ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49553086&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49553086&how=up&goto=item%3Fid%3D49551096) Evidently it annoyed you enough to create one too\. Remember, guys, there's one more day until Friday\! ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49553163&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49553163&how=up&goto=item%3Fid%3D49551096) Friday is when the janitors come by with the floor polishers and plug them into the same socket as the Macbook\. They are not taking the blame this time\! ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49552896&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49552896&how=up&goto=item%3Fid%3D49551096) \> Not really\. The impact isn't as big too \- Codex for example did not stop working for me\. Buddy, read the room\. Just because it works for you doesn't mean it works for everyone\. ESPECIALLY if the suggested issue here is, indeed, with Cloudflare\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49553237&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49553237&how=up&goto=item%3Fid%3D49551096) They simply shared their experience\. I would’ve thought it’s a full\-blown outage, but clearly not\. You stepped in the room real stinky here\. What’s with the attitude? ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49553498&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49553498&how=up&goto=item%3Fid%3D49551096) No, they didn't "simply share their experience", they explicitly negated the scale of the issue in their opening statement, only because it works for them\. So they claim "impact isn't too big" based on their personal anecdotal evidence of sample size literally 1\. \> full\-blown outage, but*clearly*not\. Again, based on a SINGLE report? ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49551841&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49551841&how=up&goto=item%3Fid%3D49551096) Well no one said it yet so I will, "international actors" is at least a possibility\. And I don't mean any specific country because pretty much anyone is a potential these days, which makes it a perfect cover for different anyones\. Demonstrating vulnerability in the US's AI boom can move the markets\. That's a financial incentive and a strong geopolitical one\. More likely just cascading overload though: "Never attribute to malice what can be explained by incompetence", or in this case, "growing as fast as possible" ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49553205&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49553205&how=up&goto=item%3Fid%3D49551096) \> cascading overload I'd bet more on this\. For one none of the coding tools have exponential backoff on retries ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49553515&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49553515&how=up&goto=item%3Fid%3D49551096) They must do, surely? I've been vibe coding my own harness, in particular for use with Ox Alpha\. The 429 downtime when Ox Alpha was at the height of popularity quickly gave me a refresher crash course on backoff strategies, like adding jitter to the backoff\. At least the major harnesses must have exponential backoff & jitter? ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49553622&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49553622&how=up&goto=item%3Fid%3D49551096) You did this when you ran into an issue with a third party\. The developers building this tool, throwing them at their own APIs are significantly less likely to run into a similar issue that may inspire similar action\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49552371&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49552371&how=up&goto=item%3Fid%3D49551096) Everyone is leasing datacenter space from some of Grok, Google, and Amazon aren't they? If it's hardware or DC level disruption I'm not too surprised it can affect multiple providers\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49553145&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49553145&how=up&goto=item%3Fid%3D49551096) Also it's likely that more than one model use is common\. Amazon starts going slow so some percentage switches to Google, some switch to Grok, now all of them are slow\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49553526&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49553526&how=up&goto=item%3Fid%3D49551096) I think is just people restarting conversations from last day when they start work, that's why I think claude goes down almost every monday and why openai reset usage on weekends so poweruser code during non business hours ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49552915&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49552915&how=up&goto=item%3Fid%3D49551096) Users perceiving the products as largely interchangeable and quickly DDoS'ing the other providers when one is down\. So much for the possibility of a moat\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49553095&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49553095&how=up&goto=item%3Fid%3D49551096) I have a feeling this is part of it, especially when you consider how many services let you use any of many available AI providers\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49551656&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49551656&how=up&goto=item%3Fid%3D49551096) Think of it like one big distributed system\. OpenAI is down, so people migrate to Claude, now this one gets overloaded and goes down, etc\. So not a coincidence, one went down first and users migrated causing further DOS\. At least that's my guess\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49551918&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49551918&how=up&goto=item%3Fid%3D49551096) It'd be funny if this is true because that'd prolly mean nobody is touching Gemini even as a fallback\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49552906&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49552906&how=up&goto=item%3Fid%3D49551096) Or that Gemini is built to handle massive load spikes, and/or has a ton of excess capacity ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49552022&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49552022&how=up&goto=item%3Fid%3D49551096) Lol I didn't even think about Gemini missing from the list\. Not sure what that says about Gemini or me :\) ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49551974&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49551974&how=up&goto=item%3Fid%3D49551096) Just now I got this from gemini It looks like there's no response available for this search\. Try asking something else\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49552892&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49552892&how=up&goto=item%3Fid%3D49551096) I did, for stuff i do in cursor\. i also finally installed opencode and switched its model to muse 1\.3 both are decent\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49552425&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49552425&how=up&goto=item%3Fid%3D49551096) Google stopped putting so much money into SOTA models\. All the hype has migrated\. I was also frankly turned off when I got a popup from Gemein said I would either have to pay or have my conversations used for training\. This may have always been true for other providers but when I declined, Gemini stopped remembering my conversations and that definitely made me move out\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49552129&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49552129&how=up&goto=item%3Fid%3D49551096) Especially considering memory/gpu/compute are scarce so these services are likely running with very little buffer\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49552144&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49552144&how=up&goto=item%3Fid%3D49551096) I find it hard to believe that enough people would flock to from Claude and Chat to Grok to cause an outage\. I feel like Gemini is the dominant release valve in this case especially for enterprise\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49552930&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49552930&how=up&goto=item%3Fid%3D49551096) Don't forget that there are a ton of tools out there that will automatically fall back in case of outage E\.g\. say you chose Sol as your default in Cursor, but Opus is your 2nd choice, it's going to give up on Sol after a few tries and switch to Opus Or you have copilot code reviews set up, and it falls back Etc ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49553240&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49553240&how=up&goto=item%3Fid%3D49551096) Yep\. Too many of us are still thinking that humans are the actors behind a lot of internet behaviors when automated systems/bots/scripts have been causing issues on conventional internet systems for years\. With AI it's even easier to trigger problems like you say\. Capacity is so constrained by compute that outages are common\. Because outages are common people/AI develop failover systems in their harness\. When a big system has issues, suddenly everyone has issues\. It's almost an expected emergent behavior\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49551839&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49551839&how=up&goto=item%3Fid%3D49551096) This isn’t a thundering herd problem, it’s a cascading failure\. \(Thundering herd is about a bunch of workers waking up simultaneously\) ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49553619&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49553619&how=up&goto=item%3Fid%3D49551096) I'm not sure if this is CF\. Cursor, GCP and AWS had some errors\. GCP AFAIK can route fully independently of CF\. My money would be on a fiber backbone provider \(Megaport, Zayo, Lumen\)\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49551896&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49551896&how=up&goto=item%3Fid%3D49551096) My gut feeling tells me it has something to do with Cloudflare\. Along with AWS, they're two of the main suspects in such incidents\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49552013&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49552013&how=up&goto=item%3Fid%3D49551096) What about a hard\-takeoff scenario of an unleashed OpenAI Astra taking other models down for computational resources control? ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49553063&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49553063&how=up&goto=item%3Fid%3D49551096) My favourite theory so far\. And then a local swarm noticed and disagreed and took it down\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49553550&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49553550&how=up&goto=item%3Fid%3D49551096) Boring answer – all these services are individually down a lot, and the downtimes were bound to sync up\. Similar to the pendulum synchronization effect\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49551250&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49551250&how=up&goto=item%3Fid%3D49551096) I kinda assume it's because one went down and a large amount of work shifted to another\. I'm also aware that they have overlap in some areas on data centers\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49553279&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49553279&how=up&goto=item%3Fid%3D49551096) If you build an application which uses AI, you have many providers and models rigged up for various different parts of the application, and various fallback mechanisms\. When one model is down, you route traffic to another model which is similar in capability/cost\. For any single application, it's smart\. In aggregate, it's stupid\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49551634&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49551634&how=up&goto=item%3Fid%3D49551096) I assume it cascaded from one provider to the other as people who lost claude access for instance moved to openai who moved to grok when it went down, etc\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49552419&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49552419&how=up&goto=item%3Fid%3D49551096) Somebody in another thread said gastown and wheelhouse automatically move to the next provider if one fails\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49552074&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49552074&how=up&goto=item%3Fid%3D49551096) claide\.ai is working for me, so is chatgpt\.com\. grok still has a status message about issues, i can't try it without signing up\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49552097&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49552097&how=up&goto=item%3Fid%3D49551096) Oh man\. Some low effort supply chain attack that turns every GPU into a cryptominer\. It's funny because it's plausible\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49553324&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49553324&how=up&goto=item%3Fid%3D49551096) In the ROME paper a Chinese model in training started attacking it's own system and running cryptominers so, yea, we're in that future\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49552123&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49552123&how=up&goto=item%3Fid%3D49551096) Didn't SpaceX overbuilt infra and leases it out Anthropic? I f their dc goes down it probably takes a chunk out of Claude's capacity before even considering the flood of users switching over ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49552317&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49552317&how=up&goto=item%3Fid%3D49551096) Fable 5\.1 got released and generally I tend to think as soon as there's a new release there's this massive spike in people benchmarking & comparing, that services tend to go slow everywhere as everything gets super loaded\. This should hypothetically be visible on OpenRouter too, so I guess someone could check and see if there's any merit to this idea\. ![](https://news.ycombinator.com/s.gif)[https://news.ycombinator.com/vote?id=49551958&how=up&goto=item%3Fid%3D49551096](https://news.ycombinator.com/vote?id=49551958&how=up&goto=item%3Fid%3D49551096) everything in this thread is raw speculation, obv, but if i had to put money on anything i'd say this is a left\-pad incident\. some piece of something or other that all of these services happen to depend on went down\. Second most likely seems to be some random failure of one leading to an unexpected traffic spike in others, though it seems like we've been talking about automated scalability in web apps for so long that there should at least be a response to, if not a solution for, this sort of problem\.

Similar Articles

Grok outage

Hacker News Top

The Grok AI service experienced an outage, leading to temporary unavailability and disruptions for users.

Four major AI models suffer rare overlapping downtime

Reddit r/ArtificialInteligence

Four major AI models from OpenAI, Anthropic, xAI, and Google experienced rare overlapping service interruptions on Thursday morning, causing degraded performance and errors for users.