@plantegg: I've been away from the company for a long time, so I can talk about this incident. The November 12, 2023 incident was arguably the most severe in Fubao Cloud's history (the person directly responsible was someone I interviewed and hired back then; later I specifically met him in person for a retrospective, so I know it very well). Even if this incident were multiplied by 10, it still wouldn't be enough for the July 26, 2026 Chrysanthemum Cloud incident...

X AI KOLs Timeline News

Summary

A former employee looks back on the severe outage at Fubao Cloud (Alibaba Cloud) on November 12, 2023, and says the outage at Chrysanthemum Cloud (Huawei Cloud) on July 26, 2026 was more than a hundred times worse, involving millions of instances locked and stopped, underscoring the absence of circuit breakers and manual confirmation mechanisms at cloud providers.

I've been away from the company for a long time, so I can talk about this incident. The November 12, 2023 incident was arguably the most severe in Fubao Cloud's history (the person directly responsible was someone I interviewed and hired back then; later I specifically went to him for a face-to-face retrospective, so I know it very well). Even multiplying this incident by 10 wouldn't be enough for the July 26, 2026 Chrysanthemum Cloud incident. Let me go through them one by one. The 2023 Fubao Cloud incident is simple. There was a routine script that refreshes the whitelist (it had been running for years). When pulling the whitelist, the API timed out, resulting in an incomplete whitelist. Without validation or error handling, it was pushed out as the full whitelist — folks, this is a whitelist; if something is missing, access won't work. So at that time, many services like OSS and image services were down. It wasn't that the services were down, but rather connection authentication failed and access couldn't get through—which is naturally no different from being down. ECS instances and RDS were running normally, but if the network relied on the whitelist, access would fail. However, ECS and RDS generally didn't depend on it. The Fubao Cloud 20231112 incident made the news and topped Weibo's trending list. But if you multiply this incident by 100, it just barely reaches the level of the Chrysanthemum Cloud 20260726 incident. Chrysanthemum Cloud 20260726: At 2 AM, a control-plane bug locked all major customer accounts (regarding them as in arrears), and then began to lock and shut down all, all, all instances of these customers (ECS/RDS/network/load balancers, etc.), covering all regions. My personal estimate is that the number of halted instances was over a million. It happened to be daytime in the Americas, so the impact on the other side of the earth was enormous. How long did it last? Recovery only started gradually after 7 AM, and instances were still being restored until 11 AM; at that point the control plane was queued up to pull instances back up. This incident can be called the most shocking and unbelievable incident in the world. There wasn't even the most basic circuit breaker, and just like that, over a million instances were shut down within half an hour. But there was no news about it at all—who cares? As a netizen said: Which dumbass customer is still using Chrysanthemum Cloud? Chrysanthemum is still an amateur in cloud products. Back when I did databases, there was a case where delayed monitoring data caused downstream to trigger automatic locking and upgrading of user instances. For this kind of operation, you must have a circuit breaker and manual intervention for secondary confirmation. Hope you enjoy reading this. Both of these incidents can only happen in a cloud environment; you couldn't reproduce them in your own IDC even if you wanted to.
Original Article
View Cached Full Text

Cached at: 08/14/26, 05:29 AM

I’ve been away from the job for a long time now, so I can talk about this incident. The outage on November 12, 2023 was probably the most severe in Fubao Cloud’s history (the person directly responsible was someone I interviewed and hired back in the day — I specifically met him in person afterward for a debrief, so I know the details very well).

That outage, even multiplied by 10, still wouldn’t match the Juhua Cloud outage on July 26, 2026. Let me go through them one by one.

The Fubao Cloud 2023 incident is straightforward. There was a routine whitelist-update script (which had been running for years). When it pulled the whitelist, the API timed out, resulting in an incomplete whitelist. No validation or error handling — it just pushed the partial whitelist out as if it were the full one. Folks, this is a whitelist — anything missing means no access.

So OSS, image services, and many other services went down. Well, not exactly down — connection authentication failed, so things were unreachable. Which is basically the same as being down.

ECS instances and RDS kept running fine, but anything that depended on the network whitelist became unreachable. ECS and RDS generally don’t depend on it.

The Fubao Cloud 20231112 incident made the news and topped Weibo’s trending list. But multiply that incident by 100 and you just barely reach the severity of the Juhua Cloud 20260726 incident.

Juhua Cloud 20260726: At 2 AM, a control-plane bug in Juhua Cloud locked down all major customer accounts (determining they were in arrears), and then proceeded to lock and shut down ALL, ALL, ALL of these customers’ instances — ECS/RDS/networking/load balancers, everything — across all regions.

My personal guess: the number of powered-down instances was in the millions. It happened to be daytime in the Americas, so the impact on the other side of the planet was enormous. How long did it last? Recovery only started around 7 AM, and instances were still being restored at 11 AM, with the control plane queued up the whole time trying to bring instances back up.

This is arguably the most alarming, most unbelievable cloud outage in the world. There wasn’t even a basic circuit breaker — over a million instances were shut down within half an hour.

But there was zero news coverage. Who cares? As one netizen put it: “What dumb customer is still using Juhua Cloud anyway?”

Juhua hasn’t even gotten started when it comes to cloud products. Back when I was working on databases, we had a case where monitoring-data latency triggered an automated account lock and instance upgrade for downstream users. For this kind of operation, you absolutely need a circuit breaker plus manual intervention for a second confirmation.

Hope you enjoyed the read. Both of these incidents are unique to cloud environments — you couldn’t recreate them with your own IDC infrastructure even if you tried.

Similar Articles

@dotey: Is OpenClaw Really Defunct? > @IFengTech [Meituan Executive Reflects on Company-wide Shrimp Farming] According to media reports, Wang Puzhong, CEO of Meituan's Core Local Business, stated in a public speech that he completely reviewed the process of AI transformation within Meituan. > When discussing from February to March this year, the company launched a company-wide 'Shrimp Farming Campaign', Wang Puzhong mentioned but the result...

X AI KOLs Following

Meituan's Core Local Business CEO Wang Puzhong reflected in a public speech on the failure of the company's AI transformation, leading to skyrocketing bills and operational disruptions.

@seclink: Internet companies release over 200,000 jobs, demand for 'AI+' talent rises. On July 2, data from the Ministry of Human Resources and Social Security showed that a month-long cloud recruitment month for internet companies has recently been launched, with over 5,000 internet companies collectively offering more than 200,000 job positions. Journalists found that the proportion of AI-related positions is continuously increasing...

X AI KOLs Timeline

Data from the Ministry of Human Resources and Social Security shows that over 5,000 internet companies released more than 200,000 jobs during the cloud recruitment month, with the proportion of AI-related positions continuously rising, reflecting that AI technology is accelerating its empowerment of various industries.

@seclink: Accused of 'Copying Code': Li Bojie (former Huawei 'Genius Youth', now Chief Scientist at Pine AI) was taking a remote interview with DeepSeek. Due to his habit of using a dual-screen setup, he occasionally glanced at the secondary screen on the left. The interviewer directly accused him of 'copying code' and demanded he prove his innocence, or the interview could not continue. The candidate's response…

X AI KOLs Timeline

Li Bojie, former Huawei 'Genius Youth', was accused of plagiarism by the interviewer for using a dual-screen setup during a remote interview with DeepSeek. He immediately terminated the interview and publicly criticized the interview process and lack of respect.

@JinYu762: Don't just focus on single tech stocks. What the July earnings truly reveal is the AI supply chain. Is it still burning hot, or is the market starting to cool? 7/7 Samsung Q2 preliminary results: First, check the temperature of memory, HBM and DRAM prices, AI server demand – all will indicate some trends. 7/15 $ASML: Watch equipment orders, EUV/DUV, advanced process demand, revenue from China region, and impact of export restrictions…

X AI KOLs Following

This article outlines key points of the AI supply chain in July earnings reports from multiple tech companies, including trends in memory, equipment, foundry, AI chips, and cloud services, emphasizing TSMC's central role in the AI supply chain.