On October 20, 2025 (Beijing time: evening of October 20 to early morning of October 21), Amazon Web Services (AWS) experienced a massive outage in its US-EAST-1 region (Northern Virginia). The incident lasted over 15 hours for some services (with full recovery taking even longer in some instances) and is regarded as one of the most severe cloud infrastructure failures in recent years. Monitoring platforms like Downdetector received more than 17 million outage reports, impacting thousands of websites, apps, and online services worldwide.
Timeline of the Outage
- Start: October 19, 23:49 PDT (approx. 14:49 Beijing time on October 20) – Error rates began spiking.
- Peak impact: Early morning, October 20 UTC (around 06:49 UTC) – Global user reports surged.
- Core issue mitigated: October 20, 02:24 PDT (17:24 Beijing time) – DNS resolution problems alleviated.
- Full recovery: Most services returned to normal by late afternoon on October 20, although backlog processing continued for several hours. Total duration from start to complete stabilization exceeded 15 hours.
Root Cause
This was not a cyberattack, but a cascading failure triggered by a latent bug in an internal automation system:
- A race condition occurred in DynamoDB’s automated DNS management system: two automated processes simultaneously attempted to update the same DNS record, resulting in the record being overwritten with an empty state.
- Consequence: Applications could no longer resolve DynamoDB API endpoints via DNS (akin to a phonebook entry being deleted).
- This triggered a chain reaction: Services dependent on DynamoDB — including EC2 instances, load balancers, and network health propagation — were all affected. Even after DynamoDB was fixed, EC2’s Droplet Workflow Manager became overloaded, handling a massive backlog of expired leases, further delaying recovery.
AWS’s official post-incident summary described it as “what appeared to be a routine DNS resolution issue” that exposed critical shortcomings in fault isolation at hyperscale.
Affected Services and Platforms (Partial List)
Because US-EAST-1 is AWS’s default and busiest region, many global services (even those claiming multi-region redundancy) still relied on it for authentication, metadata, or databases, leading to worldwide disruptions:
- Social & Entertainment: Snapchat, Roblox, Fortnite, Reddit, Disney+, Hulu
- Finance & Payments: Coinbase, Venmo, several UK banks (Lloyds, Halifax, etc.)
- Smart Home: Amazon Ring, Alexa, Eight Sleep innovative mattresses
- Productivity: Slack, Zoom, Canva, Duolingo
- Others: Signal, Perplexity AI, Pokémon GO, UK tax authority HMRC, airline apps (United, Delta, etc.)
Even Amazon’s own e-commerce platform, Prime Video, and customer support systems were briefly paralyzed.
Scale of Impact and Aftermath
- Global reach: European morning commuters and North American daytime users were hit in two successive waves.
- Economic losses: Mid-sized e-commerce sites reported losses in the tens of thousands of dollars per day; total internet-wide damage remains incalculable.
- AWS response: The automated DNS tool for DynamoDB has been turned off globally, with additional testing and safeguards implemented. Affected customers qualifying under SLA terms can apply for a 30% service credit (manual application required).
- Industry reflection: This marks the third major incident in the US-EAST-1 region in recent years (previous ones in 2021 and 2020). Experts highlight the dangers of over-centralization on a handful of cloud giants (AWS ≈30% market share, Azure ≈24%, GCP, and others trailing).
Comparison with Other Recent Incidents
- October 2025: AWS (DNS configuration error, >15 hours)
- November 2025: Large-scale Microsoft Azure outage (details limited, but occurred shortly after AWS)
- November 18, 2025: Cloudflare bot-protection collapse (unrelated to AWS but similar global impact)
These consecutive events serve as a stark reminder: multi-cloud/hybrid architectures, stronger fault isolation, and genuine regional redundancy are now essential survival strategies for any serious online business.
If you need the official AWS post-incident report link or recovery details for specific services, please let me know — I can retrieve them for you.
The above content is compiled by ModeZone, a fashion and entertainment magazine.