← Back to the 309 Software Failures Timeline
Case Study #308 — 309 Software Failure Events Timeline

AWS US-EAST-1 DynamoDB DNS outage

A roughly 15-hour outage in Amazon's largest cloud region (US-EAST-1) took down or degraded over 100 AWS services and rippled out to major platforms including Snapchat, Fortnite, Roblox, Signal, Coinbase and Venmo, among many others.

2025-10-19
When it happened
Technology
Sector
#308 of 309
Ranked by documented impact
01 — What Happened

The damage was public. The root cause was preventable.

A roughly 15-hour outage in Amazon's largest cloud region (US-EAST-1) took down or degraded over 100 AWS services and rippled out to major platforms including Snapchat, Fortnite, Roblox, Signal, Coinbase and Venmo, among many others.

Technology
Sector affected
2025-10-19
Date of the event
#308
Rank in the 309 Software Failure Events Timeline
3
Distinct root-cause clauses identified below
02 — The Root Cause

What the software actually got wrong

AWS traced the root cause to a race condition in the automated system that manages DNS records for the DynamoDB database service's regional endpoint, which caused that endpoint's DNS record to be emptied out; every service that needed a fresh connection to DynamoDB in that region began failing DNS resolution, and the failure then cascaded into EC2's instance-launch subsystem and Network Load Balancer health checks, prolonging the recovery.

01Root Cause

AWS traced the root cause to a race condition...

What Happened

AWS traced the root cause to a race condition in the automated system that manages DNS records for the DynamoDB database service's regional endpoint, which caused that endpoint's DNS record to be emptied out

How Requs AI Catches This

Requs AI Edge Case flags this exact pattern at the requirements and architecture stage — before a single line of code implementing it exists — so the assumption behind it gets challenged while it is still cheap to fix.

02Root Cause

Every service that needed a fresh connection to DynamoDB...

What Happened

every service that needed a fresh connection to DynamoDB in that region began failing DNS resolution

How Requs AI Catches This

Requs AI Software FMEA traces this failure mode back to the system-level hazard it feeds, tagging it against the Common Defect Enumeration so it surfaces in review instead of in the field.

03Root Cause

The failure then cascaded into EC2's instance-launch subsystem and...

What Happened

the failure then cascaded into EC2's instance-launch subsystem and Network Load Balancer health checks, prolonging the recovery

How Requs AI Catches This

Requs AI Edge Case and Software FMEA together treat this as a known, catalogued defect pattern — not a novel surprise — so it gets tested for deliberately rather than discovered after deployment.

03 — How This Gets Caught Before It Happens

Beyond code coverage and "shall" testing

Root causes like this one rarely show up in code coverage or requirements-compliance testing, because nobody wrote a requirement anticipating the specific edge case that broke. Requs AI Edge Case and Requs AI Software FMEA are built to surface exactly this class of overlooked failure mode — before the software is written.

Requs AI Edge Case

Surfaces this before code exists

Identifies edge cases like this one at requirements and architecture time, using the Common Defect Enumeration to catalog failure patterns seen across hundreds of real-world software failures — including this one.

Requs AI Software FMEA

Connects the failure mode to the hazard

Traces this class of root cause directly to the system-level hazard it can produce, so a defect pattern like this one gets flagged during design review instead of after it ships.

Find the overlooked root causes before they ship.

Schedule a demonstration, or explore how Requs AI Edge Case and Software FMEA use the Common Defect Enumeration.