GitHub 24-hour database consistency incident
GitHub experienced its longest-ever incident, with degraded service for over 24 hours, as engineers worked to safely reconcile inconsistent data between two data centers.
The damage was public. The root cause was preventable.
GitHub experienced its longest-ever incident, with degraded service for over 24 hours, as engineers worked to safely reconcile inconsistent data between two data centers.
What the software actually got wrong
A brief network partition caused GitHub's database clusters in two data centers to each continue accepting writes independently; reconciling the resulting conflicting data without losing any user information required a lengthy, careful manual process rather than an automatic recovery.
A brief network partition caused GitHub's database clusters in...
A brief network partition caused GitHub's database clusters in two data centers to each continue accepting writes independently
Requs AI Edge Case flags this exact pattern at the requirements and architecture stage — before a single line of code implementing it exists — so the assumption behind it gets challenged while it is still cheap to fix.
Reconciling the resulting conflicting data without losing any user...
reconciling the resulting conflicting data without losing any user information required a lengthy, careful manual process rather than an automatic recovery
Requs AI Software FMEA traces this failure mode back to the system-level hazard it feeds, tagging it against the Common Defect Enumeration so it surfaces in review instead of in the field.
Beyond code coverage and "shall" testing
Root causes like this one rarely show up in code coverage or requirements-compliance testing, because nobody wrote a requirement anticipating the specific edge case that broke. Requs AI Edge Case and Requs AI Software FMEA are built to surface exactly this class of overlooked failure mode — before the software is written.
Surfaces this before code exists
Identifies edge cases like this one at requirements and architecture time, using the Common Defect Enumeration to catalog failure patterns seen across hundreds of real-world software failures — including this one.
Connects the failure mode to the hazard
Traces this class of root cause directly to the system-level hazard it can produce, so a defect pattern like this one gets flagged during design review instead of after it ships.
Find the overlooked root causes before they ship.
Schedule a demonstration, or explore how Requs AI Edge Case and Software FMEA use the Common Defect Enumeration.