|
||||||||||||||||
|
Software Reliability Benchmarking The Reliability Myths: Five Common Practices, Zero Measurable Effect679 factors were tracked. Some of the most popular ones turned out to be noise. |
||||||||||||||||
|
Last issue was about what the top 3% of software teams do that everyone else doesn't. This issue is the mirror image: the practices that organizations pour time, budget, and headcount into — that the 679-factor benchmarking study found little to no relationship to defect density. None of these are bad ideas on their own. The data just doesn't support treating them as the lever that moves you from mediocre to elite. If you're resourcing-constrained, this is where you can probably stop over-investing. 1. Chasing CMMi Level 4 or 5There's a real, visible climb from Unrated through Level 3 — average defect density drops from 18–22 down to 8.58. Past that, the curve goes flat. Level 4 organizations averaged 10.23 defects per work year; Level 5 averaged the same 10.23. The most reliable projects in the entire database (3rd percentile) averaged only a Level 2.08 rating. A higher maturity rating past Level 3 doesn't buy you anything the data can detect.
2. The 2-week sprint as a hard ruleShorter cycles genuinely correlate with lower defect density — that part of the Agile playbook holds up. But the study found no evidence that the cycle specifically needs to be two weeks. Several successful, reliable organizations in the database ran 9- to 10-week sprints. The reliability driver is granularity and feedback frequency, not a specific sprint length. Treating "two weeks" as gospel is optimizing for a number instead of the underlying behavior.
3. Hiring more testersThe fraction of testers to engineers stays remarkably flat — between 10% and 18% — across every single percentile group, from the elite 3% to the distressed 97%. Padding the tester headcount doesn't move a team up the reliability curve. The one place the ratio does matter: when it drops below 10%, distress becomes more likely. So there's a floor, not a target to maximize. 4. Screening for language seniority"Years of experience with the selected language" is high everywhere — 90% in the top 3% tier, 100% in the 50th percentile, 100% in the worst-performing 97th percentile. Every tier is stacked with language veterans. It's a baseline hiring filter, not a differentiator. (Last issue's finding still holds: domain and industry expertise is what actually separates the tiers — 1.25 experts per team at the top vs. 0.33 at the bottom.) 5. Formal SQA audits and conformity reviewsThis is the strangest finding in the whole study. If SQA activity drove reliability, you'd expect it to climb steadily toward the elite 3% tier. Instead, it's non-linear — the mediocre 50th percentile actually out-audits the elite 3%.
The 3% tier isn't winning because of SQA oversight — it's winning on technical engineering rigor earlier in the lifecycle. SQA verifies that a document exists; it doesn't verify that the requirements cover boundary conditions or that the unit tests hit MCDC. It's a floor that keeps the worst outcomes from happening, not a ceiling that produces the best ones. The common threadEvery item on this list is an administrative or proxy metric — a maturity rating, a ceremony, a headcount ratio, a resume line, an audit checkbox. None of them require anyone to look at a specific failure mode and decide how the software should handle it. That's the same theme from last issue, just from the other direction: technical engineering rigor differentiates; administrative process is a floor everyone should clear and then stop optimizing.
|
||||||||||||||||
|
Predict Earlier. Prevent More. Be Mission Ready. Source: Industry Benchmark on Software Defects, Revision 8, 2026, © Mission Ready Software. 149 projects, 89 high-fidelity datasets, 679 factors tracked over a 33-year study period. © 2026 Mission Ready · missionreadysoftware.com |