← Back to Articles
CMMi Level 5 doesn't beat Level 3. More testers doesn't help. Here's what the data actually says is overkill.
RELIABLE SOFTWARE PULSE Mission Ready | Requs AI

Software Reliability Benchmarking

The Reliability Myths: Five Common Practices, Zero Measurable Effect

679 factors were tracked. Some of the most popular ones turned out to be noise.

Last issue was about what the top 3% of software teams do that everyone else doesn't. This issue is the mirror image: the practices that organizations pour time, budget, and headcount into — that the 679-factor benchmarking study found little to no relationship to defect density.

None of these are bad ideas on their own. The data just doesn't support treating them as the lever that moves you from mediocre to elite. If you're resourcing-constrained, this is where you can probably stop over-investing.

1. Chasing CMMi Level 4 or 5

There's a real, visible climb from Unrated through Level 3 — average defect density drops from 18–22 down to 8.58. Past that, the curve goes flat. Level 4 organizations averaged 10.23 defects per work year; Level 5 averaged the same 10.23. The most reliable projects in the entire database (3rd percentile) averaged only a Level 2.08 rating. A higher maturity rating past Level 3 doesn't buy you anything the data can detect.

8.58 → 10.23
avg. defect density, CMMi Level 3 vs. Level 5
2.08
avg. CMMi level of the most reliable (3rd percentile) projects
Bar chart: average defect density by SEI CMMi level. Unrated 22.77, Level 1 18.48, Level 2 11.88, Level 3 8.58, Level 4 10.23, Level 5 10.23. The curve drops steeply through Level 3, then flattens.

2. The 2-week sprint as a hard rule

Shorter cycles genuinely correlate with lower defect density — that part of the Agile playbook holds up. But the study found no evidence that the cycle specifically needs to be two weeks. Several successful, reliable organizations in the database ran 9- to 10-week sprints. The reliability driver is granularity and feedback frequency, not a specific sprint length. Treating "two weeks" as gospel is optimizing for a number instead of the underlying behavior.

Contrary to claims by Agile zealots, there is no correlation that cycle times must be 2 weeks. The shortest cycle that lets the work actually get done is the one that matters.

3. Hiring more testers

The fraction of testers to engineers stays remarkably flat — between 10% and 18% — across every single percentile group, from the elite 3% to the distressed 97%. Padding the tester headcount doesn't move a team up the reliability curve. The one place the ratio does matter: when it drops below 10%, distress becomes more likely. So there's a floor, not a target to maximize.

4. Screening for language seniority

"Years of experience with the selected language" is high everywhere — 90% in the top 3% tier, 100% in the 50th percentile, 100% in the worst-performing 97th percentile. Every tier is stacked with language veterans. It's a baseline hiring filter, not a differentiator. (Last issue's finding still holds: domain and industry expertise is what actually separates the tiers — 1.25 experts per team at the top vs. 0.33 at the bottom.)

5. Formal SQA audits and conformity reviews

This is the strangest finding in the whole study. If SQA activity drove reliability, you'd expect it to climb steadily toward the elite 3% tier. Instead, it's non-linear — the mediocre 50th percentile actually out-audits the elite 3%.

SQA Activity Top 3% 50th %ile
Conformity review before delivery 25% 46%
Formal audit/record results 25% 31%

The 3% tier isn't winning because of SQA oversight — it's winning on technical engineering rigor earlier in the lifecycle. SQA verifies that a document exists; it doesn't verify that the requirements cover boundary conditions or that the unit tests hit MCDC. It's a floor that keeps the worst outcomes from happening, not a ceiling that produces the best ones.

The common thread

Every item on this list is an administrative or proxy metric — a maturity rating, a ceremony, a headcount ratio, a resume line, an audit checkbox. None of them require anyone to look at a specific failure mode and decide how the software should handle it. That's the same theme from last issue, just from the other direction: technical engineering rigor differentiates; administrative process is a floor everyone should clear and then stop optimizing.

Not sure which of your practices are actually working?

Requs AI Predict scores your project against all 679 benchmarked factors — so you know where to keep investing and where to stop.

Explore Requs AI Predict

Predict Earlier. Prevent More. Be Mission Ready.

Source: Industry Benchmark on Software Defects, Revision 8, 2026, © Mission Ready Software. 149 projects, 89 high-fidelity datasets, 679 factors tracked over a 33-year study period.

© 2026 Mission Ready · missionreadysoftware.com