Schedule a Software Walkthrough
Reliable Software SOW Overview
DoW Reliable Software SOW — Task 6 of 9

Reliable Software Evaluations

A prediction is only useful if someone checks it against reality. Below are the core elements of periodically evaluating actual software reliability as the program progresses.

5 Elements
From metrics collection through review reporting
Every Build
Evaluated on a regular cadence tied to builds and increments
3-Way Check
Actual results compared against both prediction and allocation
Overview

Why software reliability evaluation is the reality check on prediction and allocation

Reliable allocation sets the target and prediction estimates whether it will be met, but software reliability evaluation is what confirms it using actual metrics gathered from development and test. Without it, a program has no way to know whether reliability is on track until it's too late.

Objective

Measured, not assumed

Reliable software evaluations are based on actual metrics collected from real builds, not status-report optimism.

Regular

Checked on a cadence

Reliable software evaluation happens on a defined schedule tied to builds and increments, not only when something goes wrong.

Comparative

Measured against something

Every reliable software evaluation compares actual results to both the predicted and allocated values.

01 — Core Elements

The core elements of reliable software evaluations

Each element below is a step in turning raw development data into a periodic, objective read on actual software reliability.

§1Metrics

Metrics Collection

Requirement

The developer shall collect defect and failure discovery times from testing for the relevant program increments.

Typical Contents
  • Defects are collected from problem reporting systems
  • Failures are events. Some defects lead to one failure, others lead to many failures.
  • Software reliability growth models such as Weibull, NHPP, Logarithmic, etc. are used. Each has unique assumptions about the inherent number of defects and their discovery trend. Most of the models assume that defect discoveries aren't increasing over the test effort. Requs Trend has the core models in the IEEE 1633. It shows you which models aren't trending to the data.
  • As per the IEEE 1633 2026, ensemble averaging is preferred. Contrary to popular myth, there is no "one" model that works on all data. The reliable software evaluation toolkit should have a selection of models with different assumptions. The models that have assumptions that differ from the data should be discarded. The remaining models should be "ensemble averaged". Requs Trend is the only software reliability forecast software that identifies the models that aren't fitting the data and ensemble averages the ones that are.
§2Cadence

Evaluation Cadence

Requirement

Evaluations shall occur on a defined, regular cadence tied to builds, increments, or program milestones.

Typical Contents
  • Cadence documented in the Reliable Software Program Plan
  • Evaluation timed to precede major technical reviews
  • Ad hoc evaluations triggered by significant changes in the software reliabiity forecast
§3Comparison

Trend Comparison

Requirement

Actual reliability metrics shall be compared against both the predicted values and the allocated targets.

Typical Contents
  • Actual-versus-predicted comparison to check prediction accuracy. Requs AI Predict Enterprise supports the comparison of the predicted versus observed.
  • Actual-versus-allocated comparison to check target achievement
  • Trend lines and software reliability forecast maintained across the program, not just point-in-time snapshots. Requs Trend uses various software reliability growth models to establish the software reliability forecasts and trend lines.
§4Outliers

Out-of-Family Identification

Requirement

If the defect discovery rate is increasing, more testing and defect reduction are in order.

Typical Contents
  • Increasing defect discovery means most of the models cannot be used.
  • This is typical early in testing because of "blocking" defects.
  • However, the software program increment should not be released as long as the trend is increasing.
§5Reporting

Reporting

Requirement

Evaluation results shall be reported at technical reviews and feed the program's risk management process.

Typical Contents
  • Results briefed at scheduled technical reviews. Requs Trend and Requs AI Predict were designed entirely for these reviews.
  • Findings documented for traceability across the program
  • Escalation path defined for evaluation results that miss targets. Requs Trend provides software reliability forecasting scenario based on different levels of defect removal and test effort.
02 — Why It Matters

What regular reliable software evaluation buys the program

01 — GROUND TRUTH

Status, not guesswork

Program management gets an objective read on reliability instead of relying on subjective confidence.

02 — EARLY WARNING

Catch drift before it's a crisis

Regular checks whether the software is trending towards the goal in time to react.

03 — ACCURACY FEEDBACK

Sharper predictions over time

Comparing actual results to predictions improves the accuracy of future prediction efforts.

04 — CONTINUOUS IMPROVEMENT

A record that pays off later

Trend data collected here strengthens test and corrective action planning downstream.

Check whether the software is actually meeting its reliability targets.

Start with the Requs Trend online demo or a discussion of your evaluation cadence.