The component does not qualify
An accelerated life test of 60 units. The point estimate of reliability clears the criterion. The 90% lower confidence bound does not. Qualification turns on the bound.
The primary data is simulated. A known Weibull and Arrhenius model generates the failure times. The injected values are beta = 2.5, Ea = 0.70 eV, and eta = 8000 h at the use condition.
This is a deliberate choice, and it is a strength here. You know the true answer, so you can measure what the estimators recover. A borrowed dataset cannot offer that check. The estimators are also validated against a published external dataset.
The numbers on this page are computed. They are read from the qualification artifact at build time.
The engineering problem
A component must survive a 3000 hour mission in the field. The use condition is 40 C.
A test at the use condition would take years. So the test applies raised temperature instead. Raised temperature makes the component wear out faster.
The test runs at 85, 105, and 125 C. An Arrhenius-Weibull model fits the failure times. The model then extrapolates the life back down to 40 C.
The acceptance criterion is fixed before the test. Reliability must be at least 0.90 at 3000 hours. You must demonstrate this on the 90% lower confidence bound, not on the point estimate.
Qualification result
NOT DEMONSTRATED
The point estimate and the lower bound fall on opposite sides of the criterion. The decision follows the bound. So the component does not qualify.
B10 life is the time by which 10% of units fail. B10 agrees with the reliability result, and B10 is an independent view of the same fit. Both point estimates clear their target. Both lower bounds miss their target.
Why the verdict says NOT DEMONSTRATED and not FAIL
The wording is deliberate. Qualification is a demonstration.
The requirement is to SHOW that the lower bound clears the criterion. This evidence does not show that. So the correct report is NOT DEMONSTRATED.
FAIL would claim more than the evidence supports. FAIL says the component is bad. The data does not say that. The point estimate of R(3000 h) is 0.9165, and the true reliability could be above the criterion.
What the data does say is narrower. At 90% confidence, this test does not establish the claim. The component does not qualify on this evidence.
This is a clear miss, not a close call
A reader can reasonably ask if the shortfall is only noise in the bootstrap. It is not.
The bootstrap bound varies by about 0.013 across seeds. That variation is roughly 10% of the gap to the criterion. Every seed puts the lower bound on the same side of the criterion.
B10 gives the same answer, and B10 misses by a wider relative margin. The B10 lower bound is 2125 h against a 3000 h mission.
The two metrics agree, and the seeds agree. So the verdict does not depend on the seed.
The qualification matrix
2 of 7 requirements are not demonstrated. 5 requirements pass. Every row is read from the qualification artifact.
| Requirement | Method | Stress | Sample | Criterion | Result | Verdict |
|---|---|---|---|---|---|---|
| R >= 0.90 at 3000 h, 40 C | Arrhenius-Weibull ALT, censored MLE, bootstrap bounds | 85 / 105 / 125 C, extrapolated to 40 C | 60 (18 censored) | 90% lower bound >= 0.90 | lower bound 0.7771 (point 0.9165); short of the criterion by 0.1229 (13.7%) | NOT DEMONSTRATED |
| B10 life >= mission time (3000 h) | Weibull B10 from the extrapolated use-condition fit | extrapolated to 40 C | 60 (18 censored) | 90% lower bound >= 3000 h | lower bound 2125 h (point 3229 h); short of the 3000 h mission by 875 h (29.2%) | NOT DEMONSTRATED |
| Failure mechanism unchanged across stress levels | Likelihood ratio test, common vs free shape, scales free | 85 / 105 / 125 C | 60 (18 censored) | p >= 0.05 (do not reject common shape) | p = 0.3286 | PASS |
| Activation energy physically plausible | Ea fitted from the ALT model, Ea = a * k | 85 / 105 / 125 C | 60 (18 censored) | 0.3 to 1.2 eV | Ea = 0.7017 eV, 90% CI [0.6204, 0.7795] | PASS |
| Wear-out mechanism identified for maintenance policy | Weibull shape parameter from censored MLE | all cells, common shape | 60 (18 censored) | beta > 1 indicates wear-out | beta = 2.576, 90% CI [2.128, 3.283] | PASS |
| Estimators validated against an external reference | Censored Weibull MLE and KM vs published dataset | n/a (NCCTG lung, Loprinzi et al. 1994) | 228 (63 censored) | two independent implementations agree | shape agrees to 1.5e-07, scale to 4.4e-06 | PASS |
| Censored units retained in every fit | Likelihood contribution S(t) for censored units | all cells | 18 censored of 60 | no censored unit discarded | all 18 retained; discarding the pooled data understates eta by 30.5% | PASS |
The two rows that are not demonstrated are the reliability claim and its B10 equivalent. Both turn on the lower confidence bound. Both point estimates clear their target.
The model passes its own validity checks
These checks support the MODEL. These checks do not support the PRODUCT. The distinction matters. A sound model that reports a bound below the criterion is still a component that does not qualify.
Common shape across the stress cells
An Arrhenius model assumes one failure mechanism at every temperature. If the Weibull shape changes with temperature, the mechanism changed. Then the extrapolation is not valid.
The likelihood ratio test gives p = 0.3286. The test does not reject the common shape.
Read that result carefully. The test does not reject, and the test does not prove. With 20 units per cell the test has limited power.
Activation energy
The fitted activation energy is 0.7017 eV. The physically plausible band is 0.3 to 1.2 eV for a single thermally activated mechanism. The fitted value sits inside that band.
Wear-out
The Weibull shape is beta = 2.576. A beta above 1 means the hazard rate rises with time, which identifies wear-out.
This result has a direct maintenance consequence. The correct mitigation is scheduled replacement before the wear-out knee. Burn-in is the wrong mitigation here. A beta below 1 would mean infant mortality, and then burn-in would be correct.
Estimators validated against a published dataset
Two independent implementations fit the same published dataset. The dataset is the NCCTG lung dataset, from Loprinzi et al. 1994, with 228 units, 63 of them censored.
Shape agrees to 1.5e-07, scale to 4.4e-06. This checks the estimator against an outside reference, rather than against its own output.
Censoring
A censored unit is a unit that did not fail before the test stopped. You know it survived to a certain time, and you do not know when it fails.
This test has 18 censored units of 60, which is 30% of the sample. Every censored unit is retained in every production fit. Each censored unit contributes its survival function to the likelihood.
The project measured what the alternative costs. When you discard the pooled censored data, the fit understates eta by 30.5%.
The bias runs in one direction for eta. The bias does not run in one direction for every derived metric. When you discard censored units, you also corrupt the Weibull shape, so a derived metric can move either way.
What the estimators recover from the injected truth
The data comes from a known model. So you can compare the fit against the values that generated the data.
| Parameter | Injected truth | Fitted | Difference |
|---|---|---|---|
| Weibull shape, beta | 2.500 | 2.576 | +3.0% |
| Activation energy, Ea (eV) | 0.7000 | 0.7017 | +0.2% |
| Characteristic life at use (h) | 8000 | 7736 | -3.3% |
Injected values read from the recorded ground truth. Fitted values read from the qualification artifact.
A borrowed dataset cannot offer this check, because a borrowed dataset has no known truth behind it. This is the reason to simulate the primary data.
One caution applies to this table. The project retired its claim that the fit was blind to the injected values. The harness could not prove that the fitting agent never saw the file. Treat this table as evidence that the pipeline works. Do not treat this table as blind validation.
Figures
Limitations
These limitations are part of the result. They are not a footnote.
- The primary data is simulated. A known model generates the failure times. No physical component was tested. The value of the simulation is the known truth, which lets you measure the estimators.
- The bounds are bootstrap percentiles. They come from 2,000 replicates. They are not likelihood-ratio bounds. The two methods would not agree exactly. A profile-likelihood bound on the same likelihood gives R = 0.708 against the bootstrap's 0.7771. Both bounds miss the criterion, so the verdict does not change. But the coverage of the percentile bootstrap is not established for this problem.
- The extrapolation is long, and the extrapolation carries the result. The test cells are at 85, 105, and 125 C. The use condition is 40 C. A small error in the activation energy compounds over that distance. The bound width at the use condition is 107% of the point estimate. At the test cells the bound width is 19% to 36%. This is the load-bearing assumption of the whole study.
- A common mechanism is supported by a non-rejection. The likelihood ratio test does not reject the common shape at p = 0.3286. A non-rejection is not proof. With 20 units per cell the test has limited power.
- One acceleration variable only. The test raises temperature. Real qualification often needs humidity or voltage as well. This is a scope choice.
- Weibull is selected on physical evidence, not on AIC. Lognormal has the lower AIC on the pooled data and in two of the three cells. The per-cell shapes, the common-shape test, and the physics of a thermally activated wear-out mechanism support Weibull. The qualification verdict does not change either way.