A 501(c)(3) non-profit organization Applied AI research for public benefit
CheAI Research Inc.

The component does not qualify

An accelerated life test of 60 units. The point estimate of reliability clears the criterion. The 90% lower confidence bound does not. Qualification turns on the bound.

The primary data is simulated. A known Weibull and Arrhenius model generates the failure times. The injected values are beta = 2.5, Ea = 0.70 eV, and eta = 8000 h at the use condition.

This is a deliberate choice, and it is a strength here. You know the true answer, so you can measure what the estimators recover. A borrowed dataset cannot offer that check. The estimators are also validated against a published external dataset.

The numbers on this page are computed. They are read from the qualification artifact at build time.

The engineering problem

A component must survive a 3000 hour mission in the field. The use condition is 40 C.

A test at the use condition would take years. So the test applies raised temperature instead. Raised temperature makes the component wear out faster.

The test runs at 85, 105, and 125 C. An Arrhenius-Weibull model fits the failure times. The model then extrapolates the life back down to 40 C.

The acceptance criterion is fixed before the test. Reliability must be at least 0.90 at 3000 hours. You must demonstrate this on the 90% lower confidence bound, not on the point estimate.

Qualification result

NOT DEMONSTRATED

The point estimate and the lower bound fall on opposite sides of the criterion. The decision follows the bound. So the component does not qualify.

Point estimate, R(3000 h) 0.9165 Clears the 0.90 criterion.
90% lower bound, R(3000 h) 0.7771 Short of the 0.90 criterion by 0.1229, or 13.7%.
Point estimate, B10 life 3229 h Clears the 3000 h mission.
90% lower bound, B10 life 2125 h Short of the 3000 h mission by 875 h, or 29.2%.

B10 life is the time by which 10% of units fail. B10 agrees with the reliability result, and B10 is an independent view of the same fit. Both point estimates clear their target. Both lower bounds miss their target.

Why the verdict says NOT DEMONSTRATED and not FAIL

The wording is deliberate. Qualification is a demonstration.

The requirement is to SHOW that the lower bound clears the criterion. This evidence does not show that. So the correct report is NOT DEMONSTRATED.

FAIL would claim more than the evidence supports. FAIL says the component is bad. The data does not say that. The point estimate of R(3000 h) is 0.9165, and the true reliability could be above the criterion.

What the data does say is narrower. At 90% confidence, this test does not establish the claim. The component does not qualify on this evidence.

This is a clear miss, not a close call

A reader can reasonably ask if the shortfall is only noise in the bootstrap. It is not.

The bootstrap bound varies by about 0.013 across seeds. That variation is roughly 10% of the gap to the criterion. Every seed puts the lower bound on the same side of the criterion.

B10 gives the same answer, and B10 misses by a wider relative margin. The B10 lower bound is 2125 h against a 3000 h mission.

The two metrics agree, and the seeds agree. So the verdict does not depend on the seed.

The qualification matrix

2 of 7 requirements are not demonstrated. 5 requirements pass. Every row is read from the qualification artifact.

Requirement Method Stress Sample Criterion Result Verdict
R >= 0.90 at 3000 h, 40 CArrhenius-Weibull ALT, censored MLE, bootstrap bounds85 / 105 / 125 C, extrapolated to 40 C60 (18 censored)90% lower bound >= 0.90lower bound 0.7771 (point 0.9165); short of the criterion by 0.1229 (13.7%)NOT DEMONSTRATED
B10 life >= mission time (3000 h)Weibull B10 from the extrapolated use-condition fitextrapolated to 40 C60 (18 censored)90% lower bound >= 3000 hlower bound 2125 h (point 3229 h); short of the 3000 h mission by 875 h (29.2%)NOT DEMONSTRATED
Failure mechanism unchanged across stress levelsLikelihood ratio test, common vs free shape, scales free85 / 105 / 125 C60 (18 censored)p >= 0.05 (do not reject common shape)p = 0.3286PASS
Activation energy physically plausibleEa fitted from the ALT model, Ea = a * k85 / 105 / 125 C60 (18 censored)0.3 to 1.2 eVEa = 0.7017 eV, 90% CI [0.6204, 0.7795]PASS
Wear-out mechanism identified for maintenance policyWeibull shape parameter from censored MLEall cells, common shape60 (18 censored)beta > 1 indicates wear-outbeta = 2.576, 90% CI [2.128, 3.283]PASS
Estimators validated against an external referenceCensored Weibull MLE and KM vs published datasetn/a (NCCTG lung, Loprinzi et al. 1994)228 (63 censored)two independent implementations agreeshape agrees to 1.5e-07, scale to 4.4e-06PASS
Censored units retained in every fitLikelihood contribution S(t) for censored unitsall cells18 censored of 60no censored unit discardedall 18 retained; discarding the pooled data understates eta by 30.5%PASS

The two rows that are not demonstrated are the reliability claim and its B10 equivalent. Both turn on the lower confidence bound. Both point estimates clear their target.

The model passes its own validity checks

These checks support the MODEL. These checks do not support the PRODUCT. The distinction matters. A sound model that reports a bound below the criterion is still a component that does not qualify.

Common shape, LRT p 0.3286
Activation energy 0.7017 eV
Weibull shape, beta 2.576
Char. life at 40 C 7736 h

Common shape across the stress cells

An Arrhenius model assumes one failure mechanism at every temperature. If the Weibull shape changes with temperature, the mechanism changed. Then the extrapolation is not valid.

The likelihood ratio test gives p = 0.3286. The test does not reject the common shape.

Read that result carefully. The test does not reject, and the test does not prove. With 20 units per cell the test has limited power.

Activation energy

The fitted activation energy is 0.7017 eV. The physically plausible band is 0.3 to 1.2 eV for a single thermally activated mechanism. The fitted value sits inside that band.

Wear-out

The Weibull shape is beta = 2.576. A beta above 1 means the hazard rate rises with time, which identifies wear-out.

This result has a direct maintenance consequence. The correct mitigation is scheduled replacement before the wear-out knee. Burn-in is the wrong mitigation here. A beta below 1 would mean infant mortality, and then burn-in would be correct.

Estimators validated against a published dataset

Two independent implementations fit the same published dataset. The dataset is the NCCTG lung dataset, from Loprinzi et al. 1994, with 228 units, 63 of them censored.

Shape agrees to 1.5e-07, scale to 4.4e-06. This checks the estimator against an outside reference, rather than against its own output.

Censoring

A censored unit is a unit that did not fail before the test stopped. You know it survived to a certain time, and you do not know when it fails.

This test has 18 censored units of 60, which is 30% of the sample. Every censored unit is retained in every production fit. Each censored unit contributes its survival function to the likelihood.

The project measured what the alternative costs. When you discard the pooled censored data, the fit understates eta by 30.5%.

The bias runs in one direction for eta. The bias does not run in one direction for every derived metric. When you discard censored units, you also corrupt the Weibull shape, so a derived metric can move either way.

What the estimators recover from the injected truth

The data comes from a known model. So you can compare the fit against the values that generated the data.

Parameter Injected truth Fitted Difference
Weibull shape, beta2.5002.576+3.0%
Activation energy, Ea (eV)0.70000.7017+0.2%
Characteristic life at use (h)80007736-3.3%

Injected values read from the recorded ground truth. Fitted values read from the qualification artifact.

A borrowed dataset cannot offer this check, because a borrowed dataset has no known truth behind it. This is the reason to simulate the primary data.

One caution applies to this table. The project retired its claim that the fit was blind to the injected values. The harness could not prove that the fitting agent never saw the file. Treat this table as evidence that the pipeline works. Do not treat this table as blind validation.

Figures

The Arrhenius fit and the extrapolation
The Arrhenius fit and the extrapolation. Log life against reciprocal temperature. The three test cells sit on the right. The use condition sits far to the left. Look at how far the fitted line must travel past the last data point to reach the use condition. Look also at how the bootstrap bounds open out over that distance. That widening is why the lower bound misses the criterion.
Weibull probability plot by cell
Weibull probability plot by cell. Each cell plots as a near-straight line, which supports the Weibull shape. The plotting positions are adjusted for censoring. The three lines run roughly parallel, and parallel lines indicate a common shape across the cells.
Kaplan-Meier survival by cell
Kaplan-Meier survival by cell. A distribution-free view of the same data, with censored units marked as ticks. This curve assumes no life distribution at all. Compare it against the fitted curves to check that the parametric model does not contradict the raw survival.
Weibull hazard rate by cell
Weibull hazard rate by cell. The hazard rate rises with time in every cell, because beta exceeds 1 in every cell. This is the wear-out arm of the bathtub curve. This data does not evidence an infant-mortality region or a constant-rate region.
Weibull against lognormal and exponential
Weibull against lognormal and exponential. The three candidate distributions fitted to the same data. They are hard to separate by eye, and AIC does not separate them decisively either. Lognormal has the lower AIC on the pooled data. Weibull is selected on the physical evidence instead, as the limitations state.
What discarding censored units would cost
What discarding censored units would cost. The same data fitted twice. The left bar of each pair handles censoring. The right bar discards the censored units. Characteristic life falls in every group when you discard, and it falls most where the censored fraction is highest.
Bound width against extrapolation distance
Bound width against extrapolation distance. The confidence bound is narrow at the test cells and wide at the use condition. This figure is the mechanism behind the verdict. The component is not shown to be unreliable. The evidence is shown to be too thin at the use condition to demonstrate the claim.

Limitations

These limitations are part of the result. They are not a footnote.