A 501(c)(3) non-profit organization Applied AI research for public benefit
CheAI Research Inc.

Can the measurement system tell good parts from bad data?

A demonstration of the Gage R&R method on simulated data. It shows how measurement variation can obscure real part differences and distort every capability number calculated downstream.

ANOVA (crossed design) Canonical seed 20260723 Published January 8, 2022

Separate part variation from measurement variation.

Before you trust any process data, ask how much of the observed variation is the process. Then ask how much of it is the measurement system itself. If the gage cannot discriminate between parts, every capability number computed from it is fiction.

A Gage R&R study divides the total measurement variation into three parts. Part-to-part variation is the variation you want to see. Repeatability is the variation of one operator who measures the same part again. Reproducibility is the difference between operators. The ANOVA method also tests for a part-by-operator interaction. An interaction shows that some operators measure certain parts differently.

One measurement system passes. One fails for a clear reason.

PASS case | %GRR 7.16%

ndc 19 | Acceptable

FAIL case | %GRR 43.55%

ndc 2 | Unacceptable

MetricPASS caseFAIL case
%GRR, study variation7.16%43.55%
ndc192
VerdictAcceptableUnacceptable
Primary causeMeasurement system acceptableReproducibility, 95.9% of GRR

The failing gage drags apparent Cpk to the acceptance boundary.

1.5000 True process Cpk
1.3337 Measured through FAIL gage

At canonical seed 20260723, the observed result lands at the 1.33 boundary, not below it. The reviewed population expectation from the injected components is 1.2815, or about 1.28. The observed Cpk loss is 11.08%.

ScenarioApparent total sd (mm)Apparent CpkCpk loss
True process (no measurement error)1.00001.50000.00%
Measured through the PASS gage1.00961.48570.95%
Measured through the FAIL gage1.12471.333711.08%

The same diagnostic views for both systems.

Each chart is embedded directly from the generated PNG artifact.

PASS case

Acceptable
Components of variation for the PASS case
Components of variation Compares contribution, study variation, and tolerance.
R chart by operator for the PASS case
R chart by operator Checks within-operator repeatability against control limits.
X-bar chart by operator for the PASS case
X-bar chart by operator Points falling OUTSIDE the control limits are the DESIRED outcome. They show that the gage can distinguish parts.
Measurement by part for the PASS case
Measurement by part Shows individual readings and part means across the study.
Measurement by operator for the PASS case
Measurement by operator Shows operator location shifts and overall measurement spread.
Part-by-operator interaction for the PASS case
Part-by-operator interaction Shows whether operator profiles remain broadly parallel across parts.

FAIL case

Unacceptable
Components of variation for the FAIL case
Components of variation Compares contribution, study variation, and tolerance.
R chart by operator for the FAIL case
R chart by operator Checks within-operator repeatability against control limits.
X-bar chart by operator for the FAIL case
X-bar chart by operator Points falling OUTSIDE the control limits are the DESIRED outcome. They show that the gage can distinguish parts.
Measurement by part for the FAIL case
Measurement by part Shows individual readings and part means across the study.
Measurement by operator for the FAIL case
Measurement by operator Shows operator location shifts and overall measurement spread.
Part-by-operator interaction for the FAIL case
Part-by-operator interaction Shows whether operator profiles remain broadly parallel across parts.

A point estimate is not a confidence bound.

With 3 operators, reproducibility is estimated with only 2 degrees of freedom. A single study cannot pin %GRR precisely, even when the verdict remains robust for a clearly failing case.

48%Mean %GRR
48%Median %GRR
18 percentage pointsStandard deviation
17.8% to 77.5%5th to 95th percentile

The point estimate can swing about 20 percentage points from run to run. The canonical FAIL verdict is still robust because the designed case is far from the marginal band.

A crossed ANOVA study.

Study design

ParameterValue
DesignCrossed (every operator measures every part)
Parts10, spanning the process range (part sd = 1.0 mm)
Operators3
Replicates3 per operator per part
Total measurements90
MethodANOVA with part-by-operator interaction
Interaction testF-test at alpha = 0.25 (AIAG convention)

Decision thresholds

%GRR, study variationVerdict
Under 10%Acceptable
10% to 30%Marginal
Over 30%Unacceptable

Every dataset, number, chart, and test result on this page is generated from the frozen plan and the canonical seed.

This validates the method, not a physical gage.

  • Simulated primary data: This validates the method against known truth. It is not a substitute for a real gage study on physical hardware.
  • Crossed design only: Every operator measures every part (non-destructive). Nested or destructive designs require a different ANOVA structure and are out of scope.
  • No time-dependent effects: The simulation does not capture day-to-day or setup-to-setup variation, which in real studies can exceed short-term repeatability.
  • The two cases are deliberately unambiguous: Both sit far from the 10-30% marginal band. The method therefore shows clearly. This is a teaching demonstration, not a claim about a specific real gage. Real gages often fall in the marginal band where the decision depends on application context.