TB-3174 · REV D · Technical newsheet
Metrology & CalibrationDevice profile
Simulation Study Validates Wilson Interval for Calibration QA Sampling
A 3,000-month simulation shows the Wald 95% interval covers the true rate only 73% of the time at 1% error rates, while the Wilson interval holds 96.5–97%. Four checks before trusting any sampling QA plan.
By Grace Kim4 min read815 words
Features
- Wald 95% intervals achieved only 73% actual coverage (2,190 of 3,000 simulated months) at a 1% error rate; Wilson score intervals achieved 96.5–97.0% at all tested rates.
- Per-technician flagging requires roughly 20 cumulative reviews (~2 months); lab-wide confidence intervals need ~6 months of accumulated data to detect a 0.5-point rate shift.
- For 2 errors in 170 reviews, the Wald formula yields [−0.44%, 2.80%] with silent zero truncation; Wilson yields [0.32%, 4.19%] with no truncation.
A 3,000-simulated-month study of confidence-interval methods at calibration-QA error rates found that the textbook Wald interval delivers 73% actual coverage when it claims 95% — while the Wilson score interval delivers 96.5% to 97.0% across every tested error rate.
The study is the third article in a trilogy by Greg Cenker on statistical QA for calibration laboratories. The first article argued that 100% review of outgoing calibration work is arithmetically impossible at lab scale. The second proposed risk-weighted stratified sampling covering roughly 10% of work. This one asks the question an auditor will eventually ask: does the confidence interval behind the sampling plan actually contain the true error rate as often as its label promises?
That property — coverage — is an empirical question, not a formulaic one. The validation method is standard practice from decades of survey research: generate synthetic months of calibration data at a known true error rate, compute the interval under test for each month, and count how often the interval contains the truth. The study varied the ground-truth error rate across 1%, 3%, 5%, and 10%, with monthly review volume, technician mix, and error-rate range mirroring a mid-sized calibration laboratory.
At a 1% true error rate, the Wald method's 95% intervals contained the true rate in only 2,190 of 3,000 simulated months — 73% coverage. The Wilson score interval held 96.5% to 97.0% coverage at all settings, slightly conservative relative to nominal, which the author considers the correct bias for an auditor-facing metric: an interval that slightly overstates uncertainty beats one that understates it.
Why the Wald interval fails here
The Wald construction assumes the observed error count behaves like a draw from a smooth, symmetric bell curve. That approximation holds with roughly 30 or more observed errors in a sample. Calibration QA does not operate there. At a 1% error rate and monthly review samples of roughly 170 items, the expected error count per month is under two; observed counts of zero, one, or two are routine. The Wald interval must be symmetric around the observed rate, so at low rates its lower bound falls below zero and gets silently truncated — clipping off probability mass without any flag in the reporting. For the representative case of 2 errors in 170 reviews (observed rate 1.176%), the Wald formula produces [−0.44%, 2.80%], truncated at zero before reporting. The Wilson formula produces [0.32%, 4.19%] with no truncation required, because it inverts the question: it asks which true rates are statistically consistent with the observed data, respecting the asymmetric shape of the count distribution.
Two warm-up timescales
Even a validated interval needs accumulation time. Per-technician flagging matures at roughly 20 cumulative reviews — about two months — a threshold that follows from binomial precision and Bayesian posterior convergence. Below it, a HIGH-RISK flag is, in the author's words, "a guess dressed up in statistical language." Laboratory-wide intervals mature at roughly six months, because sampling precision narrows with the square root of accumulated sample size; detecting a genuine 0.5-percentage-point shift in the underlying error rate requires that much data to separate it from month-to-month noise. Month-over-month comparisons are defensible from month one. Year-over-year trend analysis needs the longer horizon.
The author's testing position is blunt: a vendor claiming reliable per-technician flags in week one, or rate-shift detection after a month, "is either misunderstanding the statistics or misrepresenting them."
Four verification questions
For any sampling-based QA approach, the article prescribes four checks. Ask for the coverage study — an unsupported "95% confidence" claim is an assertion, not a claim. Ask for the audit log format: a defensible plan records the pseudo-random seed per daily selection, allocation per technician, and timestamped decision points, so an auditor can re-run any day's selection bit-for-bit. Ask which published methods underlie the approach — Wilson, Cochran, Agresti-Coull, and Kish are the names that should appear in the documentation. Ask how the tool handles a cold start; a HIGH-RISK flag before roughly 20 cumulative reviews signals the tool ignores the math.
The trilogy leaves one problem open: data staleness after roughly 18 months, when a technician's risk score stays anchored to a historical error rate that no longer reflects current performance. The author is developing a second series on recalibrating the plan to a lab's evolving reality without sacrificing reproducibility. The referenced literature spans Wilson (1927), Agresti and Coull (1998), Brown, Cai, and DasGupta (2001), and ISO/IEC 17025:2017 clause 7.7.
For labs accredited under ISO/IEC 17025, the question this work raises is direct: if your QA metrics quote 95% confidence intervals computed by default statistical software, what does your coverage study say that number actually means?
via bnpmedia.com (Original)
Filed under
- wilson-interval
- calibration-qa
- iso-iec-17025
- statistical-sampling
- confidence-interval-coverage
More from Grace Kim
Show full bio
Correspondent covering consumer brands and retail at Testbench Report.
21 articles