MyQCLab LogoMyQCLab
← Back to Blog
monte-carlo-simulationpower-functionprobability-error-detectionfalse-rejection-ratesigma-metricswestgard-rulesinstrumen-mutu-lanjutanrisk-based-QC

Monte Carlo Simulation: Testing the Reliability of Your QC Rules Before Problems Occur

MyQCLabSeptember 27, 20264 views
Monte Carlo Simulation: Testing the Reliability of Your QC Rules Before Problems Occur

Learn how Monte Carlo simulation tests the reliability of your Westgard rules in detecting errors, complete with a power function calculator.

Category: Advanced Quality Instruments | Estimated reading time: 11 minutes

The Rarely Asked Questions

All the methods discussed previously — Westgard Rules, CUSUM, Moving Average, EWMA — focus on one thing: detecting problems that have already occurred (or are starting to occur) in QC data. But there is another question that is rarely asked explicitly: how reliable is the QC system I am actually using right now?

If a 1 SD error appears in my method, what is the percentage probability that the Westgard rules I use will catch it? What if the error is 2 SD? 3 SD? This is not a question that can be answered by looking at daily QC data — this is a question about the design of the QC system itself, and this is where Monte Carlo Simulation (MCS) plays a role.

Definition

Monte Carlo Simulation is a computational statistical method that uses random sampling to simulate various possible outcomes of a process or system. In the laboratory context, MCS is used to predict variations in test results based on the statistical distribution of parameters that affect results — such as precision (CV), bias, and target values — then testing how the QC system responds to those various scenarios.

A Brief History

This method was developed in the 1940s by scientists at the Manhattan Project in Los Alamos — including John von Neumann and Stanislaw Ulam — to solve neutron diffusion problems in fissile materials that were too complex to calculate analytically. The results were formally published by Nicholas Metropolis and Stanislaw Ulam in the paper "The Monte Carlo Method" in the Journal of the American Statistical Association in 1949. The name "Monte Carlo" was inspired by a casino in Monaco, supposedly referring to Ulam's uncle's penchant for gambling there — a name chosen by Metropolis because it fit the approach based on probability and randomization.

In the medical laboratory context, the same principle has been applied by Westgard since the early 1980s to design and test the performance of QC rules — including the simulations that underlie the classic Westgard multirule paper (1981). Decades later, Curtis Parvin (2008) published a study that is considered a breakthrough in connecting QC frequency with patient risk using a similar simulation approach — work that remains a primary reference in the design of modern risk-based statistical quality control.

Reasons for Use

Monte Carlo Simulation is used to test the uncertainty and variability of a laboratory system or process. This method is very useful when direct physical testing is difficult, expensive, or time-consuming — imagine if you wanted to know "how often this QC rule will fail to detect a problem", the only physical way to do so would be to wait for the problem to actually occur many times in the real world, which is clearly impractical (and dangerous for patients). In the context of laboratory quality, MCS is used to predict the distribution of Z-scores, Total Error, or the QC rule performance of a test parameter — without having to wait for real events.

Advantages

  • Provides a realistic picture of the distribution of results and risk of error.
  • Can model complex and non-linear conditions — such as combinations of several Westgard rules at once, which are difficult to calculate analytically because the rules are correlated.
  • Suitable for evaluating the performance of QC systems and analytical methods before they are actually implemented.

Disadvantages

  • Requires an understanding of statistics and programming/simulation skills to build from scratch (though for users, it is enough to understand how to read the results).
  • The results are probabilistic, not deterministic — there is no absolute guarantee, only an estimation of probability.
  • Dependent on input quality (CV, bias, target) — garbage in, garbage out: if the CV/bias numbers entered are not representative of actual laboratory conditions, the simulation results will be misleading.

Required Data

  • CV (Coefficient of Variation) of the test method — process precision, from preliminary testing or monthly evaluations.
  • Bias from the reference value — process accuracy, from EQA (PME) results/monthly evaluations.
  • N — the number of control levels measured per run.
  • Combination of QC rules for which you want to test performance.
  • Number of simulation iterations — the higher, the more precise the result (this calculator uses 1,000 iterations per point, which is precise enough for illustrative purposes and relative comparisons).
Note that the Target Value does not need to be explicitly input. Because CV and Bias are already expressed as percentages relative to the target, the simulation can be run entirely in Z-score (SD unit) without affecting the final result — a valid technical simplification, not missing data.

How to Use the Tools

  1. Enter the CV and Bias of your method — CV measures precision, Bias measures the accuracy of the current existing method.
  2. Enter N — the number of control levels measured each run (typically 2).
  3. Select the combination of QC rules you want to test — click the badge to enable/disable. Combinations are evaluated using OR logic (a rejection occurs if any one rule is triggered), in accordance with the actual operation of Westgard Multirules.
  4. Click "Run Simulation" — the computer will generate thousands of random scenarios for various error sizes (from 0 to 4 SD), and calculate the percentage successfully detected by the selected rule combination.
  5. Read the results from the three statistical cards and the power function curve — a full explanation of how to read them is in the next section.
Important note: The bias input here is the bias that already exists in your current method (the existing condition). The simulation tests how much additional error on top of that condition is needed before the QC rules actually detect it — it does not assume the method is in a perfect state without any bias at all.

Try the Monte Carlo Calculator

[tool:monte_carlo]

How to Read Simulation Results

  1. There are three key numbers and one curve to understand:
  2. Pfr (Probability of False Rejection) — target: as low as possible. This is the percentage of "false alarms" at shift = 0 (no additional error, only the existing bias). Ideally below 5%. If Pfr is high (e.g., 15-20%), it means the chosen rule combination is too sensitive for the current method conditions — QC will be rejected frequently without any reason that is truly clinically significant, wasting time and reagents, and risking normalizing a culture of "oh, rejected again, probably just a false alarm" as discussed in the article on QC Reject.
  3. Ped (Probability of Error Detection) at a specific shift (e.g., 2 SD) — the higher the better. This number answers: "if there is indeed a 2 SD error, what is the percentage probability this rule combination will catch it?" There is no universal standard number for "how much Ped is enough" at a specific shift point — this value is more meaningful when viewed as part of the curve as a whole, not as a single standalone number.
  4. Power Function Curve — the shape is the most important thing to read, not just one point. A curve that rises quickly from the left (Ped is already high at a small shift) shows a rule combination that is sensitive to detecting even small problems. A curve that is flat and only rises at large shifts means the rule only "realizes" when the problem is already quite severe — which is less ideal for patient safety, as it means there is a range of medium errors that could go undetected for quite some time.
  5. "Ped reaches 90%" point — practical reference from literature. 90% is the Ped target commonly used in SQC design literature (including the Parvin framework) as a standard for being "sufficiently reliable" to detect clinically significant errors. The SD number at this point answers: "how big of a problem must occur before I am sufficiently confident (90%) that this QC system will catch it?" — the smaller the SD number, the more sensitive and better the rule combination.

Monte Carlo Simulator vs. Westgard Sigma Nomogram

  1. Readers already familiar with Sigma Metrics articles may realize: there is a close relationship between this simulation and the rule recommendation table based on Sigma categories ("Sigma ≥6 is sufficient to use 1₃s only", etc.). Both indeed come from the same roots — the Westgard Sigma nomogram is essentially a summary of similar simulation results, condensed into a quick reference table so that you do not need to re-simulate every time. But there are important differences between the two:
  2. AspectWestgard Sigma NomogramMonte Carlo SimulatorNatureGeneral table/graph, ready-made — quick referenceCalculated specifically for your method's CV, Bias, and NOutputCategory (World Class–Poor) + standard rule recommendationPed vs. shift curve + precise Pfr numberRule combination flexibilityGeneral recommendations (e.g., "1₃s only" or "full multirule")Can test any rule combination, including non-standard onesUsage SpeedInstant — just look at the tableRequires running a simulation (still instant on modern browsers)
  3. Both complement each other, rather than replacing one another. The nomogram is suitable for quick checks and rough estimates at the beginning. The Monte Carlo Simulator is suitable when you need a more precise and specific answer — especially for critical parameters in your laboratory that deserve more detailed attention than simply following a general table.

Suitable for Use When

  • Assessing the risk of test method failure in the long term.
  • Developing or evaluating a laboratory quality control system — especially when considering changing QC rule combinations for specific parameters.
  • Creating data-driven plans for resource allocation (e.g., determining if QC frequency needs to be increased) or changes in analytical methods.
  • Validating whether a long-used QC rule combination is still in line with current method performance — especially after reagent changes, instrument changes, or after Sigma metric evaluations show significant performance changes.

Conclusion

  1. Monte Carlo Simulation is a very powerful statistical tool in laboratory quality management, because it is able to model various scenarios and test result uncertainties realistically — without having to wait for real problems to occur repeatedly in the real world. By simulating thousands of possible results based on precision and accuracy data, laboratories can predict method performance, determine the probability of error detection, and make decisions about the design of a QC system that is truly risk-based — not just following habits or general tables without understanding the specific context of the method being run.
  2. Although it requires an understanding of statistics and supporting software to build, its benefits in strengthening quality control and validating QC design are significant for modern laboratories — especially as a complement to the more widely known Sigma nomograms, providing more precise and specific answers for your own laboratory's conditions.

References

  1. 1.Metropolis N, Ulam S. The Monte Carlo method. J Am Stat Assoc. 1949;44(247):335–341. https://doi.org/10.1080/01621459.1949.10483310
  2. 2.Westgard JO, Barry PL, Hunt MR, Groth T. A multi-rule Shewhart chart for quality control in clinical chemistry. Clin Chem. 1981;27(3):493–501.
  3. 3.Parvin CA. Assessing the impact of the frequency of quality control testing on the quality of reported patient results. Clin Chem. 2008;54(12):2049–2054. https://doi.org/10.1373/clinchem.2008.113639

© 2026 MyQCLab. All rights reserved.