
An EQA report is only useful if someone reads it properly and acts on it. This page explains the scores in a typical report, why the peer group matters, how to see a problem coming before a result fails, and what to do when one does.
Most reports for quantitative tests show the laboratory’s result against a target value and express the difference in one or more scores. The common ones:
| Score | What it measures | How to read it |
|---|---|---|
| z-score | Difference from the assigned value in units of the standard deviation for proficiency assessment | |z| ≤ 2 satisfactory, 2 < |z| < 3 questionable, |z| ≥ 3 unsatisfactory |
| SDI (standard deviation index) | Difference from the peer-group mean in units of the peer-group SD | Close to 0 is good; beyond ±2 needs an investigation |
| Deviation in % | Difference from the target in percent | Compared with fixed acceptance limits (for example CLIA limits or national rules) |
| CVR (CV ratio) | The laboratory’s CV divided by the peer-group CV | Above 1.5 points to higher imprecision than comparable laboratories |
| Grade or score per round | Share of acceptable results across all analytes of a distribution | Shows the overall picture; the single analytes still need review |
z-score limits follow ISO 13528:2022, the standard for statistical methods in proficiency testing; providers may use their own limits and name them in the report.
Reading a z-score
Limits after ISO 13528: |z| ≤ 2 satisfactory, 2 < |z| < 3 questionable, |z| ≥ 3 unsatisfactory
Because many methods do not agree with each other, and a comparison across all methods would flag a laboratory for using a different platform rather than for an error. Peer groups compare results only with laboratories using the same method, analyser and often reagent. That is essential for immunoassays and PT/INR, and useful almost everywhere else.
Peer grouping has one limit: if a whole method is biased, the peer group shares the bias and every member looks fine. Accuracy-based surveys with commutable samples and reference-method targets close that gap and show trueness, not only agreement.
By reading the reports as a series rather than one by one. A single result outside the limits can be chance; a bias that grows in one direction over three rounds is a method moving away from the truth, even while every result is still acceptable.
A bias that grows over rounds
Illustrative z-scores of one analyte across eight EQA rounds
Investigate before correcting anything, and always ask whether patient results were affected. A proven sequence:
Most unacceptable results fall into a small number of categories, and the category decides the corrective action. Classifying every case also shows patterns over time.
| Category | Examples | Typical action |
|---|---|---|
| Clerical | Transcription error, wrong unit, wrong method code, results for two samples swapped | Second-person check or electronic transfer from the LIS |
| Sample handling | Wrong reconstitution volume, sample analysed too late, wrong storage temperature | Work instruction for EQA samples, training |
| Method | Calibration bias, reagent lot, method not fit for the concentration range | Recalibration, lot verification, method review |
| Equipment | Maintenance overdue, worn component, temperature fault | Service, maintenance plan, instrument log review |
| EQA material or evaluation | Non-commutable sample, wrong or too small peer group, wrong target | Query to the provider, change of peer group or scheme |
| No cause found | Investigation complete without a finding | Document the investigation, increase monitoring for the next rounds |
From an unacceptable result to a closed CAPA
Not by scores but by agreement with the expected answer, often weighted by clinical relevance. The criteria differ by discipline:
Everything an assessor needs to follow a sample from arrival to the closed CAPA. In practice that is a short, repeatable set of records: