ACR–SIIM Practice Parameter · Adopted 5 May 2026
How many cases does your acceptance test need?
The practice parameter requires local acceptance testing before an imaging AI tool goes live. It specifies the process and leaves the sample sizes to you. Fifty cases can only detect a drop of about eight points — which is worth knowing before you collect them.
Counts only · No images or patient data · No login
Three controls, assessed before you collect a single case
The test set must span multiple manufacturers, models and protocols. We check the spread and flag where a manufacturer has too few cases to support a performance claim.
Dose-optimised, ultra-low-dose and AI-postprocessed images must be represented, because the parameter notes these may alter performance.
Metrics must be agreed before evaluation begins. Every assessment produces a dated record you can sign and file as that evidence.
Most facilities size an acceptance test by pulling fifty to a hundred past scans and seeing whether the AI agrees with the report.
For a tool claiming 95% sensitivity, fifty cases can only reliably detect a drop of about eight points. A five point drop — the kind that matters clinically and would not be visible by eye — passes unnoticed. Detecting it would take around 118 positive cases.
A record for the governance file
The output is a dated, pre-specified metric record: the tool, its claimed performance, the margin you set, the case counts by manufacturer, what the test can and cannot detect, and each control assessed against the parameter. Copy it into your quality system or print it as a memo.
Common questions
What does the practice parameter require for local acceptance testing?
That qualified end-users evaluate initial performance on local data before go-live, that the test set spans multiple scanner manufacturers, models and protocols, that it includes dose-optimised and AI-postprocessed images where applicable, and that acceptance metrics are agreed before evaluation begins. It does not specify sample sizes or cut-offs — those are left to the facility.
How many cases do we actually need?
It depends on the tool's claimed performance and the drop you would consider clinically meaningful. For a tool claiming 95% sensitivity, detecting a 10 point drop needs around 30 positive cases; detecting a 5 point drop needs around 118. A set of 50 cases can only reliably detect a drop of about 8 points or larger.
Does this replace our medical physicist?
No. It sizes the test and produces the pre-specified metric record. Judging whether a tool is fit for local deployment rests with the governance group and the qualified medical physicist — the parameter is explicit about that.
Do we have to upload images or patient data?
No. It takes counts and claimed performance figures only. Nothing leaves your browser except those numbers, and nothing is stored.
