AI can reliably restate a diagnostic report in language a patient can read, and the published evidence supports that. It cannot safely decide whether a finding is serious, what should happen next, or what a result means for a particular person. The distinction is not a marketing nuance. It is the line between a communication tool and a clinical decision tool, and it determines what oversight your organisation must keep in place.
Clinicians are right to be sceptical of AI in patient-facing settings, and vendors have earned that scepticism. This piece sets out what the evidence actually shows, including the failure rate, and what a provider should require before putting any such system in front of patients.
What the evidence shows
A systematic review and meta-analysis in The Lancet Digital Health examined 38 studies published between 2022 and 2025, covering 12,922 simplified reports evaluated by 508 assessors. It found that large language model simplification improved patient understanding and readability while largely preserving accuracy and completeness. In practical terms, readability shifted from university-level text to roughly school-level, around ages 11 to 13, for CT and X-ray reports.
That is a meaningful finding. Reports are typically written far above the reading level of the average adult, so a shift of that size is the difference between a document a patient can attempt and one they cannot.
The error rate, stated plainly
The same meta-analysis reported an error rate of 7.2% in LLM-rewritten reports, with clinically significant errors at 0.9%. Both figures deserve to be quoted by anyone selling or buying this technology, and most vendors quote neither.
Nought point nine percent is small. It is not zero. At meaningful volume it is not a rounding error either, which is precisely why the authors concluded that human oversight remains necessary. The correct response to that number is not to dismiss the technology, nor to pretend the number does not exist, but to design around it: keep the original clinical report authoritative and unaltered, position the simplified version as an accompaniment rather than a replacement, and preserve a clear route to a clinician.
| Task | Suitable for AI assistance? | Why |
|---|---|---|
| Restating clinical terms in plain language | Yes, with oversight | High volume, repeatable, evidence supports readability gains |
| Explaining what a test measures | Yes | Educational content, not patient-specific judgement |
| Multilingual delivery of the above | Yes, with per-language review | Extends access, but mistranslated hedges are clinical errors |
| Deciding whether a finding is urgent | No | Clinical judgement requiring full context and accountability |
| Recommending next steps or treatment | No | Clinical decision-making |
| Detecting disease or offering a diagnosis | No | A different product category with different regulatory obligations |
Simplification is not diagnosis
The boundary is easy to state and easy to erode. A simplification layer takes content that already exists in the report and expresses it differently. It adds no new clinical assertion. The moment a system tells a patient that a finding is probably harmless, or that they should be seen urgently, it has made a clinical claim the report did not make, and it has become something other than a communication tool.
If a system says anything the report did not already say, it is no longer simplifying. It is diagnosing, and it should be evaluated as a clinical device rather than a communication tool.
What this means for your centre
The operational case does not rest on the technology being impressive. It rests on the fact that patients now receive reports before anyone explains them, and that a large share of the resulting questions are vocabulary questions rather than clinical ones. Those are exactly the tasks the evidence supports automating, under oversight. The clinical questions stay where they belong.
Six questions to ask any vendor
- What does your system do when a patient asks whether their result is dangerous? The correct answer is that it directs them to a clinician.
- What is your measured error rate, how was it measured, and by whom?
- Does the original clinical report remain unaltered and authoritative?
- What clinical oversight exists, and what happens when the underlying model version changes?
- Is patient data used for model training, and can that be excluded contractually?
- What regulatory claims are you making, and can you evidence each one? Be sceptical of implied approvals.
A vendor who answers these plainly, including where the answer is unflattering, is more trustworthy than one with a more impressive pitch. That applies to us as much as anyone.



