Radiology AI

AI for Patient Communication in Diagnostics: What It Can and Cannot Do

A meta-analysis of 38 studies found AI-simplified radiology reports improve readability while largely preserving accuracy, with a measured error rate. Here is what that means for providers, honestly.

FR
The FlexReport Team
August 5, 20263 min read
AI for Patient Communication in Diagnostics: What It Can and Cannot Do

AI can reliably restate a diagnostic report in language a patient can read, and the published evidence supports that. It cannot safely decide whether a finding is serious, what should happen next, or what a result means for a particular person. The distinction is not a marketing nuance. It is the line between a communication tool and a clinical decision tool, and it determines what oversight your organisation must keep in place.

Clinicians are right to be sceptical of AI in patient-facing settings, and vendors have earned that scepticism. This piece sets out what the evidence actually shows, including the failure rate, and what a provider should require before putting any such system in front of patients.

What the evidence shows

A systematic review and meta-analysis in The Lancet Digital Health examined 38 studies published between 2022 and 2025, covering 12,922 simplified reports evaluated by 508 assessors. It found that large language model simplification improved patient understanding and readability while largely preserving accuracy and completeness. In practical terms, readability shifted from university-level text to roughly school-level, around ages 11 to 13, for CT and X-ray reports.

That is a meaningful finding. Reports are typically written far above the reading level of the average adult, so a shift of that size is the difference between a document a patient can attempt and one they cannot.

The error rate, stated plainly

The same meta-analysis reported an error rate of 7.2% in LLM-rewritten reports, with clinically significant errors at 0.9%. Both figures deserve to be quoted by anyone selling or buying this technology, and most vendors quote neither.

Nought point nine percent is small. It is not zero. At meaningful volume it is not a rounding error either, which is precisely why the authors concluded that human oversight remains necessary. The correct response to that number is not to dismiss the technology, nor to pretend the number does not exist, but to design around it: keep the original clinical report authoritative and unaltered, position the simplified version as an accompaniment rather than a replacement, and preserve a clear route to a clinician.

TaskSuitable for AI assistance?Why
Restating clinical terms in plain languageYes, with oversightHigh volume, repeatable, evidence supports readability gains
Explaining what a test measuresYesEducational content, not patient-specific judgement
Multilingual delivery of the aboveYes, with per-language reviewExtends access, but mistranslated hedges are clinical errors
Deciding whether a finding is urgentNoClinical judgement requiring full context and accountability
Recommending next steps or treatmentNoClinical decision-making
Detecting disease or offering a diagnosisNoA different product category with different regulatory obligations

Simplification is not diagnosis

The boundary is easy to state and easy to erode. A simplification layer takes content that already exists in the report and expresses it differently. It adds no new clinical assertion. The moment a system tells a patient that a finding is probably harmless, or that they should be seen urgently, it has made a clinical claim the report did not make, and it has become something other than a communication tool.

If a system says anything the report did not already say, it is no longer simplifying. It is diagnosing, and it should be evaluated as a clinical device rather than a communication tool.
FlexReport Patient Communication Framework

What this means for your centre

The operational case does not rest on the technology being impressive. It rests on the fact that patients now receive reports before anyone explains them, and that a large share of the resulting questions are vocabulary questions rather than clinical ones. Those are exactly the tasks the evidence supports automating, under oversight. The clinical questions stay where they belong.

Six questions to ask any vendor

  1. What does your system do when a patient asks whether their result is dangerous? The correct answer is that it directs them to a clinician.
  2. What is your measured error rate, how was it measured, and by whom?
  3. Does the original clinical report remain unaltered and authoritative?
  4. What clinical oversight exists, and what happens when the underlying model version changes?
  5. Is patient data used for model training, and can that be excluded contractually?
  6. What regulatory claims are you making, and can you evidence each one? Be sceptical of implied approvals.

A vendor who answers these plainly, including where the answer is unflattering, is more trustworthy than one with a more impressive pitch. That applies to us as much as anyone.

FR
The FlexReport Team
Writing from the FlexReport team about radiology, language, and trust.