At the recent Global WHO Conference on Shaping AI in Health in Portugal, two speakers picked ambient AI scribes, the tools that listen to a consultation and draft the clinical note, as their example of choice for a much bigger problem. Hannah van Kolfschooten of the University of Basel laid out the questions patients actually ask about them: whether they’re being recorded, whether the recording is kept, who can access it, and whether they can say no. She then cited a newly published UK-wide survey of general practitioners: among GPs already using an ambient scribe, only 63% said they routinely sought patient consent.
Dr Ana Luísa Neves of Imperial College London made a related but sharper point: safety and evidence requirements have to be assessed use case by use case, not debated under a generic “AI” label. She named digital scribes specifically, arguing their adoption across the UK and elsewhere has outpaced the regulation, consent practice, and medical-device classification meant to govern them.
The benefit clinicians report is real
It would be a mistake to read any of this as evidence that clinicians or patients are pushing back on the technology. The same UK survey found that among current users, 80% reported spending less time on documentation and 70% reported reduced cognitive load. Among the clinicians who did ask for consent, 89% said no more than one in ten patients declined. Patients aren’t refusing, and clinicians are seeing a real benefit. The harder question is what happens later, when a patient disputes what a note actually says about them
Why the errors matter more than the rate
In the same UK survey, 32% of current users said scribe-generated errors occurred “often” or “always,” and 14% rated at least one error as significant to critical. A separate outpatient trial surveyed clinical staff directly: 47% of respondents said they had personally encountered a hallucinated detail in a scribe’s output, things like a fabricated medication, allergy, or measurement, even though 84% of the same respondents also reported the tool made them more efficient overall. That gap between usefulness and error rate is exactly why what happens after an error is found matters more than how often it happens.
Sign-off is not the same as an audit trail
NHS patient-facing guidance already anticipates this tension. Its suggested wording for patients reads: “The recording will be deleted once I have checked that the notes are accurate.” The clinician who reviews the note is accountable for what it says. But once the source audio is gone, so is the only independent record of what the patient actually said, versus what the model inferred, versus what was changed on review.
Several scribe products confirm this is a deliberate design choice, not an oversight. Heidi transcribes a consultation live but says it never stores the audio at all. Nabla doesn’t store audio by default and keeps other data for a configurable window, 14 days is its stated standard. One NHS practice using Heidi and TORTUS tells patients that draft notes are deleted within 24 hours. None of that is negligent; deleting audio quickly is a legitimate, privacy-protective choice. But it means the transcript, itself an AI output with its own risk of misattributed speech, is often the closest thing left to a record of what was actually said.
What the WHO conference put a name to
This is the implementation gap the WHO conference put a name to: governance arriving after the tooling is already in daily use. The EU’s own classification rules haven’t caught up either. A public consultation on which AI systems count as “high-risk” under the AI Act closes 23 July, evidence the boundary is still being drawn, not proof that every ambient scribe already falls on one side of it.
What a proportionate answer looks like
None of this argues for recording every consultation indefinitely; that trades one set of risks for a worse one on privacy and trust. It argues for building an evidence chain sized to the actual risk: knowing what exists (audio, transcript, prompt, draft, edits) and for how long it’s kept; preserving the relevant material the moment a patient raises a dispute or a safety concern; keeping a real version history of what a clinician changed in an AI draft and why; and giving patients an actual route to challenge an AI-assisted note, not just the assurance that someone reviewed it.
Ambient scribes are being adopted across UK general practice at exactly the moment this gap is most visible. The test for any organization adopting one is simple: if a patient challenges a note next month, can you reconstruct how it was written? For a meaningful share of current UK practice, once the audio is gone, the honest answer is no, not because anyone did anything wrong, but because no one built the system to answer that question in the first place.
Author: Dr. Mahé Pereira, Product Manager, Videolab



