← Back to insights AI in Medicine

Claude in Clinical Documentation: What Ambient AI Evidence Actually Shows

Kamna Singh Thakur 20 min read

Claude can be a useful drafting and summarisation component in clinical documentation workflows. But the strongest peer-reviewed evidence available today comes from ambient documentation systems generally, not from clinical trials of Claude itself. That evidence supports clinician-reviewed drafting; it does not support autonomous note completion or treating fluent output as clinical truth. A BAA can enable a compliant architecture, but it does not make the finished workflow compliant by itself.


Key takeaways

  • The safe question is not “is Claude medical?” It is: what precisely is this workflow asking the model to do, what evidence does it receive, and who checks the result?
  • Ambient documentation has real supporting evidence. In a 263-participant quality improvement study across six US health systems, burnout among the 186 clinicians included in the burnout models fell from 51.9% to 38.8% after 30 days — an adjusted odds ratio of 0.26.
  • That study was not a randomised trial. It was a voluntary pre/post evaluation of a single commercial platform (Abridge), with self-reported outcomes, no control group, incomplete follow-up, and an intervention funded by the participating institutions.
  • Note quality is close, but the failure mode differs. Across 97 note pairs in five specialties, physician notes scored 4.25 out of 5 and ambient notes 4.20 — but hallucinations appeared in 31% of ambient notes against 20% of physician notes.
  • The most uncomfortable finding: reviewers preferred the ambient notes anyway, 47% to 39%. More thorough, better organised, more likely to contain hallucinated content — and preferred. That result is consistent with the kind of automation bias safety frameworks warn about.
  • Ceremonial review is measurable, not theoretical. A separate pilot of 7,545 AI-generated notes found 14.9% were left entirely unedited, with a median of 9% of AI-generated words changed.
  • The same pilot moderates the 31% figure: hallucinations appeared in 11.5% of formally evaluated notes, and only 5.3% contained errors rated as posing serious or imminent risk.
  • Fluency conceals omission as well as fabrication. NIST names the healthcare case explicitly: a confabulated patient summary could lead a doctor to an incorrect diagnosis or the wrong treatment.
  • HIPAA-ready is not compliant. A Primary Owner must activate HIPAA compliance before a standard Claude Enterprise organisation receives BAA coverage, and Anthropic’s BAA does not cover data sent to third parties through connectors.
  • Drafting is not deciding. A system that organises supplied facts has a different regulatory profile from one that recommends a diagnosis, treatment or coverage determination.

What ambient AI evidence can actually tell us about Claude

Anthropic introduced Claude for Healthcare on 11 January 2026 as a set of tools, connectors and resources for providers, payers and health-technology companies. The announcement included access to the CMS Coverage Database, ICD-10 data, the NPI Registry and PubMed, plus a FHIR development skill. Anthropic also identified ambient clinical documentation as a use case for startups building on its developer platform.

That is a product announcement, not independent clinical validation. Claude remains a general-purpose generative model with healthcare data plumbing attached.

So the useful framing is not “is Claude medical.” It is three narrower questions: what is this workflow asking the model to do, what evidence does the model receive, and who checks the result before it reaches a patient record?

Supported by current evidenceNot established by current evidence
Ambient documentation can reduce perceived burden and after-hours work in some health-system deploymentsThat every product, specialty, accent, visit type or note template achieves the same result
LLM-generated notes can be organised and thorough enough to be useful draftsThat a fluent note is factually complete, free of omissions, or safe to sign without review
Claude can support healthcare administration and documentation with appropriate controlsThat Claude is a validated autonomous diagnostician or a substitute for professional judgement

Where Claude can add value in clinical documentation

Healthcare produces large volumes of repetitive, text-heavy work. That is where language models are most credible: turning supplied information into a first draft a qualified person inspects.

Suitable tasks include organising an encounter transcript into a structured draft note; rewriting clinician-approved content into a patient-friendly summary; extracting facts from a defined record set for chart review; preparing draft correspondence, handoffs or prior-authorisation material from verified sources; and helping developers work with FHIR resources and healthcare terminology, with authoritative code systems remaining the source of record.

The distinction that matters: drafting is not deciding. A system that organises supplied facts into a note has a different risk profile from one that recommends a diagnosis, treatment, disposition or coverage determination. Everything downstream — regulatory classification, oversight design, liability — follows from which side of that line the product sits on. We cover the regulatory consequences in agentic AI regulation.


What the ambient-scribe evidence actually shows

The strongest public evidence for generative AI in routine clinical work concerns documentation, not diagnosis. The studies below did not evaluate Claude directly; they evaluated ambient documentation systems and therefore inform the risk and workflow design of this class of product rather than establish a Claude-specific clinical error rate.

Two studies carry most of the weight, and they point in different directions.

The burnout study

A 2025 multicentre quality improvement study led from Yale, conducted between February and October 2024, enrolled 263 physicians and advanced practice practitioners across six US health systems. After 30 days of ambient scribe use, burnout among the 186 participants included in the burnout models fell from 51.9% to 38.8% — a difference of 13.1 percentage points, corresponding to an adjusted odds ratio of 0.26 (95% CI 0.13–0.54, P<.001). Cognitive task load, after-hours documentation, focused attention on patients and urgent access to care all improved significantly.

That is a substantial result and it should be taken seriously. It should not be converted into a universal performance promise, and the study’s own authors do not do so.

Five limitations belong with the figure whenever it is quoted. It was a voluntary pre/post evaluation, not a randomised comparison. There was no control group, so temporal trends could not be adjusted for. Outcomes were self-reported, with no quantitative EHR data. The dataset included only complete survey responses, so non-completers could not be characterised. The study evaluated one commercial ambient documentation platform, Abridge, so its findings should not be treated as evidence that every ambient-scribe product will produce the same result. The intervention was funded by the participating institutions. 

Note also the participant arithmetic: 263 clinicians completed the study, but the burnout analysis rests on 186. Both numbers are correct; only one is the denominator for the headline figure.

The note quality study

A separate 2025 study compared 97 pairs of ambient and physician-authored notes across five specialties using the validated PDQI-9 instrument.

Overall quality was close: physician “gold” notes scored 4.25 out of 5, ambient notes 4.20 (P=0.04). Physician notes were more succinct, more accurate and more internally consistent. Ambient notes were more thorough and better organised.

And hallucinations were detected in 31% of ambient notes against 20% of gold notes (P=0.01).

Two qualifications. Several authors were affiliated with Suki AI, the vendor whose scribe was evaluated, and with Hippocratic AI. The sample was modest and assessed a single pipeline. This is a warning signal, not a market-wide error rate.

The finding the source coverage keeps omitting

Here is the result that should change how you design review workflows, and it is in the same paper’s abstract: reviewers overall preferred the ambient notes, 47% against 39% for the physician-authored ones.

More thorough. Better organised. More likely to contain fabricated content. And preferred by the clinicians reading them.

That result is consistent with the kind of automation bias safety frameworks warn about. The qualities that make an AI note feel better — completeness, structure and readability — are not the same as factual accuracy. A review step that relies only on a clinician noticing something wrong in a well-organised document may therefore be weaker than it appears.

The third study, which moderates both

A pragmatic prospective pilot ran 31 physicians across multiple specialties through 7,545 AI-generated clinic notes, formally evaluating 356 of them against four error types.

Accidental omissions were the most frequent error at 18%, followed by hallucinations at 11.5% and accidental inclusions at 9.3%. Bias was rare at 1.1%. Most errors — 83.8% — were rated mild to moderate, with only 5.3% of notes containing errors rated as posing serious or imminent risk.

That moderates the 31% headline considerably. Different instrument, different threshold for what counts as a hallucination, much larger sample. Anyone quoting a single ambient-scribe error rate is quoting a measurement methodology as much as a product.

But the same study contains the most operationally alarming number in this article. Physician editing varied enormously — average AI-generated words changed ranged from 1.9% to 69.3%, median 9% — and 14.9% of notes were left entirely unedited.

Roughly one in seven generated notes received no edits. That does not by itself prove the note was signed or filed without review, but it shows why teams should measure review behaviour rather than assume that a review step is meaningful.

What follows from all three studies together: the right operating model is clinician-reviewed drafting with measured quality, not automatic chart completion based on fluent output. And the review step itself needs designing, monitoring and measuring — because left alone, a material fraction of it does not happen.


The five limitations that matter in real workflows

1. The model inherits upstream errors

In an ambient workflow the model sits downstream of microphones, speech recognition, segmentation, speaker attribution and EHR retrieval. If a dose is transcribed incorrectly, a patient statement is attributed to the clinician, or the wrong encounter context is retrieved, a polished summary preserves or amplifies the error.

Language-model quality cannot reliably repair missing evidence. This is why ASR selection is a clinical safety decision rather than a procurement detail — we cover the specific failure modes in open medical AI models.

2. Fluency conceals fabrication and omission

NIST defines generative-AI confabulation as the production of confidently stated but erroneous content, and names the healthcare case explicitly: a confabulated summary of patient information could cause doctors to make incorrect diagnoses or recommend the wrong treatment.

Omission matters equally and gets far less attention — note that in the 7,545-note pilot, omissions were the most frequent error type at 18%, well ahead of hallucinations. A note can read perfectly while excluding a negative finding, a stated uncertainty, a medication detail or a follow-up condition that changes its clinical meaning. Fabrication is easier to spot because something is there that should not be. Omission leaves no trace to notice.

3. Output is sensitive to context and instructions

Template wording, transcript length, specialty, language, local abbreviations and the ordering of EHR facts all change the response. A prompt that works for primary care may perform poorly in mental health, obstetrics or emergency medicine.

Evaluation therefore has to be stratified by real use case rather than reported as a single average accuracy number. One figure across all specialties conceals exactly the variation you need to see.

4. Human review can become ceremonial

WHO warns that plausible generative output creates automation bias, in which professionals overlook errors or delegate difficult choices to the model. The findings above make that risk operationally relevant: reviewers preferred the notes that hallucinated more, and roughly one in seven generated notes received no edits.

A review button is not a safety control if time pressure makes approval automatic. Review must expose the source evidence, make edits easy, and — critically — measure what clinicians actually change. Edit rate is a safety metric, not a usage statistic.

5. Models and product surfaces change

Model updates alter verbosity, extraction behaviour, refusal patterns and structured-output reliability. Production teams need model and prompt versioning, regression cases and a rollback path. A one-time validation does not cover future model releases, and API model versions change on the provider’s schedule rather than yours.


HIPAA-ready does not mean automatically compliant

Anthropic offers a BAA for qualifying HIPAA-ready services. For Claude Enterprise, a Primary Owner must activate HIPAA compliance and accept the BAA — a standard Enterprise organisation is not covered automatically. For first-party API use with PHI, the organisation must sign a BAA and have the capability enabled. Coverage is organisation- and feature-specific, and several API features are excluded or unavailable entirely within a HIPAA-ready organisation.

This matters because “Claude” is not one legal surface. Consumer accounts, enterprise chat, the first-party API, cloud-platform access, connectors and third-party tools carry different contracts, retention settings and data flows. Anthropic also states that data sent to third parties through connectors is not covered by its BAA — which is significant given that connectors are the core of the healthcare offering. We map the full coverage boundary in what the BAA actually covers.

HHS guidance is equally clear that a BAA is one control among many. A covered entity or business associate must understand the cloud environment, perform its own risk analysis, and implement appropriate administrative, physical and technical safeguards. A vendor’s willingness to sign a BAA is not a government certification of your finished application.

Procurement rule: map every PHI flow, subprocessor, storage location and enabled feature, and confirm contractual coverage for the exact configuration before sending PHI — not after a pilot has started.

For UK and Australian deployments, the analysis is different

HIPAA is US law and has no application to an NHS trust or an Australian practice. The equivalent questions are:

In the UK, whether patient data can lawfully leave the estate at all, and where inference happens. Anthropic’s first-party API offers US or global inference geographies with no UK-only option; in-region processing requires AWS Bedrock or Google Vertex AI with explicit region configuration. Beyond that sit DCB0129 clinical risk management for the supplier, DCB0160 for the deploying organisation, DTAC for procurement, DSPT for data security, and a DPIA under UK GDPR. NHS England’s position is that safety requirements apply to all digital products regardless of whether they are considered a medical device.

In Australia, the TGA regulates by intended purpose under section 41BD of the Therapeutic Goods Act 1989, and AHPRA holds the practitioner responsible for clinical decisions regardless of the technology used, with an expectation that patients are told when a digital scribe is in use.

The underlying discipline is the same in all three jurisdictions. The instruments are not, and a BAA satisfies none of the UK or Australian ones.


Documentation support and clinical decision support are not the same thing

Whether software is regulated as a medical device depends on its intended function and claims, not on the model vendor underneath it.

The FDA’s January 2026 Clinical Decision Support guidance explains that some CDS functions fall outside the device definition when statutory criteria are met, while functions meeting the device definition remain subject to FDA digital-health policies. Patient- or caregiver-facing functions receive separate scrutiny.

A product that drafts a note from an encounter is not equivalent to one that ranks diagnoses or recommends treatment. Once a workflow begins influencing diagnosis, treatment, triage or another consequential decision, the team should obtain specialised regulatory advice and document how a clinician can independently review the basis of any recommendation.

Product labels such as “copilot” or “clinical intelligence” do not decide regulatory status. Marketing claims can, however, change it — which is why the intended purpose statement is a product artefact rather than a submission document. Our breakdown of is an AI scribe a medical device covers where the line sits in each market, including the nine worked examples the MHRA published in July 2026.

For certified US health IT, ONC’s HTI-1 rule adds transparency expectations for predictive decision support interventions, including information helping users assess fairness, appropriateness, validity, effectiveness and safety.


A worked example: Claude as a drafting layer

Experience, not independent evidence. At a high level, KH Scribe uses Claude after transcription to organise encounter evidence into a clinician-reviewable draft note and a plain-language patient summary. The design intent is documentation assistance, not autonomous diagnosis.

The controls that matter: traceability from every generated statement back to the source transcript or selected EHR facts; explicit handling of missing evidence rather than silent gap-filling; clinician editing and sign-off before anything enters the record; and separate validation of any coding suggestion.


A defensible deployment checklist

Eleven items, ordered so nothing blocks anything after it.

  1. Define one narrow intended use, and write down explicitly what the system must never decide.
  2. Use the minimum necessary patient data, and document every third-party data flow including connectors.
  3. Confirm the BAA and eligible features for the exact organisation and product surface — or, outside the US, confirm the residency position and DPIA before any data moves.
  4. Obtain patient disclosure or consent for ambient recording under applicable law and organisational policy. AMA educational guidance emphasises proactive disclosure, consent and manual verification; AHPRA expects patients to be informed in Australia.
  5. Keep generated notes in draft status until an authorised clinician reviews and accepts them.
  6. Show supporting transcript or EHR evidence adjacent to the generated statement, not one click away.
  7. Fail visibly when audio, transcript, patient identity or evidence is incomplete. Do not silently fill gaps.
  8. Test by specialty, language, accent, audio quality and encounter type, including deliberately difficult negative cases.
  9. Measure omissions, unsupported additions, laterality, medication details, negation, speaker attribution and clinician edits — not readability. Track edit rate as a safety metric and investigate reviewers whose rate approaches zero.
  10. Version the model, prompt, template and terminology sources, and rerun regression tests before any change reaches clinicians.
  11. Monitor post-deployment errors, maintain an incident path, and sample signed notes for continuing quality review.

What this means for using Claude in clinical documentation

The most credible role for a general-purpose model in healthcare today is augmentation: helping qualified people turn supplied information into a useful first draft. That is not a small thing. The burnout evidence is real, and reducing after-hours documentation is a genuine improvement in the experience of clinical work.

But the same technology omits, misattributes and invents clinically important content, and the note-quality research suggests something more uncomfortable than a simple error rate: clinicians preferred the notes that hallucinated more often. Thoroughness and organisation are the qualities a reader notices. Accuracy is the quality a reader has to actively check for. Those are not the same thing, and a well-structured wrong note is harder to catch than a scrappy one.

Trust therefore cannot come from the model. It comes from the system around it — narrow intended use, evidence traceability, appropriate contracts, patient transparency, clinician authority, measured performance, and monitoring that continues after go-live.

In healthcare, a polished answer is the beginning of review, not the end of it. The design question is whether your workflow makes that review real or merely available.


Frequently asked questions

Can Claude be used safely in healthcare?
As a drafting and summarisation component under clinician review, with appropriate contracts and controls, yes. As a source of clinical truth or a substitute for professional judgement, no. Claude is a general-purpose generative model, not a validated clinical device, and Anthropic does not market it as one.

Does ambient AI documentation actually reduce burnout?
There is real evidence that it can. A 2025 quality improvement study across six US health systems found burnout fell from 51.9% to 38.8% after 30 days among the 186 clinicians included in the burnout analysis, with an adjusted odds ratio of 0.26. However, it was a voluntary pre/post evaluation with no control group, self-reported outcomes, a single commercial platform, and an intervention funded by the participating institutions. It shows documentation support can help in some settings; it does not establish a universal result.

How often do AI-generated clinical notes contain hallucinations?
Estimates vary substantially with methodology. One validated PDQI-9 comparison of 97 note pairs found hallucinations in 31% of ambient notes against 20% of physician-authored notes. A larger pragmatic pilot evaluating 356 notes drawn from 7,545 found hallucinations in 11.5%, with accidental omissions more frequent at 18%, and only 5.3% of notes containing errors rated as posing serious or imminent risk. Any single quoted rate reflects a measurement method as much as a product.

Are AI-generated notes worse than physician notes?
Not straightforwardly. In the PDQI-9 comparison, overall quality was near-identical — 4.25 out of 5 for physician notes against 4.20 for ambient. Physician notes were more succinct, accurate and internally consistent; ambient notes were more thorough and better organised. Ambient notes hallucinated more often. Notably, reviewers preferred the ambient notes 47% to 39% despite the higher hallucination rate.

Why does it matter that reviewers preferred the AI notes?
Because the result is consistent with automation-bias concerns. The qualities that make a note feel better to read — completeness, structure and clarity — are not the same as factual accuracy. A review process that depends only on a clinician noticing an error inside a polished document may therefore be weaker than it appears.

What is the biggest risk in an ambient scribe workflow?
Omission rather than fabrication, on the available evidence. In the largest sample evaluated, accidental omissions were the most frequent error type at 18%. Fabrication is easier to detect because something is present that should not be; omission leaves nothing to notice. A note can read perfectly while excluding a negative finding, an uncertainty or a medication detail that changes its clinical meaning.

Do clinicians actually review AI-generated notes?
Not always. In a pilot covering 7,545 AI-generated notes, 14.9% were left entirely unedited, and the average proportion of AI-generated words changed ranged from 1.9% to 69.3% across physicians, with a median of 9%. Edit rate should be treated as a safety metric and monitored, not assumed.

Is Claude HIPAA compliant?
Not automatically. Anthropic offers a BAA for qualifying HIPAA-ready services, but a standard Claude Enterprise organisation is not covered until a Primary Owner activates HIPAA compliance and accepts the BAA. For first-party API use with PHI, the organisation must sign a BAA and have the capability enabled. Coverage is organisation- and feature-specific, and data sent to third parties through connectors is not covered.

Does a signed BAA mean our application is compliant?
No. HHS guidance is explicit that a covered entity or business associate must understand its cloud environment, perform its own risk analysis, and implement appropriate administrative, physical and technical safeguards. A vendor’s willingness to sign a BAA is not a certification of your finished application.

Does any of this apply to the NHS?
The discipline does; the instruments do not. HIPAA has no application in the UK. The relevant questions are whether data may leave the estate, where inference occurs, and compliance with DCB0129, DCB0160, DTAC, DSPT and UK GDPR including a DPIA. Anthropic’s first-party API offers no UK-only inference region, so in-region processing requires AWS Bedrock or Google Vertex AI with explicit region configuration.

Is an AI scribe a medical device?
It depends on intended purpose, claims and jurisdiction. A product that only transcribes or summarises information already present in the consultation for clinician review may fall outside medical-device regulation in some circumstances. Adding diagnostic, treatment, triage or autonomous clinical functions can materially change that position. Marketing claims can also affect classification.

What is the difference between documentation support and clinical decision support?
Documentation support organises supplied facts into a draft a clinician reviews. Clinical decision support influences a diagnosis, treatment, triage or coverage decision. The FDA’s January 2026 guidance sets out statutory criteria under which some CDS functions fall outside the device definition. Once a workflow influences a consequential clinical decision, specialised regulatory advice is warranted.

How should we evaluate an ambient scribe before deploying it?
Stratified by real use case rather than as a single average. Test by specialty, language, accent, audio quality and encounter type, including difficult negative cases. Measure omissions, unsupported additions, laterality, medication details, negation, speaker attribution and clinician edit rates — not readability or general accuracy.

What happens when the underlying model is updated?
Behaviour can change, including verbosity, extraction reliability, refusal patterns and structured-output stability. A one-time validation does not cover future releases. Production deployments need model and prompt versioning, a regression suite, and a rollback path, with a documented revalidation trigger in the clinical safety case.

Does using Claude make our product a medical device?
No. Classification depends on your product’s intended purpose, not on which model powers it. But if your product does meet the device definition, you are the manufacturer, not Anthropic — the TGA states this explicitly, and the MHRA and FDA frameworks assign obligations the same way.


Deciding whether an ambient scribe is right for your setting?

The evidence supports documentation assistance under genuine clinician review. It says nothing about your specialty, your accents or your governance position — and those decide whether a deployment works.

That is what the Discovery Workshop covers. Sometimes the answer is that the documentation burden is a workflow problem, and adding AI would accelerate a process that needs redesigning first. We would rather say so before you buy something than after.


Sources

Peer-reviewed studies

  1. Olson KD, Meeker D, Troup M, Barker TD, Nguyen VH, Manders JB, Stults CD, Jones VG, Shah SD, Shah T, Schwamm LH. Use of Ambient AI Scribes to Reduce Administrative Burden and Professional Burnout. JAMA Network Open. 2025;8(10). doi:10.1001/jamanetworkopen.2025.34976
  2. Palm E, Manikantan A, Mahal H, Belwadi SS, Pepin ME. Assessing the quality of AI-generated clinical notes: validated evaluation of a large language model ambient scribe. Frontiers in Artificial Intelligence. 2025;8:1691499. doi:10.3389/frai.2025.1691499
  3. Quality of Clinical Notes Created by Ambient Listening Generative AI: Pragmatic Prospective Pilot Study. JMIR Medical Informatics, 2026. (31 physicians, 7,545 notes, 356 evaluated)

Standards and safety frameworks

  1. NIST. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1, July 2024. PDF
  2. World Health Organization. Ethics and Governance of Artificial Intelligence for Health: Guidance on Large Multi-Modal Models, 2024. Source

Regulatory and legal

  1. US FDA. Clinical Decision Support Software: Guidance for Industry and FDA Staff, January 2026. Source
  2. US HHS Office for Civil Rights. Guidance on HIPAA & Cloud Computing. Source
  3. ONC. HTI-1 Final Rule overview, last updated 7 November 2025. Source
  4. AMA Code of Medical Ethics. A Structured Decision Simulator for Using an Ambient AI Scribe, 16 June 2026. Source
  5. MHRA. Ambient voice technology-enabled products, 29 July 2026.
  6. NHS England. Guidance on the use of AI-enabled ambient scribing products in health and care settings, Version 3, 29 July 2026.
  7. TGA. Digital scribes; Artificial intelligence (AI) and medical device software regulation, 5 February 2026.