AI in healthcare
AI Scribes Are Quietly Reshaping Hospital Bills, and Insurers Are Alarmed
A new Blue Cross Blue Shield analysis links AI coding and ambient scribe tools to $942 million in added hospital costs. Here is what the evidence actually shows, and where it falls short.
A quiet change in the medical record, and a billion-dollar argument about it#
Something odd started showing up in the records of new mothers. More of them were being logged with acute posthemorrhagic anaemia, a diagnosis that signals dangerous blood loss and usually calls for a transfusion. The transfusions, though, were not keeping pace with the diagnoses. The label was climbing. The treatment was not.
That mismatch sits at the centre of a report the Blue Cross Blue Shield Association (BCBSA) released on 24 September 2026, and it has turned into one of the sharper fights in health policy this autumn. The insurer group says AI tools that write clinical notes and assign billing codes are pushing patients into more complex, higher-paying categories without any matching change in the care they receive. It puts the added cost to Blue plans at roughly $942 million over 2024 and 2025 compared with a 2023 baseline.
Hospitals reject the reading. They say patients really are older and sicker, and that better software is finally capturing conditions that rushed human documentation used to miss. Both sides are partly describing the same technology. What they disagree about is what it is doing to the truth of a medical record, and who pays when the record and the care drift apart.
Background: what an AI medical scribe actually does#
Two related tools sit behind this story, and it helps to separate them.
The first is the ambient AI medical scribe. It listens to a clinical visit through a microphone, transcribes the conversation, and drafts a structured note for the doctor to review and sign. The pitch is straightforward: doctors spend hours every week on paperwork, and an AI-powered medical scribe can hand some of that time back. Companies such as Abridge, Nuance, Suki, Nabla and Heidi compete in this space, and adoption has moved fast.
The second is AI medical coding. Coding is the process of translating a visit into standardised codes that hospitals use to bill insurers. Every diagnosis and procedure maps to a code, and secondary diagnoses, the conditions listed alongside the main reason for a stay, can raise how much a hospital is paid. AI coding tools scan notes, lab results and the ambient transcript, then suggest codes automatically.
The connection matters. A fuller note supports more codes. When an ambient scribe captures a passing mention of, say, kidney disease during a stroke admission, that mention can become a billable secondary diagnosis. This is where efficiency and revenue start to blur, and where the term upcoding enters. Upcoding means billing for a more severe or complex case than the care delivered supports. It can be deliberate fraud, but it can also emerge quietly from documentation that is simply more complete than it used to be.
What the Blue Cross analysis found#
The analysis came from Blue Health Intelligence, the data arm of the Blue system, working with de-identified claims from tens of thousands of admissions. Its headline number is the share of inpatient claims coded as medically complex, which rose from 37% at the start of 2023 to 40% by the end of 2025. On its own a three-point rise sounds small. Spread across millions of hospital stays, it moves a lot of money.
About $653 million of the $942 million came from secondary diagnoses that pushed more than 55,000 claims into higher-paying tiers, roughly $11,000 per extra-complex case. In major bowel procedures, the proportion of claims coded at the highest complexity level jumped from 10.2% to 22.7%. And in the maternity example that opened this piece, hospitals with the fastest rise in anaemia diagnoses gave transfusions less often than their peers, 16.9% against 19.3%, which is the opposite of what you would expect if these patients were genuinely bleeding more.
The group also tried to isolate the AI signal. Reviewing facilities that had publicly disclosed AI adoption, it found one hospital whose complexity rating climbed 6.7% after it announced a switch to AI, against 0.9% for other facilities in the same state. An earlier BCBSA analysis in March had focused on the postpartum anaemia pattern and estimated that roughly $2.3 billion in inpatient and outpatient spending nationwide might be tied to AI-enabled coding.
The people behind the report were careful to disagree with each other about how hard to push the claim. "The disconnect between diagnoses and treatment suggests that AI is identifying more billable conditions, not sicker patients," said Luke Chalker, the association's senior vice president of product and data science. Dr Razia Hashmi, its vice president of clinical affairs, was more cautious: "There may be an element of correct coding there, but the likelihood that this is technology-enabled upcoding is higher, in my view," she told reporters.
The hospitals' rebuttal, and why it is not easy to dismiss#
Read the source before you read the conclusion. The analysis came from an insurer trade group that has a direct financial interest in lowering hospital payments, and it was published as a white paper rather than a peer-reviewed study. It rests on claims data, not on medical charts, and the association acknowledged that reviewing actual charts would be a more direct test of whether patients were sicker.
The American Hospital Association has been making the counterargument since the summer. In a July fact sheet it argued that an ageing population and rising chronic disease are genuinely increasing patient complexity, and it cited an AHA and Vizient analysis showing the hospital case-mix index, a standard measure of how sick admitted patients are, rose about 5% from 2019 to 2024. Its sharpest point turns the accusation around: a 2025 finding from the Medicare Payment Advisory Commission that upcoding contributed to about $40 billion in overpayments to Medicare Advantage plans, which are run by insurers. If coding intensity is a sin, it is not one committed only by hospitals.
Both accounts can hold at once. Some of the new diagnoses probably reflect real conditions that hurried notes used to miss. Others are likely thin. A population-level claims analysis cannot tell any individual patient which category their own record falls into, and that is the honest limit of what has been published so far.
What the peer-reviewed evidence says about coding intensity#
Set the insurer white paper aside and the pattern still appears in more careful work, which is what makes the debate worth taking seriously rather than dismissing as a payer talking point.
A study from UCSF published in January 2026 tracked ambulatory visits and found that ambient scribes raised relative value units and the number of patient encounters, with no change in claim denials. Relative value units are the standard yardstick for the intensity of a service. The measured effect worked out to about $3,044 more per physician per year on the 2025 Medicare fee schedule. Riverside Health, a nine-hospital system in Virginia, reported an 11% rise in physician relative value units after adopting Abridge's scribe. A quasi-experimental study posted to medRxiv examined the longitudinal effect of ambient AI scribe use on documentation burden and financial productivity and found gains on both, including a 16% drop in note-writing time by day 150 across 220 clinicians and more than 314,000 encounters, though it is a preprint (not peer reviewed) and should be read as preliminary. Separate market research from Trilliant Health likewise found outpatient coding intensity rising as hospitals adopt AI-enabled scribing, though it read the trend as largely reflecting more thorough, rules-based documentation rather than abuse.
The consistent finding across these sources is not fraud. It is drift. Give a physician a note that is more complete than the one they would have typed at the end of a long shift, and the codes attached to that note tend to rise. Whether that is accuracy finally catching up or billing quietly inflating is exactly the question the numbers alone cannot settle.
The deeper problem: deployment is outrunning evaluation#
The billing fight is really a symptom of a wider issue in AI in healthcare. These systems are being rolled out at scale faster than anyone is measuring what they do. Federal data shows seven in ten US hospitals used predictive AI in 2024, and the use of AI for billing jumped sharply year on year. Ambient scribes went from pilots to enterprise contracts in the space of a couple of years.
The independent evidence has not kept up, and where it exists it complicates the marketing. A June 2026 study in Nature Medicine found that general-purpose frontier models outperformed specialised clinical AI tools on medical knowledge benchmarks and on real clinical questions, which suggests that a "clinical" label does not guarantee a better or safer product. Other work has looked at how clinicians edit AI drafts, including a preprint examining whether doctors soften or remove the hedging language an AI note generates, a small choice that can change what a record implies about certainty. When systems are documenting and coding for millions of people, small systematic choices become large systematic effects.
For patients, the most concrete stake is not the premium line on a spreadsheet. It is accuracy. A secondary diagnosis added during a hospital stay does not stay on the bill. It follows you into the records future clinicians read. If a condition was recorded but never discussed or treated, it can confuse later care. Under HIPAA, people can request copies of their medical and billing records and ask for corrections, which is worth remembering after any hospital stay where the explanation of benefits lists something that does not match what happened in the room. Any new or worsening symptom, of course, belongs with a clinician rather than a billing review.
What a fix would look like, and why it matters beyond the US#
The obvious way to end this argument is also the one nobody has done: read the charts. A patient-level review, comparing coded diagnoses against the notes and the treatment actually delivered, would show how much of the rise is real illness and how much is inflation. The BCBSA says it plans further analyses of outpatient care and other diagnosis categories, and hospitals have their own audit data. Until one side publishes chart-level results, the public is left choosing between two motivated readings of the same claims.
There are practical guardrails short of that. Insurers and health systems can flag facilities whose coding jumps out of step with peers, as the Blue analysis did with the hospital whose complexity rose nearly seven times faster than others in its state. Coding tools can be required to link each suggested secondary diagnosis to supporting evidence in the record. And independent testing, of the kind the Nature Medicine benchmark study modelled for clinical chatbots, could be extended to documentation and coding systems before they are trusted at scale.
The stakes are not only American. Health systems from the UK to India are piloting ambient scribes, and researchers have already flagged that most of these tools are trained and validated on English-language, Western clinical conversations. A recent study on evaluating ambient scribes in India argued that multilingual, real-world clinical data is missing from the evidence base entirely, which is a preprint (not peer reviewed) but points at a real gap. Whatever the United States decides about AI coding will shape how the technology is regulated, reimbursed and trusted across the global scientific and clinical community. Getting the evidence right, rather than fast, is the part that travels.
Key takeaways#
The BCBSA links AI coding and ambient scribe tools to about $942 million in added inpatient costs for its plans across 2024 and 2025, driven mainly by secondary diagnoses that raised claim complexity without a matching rise in treatment.
The strongest single data point is the treatment gap: at hospitals diagnosing more blood-loss anaemia, patients received transfusions less often, not more, which is hard to square with the idea that they were sicker.
The analysis is not independent or peer reviewed. It comes from an insurer group, relies on claims rather than charts, and the hospital sector counters that patients are genuinely more complex and that insurers upcode too.
Peer-reviewed and preprint studies separately confirm that ambient scribes raise coding intensity and physician productivity. They cannot yet say how much of that reflects real illness versus fuller documentation.
The underlying problem is an evaluation gap. AI documentation and coding are being deployed across most US hospitals faster than independent research can measure their effects on cost, accuracy and patient records.
Frequently asked questions#
What did the Blue Cross Blue Shield analysis actually claim? That rising coding intensity, which it attributes largely to AI tools, added an estimated $942 million in costs for Blue plans over 2024 and 2025 versus 2023, as complex inpatient cases rose from 37% to 40% of claims with no matching change in treatment.
What is an ambient AI scribe? Software that listens to a doctor and patient during a visit, transcribes the conversation, and drafts a clinical note for the doctor to review and sign. It is meant to cut paperwork and reduce burnout.
What is upcoding? Billing for a more severe or complex case than the care delivered supports. It can be deliberate, but it can also emerge from documentation that is more complete than before, which is why this case is genuinely contested.
Do hospitals accept the insurer's conclusion? No. The American Hospital Association argues that patients are older and sicker, points to a roughly 5% rise in case-mix index from 2019 to 2024, and notes that insurers themselves have been found to upcode, citing about $40 billion in Medicare Advantage overpayments.
Is there independent evidence that AI scribes raise billing? Yes. A January 2026 UCSF study and other research found that scribes increased relative value units and encounters. What that research cannot settle is whether the higher codes reflect real illness or simply better notes.
How does this affect me as a patient? Diagnosis codes enter your permanent record, not just your bill. It is reasonable to compare your insurer's explanation of benefits with your discharge summary and to ask about any diagnosis that was never discussed or treated.
Should I trust AI-written medical notes? The technology can reduce clerical burden and capture detail, but the note is only as reliable as the clinician's review. You are entitled to see your records and request corrections under HIPAA.
Glossary#
Ambient AI scribe: A tool that captures a clinical conversation through a microphone and drafts a structured medical note for the clinician to approve.
AI medical coding: Software that reads clinical notes and automatically suggests the standardised codes hospitals use to bill insurers.
Upcoding: Billing for a more complex or severe condition than the delivered care supports, whether intentionally or through documentation drift.
Secondary diagnosis: A condition listed alongside the main reason for a hospital stay. Additional secondary diagnoses can raise how much a hospital is paid.
Case-mix index: A standard measure of how sick a hospital's admitted patients are on average. A higher index generally means more complex, costlier care.
Relative value unit (RVU): A metric that quantifies the time, skill and intensity involved in a medical service, widely used to measure physician productivity and set payment.
Claims data: Billing records submitted to insurers. Useful for spotting patterns at scale, but weaker than medical charts for judging whether a patient was truly sick.
Preprint: A research paper shared publicly before peer review. It can signal early findings but has not yet been vetted by independent experts.
References#
Blue Cross Blue Shield Association, "Study suggests AI is boosting hospital billing," Blue Health Intelligence analysis. https://www.bcbs.com/news-and-insights/report/ai-boosting-hospital-billing
Medical Daily, "Blue Cross Ties $942 Million in Added Hospital Costs to AI Coding, but Hospitals Say Patients Are Sicker," 25 September 2026. https://www.medicaldaily.com/blue-cross-ai-coding-medically-complex-hospital-records-479102
Fierce Healthcare, "Hospitals' use of AI coding tools cost BCBSA plans $942M more for similar care, analysis finds." https://www.fiercehealthcare.com/finance/hospitals-use-ai-coding-tools-cost-bcbsa-plans-942m-more-similar-care-analysis
TechCrunch, "Insurers claim AI is already increasing healthcare costs," 26 September 2026. https://techcrunch.com/2026/09/26/insurers-claim-ai-is-already-increasing-healthcare-costs/
American Hospital Association, "Fact Sheet: Artificial Intelligence and Coding Intensity," 31 July 2026. https://www.aha.org/fact-sheets/2026-07-31-fact-sheet-artificial-intelligence-and-coding-intensity
Healthcare Brew, "The benefits and hidden costs of AI scribes" (UCSF and Riverside Health figures). https://www.healthcare-brew.com/stories/the-benefits-and-hidden-costs-of-ai-scribes
Trilliant Health, "Outpatient Coding Intensity Increases as Hospitals Adopt AI-Enabled Scribing." https://www.trillianthealth.com/market-research/studies/outpatient-coding-intensity-increases-as-hospitals-adopt-ai-enabled-scribing
medRxiv, "Longitudinal effects of ambient AI scribe use on documentation burden and financial productivity: a quasi-experimental study." Preprint (not peer reviewed). https://www.medrxiv.org/content/10.64898/2026.01.12.26343538v3
"General-purpose large language models outperform specialized clinical AI tools on medical benchmarks," Nature Medicine, June 2026. https://www.nature.com/articles/s41591-026-04431-5
US Department of Health and Human Services, "HIPAA guidance for individuals: access to and correction of records." https://www.hhs.gov/hipaa/for-individuals/guidance-materials-for-consumers/index.html