Est.
AI in ClaimsLong read

Large Language Models in Medical Coding Assistance and Payer Response

Senior Writer · · 10 min read
Cover illustration for “Large Language Models in Medical Coding Assistance and Payer Response”
AI in Claims · September 30, 2026 · 10 min read · 2,183 words

Claim denials are climbing at the same time medical coding software is getting better at its job, and that pairing is the whole story here. Initial denial rates at Medicare Advantage and commercial payers are both in the double digits now, with Medicare Advantage running a bit hotter than commercial plans. Private payer denials moved from 8% to 11% between 2021 and 2023, and the climb hasn't leveled off since. None of this is happening in spite of AI adoption in billing offices. It's happening alongside it, and that timing is not a coincidence.

What LLMs do in a medical coding workflow

Medical coding is translation work. Someone has to take a diagnosis, a set of symptoms, a treatment plan, and a procedure note, and turn all of it into the standardized codes a payer's system will recognize: CPT, ICD, HCPCS, and a handful of adjacent sets. ICD-10 alone runs to 69,000 codes, and any coder working in the US has to move across six HIPAA-mandated code sets. That's the raw scale of the number of codes to look up before anyone touches a keyboard.

Large language models entered this workflow to do something narrower and more useful than "read the chart and pick a code." Their real contribution sits at the semantic layer: judging whether a clinical note actually describes the service being billed, with the specific attributes a payer's fee schedule demandsc5. That used to mean a coder reading dense free text line by line, matching phrasing to code definitions by hand. An LLM can surface that match in seconds instead of minutes because matching is a semantic read the model performs directly on the text, not a manual line-by-line search.

The stakes this produces are not small. One estimate puts incorrect billing losses to US taxpayers at $935 million a week, which gives some sense of why vendors keep building tools aimed at this exact failure point. Platforms in this space, MediMobile's Genesis among them, describe their systems as handling the routine cases without a human touching them and routing only the genuine exceptions to a person for review. That's the pitch across the category: less manual work, fewer inconsistent judgment calls, faster turnaround. A pitch is not a guarantee, and the next section gets into why.

Vendors lean hard on benchmark scores to sell these tools, and the benchmarks themselves deserve more scrutiny than they usually get. A June 2026 paper out of New York University, from researchers Perrett, Elliott, Hill, and Scott, makes the case that LLM benchmark evaluations systematically overstate how reliable these models actually are. Part of the problem is mundane: a lot of benchmark content already sits inside the training data the model learned from, so the model isn't really being tested, it's being asked to recall. Benchmarks also tend to skip over variance and the size of the errors when they happen, which matters more in billing than in most domains.

The same researchers ran a cleaner test. When benchmark questions and their answers were deliberately excluded from training data, frontier model accuracy on those same tasks fell off sharply, while human experts working the identical novel material got the large majority right. That gap is the whole argument in miniature: models look expert-level right up until the material is genuinely new to them, and then the gap opens fast.

Layer onto that the fact that LLM output is stochastic by design. In most applications that's a tolerable quirk. In medical coding, it isn't a curiosity, it's a liability: a code that's correct most of the time will still generate wrong codes on a real share of claims, and each of those is a denial with a name and a date attached.

A proof-of-concept study out of Aschaffenburg University, from Holter, Geis, Raab, and Bauke, tested LLM-assisted billing verification directly and found outcome agreement bounced around across different models, with no model showing a consistent edge over a plain end-to-end baseline. More telling: some of the raw judgments the models produced were correct, but lacked any real supporting evidence, and had to be downgraded to incomplete decision-basis entries rather than trusted outright. A right answer with no defensible reasoning behind it isn't something a billing office can stand behind in an appeal.

None of this means the models are useless. It means the benchmark hype creates a kind of false confidence that a careful operation can't afford. The same study found that when documentation requirements were spelled out explicitly, all four models tested improved at identifying missing information. Keep the deterministic rule checks separate from the LLM's semantic read, and let the model say "I don't have enough here" instead of forcing it to guess.

The five denial categories that coding accuracy alone cannot prevent

Diagram: Where Denials Actually Come From. Visualizes: Visualize the breakdown of denial root causes to show that coding errors are only one of five major categories.

Coding is one failure point among several, and it isn't even the biggest one. Five categories account for roughly 75% of all denials: eligibility issues, missing information, missing prior authorization, coding errors, and timely filing misses. Coding errors sit in that list as a single entry. Roughly half of all denials trace back to front-end problems instead: eligibility checks that never happened, demographic data entered wrong, authorizations that were never secured. LLM coding tools have no reach into any of that.

Follow that through to its practical conclusion. A practice that installs an LLM coding assistant and stops there, without touching its prior authorization process, its eligibility verification step, or its registration accuracy, will not see the denial rate move the way the sales materials implied it would. Coding accuracy fixes coding errors. It doesn't fix a front desk that didn't check the patient's coverage before the visit.

Payer AI systems reviewing claims LLM tools help produce

The same wave of AI adoption reshaping how practices code claims is reshaping how payers review them, and that second half of the story is the one most practices haven't caught up to yet. Private insurers have rolled out algorithm-based claim review systems of their own, and Medicare introduced its WISeR pilot in 2026, bringing AI directly into the prior authorization decision. A practice that spent a year tightening its coding accuracy is now submitting those better-coded claims into a review process that has also changed, and not necessarily in its favor.

The overlap is not flattering. The tool built to solve coding hands the payer's AI a slightly different version of the same lookup to find.

There's also a visibility gap that makes all of this harder to size up. No federal rule requires commercial payers to disclose their denial rates. Medicare Advantage, Medicaid managed care, CHIP, and ACA exchange plans have to report; commercial fully-insured and self-insured employer plans do not. Practices are, in effect, negotiating with a system whose behavior they can only partly see. Payers' common denial grounds (incomplete documentation, lack of medical necessity, coding discrepancies) map directly onto the weaknesses in LLM coding output: stochastic variance, missing evidence spans, and documentation gaps the model did not flag.

The 2026 Regulatory Environment for Both Sides of That Dynamic

Regulators started closing some of that gap in 2026, though not all of it. CMS-0057-F took operational effect January 1, 2026, and it cuts the standard prior authorization decision window in half: covered payers now have to decide within 72 hours for urgent requests and 7 calendar days for standard ones (for most covered payers, excluding QHP issuers on Federally-Facilitated Exchanges), down from the old 14-day standard. The rule reaches Medicare Advantage organizations, state Medicaid and CHIP fee-for-service programs, Medicaid managed care plans, CHIP managed care entities, and QHP issuers on the federal exchanges. Drugs are carved out of it entirely. A companion rule, CMS-0062-P, released April 10, 2026, would extend similar requirements to drug authorizations, but as of the most recent reporting it remains a proposed rule, not a finalized one.

Transparency picked up a bit too. Covered payers now have to publish their prior authorization metrics once a year, starting March 31, 2026, which is the first time anything resembling a public reporting requirement has touched this population, even though commercial plans still sit outside it. Four FHIR APIs, covering patient access, provider access, payer-to-payer data exchange, and prior authorization, are required under the same rule, but payers have until January 1, 2027 to get the APIs running, even as the decision timeframes, denial-reason requirements, and public reporting obligations have already been in force since January 1, 2026.

States are moving on a separate track, and a few have gone further than the federal rule does. States including Texas, Arizona, and Maryland now require human review of adverse prior authorization decisions that an AI system generates, which puts a direct regulatory check on algorithmic denial.

On the coding side of the ledger, the AMA's 2026 CPT code set made a structural expansion of codes that explicitly recognize AI-augmented services, sitting alongside the existing codes for services a clinician performs alone. That's a structural shift. Revenue cycle teams need to fold it into their coding logic to capture the reimbursement it makes available, and to avoid a denial category that didn't exist the year before.

The clean claim rate is still the number that determines whether any of this matters

The clean claim rate is the number that cuts through all of the above, and it's worth starting with what "good" looks like. Industry benchmarks put a solid clean claim rate at 95% or higher, with the strongest performers going even higher than that. Fall below that line and it's a sign of a systemic upstream problem, usually sitting in coding or in eligibility.

A practice sitting below that threshold is generating a heavy volume of rework every single billing cycle: staff hours spent chasing corrections, payment delayed while claims sit in limbo, and mounting exposure to denials that compound month over month. MGMA data shows top-quartile practices holding denial rates well under the industry average, while the broader industry saw initial denials reach roughly 11.8% in 2024, with projections pointing higher still for 2025.

LLM coding assistance only moves the needle on the coding-error slice of clean claim rate when treated as a single aggregate figure. The other four denial categories, eligibility, missing information, missing authorization, timely filing, need their own fixes entirely. A practice tracking clean claim rate only in aggregate can watch the topline number barely budge after a coding tool rollout and wrongly conclude the tool failed, when really it did exactly the narrow job it was built for. Days in AR under 30 is strong; 31–40 is tolerable; over 50 is a red flag, and every day a denied claim sits unworked is a day the window to recover revenue narrows.

Good LLM-assisted billing operations with both AI layers accounted for

The Aschaffenburg findings point toward an architecture that holds up well beyond that one proof-of-concept study. Keep the deterministic rule checks, code availability, quantity limits, and exclusion criteria running separately from the LLM's semantic judgment: the documentation must actually support the claimed service. Require the model to abstain when the evidence isn't there instead of filling the gap with a guess. That separation is what makes the model's contribution something a human can actually inspect and audit, rather than a black box that occasionally gets lucky.

Fail-closed abstention sounds like a small engineering choice, but it changes what happens downstream. A system that flags "insufficient documentation" hands a billing team a workable exception queue. A system that guesses instead hands them a denial they'll have to unwind later, usually with less information than they had going in.

None of this replaces the people doing the work. LLMs are good at volume and pattern recognition across thousands of claims a week. The exceptions those systems surface, the appeals a payer's own AI generates, and the policy changes that appear without warning still need in-house billing specialists who carry payer-specific institutional knowledge. Payer rules shift without notice: the same payer AI that denied a claim last week may be running a different rule next week, and a practice with no one tracking that pattern is always finding out after the money is already gone.

The 2026 CPT changes add one more layer that can't be skipped. Practices now have to correctly separate AI-assisted services from clinician-only services in both documentation and coding, or they'll run into a denial category that didn't exist before, one that an LLM tool trained on last year's code set has no way of catching on its own. That's not a hypothetical risk. It's a direct consequence of the code set change already sitting in front of every RCM team.

The practices that come out ahead in this environment tend to share one habit: they integrate AI assistance with existing EMRs and clearinghouses rather than ripping out workflows, which lets them apply AI assistance without retraining front-desk staff and preserves the operational continuity that makes adoption sustainable. That approach lets a coding tool do its job without forcing front-desk staff through a retraining cycle, and that continuity, more than any single feature of the software, is what makes adoption actually stick.

Sources

  1. Aligning AI-Powered Coding With 2026’s New Season
  2. AI in Medical Coding and Billing: Use Cases, Risks and Opportunities | Uptech
  3. Flaws in the LLM Automation Narrative
  4. A decision-basis contract for auditable LLM-assisted medical billing verification: deterministic rules, verbatim evidence, and fail-closed abstention
  5. AI CPT Codes 2026: Updates for Medical Practices
  6. Prior Authorization in 2026: IRA, CMS Reform & AI Automation
Filed underAI in Claims

More in AI in Claims