AI Denial Prediction Models and Real-Time Claim Routing in RCM Platforms
Payers now deny claims at machine speed; practices need AI to keep up.

The revenue cycle's denial problem has changed in kind, not just in size: payers now adjudicate claims through automated systems that operate at machine speed, while most practices still respond to denials one at a time, after the fact. That mismatch in velocity, not any single payer policy, is the structural reason predictive AI has moved from a competitive advantage to an operational necessity.
The revenue cycle's denial problem is getting structurally worse, not just bigger
Commercial payers including Aetna, UnitedHealthcare, and several Blue Cross Blue Shield affiliates have rolled out automated retrospective reviews that pull back payments weeks to months after a claim was originally adjudicated. Practices that treat the moment of submission as the end of the denial risk window are mistaken, because the window now extends well past the date a claim clears. The scale one payer's system can operate at illustrates the gap: EngineerBabu documented a single payer system rejecting hundreds of thousands of claims within two months, a volume that no manual denial workflow can answer on equal terms. A billing team built to review claims by hand, one at a time, cannot keep pace with an adjudication engine built to reject at that rate.
The time allotted for practices to respond has also compressed. Shorter response windows mean a practice working denials manually has less time to catch an error before an appeal deadline passes, regardless of whether the underlying clinical case was sound.
Prior authorization denial volume has risen sharply year over year across both commercial and Medicare Advantage payers, a trend driven by longer lists of services requiring prior authorization and by AI-assisted adjudication on the payer side. The 2026 CMS WISeR Model expanded prior authorization requirements into Traditional Medicare, building on the narrower prior authorization programs that already existed there and opening a denial surface that simply did not exist for those claims before.
None of this reflects bad faith on the part of payers so much as a straightforward asymmetry in tooling. Payers appeal to practices that don't contest denials, and most practices don't: fewer than 0.2% of denied ACA marketplace claims went to internal appeal, even though policyholders who did appeal won the majority of the time. Among Medicare Advantage payers, standard appeal overturn rates vary enormously from one payer to the next, a spread that shows how much the outcome of an appeal depends on which payer issued the denial in the first place and whether anyone worked the appeal at all. Practices that don't appeal aren't losing because their claims lacked merit. They are losing because nobody on the payer side is incentivized to volunteer money back, and few practices have the staff time to ask for it systematically.
This pressure compounds against a tightening margin environment. Operating expenses climbed significantly from 2024 to 2025 while the Medicare conversion factor fell, so every preventable denial now costs more in lost revenue and in the labor required to rework it than it did two years earlier. Commure's analysis of one healthcare organization found millions of dollars in annual revenue lost, representing a meaningful share of that organization's annual recurring revenue, traced simply to the absence of infrastructure capable of triaging, correcting, and resubmitting denied claims efficiently. That figure describes one organization's experience rather than an industry average, but it illustrates the order of magnitude at stake when denial management remains a manual, reactive function.
Where denials originate
The scale of the problem matters less than its origin, and the data on origin is reassuring in one specific sense: most denials are not clinical disputes. They are administrative and documentation errors that could, in principle, have been caught before the claim ever left the practice. That distinction turns denial management from a legal or clinical argument into a process engineering problem, which is the premise every AI intervention described later in this piece depends on.
Experian Health's 2025 State of Claims survey, which polled 250 revenue cycle management leaders, identifies the top three denial causes as missing or inaccurate claim data, authorization complexity, and incomplete patient registration. Three out of every four denials trace back to paperwork or plan design rather than to a disagreement over medical necessity. Missing or inaccurate claim data was cited as a top cause by the largest share of respondents in that survey, and that share grew from the prior year, a trend the survey connects to the constant churn of annual CPT and ICD-10 code updates and to ongoing Medicaid redeterminations that destabilize patient eligibility data.
The error patterns behind those denials vary across practices. Commure's site-level data found one practice with nearly half its denials attributed to CARC 59, the code covering multiple or concurrent procedure rules such as multiple surgery, diagnostic imaging, or concurrent anesthesia billing. A second practice in the same dataset saw nearly three-quarters of its denials attributed to CARC 104, which covers managed care withholding typically tied to inadequate or incomplete supporting documentation, and that single denial category cost the second practice a substantial sum in just one quarter. Two practices, two almost entirely different denial profiles. That variation explains why a single, generic set of claim-scrubbing rules applied uniformly across a specialty or a region will always miss a large share of preventable denials: the error pattern is specific to the practice and specific to the payer, not something that holds true industry-wide.
Specialty compounds the variation further. Among common specialty types, behavioral health carries the highest denial rate and the longest accounts receivable days, while orthopedics and ambulatory surgery centers occupy a middle tier. Each specialty's particular combination of denial rate and days in AR defines its own recovery challenge. A prediction model tuned for one specialty's claim patterns will not transfer cleanly to another. Practices that perform best in class hold their denial rates well below the broader practice average and their clean claim rates well above it, and the space between that benchmark and the average is, by definition, the preventable portion of the problem.
Part of what keeps that preventable portion so persistent is the annual churn in coding rules itself. The 2026 CPT and ICD-10-CM updates introduced hundreds of new code changes, and the 2026 Medicare Physician Fee Schedule alone added more than 400 new CPT code changes along with a negative efficiency adjustment affecting thousands of specialty services. Each update shifts the logic payers use to adjudicate claims and opens new opportunities for mismatch between what a practice bills and what a payer's system expects to see. For an orthopedic practice with substantial annual Medicare revenue, the efficiency adjustment alone can mean a loss of tens of thousands of dollars a year, before any cost of reworking the resulting denials gets added on top. Code updates arrive on a predictable annual calendar, yet a predictable event still produces a denial spike every year in practices without automated validation in place to catch the mismatch before submission.
What AI denial prediction models score
Denial prediction, done properly, is a supervised learning system trained on a specific practice's own claim history and its own record of how specific payers have responded to those claims, producing a score for each new claim based on the variables known to predict denial for that payer and that practice.
The architecture behind that score typically relies on gradient boosting methods such as XGBoost, trained on 12 to 24 months of a practice's own historical claim and denial data instead of pooled, industry-wide data. That choice of training data is deliberate: denial patterns are payer-specific and practice-specific, as the Commure site-level numbers above demonstrate, so a model trained on another organization's claims would be scoring against the wrong distribution. The features that go into the model include procedure code, diagnosis code, payer, plan type, place of service, day of week, provider NPI, and prior authorization status, and it's the combination of these factors together, not any single field in isolation, that produces the denial probability score attached to a claim.
Named systems in production illustrate how this plays out in practice. Commure's platform builds a knowledge graph from encounter-level EHR data, remittance records, and insurance responses, and uses that graph to surface payer-specific denial patterns and recommend compliant CPT, ICD, and modifier codes based on which combinations have actually succeeded with a given payer in the past. Commure's AI agent, called Scout, flags errors and proposes corrections before a claim is ever submitted. Apexon's AgentRise takes a related but distinct approach, embedding payer-specific rules and APIs directly into the platform so that it adapts automatically as individual payers change their adjudication criteria, rather than requiring a billing team to manually rewrite scrubbing rules every time a payer's policy shifts.
Claims-data models alone cannot see into clinical documentation. Natural language processing extends the prediction layer further upstream into that territory. NLP engines suggest specific documentation language capable of satisfying a payer's current criteria, and generative AI presents that language to the provider before the clinical note is finalized, intervening before CARC 104-type denials, the kind tied to inadequate supporting documentation, ever have a chance to occur. Commure's large language model layer parses payer policy documents directly and surfaces coding recommendations with a 95% QA pass rate, testing data pulled from EHRs and giving billing staff precise, specific recommendations on coding denials and rejections rather than generic guidance. NLP also works on the back end of the process, interpreting payer Explanation of Benefits notes and remittance advice to improve the accuracy of denial reason coding after a denial has already landed. Apexon's platform applies this capability to enable faster, more precise action across millions of transactions.
Real-time claim routing across the five intervention points
Prediction only produces value once it is wired into a workflow that acts on the score, and that workflow is not a single checkpoint sitting in front of claim submission. It operates as a series of interventions positioned at every point in the revenue cycle where denial risk can enter a claim, starting at patient registration and continuing through the final scrub before the claim goes out.
The first intervention point sits at eligibility and coverage verification. Insurance information confirmed at the time of scheduling is frequently stale by the date of service, a problem made worse by post-ACA Medicaid redeterminations and by routine churn in commercial plan enrollment. AI-driven real-time eligibility checks run at every patient touchpoint to flag coverage anomalies before the appointment takes place. OhioHealth reduced registration and eligibility-related denials by 42% using Experian Health's Patient Access Curator platform, which stands as the most concrete published outcome tied to this particular intervention point.
The second intervention point addresses clinical documentation validation. NLP engines compare the content of clinical notes against current payer-specific medical necessity criteria at the moment of charge capture, before the encounter ever reaches a biller's desk. That timing matters because payer medical necessity criteria have tightened in 2026 and are now reviewed on the payer's own side, so documentation that would have satisfied a payer's standards in prior years can fail automated adjudication today.
The third intervention point audits coding accuracy. An AI coding audit validates the codes selected against the clinical documentation on file, identifies modifiers that are missing, and flags code combination edits, including Correct Coding Initiative edits and medically unlikely edits, before the claim itself is even assembled.
The fourth intervention point turns the prediction model described in the previous section into an operational decision rather than just a score. The model scores each fully assembled claim, routing high-risk claims to human review while allowing low-risk claims to pass through automatically. At this stage a claim is either cleared outright, flagged for a biller to review with specific context about what triggered the flag, or held pending prior authorization validation. Apexon's AgentRise performs this function across a full set of checks at once: real-time claim scrubbing, missing documentation detection, prior authorization validation against each payer's specific requirements, eligibility verification at the point the claim is created, and coding and modifier accuracy checks aligned to both CMS rules and commercial payer rules, all completed before the claim is submitted.
The fifth intervention point picks up after a denial has already landed. Automated categorization sorts denials by root cause and by payer, auto-generated appeals assemble the relevant supporting clinical documentation, and payer-specific appeal templates move the claim back into active workflow immediately rather than letting it sit in a queue waiting for staff attention. Commure's AI Denial Automation System automates the majority of denied claim reprocessing under this model, decreasing the labor cost of denial resubmissions, increasing the volume of claims that actually get resubmitted, and improving denial re-approval rates significantly. The more complicated cases, the ones genuinely involving a dispute over medical necessity or a question of payer policy interpretation, still route to an in-house biller or specialist who gets the full context the AI system has already assembled. The hybrid model, combining automated triage with human judgment on the hard cases, consistently outperforms either an all-AI or an all-manual approach working alone.
The data foundation that makes routing and prediction work
A specific kind of data infrastructure produces this prediction and routing, and most practices evaluating these platforms underestimate that layer. A model trained on claims history that is fragmented across disconnected systems produces fragmented, unreliable scores no matter how sophisticated its algorithm is. The data architecture a practice has in place matters as much as the model sitting on top of it.
The technical foundation that makes accurate scoring possible is a centralized data layer that consolidates EHR and EMR data into one unified data warehouse, ingesting every claim submitted, every denial received, every payer response, and every appeal outcome into a single continuous record. A model built on that kind of foundation sees the full history of how a given payer has treated a given practice's claims over time. A model built on snapshots, pulled inconsistently from whatever system happened to be queried at a given moment, sees only fragments of that history, and its predictions degrade accordingly. For a practice evaluating whether to adopt an AI denial prediction system, the first diagnostic question is whether the practice's own data infrastructure can support one. It is whether the practice's own claims, remittance, and appeal data already lives in a form a model can actually learn from, because the gradient boosting architecture, the NLP layer, and the five-point routing system described above all sit on top of that foundation, and none of them can compensate for its absence.


