Est.
AI in ClaimsLong read

AI Fraud Detection Models and False Positive Denial Rates for Independent Practices

AI models flag legitimate claims as fraud because they lack practice-specific context.

Editor at Large · · 12 min read
Cover illustration for “AI Fraud Detection Models and False Positive Denial Rates for Independent Practices”
AI in Claims · September 29, 2026 · 12 min read · 2,786 words

AI Fraud Detection Models and False Positive Denial Rates for Independent Practices.

How payer AI fraud detection flags legitimate claims

Government and commercial payers have moved toward pre-payment AI fraud detection that flags billing patterns before payment, creating denials that are technically false positives (legitimate claims caught by algorithms trained on population-level anomalies rather than practice-specific context), and understanding how these models work is the first step to fighting back. The federal posture change is instructive here. HHS published a Request for Information on February 25, 2026, seeking industry input on AI-driven fraud prevention, and CMS and the DOJ have both signaled plans to deploy AI models that flag claims before a dollar goes out the door medicaleconomics.com. That's a shift from "pay and chase," the old model where fraud got investigated after payment, to something regulators are calling "detect and deploy."

Machine learning fraud scoring works by training on aggregate claims data pulled from thousands of providers and payer populations. The model looks for statistical anomalies: procedure frequency, billing intensity, specialty benchmarks, referral patterns, all measured against a peer cohort. What the model does not have is any visibility into why a specific practice looks the way it does. It doesn't know that a practice serves an unusually complex patient panel, or that it operates in a rural market with a narrow specialty mix, or that its patient volume swings seasonally for reasons that have nothing to do with fraud. So when that practice's billing pattern deviates from the population average, the algorithm has no mechanism to distinguish a legitimate outlier from an actual bad actor.

Shannon Sumner, CPA, CHC and managing principal of PYA's Nashville office, said investigations increasingly get triggered by analytics, so practices get flagged simply because their data doesn't look like their peers'. That observation applies just as much to payer adjudication engines as it does to government audits. Two separate arenas are damaging practices here: government enforcement exposure, which moves slowly and carries legal risk, and payer-side claim denials, which hit revenue immediately. Both deserve scrutiny, and both trace back to the same structural weakness in how these models are built.

The false positive gap between rule-based and ML fraud detection systems

Diagram: False Positive Rates: Rule-Based vs. Machine Learning Fraud Detection. Visualizes: Show the contrast between two fraud detection system types on a single visual dimension: false positive rates.

The data on false positive rates tells its own story. Legacy rule-based fraud detection systems, the kind built on fixed thresholds and hardcoded logic, produce false positive rates between 30% and 50% medicaleconomics.com nirmitee.io. Machine learning-based fraud scoring, by contrast, operates under 10% medicaleconomics.com nirmitee.io. That's a meaningful gap, and on paper it looks like clear evidence that the industry should race toward ML adoption as fast as possible.

The complication is that most payers haven't finished that migration. Adjudication engines across the industry are sitting in a mixed state, some claims running through legacy rule-based logic and others through newer ML scoring, often within the same payer's systems. Practices end up absorbing high volumes of false positive denials generated by the older rule-based systems, all while the payer markets its overall claims process as AI-powered and modern medicaleconomics.com. That branding gap matters because it shapes how a practice interprets a denial. A denial from a system marketed as intelligent reads differently to office staff than a denial they know came from a blunt, rule-based filter, even when the underlying accuracy problem is identical.

Even fully migrated ML systems carry their own false positive risk, just a smaller one. They're trained on population data, not on the specific context of a given practice, and small practices with limited claim volume are statistically more likely to look anomalous than a large health system whose sheer volume smooths out the edges. Model drift adds another layer: fraud patterns evolve, and a model trained on last year's historical data can flag a legitimately growing service line as suspicious simply because it wasn't common before.

Then there's the explainability failure, arguably the most frustrating part of this system. When an AI-assisted adjudication engine denies a claim, the denial reason code the practice receives, something like CO-16 or CO-197, reflects the system's output, not its reasoning. A billing staffer sees a procedural code. They do not see a note explaining that an algorithm flagged the claim's billing pattern as statistically unusual. That opacity makes AI-generated false positives harder to fight than an old-fashioned manual-review denial: the appeal has to respond to a surface-level code whose real cause, the statistical flag, the payer never discloses.

False positive denial concentration in 2026 (Medicare Advantage and prior authorization)

Diagram: The MA Denial Funnel: Volume, Reversals, and Abandoned Appeals. Visualizes: Visualize the Medicare Advantage prior authorization denial pipeline as a three-stage funnel or flow.

Medicare Advantage is the clearest place to watch this play out. KFF reported that MA insurers made nearly 53 million prior authorization determinations in 2024, a volume so large that even a modest false positive rate translates into an enormous number of wrongly denied claims.

The appeal overturn rate is the number that should trouble anyone paying attention: 80.7% of MA prior authorization denials, when actually challenged, get reversed. Four out of five denials, once a human looks closely, turn out to have been medically necessary all along. That's systemic evidence that the initial denial decision was wrong far more often than it was right.

Layer onto that the disruption in the MA market for 2026. Roughly 2.9 million Medicare beneficiaries were forced to switch MA plans this year because of plan withdrawals and service area reductions. Prior authorizations obtained under a beneficiary's old plan don't carry over to the new one, so patients have been showing up for scheduled, medically appropriate care carrying authorizations that technically no longer apply. Practices in the markets hit hardest by these transitions saw denial spikes of 20% to 30% above their baseline, driven entirely by authorization gaps tied to plan churn, not by anything wrong with the clinical decision-making medicaleconomics.com.

Prior authorization denials across both commercial and MA payers rose 31% year over year in 2026, and now make up 34% of all first-pass claim denials, up from 22% in 2023 moneygeek.com MBC. Decisions that used to take several business days under human review now come back in hours, but at denial rates far higher than the human-reviewed process ever produced, according to the AMA's Prior Authorization Survey medicaleconomics.com. Whatever efficiency gain the faster turnaround offers gets completely erased by the accuracy loss medicaleconomics.com. MA denial rates now top 17%, more than double traditional Medicare, per Experian (the steepest year-over-year increase of any payer category at 4.8% (Experian)).

The broader denial environment's baseline for AI false positives

None of this happens in isolation. Denial rates across the industry were already climbing before AI adjudication accelerated the trend. Experian Health's State of Claims Report found that 41% of providers now report more than 10% of their claims getting denied, up from 30% in 2022 and 38% in 2024 medicaleconomics.com nirmitee.io. The industry-wide initial denial rate reached nearly 12% in 2024 and held near that level through 2025, according to HFMA Experian Health State of Claims Report medicaleconomics.com.

The dollar figures are hard to look past. $262 billion in medical claims get initially denied every year, and 65% of those denials are never resubmitted Experian Health State of Claims Report. The majority of denials, including the AI-generated false positives buried inside that pool, simply become written-off revenue that no one ever tries to recover Experian Health State of Claims Report. Commercial payers and Medicare Advantage plans drove most of the recent increase, while traditional fee-for-service Medicare stayed relatively stable, which lines up neatly with where AI adjudication has been deployed most aggressively.

ACA Marketplace transparency data from plan year 2024 gives a useful cross-section of where individual payers land. UnitedHealthcare posted roughly a 19.1% overall denial rate across 6.4 million claims moneygeek.com. Oscar Health came in at 25.3% moneygeek.com medicaleconomics.com. Molina Healthcare sat at 22% moneygeek.com. Across all of these payers, fewer than 0.2% of denied claims ever went to internal appeal moneygeek.com. That utilization rate is so low it more or less confirms what the resubmission data already suggests: practices are not systematically challenging denials, wrong or otherwise moneygeek.com.

Working a denial isn't free, either. The administrative cost of handling a single denied claim climbed to $57.23 in 2023, and at current denial volumes, that rework cost compounds on top of whatever revenue never gets recovered in the first place gomedicalbilling.com. Put together, independent practices are absorbing AI-generated false positives inside a market where the baseline denial rate already sits well above best-in-class thresholds, their staff bandwidth to work denials is limited, and most wrong denials will never see an appeal.

The billing patterns that AI fraud models flag most often at independent practices

Fraud models flag statistical outliers relative to a peer cohort, and independent practices are structurally more likely to look like outliers than large health systems, mostly because their sample sizes are smaller and their patient panels less predictable. A handful of specific billing patterns repeatedly trigger the flag.

Procedure frequency that runs above the specialty benchmark is one, a solo gastroenterologist, say, who performs more procedures per visit than the national specialty average simply because of who walks through the door. Billing intensity that spikes in a short window is another: a practice clearing a coding backlog, or adding a new service line, can look to a context-blind model like sudden upcoding. Referral network anomalies get flagged too, a practice that sends most of its referrals to one specialist or facility for reasons of geography or patient need, not financial relationship. Modifiers applied at rates above the peer average, legitimate use of modifier -25 or -59 driven by genuine patient complexity, can read as anomalous even when every claim behind them is clean. New service categories added without a payer-specific prior authorization on file get caught by AI adjudication as a missing authorization rather than what it actually is: a new clinical capability the practice is offering.

UnitedHealthcare's own history illustrates the risk well. The insurer has applied extensive prior authorization requirements in the past, though it now reports requiring prior authorization for just 2% of medical services and says it applies fewer Medicare Advantage prior auth requirements than any other insurer, actively working to reduce CO-197 exposure. Its proprietary clinical algorithms still generate CO-16 denials paired with medical necessity RARCs, and both denial types get issued faster and at higher volume under AI-assisted adjudication than they ever did under human review.

An enforcement-side parallel exists here directly. OIG and DOJ analytics flag practices using the same peer-comparison logic that payers use, which means a practice already drawing payer AI scrutiny for its billing patterns is sitting inside an environment where those same patterns could just as easily trigger a government data request. The two risks aren't separate tracks; they run on the same underlying logic. Regulators themselves have acknowledged as much. The February 2026 HHS RFI explicitly asked for stakeholder input on frameworks to ensure AI-driven fraud detection respects due process and minimizes false positives. That request for stakeholder input means the pattern-flagging practices are living through right now isn't a bug that fixes itself quickly.

The cost of false positive denials for an independent practice that does not fight back

The cost here doesn't announce itself as one clean number. It accumulates.

Start with the 65% resubmission gap: 65% of denied claims never get resubmitted at all, per Experian Health's report, and for AI-generated false positives specifically, that figure is probably higher, because the denial reason code doesn't clearly point to what needs fixing Experian Health State of Claims Report medicaleconomics.com. The average denied Medicare Advantage claim rose 22.4% between 2024 and 2025, landing around $1,000, according to MDaudit medicaleconomics.com. Each unworked false positive is a meaningful individual loss, one that repeats every time the pattern goes unaddressed MDaudit medicaleconomics.com.

Best-in-class practices run denial rates below 3% Experian Health State of Claims Report medicaleconomics.com. The industry average is 11.8%, and the gap between a practice at baseline and one at best-in-class performance represents a substantial share of annual revenue Experian Health State of Claims Report medicaleconomics.com. Independent practices face a specific set of structural disadvantages that widen this gap further. Most don't have dedicated denial management staff able to work AI-generated denials the moment they land. Most don't carry institutional memory of which payer AI flags reliably reverse on appeal versus which ones require a pre-submission fix. Appeal windows keep shrinking as payers deploy AI on their end, narrowing the recovery window faster than a manual billing process can keep pace with. And staff turnover, a near-constant reality in small practice administration, erases the payer-specific knowledge needed to write appeal letters that actually work against AI-generated denial types.

The real cost is the compounding effect of writing off an entire category of denials, not just the individual lost claim, one the practice never recognizes as false positives in the first place, because the denial code on the page looks procedural instead of algorithmic.

Operational steps for catching and responding to AI-generated false positive denials before revenue is lost

Recognizing the denial type is the entry point. AI-generated false positives cluster in specific CARC and RARC combinations, CO-16 paired with MA-130 or N-386, or CO-197 on its own, and they appear as sudden spikes in prior authorization denials rather than as scattered, isolated events medicaleconomics.com. Pattern recognition at the denial reporting level is where the response has to begin medicaleconomics.com.

Pre-submission work is the strongest defense available, because it stops the denial before it ever gets generated. Real-time eligibility verification catches plan transition gaps before a claim goes out the door, and given that 2.9 million MA beneficiaries switched plans for 2026, that specific category of false positive is one eligibility checks can surface in advance. Claim scrubbing that validates modifiers against NCCI edit tables addresses CO-4 and CO-16 denials before they ever reach the payer's AI model. Verifying prior authorization at the point of scheduling, rather than at the point of billing, turns authorization gaps into a scheduling problem instead of a denial management problem.

When a denial does land, the response discipline matters just as much. Work AI-generated denials the day they arrive rather than letting them sit in a weekly or monthly queue, because appeal windows are compressing and an 80.7% MA overturn rate means the money is recoverable if the appeal actually gets filed. Decode the RARC paired with the CARC to find the real root cause the model flagged: CO-16 alone tells you almost nothing actionable, but CO-16 paired with M-76 for a missing modifier, or MA-130 for missing medical records, tells you what to fix medicaleconomics.com. Document the clinical context the algorithm never had access to, letters of medical necessity, patient complexity notes, specialty-specific coding rationale. The appeal, in effect, is making the argument the model was never built to make.

Tracking denial patterns across payers over time turns this from a reactive exercise into a predictive one, surfacing which payer models produce recurring false positives on specific service types. That kind of payer-specific intelligence compounds, and it's what eventually lets a practice fix the problem before submission instead of appealing after the fact. The clean claim rate benchmark is 95% or higher, with top performers reaching 98% to 99% revenuesynergy.com. The distance between a practice's current rate and that benchmark is, more or less, the exact portion of AI-generated denials that better pre-submission workflow could eliminate before an appeal is ever necessary revenuesynergy.com. The AMA's 2025 Prior Authorization Survey finding that AI-driven PA decisions return faster but with dramatically higher denial rates is the evidence basis for investing in pre-authorization workflows rather than assuming faster payer responses mean cleaner adjudication medicaleconomics.com.

Closing the false positive gap with payer-specific AI knowledge in billing operations

Independent practices repeatedly run into the limitation of institutional memory. Knowing that UHC's clinical algorithms tend to generate CO-197 denials at particular rates, or that MA plan transitions produce authorization gaps that look procedural on the surface but aren't, takes accumulated, payer-specific experience. Most practice billing staff simply don't have the time to build that knowledge, and whatever they do build gets lost the moment they leave.

HFMA's analysis points to where AI actually earns its keep in this fight: not as a black box issuing denials, but as a tool inside the billing operation itself, one that recognizes payer-specific patterns fast enough to correct claims before submission rather than fight them after the fact. That's the real shape of the fix. The payers built pattern recognition into their adjudication engines first. Practices that want to hold their ground need the same discipline turned back the other way, applied to the payers' own behavior, claim by claim, denial code by denial code, until the false positives stop being a mystery and start being a known, addressable pattern.

Sources

  1. False Claims Act recoveries hit a record $6.8 billion in 2025 | Medical Economics
  2. Predict, prevent, perform: The AI evolution of denials management
  3. Healthcare Denial Trends and AI Playbook 2026
  4. kff.org
  5. ama-assn.org
Filed underAI in Claims

More in AI in Claims