A healthcare claim is built to be adjudicated as a transaction. It identifies a member, a provider, a service, a date, a diagnosis, a procedure and an amount. The claim may be complete. The codes may be valid. The service may be covered. Nothing on its face necessarily explains why the payer should hesitate.

Yet many payment integrity problems are not visible on the face of a claim. They emerge in repetition, concentration, sequence and comparison. A modifier appears at an unusual frequency. Units accumulate just below an edit threshold. A procedure mix drifts away from that of credible peers. Diagnoses repeatedly justify a more lucrative service. The same member pattern appears across related providers. Each claim can remain plausible while the provider pattern becomes increasingly difficult to explain.

This is the gap AI is beginning to close. The most important change is not a more sophisticated score applied to an isolated claim. It is the connection of two decisions that most payers still manage separately. The first asks what can be learned from the provider’s history. The second asks what should happen to the next claim before money leaves the payer.

Payment integrity improves when the answer to the first question changes the answer to the second.

The claim is designed to look ordinary

Claims adjudication has long relied on deterministic controls. Eligibility, coverage, coding, pricing, duplication and authorization rules decide whether a transaction satisfies known requirements. Correct coding edits remain essential. CMS describes Procedure to Procedure edits as controls against inappropriate payment for services that should not be reported together, while Medically Unlikely Edits address incorrect units of service.[1]

These controls are precise because they encode known relationships. Their precision is also their boundary. They can identify an invalid combination or a unit count that breaches a defined limit. They are less able to determine whether a valid modifier is being used with an improbable frequency, whether a provider’s coding intensity has changed in a clinically implausible way, or whether several individually acceptable claims form a concerning longitudinal pattern.

The distinction matters because improper payment is broader than fraud. CMS explicitly cautions that improper payment measurement is not a measure of fraud and that not every improper payment is attributable to fraud or abuse.[2] A claim may reflect a documentation gap, a coding error, a policy misunderstanding, an unusual but legitimate case mix, wasteful practice, abuse or deliberate deception. An analytical system that collapses those possibilities into a declaration of fraud will create false certainty precisely where judgment is required.

The purpose of AI is therefore not to replace rules with a black box. It is to add context that rules cannot see and to help the payer choose the next proportionate action.

The provider pattern changes what the claim means

Consider two claims containing the same high-paying modifier. On the first claim, the modifier is consistent with the provider’s specialty, the member’s recent care and the provider’s established billing profile. On the second, it follows a sudden shift in coding behavior, appears at several times the peer-group rate and is concentrated in combinations that have previously produced unsupported payments.

The claim-level facts may be nearly identical. The provider context is not. Once that context is available, the payer is no longer asking only whether the claim is technically payable. It is asking whether the claim remains credible in light of what is known now.

This provider view depends on accumulations rather than snapshots. It may include rolling rates for procedures and modifiers, changes in coding intensity, ratios between diagnoses and services, concentrations by member or location, temporal bursts, reversal and correction behavior, peer-group deviation, referral relationships, prior review outcomes and financial exposure. The objective is not to produce the longest possible feature inventory. It is to represent the few aspects of behavior that materially change the payment decision.

Official program integrity work already demonstrates the importance of pattern. HHS OIG has recommended analyzing supplier billing patterns to determine whether additional prepayment or postpayment review is warranted.[3] Its studies have also identified providers whose billing for telehealth differed sharply from broader populations.[4] These examples reinforce a practical truth. The unit of observation may be the claim, but the unit of suspicion is often the provider, network or episode.

Retrospective intelligence arrives after the economic decision

Most payers already conduct some form of retrospective analytics. They examine paid claims, compare providers with peers, rank anomalies and send cases to investigators. This can improve recoveries and reveal emerging schemes. It can also create the comforting impression that AI is already embedded in payment integrity.

The economic limitation is timing. By the time an investigator sees the pattern, the payer has already adjudicated and paid many of the claims that made the pattern visible. Recovery is slower and less certain than prevention. Investigators must reconstruct evidence, contact providers, manage disputes and pursue funds that have left the organization. The analytics may be effective while the operating model remains reactive.

CMS has used predictive analytics to generate alerts on claims and providers for program integrity review, and GAO has repeatedly examined the extent to which such analytics enable quicker investigation and more timely corrective action.[5] Prepayment review programs similarly examine selected claims before payment where historical evidence indicates elevated risk.[6] The strategic lesson is not that every claim should be delayed. It is that retrospective learning should determine where friction is worth introducing before payment.

That changes the role of the retrospective environment. It is no longer merely a place to find yesterday’s anomalies. It becomes the intelligence layer for today’s adjudication decisions.

The same feature must mean the same thing twice

Connecting the two environments is more difficult than moving a model from a notebook into a transaction system. Retrospective analytics may calculate a provider’s modifier rate across two years of paid claims. Operational adjudication may need the same measure in milliseconds while the latest claims are still being processed. One environment can scan history. The other must maintain current state.

The physical calculation can differ. The meaning cannot.

If a feature called provider modifier rate uses one population, time window, claim status or denominator during model development and another at runtime, the payer has not operationalized the same intelligence. It has created two similarly named measures that can drive materially different decisions. The model may appear healthy in retrospective testing and fail quietly in production because its operational inputs no longer represent what it learned.

Feature parity is therefore a business control as much as an engineering discipline. Each feature needs a stable semantic definition, a point-in-time rule, an accountable owner and evidence that its retrospective and operational implementations reconcile. Historical paid claims may establish accumulations. Newly adjudicated claims may update them. Corrections, reversals, late submissions and provider identity changes must be handled without leaking future information into past decisions or counting the same event twice.

This is where a common claims and feature contract becomes strategically important. Payers do not share one source schema. One may separate claim headers and lines. Another may flatten them. Identifiers, statuses, code conventions and histories vary. Reuse does not come from pretending those differences do not exist. It comes from normalizing each source into a canonical representation, preserving provenance and then generating the same decision features from that representation.

The reusable asset is not a universal fraud model. It is the hard-won scaffolding that converts payer data into trustworthy, point-in-time provider and claim context. Policies, peer groups, thresholds, calibration and workflow decisions remain payer specific. The architecture becomes portable without making the judgment generic.

A provider anomaly is not a claim denial

Anomaly detection identifies difference, not wrongdoing. A provider may be unusual because it treats a rare population, operates a specialist referral center, adopted a new therapy, acquired another practice or corrected earlier undercoding. Even a strong statistical deviation does not establish that a specific claim is unsupported.

This is why the output of provider intelligence should be an integrity disposition rather than a verdict. Most claims should continue to be released. Some should be held for a targeted documentation request. Some should be routed to specialist review. Others may be paid while the provider enters a focused postpayment investigation. The action should reflect the strength of the signal, the financial exposure, the reversibility of the decision and the likely burden imposed on legitimate providers and members.

Government practice illustrates the principle of proportionality. CMS medical review programs provide detailed reasons when a reviewed claim is denied or not affirmed, while targeted review and education programs are intended to reduce future denials and appeals through provider-specific assistance.[7] The broader lesson for payers is that payment integrity is not only about stopping the wrong payment. It is also about explaining decisions, correcting behavior and avoiding unnecessary friction for providers who are billing appropriately.

AI should therefore recommend the next best control, not simply maximize the number of claims it interrupts. A low-cost documentation request may be appropriate for one signal. A clinical review may be necessary for another. A network investigation may be more useful when the pattern spans several entities. Automation earns trust when it narrows human attention and preserves an auditable path from evidence to action.

Payment integrity needs more than one analytical method

No single model can represent the full problem. Deterministic edits remain the right instrument for explicit coding and policy violations. Supervised models can recognize patterns associated with previously confirmed outcomes. Unsupervised methods can surface emerging behavior that has no reliable label. Graph analytics can reveal relationships among providers, members, addresses, referrals and billing entities. Language models can help reviewers navigate policies and summarize supporting records, but they should not manufacture evidence or substitute fluent explanation for a defensible decision.

The strategic design is a layered decision system. A claim enters adjudication and passes through established policy controls. Current claim facts are combined with point-in-time provider, member and network context. Analytical methods contribute distinct signals. A decision layer then weighs risk, exposure, operational capacity and policy to choose release, hold or review.

This avoids two common failures. The first is asking a model to rediscover rules that the payer already knows and can apply deterministically. The second is treating every anomaly as equally urgent. The most useful AI does not compete with the claims engine. It makes the claims engine context aware.

The feedback loop must learn from uncertainty

Once provider intelligence influences live claims, every intervention becomes a potential learning event. A documentation request may confirm the service. A clinical reviewer may find a coding issue. A provider may change its billing behavior after education. An investigation may uncover a broader network. An appeal may overturn the original decision.

Those outcomes need to return to the intelligence layer. Otherwise, the system learns only from what it chose to inspect and may reinforce the biases of earlier review strategies. Providers that were never selected remain underrepresented. Cases closed for lack of evidence may be mistaken for clean claims. Confirmed errors may be labeled as fraud even when intent was never established. A mature feedback loop preserves these distinctions and records uncertainty rather than forcing every outcome into a binary label.

Governance must extend across the full loop. Payers need to know which data were used, which version of a feature and model influenced the disposition, which policy was applied, who reviewed the case and what happened after appeal or investigation. Protected health information may be used for payment and healthcare operations under HIPAA within applicable limits and safeguards, but that permission does not remove the need for purpose limitation, access control, security and accountable data handling.[8]

The objective is not merely model monitoring. It is decision monitoring. A model can remain statistically stable while review queues become unmanageable, provider abrasion rises or savings concentrate in cases that would have been caught by existing rules. Leaders should measure incremental prevented loss, investigative yield, review burden, time to resolution, appeal outcomes and unintended effects across provider and member populations.

The advantage is a learning claims decision system

Many payers have already bought analytics, built rules and established investigation teams. The next source of advantage will not come from adding another isolated model. It will come from connecting those capabilities around a common decision.

That requires a different operating ambition. Payment integrity cannot remain a retrospective department that sends findings to claims operations. Claims operations cannot remain a transaction factory that sees only the fields on the current claim. Data science cannot optimize a score without owning how that score changes workflow. Clinical, policy, investigative and engineering teams must share the same definitions of provider behavior, intervention and outcome.

When those pieces connect, the payer begins to compound what it learns. Retrospective analysis identifies a meaningful provider pattern. The pattern is expressed through governed features. Those features are made current at adjudication. The next claim receives a proportionate disposition. The outcome improves the next retrospective analysis. Each decision leaves the system better prepared for the one that follows.

This is how AI is reshaping claims and provider integrity. It is moving the payer from recovering value after a pattern becomes undeniable to acting while the next payment can still be influenced. The claim remains the transaction. The provider pattern supplies the intelligence. The advantage belongs to the payer that can join them before the money moves.

References

  1. Centers for Medicare and Medicaid Services, National Correct Coding Initiative for Medicare
  2. Centers for Medicare and Medicaid Services, Fiscal Year 2025 Improper Payments Fact Sheet
  3. US Department of Health and Human Services Office of Inspector General, Medicare remains vulnerable to improper payments for off-the-shelf orthotic braces
  4. US Department of Health and Human Services Office of Inspector General, Medicare telehealth services and program integrity risks
  5. US Government Accountability Office, CMS use of data analytics to identify and prevent fraud
  6. Centers for Medicare and Medicaid Services, Recovery Audit Program Prepayment Review Demonstration
  7. Centers for Medicare and Medicaid Services, Review Reason Codes and Statements
  8. US Department of Health and Human Services, Uses and disclosures for treatment, payment and healthcare operations