Time-to-Event Estimation with Unreliably Reported Events in Medicare Health Plan Payment

📅 2026-02-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges posed by unreliable diagnosis coding in Medicare Advantage plans, which can lead to payment inaccuracies and upcoding fraud, while traditional survival analysis methods struggle with competing risks and non-proportional hazards. Building upon the restricted mean time lost (RMTL) framework, this work extends RMTL for the first time to the context of Medicare payment policy, introducing novel estimators robust to misreported events. The authors also develop the first open-source R package that simulates realistic upcoding behavior. Through Monte Carlo simulations and validation using All of Us data, the proposed approach effectively identifies upcoding practices potentially costing up to $40 billion annually, offering regulators a reproducible and scalable analytical tool for Medicare oversight.

Technology Category

Reasoning under Uncertainty: CausalityMachine Learning: Calibration & Uncertainty QuantificationMultiagent Systems: Mechanism Design

Application Category

Economics, Online Markets and Human Computation: Trust and reliance of crowd workers and data experts on GenAIUser Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systemsResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
Time-to-event estimation (i.e., survival analysis) is common in health research, most often using methods that assume proportional hazards and no competing risks. Because both assumptions are frequently invalid, estimators more aligned with real-world settings have been proposed. An effect can be estimated as the difference in areas below the cumulative incidence functions of two groups up to a pre-specified time point. This approach, restricted mean time lost (RMTL), can be used in settings with competing risks as well. We extend RMTL estimation for use in an understudied health policy application in Medicare. Medicare currently supports healthcare payment for over 69 million beneficiaries, most of whom are enrolled in Medicare Advantage plans and receive insurance from private insurers. These insurers are prospectively paid by the federal government for each of their beneficiaries'anticipated health needs using an ordinary least squares linear regression algorithm. As all coefficients are positive and predictor variables are largely insurer-submitted health conditions, insurers are incentivized to upcode, or report more diagnoses than may be accurate. Such gaming is projected to cost the federal government $40 billion in 2025 alone without clear benefit to beneficiaries. We propose several novel estimators of coding intensity and possible upcoding in Medicare Advantage, including accounting for unreliable reporting. We demonstrate estimator performance in simulated data leveraging the National Institutes of Health's All of Us study and also develop an open source R package to simulate realistic labeled upcoding data, which were not previously available.
Problem

Research questions and friction points this paper is trying to address.

upcoding
Medicare Advantage
unreliable reporting
healthcare payment
coding intensity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Restricted Mean Time Lost
upcoding detection
competing risks
Medicare Advantage
simulation framework
O
Oana M. Enache
Stanford University School of Medicine, Department of Biomedical Data Science, Edwards Building, 300 Pasteur Drive, Stanford, CA, 94304
Sherri Rose
Sherri Rose
Professor, Stanford University
statisticshealth policymachine learningcausal inferencecomputational health economics