🤖 AI Summary
This paper addresses the challenge of predicting incurred-but-not-reported (IBNR) claim frequencies in non-life insurance. We propose a survival analysis–machine learning hybrid modeling framework leveraging individual claim-level data (e.g., accident date, reporting delay). Innovatively, we formulate a risk function estimation framework under reversed development time, unifying the Cox proportional hazards model, feedforward neural networks, and XGBoost within a single coherent structure. Our method employs partial likelihood estimation to jointly handle left truncation and tied event times, ensuring internal consistency of development factors across monthly, quarterly, and annual granularities. Unlike traditional chain-ladder methods reliant on aggregated time-series data, our approach achieves significantly improved prediction accuracy and stability on both synthetic and real-world datasets. A production-ready R implementation is provided via the open-source *ReSurv* package.
📝 Abstract
We introduce new approaches for forecasting IBNR (Incurred But Not Reported) frequencies by leveraging individual claims data, which includes accident date, reporting delay, and possibly additional features for every reported claim. A key element of our proposal involves computing development factors, which may be influenced by both the accident date and other features. These development factors serve as the basis for predictions. While we assume close to continuous observations of accident date and reporting delay, the development factors can be expressed at any level of granularity, such as months, quarters, or year and predictions across different granularity levels exhibit coherence. The calculation of development factors relies on the estimation of a hazard function in reverse development time, and we present three distinct methods for estimating this function: the Cox proportional hazard model, a feed-forward neural network, and xgboost (eXtreme gradient boosting). In all three cases, estimation is based on the same partial likelihood that accommodates left truncation and ties in the data. While the first case is a semi-parametric model that assumes in parts a log linear structure, the two machine learning approaches only assume that the baseline and the other factors are multiplicatively separable. Through an extensive simulation study and real-world data application, our approach demonstrates promising results. This paper comes with an accompanying R-package, $ exttt{ReSurv}$, which can be accessed at url{https://github.com/edhofman/ReSurv}