🤖 AI Summary
Credit time series often exhibit spurious trailing zero-balance observations—termed “temporal contamination”—arising from systemic delays in account closure, which distort the true loan termination point and impair risk modeling.
Method: This paper proposes a data-driven small-balance thresholding method to automatically identify the genuine loan termination date by filtering out non-informative tail zeros. The approach integrates empirical distribution analysis, threshold optimization, and survival model–based diagnostics, calibrated and validated on South African residential mortgage data.
Contribution/Results: Unlike ad hoc or rule-based truncation, our method ensures statistical robustness while preserving business interpretability—constituting the first systematic solution to credit time-series endpoint identification. Empirical results demonstrate substantial improvements in the accuracy of risk event timing and severity prediction, reduced bias in credit loss estimation, and enhanced accuracy and robustness of IFRS 9 expected credit loss (ECL) provisioning.
📝 Abstract
A novel procedure is presented for finding the true but latent endpoints within the repayment histories of individual loans. The monthly observations beyond these true endpoints are false, largely due to operational failures that delay account closure, thereby corrupting some loans. Detecting these false observations is difficult at scale since each affected loan history might have a different sequence of trailing zero (or very small) month-end balances. Identifying these trailing balances requires an exact definition of a"small balance", which our method informs. We demonstrate this procedure and isolate the ideal small-balance definition using South African residential mortgages. Evidently, corrupted loans are remarkably prevalent and have excess histories that are surprisingly long, which ruin the timing of risk events and compromise any subsequent time-to-event model, e.g., survival analysis. Having discarded these excess histories, we demonstrably improve the accuracy of both the predicted timing and severity of risk events, without materially impacting the portfolio. The resulting estimates of credit losses are lower and less biased, which augurs well for raising accurate credit impairments under IFRS 9. Our work therefore addresses a pernicious data error, which highlights the pivotal role of data preparation in producing credible forecasts of credit risk.