Starting Off on the Wrong Foot: Pitfalls in Data Preparation
Published in arXiv (Under Review), 2026
The Problem: Starting Off on the Wrong Foot
Actuarial modeling often starts off on the wrong foot when handling real-world, highly imbalanced insurance datasets. Conventional techniques, such as random train-test splitting, fail to properly partition rare but critical extreme claim events, leading to severe distribution shifts between datasets. Compounded by the “curse of dimensionality” and missing data, these initial missteps can completely undermine the statistical validity and reliability of downstream ML algorithms.
Our Methodology: Statistically Informed Data Preparation
To tackle these pitfalls, we developed IDPP that systematically automates and improves upstream preprocessing. We utilized SPlit, which leverages support points to guarantee the distributional consistency between train and test sets, particularly for heavy-tailed imbalanced distributions. For feature selection, we applied the Chatterjee correlation coefficient (CCC) to capture complex, non-linear dependencies without relying on specific model architectures. Finally, we handled missingness using MissForest imputation, seamlessly embedding this entire IDPP into our custom InformedAutoML framework.
An illustration of the IDPP
The Results: Robust Models and Efficient Computation
Through rigorous simulations and real-world studies, our approach demonstrated substantial improvements. The SPlit method successfully stabilized the allocation of extreme claim events, drastically reducing the variance in coefficient estimation. Furthermore, InformedAutoML achieved globally optimal predictive performance metrics while running four to five times faster than baseline automated frameworks. Ultimately, the proposed framework is both statistically robust and computationally efficient.
Train/Test RMSE and runtime on college Pell Grant dataset
Recommended citation: Guo, J., Dong, P., Quan, Z. (2026). Starting Off on the Wrong Foot: Pitfalls in DataPreparation.
Read Paper | Download Bibtex
