Bridging the Divide While Walking a Tightrope: Evidence from the Insurance Industry on Federated Data Sharing
Published in ssrn (Under Review), 2026
The Motivation: Overcoming Data Silos and Privacy Tightropes
Data in the insurance industry is severely fragmented due to competitive pressures and strict regulatory privacy constraints. Insurers often operate as data silos, facing a two bottleneck: extreme data imbalance and insufficient feature coverage. While centralized repositories or direct data purchases offer partial solution, they create significant trade-offs regarding data privacy and continuous licensing costs. To resolve this, we sought to design a privacy-preserving collaborative learning framework that balances data privacy with data utility.
Our Methodology: Hybrid Federated Learning (HyFL)
Classical Federated Learning (FL) is typically limited to either horizontal (HFL) or vertical (VFL) data partitioning. However, real-world insurance ecosystems are structurally complex, featuring overlapping samples across insurers alongside non-overlapping proprietary features from InsurTech partners. To address this, we developed HyFL, an framework simultaneously aggregate horizontal and vertical partitions. In addition, to handle the extreme sparsity and class imbalance of insurance claims, we integrated a domain-specific warm-up pre-training phase that stabilized the training process.
Structure of HyFL
The Impact: Substantial Accuracy and Multimillion-Dollar Value
Using real-world proprietary data from two mid-sized commercial insurers and an InsurTech partner, we demonstrated that HyFL substantially outperforms local training as well as standalone HFL and VFL. Economically, this reduction in prediction error translates to $57.8 million and $13.3 million in portfolio uncertainty reduction for the respective carriers. Furthermore, our interpretability analysis using ALE revealed that HyFL effectively eliminates localized feature biases inherent in single-insurer datasets, demonstrating that even large carriers with extensive local data stand to gain immensely from privacy-preserving data collaboration.
| Collaborator | Claim Portfolio Size (in million) | Improvement (in million) | ||
|---|---|---|---|---|
| HFL | VFL | HyFL | ||
| Company A | 361.2 | 32.5 | 50.6 | 57.8 |
| Company B | 102.3 | 7.1 | 7.1 | 13.3 |
Recommended citation: Dong, P., Feng, F., Quan, Z., Wang, T. (2026). Bridging the Divide While Walking a Tightrope: Evidence from the Insurance Industry on Federated Data Sharing.
Read Paper | Download Bibtex
