Sitemap
A list of all the posts and pages found on the site. For you robots out there, there is an XML version available for digesting as well.
Pages
Posts
code
Automated Machine Learning (AutoML)
https://github.com/PanyiDong/InsurAutoML
Open-source Automated Machine Learning (AutoML) framework tailored for the insurance domain, automating data preprocessing, hyperparameter optimization, and imbalance learning. Also powers the Informed Data Preparation Pipeline (IDPP) and InformedAutoML.
SuppsPlit
https://github.com/PanyiDong/SuppsPlit
A Python Wrapper for sPlit, based on Joseph, V. R., & Vakayil, A. (2022). SPlit: An optimal method for data splitting. Technometrics, 64(2), 166-176.
InsurTech NLP
https://github.com/PanyiDong/InsurTech_NLP
Natural Language Processing toolkit for insurance analytics: sentiment-based rating de-biasing, neural embeddings for high-cardinality business categories, and unsupervised NAICS industry classification with LDA and RAKE.
InsurGenAI
https://github.com/PanyiDong/InsurGenAI
An integrated interface for deploying open-source large language models (LLMs).
portfolio
publications
Privacy-Preserving Collaborative Information Sharing Through Federated Learning
Published in arXiv, 2024
Recommended citation: Dong, P., Quan, Z., Edwards, B., Wang, H., Feng, R., Wang, T., Foley, P., Shah, P. (2024). Privacy-Preserving Collaborative Information Sharing Through Federated Learning. Read Paper | Download Bibtex
Federated Learning for Insurance Companies
Published in Society of Actuaries Research Institute, 2024
Recommended citation: Dong, P., Feng, F., Quan, Z., Wang, T. (2024). Federated Learning for Insurance Companies. Society of Actuaries Research Institute. Read Paper | Download Bibtex
Improving Business Insurance Loss Models by Leveraging InsurTech Innovation
Published in North American Actuarial Journal, 2024
This paper demonstrates that enriching traditional in-house insurance datasets with real-time, personalized InsurTech data improves the predictive accuracy of business insurance loss models.
Recommended citation: Quan, Z., Hu, C., Dong, P., Valdez, E. (2025). Improving Business Insurance Loss Models by Leveraging InsurTech Innovation. North American Actuarial Journal, 29(2), 247-274. Read Paper | Download Bibtex
Bridging the Divide While Walking a Tightrope: Evidence from the Insurance Industry on Federated Data Sharing
Published in ssrn (Under Review), 2026
This paper proposes a Hybrid Federated Learning (HyFL) framework tailored for the insurance industry that bridges data silos across insurers and InsurTech partners, achieving significant predictive accuracy gains and unlocking substantial financial value while preserving data privacy.
Recommended citation: Dong, P., Feng, F., Quan, Z., Wang, T. (2026). Bridging the Divide While Walking a Tightrope: Evidence from the Insurance Industry on Federated Data Sharing. Read Paper | Download Bibtex
Efficient and Interpretable Transformer for Counterfactual Fairness
Published in arXiv, 2026
In this paper, we introduce the Feature Correlation Transformer (FCorrTransformer) and Counterfactual Attention Regularization (CAR). This efficient, attention-light framework achieves counterfactual fairness and high interpretability for tabular datasets, effectively mitigating algorithmic bias with minimal predictive performance degradation.
Recommended citation: Dong, P., Quan, Z. (2026). Efficient and Interpretable Transformer for Counterfactual Fairness. Read Paper | Download Bibtex
Hybrid Tree-based Interpretable Pricing
In preparation
With: Zhiyu Quan, Montserrat Guillen, Lluís Bermúdez.
Automated Machine Learning (AutoML) in Insurance
Published in Insurance: Mathematics and Economics, 2024
This paper introduces an open-source Automated Machine Learning (AutoML) framework tailored for the insurance domain, effectively automating data preprocessing, hyperparameter optimization, and imbalance learning, to streamline actuarial data science.
Recommended citation: Dong, P., Quan, Z. (2025). Automated Machine Learning (AutoML) in Insurance. Insurance: Mathematics and Economics, 120, 17-41. Read Paper | Download Bibtex
Starting Off on the Wrong Foot: Pitfalls in Data Preparation
Published in arXiv (Under Review), 2026
To prevent flawed actuarial modeling caused by inappropriate data preprocessing, we developed an effective and efficient Informed Data Preparation Pipeline (IDPP) utilizing various statistical tools.
Recommended citation: Guo, J., Dong, P., Quan, Z. (2026). Starting Off on the Wrong Foot: Pitfalls in DataPreparation. Read Paper | Download Bibtex
InsurTech innovation using natural language processing
Published in North American Actuarial Journal, 2026
This paper explores the transformative potential of Natural Language Processing (NLP) in modernizing insurance analytics by extracting actionable insights from unstructured InsurTech data, with applications such as feature de-biasing, high-cardinality feature representation, and automated industry classification.
Recommended citation: Dong, P., Quan, Z. (2026). InsurTech innovation using natural language processing. North American Actuarial Journal, 1-34. Read Paper | Download Bibtex
Rare Events Data and Zero Reduction Sampling
In preparation
With: Zhiyu Quan, Haiying Wang, Jing Wang, Jiaqi Liu.
talks
teaching
Teaching Assistant
Actuarial Science courses, University of Illinois Urbana-Champaign, 2022
I’ve been teaching assistant for University-Earned Credit (UEC) life contingencies (Exam FAM-L) and awarded Teaching Assistants Ranked as Excellent by Their Students for multiple semesters.
Course Instructor
ASRM 195, Foundations of Data Management, University of Illinois Urbana-Champaign, 2026
