TY - GEN
T1 - Optimizing Profit Scoring in P2P Lending Using Prepayment and IRR Analysis
AU - Hegde, Anusha
AU - Nayaka, Premkumar
AU - Bhowmik, Biswajit
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Peer-to-peer (P2P) lending platforms face significant challenges in accurately assessing borrower risk and predicting default probability due to the complex nature of lending data characterized by high dimensionality, severe class imbalance, and intricate relational dependencies. This work develops a comprehensive machine learning pipeline for credit and profit scoring in P2P lending environments. A critical innovation of this work is the identification and incorporation of prepayments as a distinct class, addressing the financial impact of early loan settlements on expected interest returns. We conduct a systematic comparison of default prediction performance between binary classification (default vs. non-default) and multi-class classification (default vs. non-default vs. prepaid) frameworks using three machine learning algorithms: Logistic Regression, Random Forest, and XGBoost. To evaluate the practical financial implications of misclassification errors, we implement Internal Rate of Return (IRR) analysis through simulated cash flow modeling. This approach quantifies the actual monetary impact of prediction errors on lending profitability. The findings demonstrate that misclassification costs are substantially higher in binary classification models compared to multi-class approaches, suggesting that incorporating prepayment prediction significantly improves both predictive accuracy and financial outcomes.
AB - Peer-to-peer (P2P) lending platforms face significant challenges in accurately assessing borrower risk and predicting default probability due to the complex nature of lending data characterized by high dimensionality, severe class imbalance, and intricate relational dependencies. This work develops a comprehensive machine learning pipeline for credit and profit scoring in P2P lending environments. A critical innovation of this work is the identification and incorporation of prepayments as a distinct class, addressing the financial impact of early loan settlements on expected interest returns. We conduct a systematic comparison of default prediction performance between binary classification (default vs. non-default) and multi-class classification (default vs. non-default vs. prepaid) frameworks using three machine learning algorithms: Logistic Regression, Random Forest, and XGBoost. To evaluate the practical financial implications of misclassification errors, we implement Internal Rate of Return (IRR) analysis through simulated cash flow modeling. This approach quantifies the actual monetary impact of prediction errors on lending profitability. The findings demonstrate that misclassification costs are substantially higher in binary classification models compared to multi-class approaches, suggesting that incorporating prepayment prediction significantly improves both predictive accuracy and financial outcomes.
UR - https://www.scopus.com/pages/publications/105042376208
UR - https://www.scopus.com/pages/publications/105042376208#tab=citedBy
U2 - 10.1109/AIDE69088.2026.11545044
DO - 10.1109/AIDE69088.2026.11545044
M3 - Conference contribution
AN - SCOPUS:105042376208
T3 - 2026 International Conference on Artificial Intelligence and Data Engineering, AIDE 2026 - Proceedings
SP - 117
EP - 122
BT - 2026 International Conference on Artificial Intelligence and Data Engineering, AIDE 2026 - Proceedings
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2026 International Conference on Artificial Intelligence and Data Engineering, AIDE 2026
Y2 - 5 February 2026 through 7 February 2026
ER -