Explainable and Fair Credit Risk Scoring with Counterfactual Explanations: A Reproducible Evaluation on the German Credit Dataset (HELOC-Motivated)
DOI:
https://doi.org/10.32664/j-intech.v14i02.2228Kata Kunci:
Algorithmic recourse, Counterfactual explanations, Credit scoring, Fairness, Explainable AI (XAI)Abstrak
Credit risk scoring requires models that are accurate, fair, and able to provide actionable explanations for adverse decisions. Motivated by the HELOC explainable lending benchmark, but using the publicly downloadable Statlog German Credit dataset as a fully reproducible proxy when direct HELOC access is constrained, this study evaluates explainable and fair credit scoring on 1,000 applicants. We train five common models—logistic regression (LR), decision tree (DT), random forest (RF), XGBoost (XGB), and LightGBM (LGBM)—on a fixed 600/200/200 train/validation/test split with a consistent preprocessing pipeline. Thresholds are selected on the validation set to maximize F1 for bad-risk detection. LR achieves the best test AUC (0.7888), LGBM the highest test accuracy (0.6550), and RF the best calibration (ECE=0.0473), showing that discrimination, thresholded accuracy, and calibration do not align. Fairness is audited by a derived sex attribute using demographic parity and equalized odds. Baseline LGBM shows an approval-rate difference of −0.0846 (female minus male) and an equalized-odds gap of 0.1000. Reweighing reduces the approval-rate difference to −0.0642 while preserving AUC (0.7804), and equal-opportunity thresholding reduces the equalized-odds gap to 0.0750. For individual explainability, we generate counterfactual recourse for rejected applicants using six actionable features. Feasible recourse is defined as the existence of at least one action-constrained counterfactual that changes the decision from reject to approve; 78.64% of rejected applicants receive such recourse, with mean cost 1.5298 measured as standardized numeric change plus categorical steps. Across five retraining seeds, LGBM AUC is stable (mean 0.7742, std 0.0022), but fairness gaps vary. The study provides a reproducible template for jointly evaluating performance, calibration, fairness, mitigation, and recourse in credit scoring
Referensi
[1] A. Bitetto, P. Cerchiello, S. Filomeni, A. Tanda, and B. Tarantino, “Machine learning and credit risk: Empirical evidence from small- and mid-sized businesses,” Socioecon. Plann. Sci., vol. 90, p. 101746, Dec. 2023, doi: 10.1016/j.seps.2023.101746.
[2] E. Dumitrescu, S. Hué, C. Hurlin, and S. Tokpavi, “Machine learning for credit scoring: Improving logistic regression with non-linear decision-tree effects,” Eur. J. Oper. Res., vol. 297, no. 3, pp. 1178–1192, Mar. 2022, doi: 10.1016/j.ejor.2021.06.053.
[3] Board of Governors of the Federal Reserve System; Office of the Comptroller of the Currency, “Supervisory Guidance on Model Risk Management (SR 11-7),” 2011.
[4] Office of the Comptroller of the Currency, “OCC Bulletin 2011-12: Sound Practices for Model Risk Management,” 2011.
[5] Consumer Financial Protection Bureau, “Consumer Financial Protection Circular 2022-03,” 2022.
[6] Consumer Financial Protection Bureau, “12 CFR §1002.9 - Notifications (Regulation B).”
[7] M. Ribeiro, S. Singh, and C. Guestrin, “‘Why Should I Trust You?’: Explaining the Predictions of Any Classifier,” in Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations, Stroudsburg, PA, USA: Association for Computational Linguistics, 2016, pp. 97–101. doi: 10.18653/v1/N16-3020.
[8] Scott M. Lundberg; and Su-In Lee, “A Unified Approach to Interpreting Model Predictions,” 2017, pp. 4765–4774.
[9] S. Wachter, B. Mittelstadt, and C. Russell, “Counterfactual Explanations Without Opening the Black Box: Automated Decisions and the GDPR,” SSRN Electronic Journal, 2017, doi: 10.2139/ssrn.3063289.
[10] B. Ustun, A. Spangher, and Y. Liu, “Actionable Recourse in Linear Classification,” in Proceedings of the Conference on Fairness, Accountability, and Transparency, New York, NY, USA: ACM, Jan. 2019, pp. 10–19. doi: 10.1145/3287560.3287566.
[11] Moritz Hardt, Eric Price, and Nathan Srebro, “Equality of Opportunity in Supervised Learning,” 2016, pp. 3315–3323.
[12] S. Verma and J. Rubin, “Fairness definitions explained,” in Proceedings of the International Workshop on Software Fairness, New York, NY, USA: ACM, May 2018, pp. 1–7. doi: 10.1145/3194770.3194776.
[13] Jon Kleinberg;, Sendhil Mullainathan;, and Manish Raghavan, “Inherent Trade-Offs in the Fair Determination of Risk Scores,” LIPIcs, 2017.
[14] UCI Machine Learning Repository, “Statlog (German Credit Data),” 1994.
[15] F. Kamiran and T. Calders, “Data preprocessing techniques for classification without discrimination,” Knowl. Inf. Syst., vol. 33, no. 1, pp. 1–33, Oct. 2012, doi: 10.1007/s10115-011-0463-8.
[19] S. Goethals, D. Martens, and T. Calders, “PreCoF: counterfactual explanations for fairness,” Mach. Learn., vol. 113, no. 5, pp. 3111–3142, 2024, doi: 10.1007/s10994-023-06319-8.
[20] I. E. Kumar, K. E. Hines, and J. P. Dickerson, “Equalizing Credit Opportunity in Algorithms: Aligning Algorithmic Fairness Research with U.S. Fair Lending Regulation,” in Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society, New York, NY, USA: ACM, Aug. 2022, pp. 357–368, doi: 10.1145/3514094.3534154.
[21] P. Ganesh, H. Chang, M. Strobel, and R. Shokri, “On The Impact of Machine Learning Randomness on Group Fairness,” in Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, New York, NY, USA: ACM, Jun. 2023, pp. 1789–1800, doi: 10.1145/3593013.3594116.
[22] A. Fabris, S. Messina, G. Silvello, and G. A. Susto, “Algorithmic fairness datasets: the story so far,” Data Min. Knowl. Discov., vol. 36, no. 6, pp. 2074–2152, 2022, doi: 10.1007/s10618-022-00854-z.
[23] A. Roy, J. Horstmann, and E. Ntoutsi, “Multi-dimensional Discrimination in Law and Machine Learning - A Comparative Overview,” in Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, New York, NY, USA: ACM, Jun. 2023, pp. 89–100, doi: 10.1145/3593013.3593979.
[24] S. Das, R. Stanton, and N. Wallace, “Algorithmic Fairness,” Annu. Rev. Financ. Econ., vol. 15, pp. 565–593, 2023, doi: 10.1146/annurev-financial-110921-125930.
[25] D. Pessach and E. Shmueli, “A Review on Fairness in Machine Learning,” ACM Comput. Surv., vol. 55, no. 3, pp. 1–44, 2022, doi: 10.1145/3494672.
Unduhan
Diterbitkan
Terbitan
Bagian
Lisensi
Hak Cipta (c) 2026 J-INTECH ( Journal of Information and Technology)

Artikel ini berlisensiCreative Commons Attribution-ShareAlike 4.0 International License.

