Vollständiger Abstract
Worum geht es in dieser Arbeit?
Deciding whether to approve a loan application is one of the highest-stakes decisions a lender makes: approving a risky applicant invites default, while turning away a sound one loses a good customer. Predicting the approval decision accurately is therefore central to comprehensive credit management. Gradient boosting has become the dominant approach for such tabular credit data, and two libraries, XGBoost and LightGBM, account for most of its practical use. Although both are widely applied to credit-risk and loan-approval prediction, they are seldom compared directly under matched conditions, which leaves little firm evidence for choosing between them; this study addresses that gap. A controlled, like-for-like comparison was carried out in which both models were trained on the same dataset of 45,000 loan applications using one pre-processing pipeline and a matched set of structural hyper-parameters, and were then evaluated with accuracy, precision, recall, F1-score and the area under the ROC curve. The comparison was re-examined with five-fold stratified cross-validation and repeated on a second, independent loan-approval dataset of a different character. On the primary dataset, XGBoost reached an accuracy of 93.5% and an ROC-AUC of 0.979, a small margin above LightGBM on every metric, and it retained the higher mean AUC under cross-validation; on the second, high-signal dataset both models exceeded 98% accuracy and XGBoost again held the higher cross-validated AUC, winning four of five folds. Across both dataset types, XGBoost therefore demonstrated a consistent, though modest, improvement in discrimination, an effect compatible with its level-wise tree growth and stronger regularization. The principal limitations are the use of a single matched configuration and two historical datasets. The findings indicate that both libraries are accurate and deployable, with XGBoost showing a marginal but reproducible edge in predictive discrimination.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Sagar Poudel, Bhoj Raj Ghimire
- Quelle
- American Journal of Data Science and Artificial Intelligence
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 3069-3632
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Sagar Poudel, Bhoj Raj Ghimire (2026). XGBoost versus LightGBM for Loan Approval Prediction: A Controlled, Like-for-Like Comparison. American Journal of Data Science and Artificial Intelligence. https://doi.org/10.54536/ajdsai.v2i2.8054
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1