International Journal For Multidisciplinary Research

E-ISSN: 2582-2160   •   Impact Factor: 9.24

A Widely Indexed Open Access Peer Reviewed Multidisciplinary Bi-monthly Scholarly International Journal

Call for Paper Volume 8, Issue 5 (September-October 2026) Submit your research before last 3 days of October to publish your research paper in the issue of September-October.

Temporal Evaluation of Gradient Boosting and Autoencoder Models for Credit Card Fraud Detection

Author(s) Arnav Thakur
Country United States
Abstract Payment card fraud losses exceeded US$33 billion worldwide in 2022, yet fraud detection models are often summarized by accuracy values that conceal how much fraud they miss. This study compares a Naive Bayes baseline, three gradient boosting libraries (LightGBM, XGBoost and CatBoost), an autoencoder anomaly detector, and a hybrid model that supplies the autoencoder's reconstruction error to LightGBM, using the IEEE-CIS Fraud Detection dataset. The 590,540 labeled transactions were ordered in time and divided into training (60%), validation (20%) and test (20%) periods, and all preprocessing was fitted on the training period only. On the later test period (118,108 transactions, 4,064 of them fraudulent), LightGBM reached an average precision (AP) of 0.501 and a ROC-AUC of 0.889. XGBoost reached an AP of 0.506 when text features were supplied as integer codes but only 0.325 with its native categorical handling, a larger gap than any observed between libraries. CatBoost reached 0.460 at its 800-iteration cap. The autoencoder alone was weaker than Naive Bayes (AP 0.111 versus 0.132), and adding its reconstruction error to LightGBM changed AP by +0.003 (95% day-bootstrap interval −0.002 to 0.009). At thresholds selected on validation data, the strongest models detected 43–46% of test fraud at 53–55% precision, whereas a classifier that flags nothing reached 96.6% accuracy. Every model performed worse on the test period than on validation. These results indicate that categorical encoding, threshold choice and temporal evaluation affect reported fraud detection performance at least as much as the choice among modern boosting libraries.
Keywords credit card fraud, IEEE-CIS, gradient boosting, autoencoder, temporal validation, class imbalance, average precision
Field Computer > Artificial Intelligence / Simulation / Virtual Reality
Published In Volume 8, Issue 5, September-October 2026
Published On 2026-09-29
DOI https://doi.org/10.36948/ijfmr.2026.v08i05.88736

Share this