Robust Fraud Detection with Ensemble Learning: A Case Study on the IEEE-CIS Dataset
preprint
OA: closed
AI-generated summary
This study evaluated ensemble learning methods, including stacking, for credit card fraud detection on the IEEE-CIS dataset, achieving 91.8% AUC-ROC and 0.891 AUC-PR with balancing and feature engineering.
One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works
Abstract
The rapid growth of digital financial transactions has led to a corresponding increase in credit card fraud, necessitating the development of sophisticated detection systems. This paper presents a comprehensive analysis of advanced ensemble learning techniques for imbalanced fraud detection using the IEEE-CIS dataset. We address the critical challenges of extreme class imbalance, concept drift, and real-time detection requirements through systematic evaluation of ensemble methods, including Random Forest, XGBoost, LightGBM, and novel stacking approaches. Our methodology incorporates advanced data balancing techniques (SMOTE, ADASYN, Borderline-SMOTE) and feature engineering strategies optimized for the IEEE-CIS dataset containing 590,540 transactions with 3.5% fraud rate. Experimental results demonstrate that our proposed ensemble stacking approach achieves superior performance with 91.8% AUC-ROC, 0.891 AUC-PR, and significant improvements in fraud detection rates while maintaining low false positive rates. The study provides empirical evidence for the effectiveness of ensemble methods in handling severely imbalanced financial fraud datasets and offers practical insights for real-world implementation.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.
Source provenance
- europepmc
- last seen: 2026-05-20T01:45:00.602351+00:00