Ijraset Journal For Research in Applied Science and Engineering Technology
Authors: Louis Bratawidjaja, Hans Philemon Limanza, Taufik Hidayat
DOI Link: https://doi.org/10.22214/ijraset.2026.84946
Certificate: View Certificate
Class imbalance and unreliable data are usually studied apart, and when they are studied together, minority label noise is treated as a single mechanism. It is not. Recording a genuine positive as a negative both depletes the positive class and contaminates the negative class with an example that contradicts the positives that remain, and a study that reports only the resulting performance drop cannot say which mechanism a method is failing to survive. We separate them with paired corruption arms that share a random draw of affected rows, relabelling the selected positives in one arm and deleting exactly the same rows in the other, so that the gap between the arms isolates the cost of the contradictory labels. Crossing six imbalance-handling strategies with three classifiers over 20 benchmark datasets and one extremely imbalanced fraud dataset, and summarising robustness by the normalised area under the corruption curve, we find that mislabelling a training positive costs 0.038 of clean PR-AUC more than deleting the same row, so a substantial part of the damage is the contradictory label rather than the lost example. Comparing methods with the Friedman test blocked on datasets — rather than on cross-validation folds, a common error that inflates the statistic in proportion to the number of folds — the imbalance strategy separates where the classifier family does not, and only 1 of five strategies separates from applying none at all. Resampling is not free: the most aggressive strategy degrades the Brier score to 4.5× that of the unmodified baseline, a cost invisible to threshold-free discrimination metrics. The decomposition costs one experimental arm, applies to any study that injects minority label noise, and turns a measurement of how much a method degrades into a statement about what it is failing to survive.
The text presents a research study on the combined problem of class imbalance and imperfect/noisy data in binary classification. These problems commonly occur in applications such as fraud detection, medical diagnosis, and defect detection, where the important class is rare and training data may contain incorrect labels or corrupted features.
The study focuses particularly on minority-label noise, where genuine positive cases are incorrectly labelled as negative. This creates two simultaneous effects:
Depletion: the positive class loses genuine examples.
Contamination: the negative class gains incorrectly labelled positive examples.
The authors argue that treating these two effects as one type of noise makes it difficult to determine why an imbalance-handling method fails and what type of solution would be appropriate.
To separate the two effects, the study introduces two paired corruption conditions:
Flip: selected minority examples are relabelled as negative, causing both depletion and contamination.
Drop: the exact same selected minority examples are removed from the training data, causing depletion but not contamination.
Because the same examples are selected in both conditions, the performance difference between flip and drop is used to estimate the “contamination cost”—the additional damage caused by keeping wrongly labelled examples in the training set.
The study also examines feature noise, created by adding Gaussian noise to continuous features.
The researchers evaluate:
21 imbalanced binary datasets, including 20 KEEL datasets and the Credit Card Fraud Detection dataset.
Imbalance ratios ranging from approximately 8.6:1 to 598.8:1.
6 imbalance-handling strategies:
No imbalance handling
Three SMOTE-based oversampling approaches
Random undersampling
Cost-sensitive class weighting
3 classifiers:
Logistic regression
Random forest
XGBoost
Corruption is introduced only into the training data, while validation and test data remain clean. This allows the researchers to measure how well each method remains effective when training data become increasingly unreliable.
Rather than evaluating performance at only one corruption level, the study uses severity curves and normalized areas to measure robustness across increasing levels of corruption.
The and incorrectly retaining them as negative examples have different consequences. By comparing matched main performance measures include:
PR-AUC for discrimination under severe class imbalance.
Brier score for probability calibration.
Friedman tests and related statistical comparisons for comparing multiple methods across datasets.
The study asks four main research questions concerning:
How different imbalance-handling methods degrade as data corruption increases.
How much performance loss comes from minority depletion versus negative-class contamination.
Whether resampling is more robust than cost-sensitive weighting or no imbalance handling.
How these methods affect probability calibration in addition to ranking performance.
Previous research has established methods such as SMOTE, Borderline-SMOTE, ADASYN, random undersampling, cost-sensitive learning, and noise-filtering techniques. However, the authors identify three important gaps:
Previous studies often use symmetric noise, whereas real applications such as fraud detection may have predominantly false-negative errors.
Many studies examine only a few corruption levels rather than complete severity curves.
Most importantly, previous work generally does not explicitly separate positive-class depletion from negative-class contamination.
We evaluated eighteen imbalance-handling pipelines under controlled training-data corruption on 20 benchmark datasets and one extremely imbalanced fraud dataset, summarising robustness by normalised area under the corruption curve and comparing methods with the Friedman test blocked on datasets. Three conclusions follow. First, minority label noise is not one mechanism but two, and they can be separated: deleting a training positive and mislabelling it differ by +0.038 of clean PR-AUC at 40% corruption, so the contradictory label carries a cost distinct from the lost example. Second, the choice of imbalance strategy separates for robustness where the classifier family does not, and only some strategies separate from applying none at all. Third, resampling is not free: it degrades probability calibration even where it leaves ranking untouched, which a PR-AUC-only evaluation cannot see. The decomposition is the part we would most like to see reused. It costs one additional experimental arm, it is applicable to any study that injects minority label noise, and it converts a statement about how much a method degrades into a statement about what it is failing to survive. The natural continuations are to test noise-aware oversamplers against it, to extend it to instance-dependent noise where contamination is concentrated near the boundary rather than spread uniformly, and to ask whether the classifier-dependence of the contamination cost reflects the locality of the decision boundary, as Section 5.1 conjectures.
[1] N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer. SMOTE: Synthetic Minority Over-sampling Technique. Journal of Artificial Intelligence Research, 16:321–357, 2002. [2] C. Elkan. The foundations of cost-sensitive learning. In Proc. International Joint Conference on Artificial Intelligence (IJCAI), pages 973–978, 2001. [3] B. Frénay and M. Verleysen. Classification in the Presence of Label Noise: A Survey. IEEE Transactions on Neural Networks and Learning Systems, 25(5):845–869, 2014. [4] J. Van Hulse and T. M. Khoshgoftaar. Knowledge discovery from imbalanced and noisy data. Data & Knowledge Engineering, 68(12):1513–1542, 2009. [5] J. A. Sáez, J. Luengo, J. Stefanowski, and F. Herrera. SMOTE–IPF: Addressing the noisy and borderline examples problem in imbalanced classification by a re-sampling method with filtering. Information Sciences, 291:184–203, 2015. [6] F. Brishti, F. Zhang, S. Mohammed, L. Bai, F. Wu, and B. Chen. Imbalanced classification with label noise: A systematic review and comparative analysis. ICT Express, 11(6):1127–1145, 2025. [7] A. Dal Pozzolo, O. Caelen, R. A. Johnson, and G. Bontempi. Calibrating Probability with Undersampling for Unbalanced Classification. In Proc. IEEE Symposium Series on Computational Intelligence (SSCI), pages 159–166, 2015. [8] D. Hendrycks and T. Dietterich. Benchmarking Neural Network Robustness to Common Corruptions and Perturbations. In Proc. International Conference on Learning Representations (ICLR), 2019. [9] J. Demšar. Statistical Comparisons of Classifiers over Multiple Data Sets. Journal of Machine Learning Research, 7:1–30, 2006. [10] N. Japkowicz and S. Stephen. The class imbalance problem: A systematic study. Intelligent Data Analysis, 6(5):429–449, 2002. [11] H. He and E. A. Garcia. Learning from Imbalanced Data. IEEE Transactions on Knowledge and Data Engineering, 21(9):1263–1284, 2009. [12] Y. Sun, A. K. C. Wong, and M. S. Kamel. Classification of imbalanced data: A review. International Journal of Pattern Recognition and Artificial Intelligence, 23(4):687–719, 2009. [13] B. Krawczyk. Learning from imbalanced data: open challenges and future directions. Progress in Artificial Intelligence, 5(4):221–232, 2016. [14] P. Branco, L. Torgo, and R. P. Ribeiro. A survey of predictive modeling on imbalanced domains. ACM Computing Surveys, 49(2):31:1–31:50, 2016. [15] H. Han, W.-Y. Wang, and B.-H. Mao. Borderline-SMOTE: A New Over-Sampling Method in Imbalanced Data Sets Learning. In Proc. International Conference on Intelligent Computing (ICIC), pages 878–887, 2005. [16] H. He, Y. Bai, E. A. Garcia, and S. Li. ADASYN: Adaptive Synthetic Sampling Approach for Imbalanced Learning. In Proc. IEEE International Joint Conference on Neural Networks (IJCNN), pages 1322–1328, 2008. [17] I. Tomek. Two modifications of CNN. IEEE Transactions on Systems, Man, and Cybernetics, SMC-6(11):769–772, 1976. [18] D. L. Wilson. Asymptotic properties of nearest neighbor rules using edited data. IEEE Transactions on Systems, Man, and Cybernetics, SMC-2(3):408–421, 1972. [19] G. E. A. P. A. Batista, R. C. Prati, and M. C. Monard. A study of the behavior of several methods for balancing machine learning training data. ACM SIGKDD Explorations Newsletter, 6(1):20–29, 2004. [20] T. Chen and C. Guestrin. XGBoost: A Scalable Tree Boosting System. In Proc. ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD), pages 785–794, 2016. [21] M. Buda, A. Maki, and M. A. Mazurowski. A systematic study of the class imbalance problem in convolutional neural networks. Neural Networks, 106:249–259, 2018. [22] R. Blagus and L. Lusa. SMOTE for high-dimensional class-imbalanced data. BMC Bioinformatics, 14:106, 2013. [23] Y. Elor and H. Averbuch-Elor. To SMOTE, or not to SMOTE? arXiv:2201.08528, 2022. [24] R. van den Goorbergh, M. van Smeden, D. Timmerman, and B. Van Calster. The harm of class imbalance corrections for risk prediction models: illustration and simulation using logistic regression. Journal of the American Medical Informatics Association, 29(9):1525–1534, 2022. [25] C. G. Northcutt, L. Jiang, and I. L. Chuang. Confident Learning: Estimating Uncertainty in Dataset Labels. Journal of Artificial Intelligence Research, 70:1373–1411, 2021. [26] H. Song, M. Kim, D. Park, Y. Shin, and J.-G. Lee. Learning from Noisy Labels with Deep Neural Networks: A Survey. IEEE Transactions on Neural Networks and Learning Systems, 34(11):8135–8153, 2023. [27] C. Seiffert, T. M. Khoshgoftaar, J. Van Hulse, and A. Folleco. An empirical study of the classification performance of learners on imbalanced and noisy software quality data. Information Sciences, 259:571–595, 2014. [28] J. Davis and M. Goadrich. The relationship between Precision-Recall and ROC curves. In Proc. International Conference on Machine Learning (ICML), pages 233–240, 2006. [29] T. Saito and M. Rehmsmeier. The Precision-Recall Plot Is More Informative than the ROC Plot When Evaluating Binary Classifiers on Imbalanced Datasets. PLOS ONE, 10(3):e0118432, 2015. [30] J. Derrac, S. García, D. Molina, and F. Herrera. A practical tutorial on the use of nonparametric statistical tests as a methodology for comparing evolutionary and swarm intelligence algorithms. Swarm and Evolutionary Computation, 1(1):3–18, 2011. [31] A. Benavoli, G. Corani, and F. Mangili. Should we really use post-hoc tests based on mean-ranks? Journal of Machine Learning Research, 17(5):1–10, 2016. [32] J. Alcalá-Fdez, A. Fernández, J. Luengo, J. Derrac, S. García, L. Sánchez, and F. Herrera. KEEL Data-Mining Software Tool: Data Set Repository, Integration of Algorithms and Experimental Analysis Framework. Journal of Multiple-Valued Logic and Soft Computing, 17(2–3):255–287, 2011. [33] L. Breiman. Random Forests. Machine Learning, 45(1):5–32, 2001. [34] F. Pedregosa et al. Scikit-learn: Machine Learning in Python. Journal of Machine Learning Research, 12:2825–2830, 2011. [35] G. Lemaître, F. Nogueira, and C. K. Aridas. Imbalanced-learn: A Python Toolbox to Tackle the Curse of Imbalanced Datasets in Machine Learning. Journal of Machine Learning Research, 18(17):1–5, 2017.
Copyright © 2026 Louis Bratawidjaja, Hans Philemon Limanza, Taufik Hidayat. This is an open access article distributed under the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
Paper Id : IJRASET84946
Publish Date : 2026-09-24
ISSN : 2321-9653
Publisher Name : IJRASET
DOI Link : Click Here
Submit Paper Online
