TY - GEN
T1 - Balancing the Imbalanced Datasets for Machine Learning Classification Problems
AU - Kumar, Rishi
AU - Gujjar, J. Praveen
AU - Prasad, M. S.Guru
AU - Devadas, Raghavendra M.
AU - Preethi, null
AU - Kumar, Utsav
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Machine learning tasks with imbalanced datasets present several difficulties, especially when trying to solve the underlying imbalance issue in the data. Insufficient data in the minority class can result in biased models that are not very good at generalizing to real-world situations. Moreover, because they must take into account class imbalance, traditional evaluation metrics like accuracy can be deceptive. It is imperative to tackle these issues with methods such as Imbalance-Learn, which offers instruments to manage unbalanced datasets efficiently through the modification of class weights, resampling techniques, and algorithmic methods. Data scientists can guarantee equitable and precise model predictions by taking into account the number of samples in the training set and managing the imbalance issue with great care. This paper shows various strategies to handle class imbalance, including data-level techniques such as resampling. This paper highlights the challenges of the imbalanced dataset and also it presents the various techniques to balance the imbalance dataset.
AB - Machine learning tasks with imbalanced datasets present several difficulties, especially when trying to solve the underlying imbalance issue in the data. Insufficient data in the minority class can result in biased models that are not very good at generalizing to real-world situations. Moreover, because they must take into account class imbalance, traditional evaluation metrics like accuracy can be deceptive. It is imperative to tackle these issues with methods such as Imbalance-Learn, which offers instruments to manage unbalanced datasets efficiently through the modification of class weights, resampling techniques, and algorithmic methods. Data scientists can guarantee equitable and precise model predictions by taking into account the number of samples in the training set and managing the imbalance issue with great care. This paper shows various strategies to handle class imbalance, including data-level techniques such as resampling. This paper highlights the challenges of the imbalanced dataset and also it presents the various techniques to balance the imbalance dataset.
UR - https://www.scopus.com/pages/publications/105010183486
UR - https://www.scopus.com/pages/publications/105010183486#tab=citedBy
U2 - 10.1109/INCIP64058.2025.11019446
DO - 10.1109/INCIP64058.2025.11019446
M3 - Conference contribution
AN - SCOPUS:105010183486
T3 - Proceedings - International Conference on Next Generation Communication and Information Processing, INCIP 2025
SP - 1
EP - 4
BT - Proceedings - International Conference on Next Generation Communication and Information Processing, INCIP 2025
A2 - Bukya, Mahipal
A2 - Kumar, Pramod
A2 - Rawat, Sanyog
A2 - Jangid, Mahesh
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2025 International Conference on Next Generation Communication and Information Processing, INCIP 2025
Y2 - 23 January 2025 through 24 January 2025
ER -