Imbalanced Data Clustering Using Improved AFCM Algorithm Based on Whale Optimization Algorithm
Abstract
Accurate clustering in data mining faces serious challenges, especially when the data is unbalanced. In unbalanced data, traditional Fuzzy Clustering Algorithms (FCM) usually do not have sufficient accuracy and do not correctly identify minority classes due to their tendency to create clusters of equal size. To reduce the effect of this limitation, in this research, a novel approach is presented in which the adaptive FCM is improved using the Whale Optimization Algorithm (WOA). By purposefully exploring the response space, the whale algorithm adjusts the parameters of the adaptive FCM optimally and therefore reduces the probability of the algorithm getting stuck in local optima. Experimental evaluations on standard Iris and Thyroid datasets show that the proposed method achieves values of 97.14% and 94.71% for F1-Score and Balanced Accuracy (BA) metrics, respectively. This improvement in evaluation criteria indicates that the proposed method has a higher ability to manage the side effects of data imbalance and maintain the true structure of clusters compared to FCM and other basic methods.
Keywords:
Clustering, Unbalanced data, Data mining, Adaptive fuzzy algorithm, Whale optimization algorithmReferences
- [1] Wang, S., Shao, C., Xu, S., Yang, X., & Yu, H. (2024). MSFSS: A whale optimization-based multiple sampling feature selection stacking ensemble algorithm for classifying imbalanced data. AIMS mathematics, 9(7), 17504–17530. https://doi.org/10.3934/math.2024851
- [2] Oskouei, A. G., Samadi, N., Khezri, S., Moghaddam, A. N., Babaei, H., Hamini, K., ... & Arasteh, B. (2025). Feature-weighted fuzzy clustering methods: An experimental review. Neurocomputing, 619, 129176. https://doi.org/10.1016/j.neucom.2024.129176
- [3] Kalaycı, T. A., & Asan, U. (2022). Improving classification performance of fully connected layers by fuzzy clustering in transformed feature space. Symmetry, 14(4), 658. https://doi.org/10.3390/sym14040658
- [4] Wang, Y., Pang, W., & Jiao, Z. (2023). An adaptive mutual K-nearest neighbors clustering algorithm based on maximizing mutual information. Pattern recognition, 137, 109273. https://doi.org/10.1016/j.patcog.2022.109273
- [5] Zeraatkar, S., & Afsari, F. (2021). Interval-valued fuzzy and intuitionistic fuzzy-KNN for imbalanced data classification. Expert systems with applications, 184, 115510. https://doi.org/10.1016/j.eswa.2021.115510
- [6] Liu, F., Wang, J., & Liu, Y. (2024). IMI2: A fuzzy clustering validity index for multiple imbalanced clusters. Expert systems with applications, 238, 122231. https://doi.org/10.1016/j.eswa.2023.122231
- [7] Mirjalili, S., & Lewis, A. (2016). The whale optimization algorithm. Advances in engineering software, 95, 51–67. https://doi.org/10.1016/j.advengsoft.2016.01.008
- [8] Liu, Y., Jiang, Y., Hou, T., & Liu, F. (2021). A new robust fuzzy clustering validity index for imbalanced data sets. Information sciences, 547, 579–591. https://doi.org/10.1016/j.ins.2020.08.041
- [9] Wang, J., Wang, S., & Zhang, Y. (2025). Deep learning on medical image analysis. CAAI transactions on intelligence technology, 10(1), 1–35. https://doi.org/10.1049/cit2.12356
- [10] Hoseini, S. S., & Abbasi, S. H. (2017). A new method for community detection in social networks based on message distribution. International journal of computer science and network security, 17(5), 298–308. https://www.researchgate.net/profile/Hamid-Abbasi-13/publication/358214918
- [11] Japkowicz, N., & Stephen, S. (2002). The class imbalance problem: A systematic study. Intelligent data analysis, 6(5), 429–449. https://doi.org/10.3233/IDA-2002-6504
- [12] He, H., & Garcia, E. A. (2009). Learning from imbalanced data. IEEE transactions on knowledge and data engineering, 21(9), 1263–1284. https://doi.org/10.1109/TKDE.2008.239
- [13] Zheng, M., Li, T., Zheng, X., Yu, Q., Chen, C., Zhou, D., ... & Yang, W. (2021). UFFDFR: Undersampling framework with denoising, fuzzy c-means clustering, and representative sample selection for imbalanced data classification. Information sciences, 576, 658-680. https://doi.org/10.1016/j.ins.2021.07.053
- [14] Yue, X. D., Miao, D. Q., Cao, L. B., Wu, Q., & Chen, Y. F. (2014). An efficient color quantization based on generic roughness measure. Pattern recognition, 47(4), 1777–1789. https://doi.org/10.1016/j.patcog.2013.11.017
- [15] Arslan, H., & Toz, M. (2019). Data clustering based on fuzzy c-means and chaotic whale optimization algorithms. Sigma journal of engineering and natural sciences, 37(4), 1107–1128. https://sigma.yildiz.edu.tr/storage/upload/pdfs/1635862524-en.pdf
- [16] Ahmed, A. M., Rashid, T. A., Hassan, B. A., Majidpour, J., Noori, K. A., Rahman, C. M., ... & Mohammed, N. B. (2024). Balancing exploration and exploitation phases in whale optimization algorithm: An insightful and empirical analysis. In Handbook of whale optimization algorithm (pp. 149-156). Academic Press. https://doi.org/10.1016/B978-0-32-395365-8.00017-8

