Resource-Efficient Optimization for Multi-Class Hematological Diagnosis: A Hybrid BPSO-Extra Trees Approach with Data Imbalance Handling

Authors

  • Dimas Chaerul Ekty Saputra Telkom University https://orcid.org/0000-0001-6978-2846
  • Zahid Abdullah Nur Mukhlishin Telkom University
  • Affifah Mutiara Pertiwi Telkom University
  • Mochammad Zulfikar Alfany Telkom University
  • Irianna Futri Ratchathani University
  • Raksmey Phann Seoul National University of Science and Technology

DOI:

https://doi.org/10.12928/mf.v8i1.15703

Keywords:

Complete Blood Count, Hematological Disorders, Feature Selection, Binary Particle Swarm Optimizatio, Computational Efficiency

Abstract

Complete Blood Count (CBC) remains the cornerstone for initial screening of hematological disorders, yet manual interpretation is often challenged by overlapping biological patterns and substantial inter-patient variability. Although machine learning approaches have demonstrated promise for automated diagnosis, many existing studies prioritize classification accuracy while neglecting computational efficiency and the persistent class imbalance inherent in medical datasets. This study develops a lightweight yet effective diagnostic framework for classifying nine hematological conditions using routine CBC parameters. Evaluated on a public dataset of 1,281 records from Kaggle, the proposed model is benchmarked against standard Random Forest, XGBoost, and Support Vector Machine (SVM) classifiers. The approach integrates the Synthetic Minority Oversampling Technique (SMOTE) to mitigate class imbalance, and Binary Particle Swarm Optimization (BPSO) to identify a compact and clinically informative feature subset of exactly 6 parameters, referred to as a clinical fingerprint, optimized for the Extra Trees classifier. Evaluated using ten-fold cross-validation, the BPSO-Extra Trees model achieved an average accuracy of 87.43 percent and an F1 score of 82.75 percent, while demonstrating superior resource efficiency with peak memory consumption of only 0.249 MB, corresponding to a 46.6 percent reduction compared with the standard Random Forest baseline. These findings confirm that swarm intelligence optimized models can effectively balance diagnostic performance with extreme computational frugality, enabling the potential deployment of accurate hematology-based decision support systems on portable devices and in resource-limited laboratory environments.

References

J. Yi, Z. Tang, L. Chen, and X. Meng, “The clinical value of complete blood count parameters in the prediction of postpartum hemorrhage severity,” Critical Public Health, vol. 35, no. 1, p. 2571873, 2025.

N. Mishra, A. Rana, P. Kumar, R. Gupta, and M. Kumar, “Complete Blood Count (CBC) Parameters as a Cost-Effective Tool for Early Diagnosis of Pediatric Sepsis: A Retrospective Cross-Sectional Study,” Cureus, vol. 17, no. 8, p. e89661, 2025.

J. Liu et al., “Development and application of machine learning models for hematological disease diagnosis using routine laboratory parameters: a user-friendly diagnostic platform,” Front. Med., vol. 12, p. 1605868, Oct. 2025, doi: 10.3389/fmed.2025.1605868.

G. Gulati, G. Uppal, and J. Gong, “Unreliable Automated Complete Blood Count Results: Causes, Recognition, and Resolution,” Ann Lab Med, vol. 42, no. 5, pp. 515–530, Sep. 2022, doi: 10.3343/alm.2022.42.5.515.

A. Kubiak, E. Ziółkowska, and A. Korycka-Wołowiec, “The diagnostic pitfalls and challenges associated with basic hematological tests,” Acta Haematol Pol., vol. 53, no. 2, pp. 104–111, Apr. 2022, doi: 10.5603/AHP.a2022.0014.

W. Hong et al., “A Comparison of XGBoost, Random Forest, and Nomograph for the Prediction of Disease Severity in Patients With COVID-19 Pneumonia: Implications of Cytokine and Immune Cell Profile,” Front. Cell. Infect. Microbiol., vol. 12, p. 819267, Apr. 2022, doi: 10.3389/fcimb.2022.819267.

H. Esmaily, M. Tayefi, H. Doosti, M. Ghayour-Mobarhan, H. Nezami, and A. Amirabadizadeh, “A Comparison between Decision Tree and Random Forest in Determining the Risk Factors Associated with Type 2 Diabetes”.

G. Airlangga, “Anemia Classification Using Hybrid Machine Learning Models: A Comparative Study of Ensemble Techniques on CBC Data,” JoSYC, vol. 5, no. 4, pp. 1108–1117, Aug. 2024, doi: 10.47065/josyc.v5i4.5848.

H. Amjad, Z. Hussain, M. Hasan, and M. Ul Hassan, “Machine learning-based models for screening of anemia and leukemia using features of complete blood count reports,” Sci Rep, vol. 15, no. 1, p. 33333, Sep. 2025, doi: 10.1038/s41598-025-21279-w.

G. Airlangga, “Leveraging Machine Learning for Accurate Anemia Diagnosis Using Complete Blood Count Data,” IJAIDM, vol. 7, no. 2, p. 318, May 2024, doi: 10.24014/ijaidm.v7i2.29869.

Y. Yang, H. A. Khorshidi, and U. Aickelin, “A review on over-sampling techniques in classification of multi-class imbalanced datasets: insights for medical problems,” Front. Digit. Health, vol. 6, p. 1430245, Jul. 2024, doi: 10.3389/fdgth.2024.1430245.

H. Hairani, T. Widiyaningtyas, and D. Dwi Prasetya, “Addressing Class Imbalance of Health Data: A Systematic Literature Review on Modified Synthetic Minority Oversampling Technique (SMOTE) Strategies,” JOIV : Int. J. Inform. Visualization, vol. 8, no. 3, p. 1310, Sep. 2024, doi: 10.62527/joiv.8.3.2283.

M. Zhu, “Solving Class Imbalance in Medical Image Classification Based on Federated Learning,” in Proceedings of the 2025 2nd International Conference on Generative Artificial Intelligence and Information Security, Hangzhou China: ACM, Feb. 2025, pp. 546–550. doi: 10.1145/3728725.3728811.

S. Dhibar, D. Gupta, G. E.A., and S. Vekkot, “An enhanced grey wolf optimization method for feature selection and explainable prediction in chronic disease analytics,” Decision Analytics Journal, vol. 17, p. 100655, Dec. 2025, doi: 10.1016/j.dajour.2025.100655.

L. K. Kumar, K. G. Suma, P. Udayaraju, V. Gundu, S. V. Mantena, and B. N. Jagadesh, “Clustering-based binary Grey Wolf Optimisation model with 6LDCNNet for prediction of heart disease using patient data,” Sci Rep, vol. 15, no. 1, p. 1270, Jan. 2025, doi: 10.1038/s41598-025-85561-7.

A. O. Salau, E. D. Markus, T. A. Assegie, C. O. Omeje, and J. N. Eneh, “Influence of Class Imbalance and Resampling on Classification Accuracy of Chronic Kidney Disease Detection,” MMEP, vol. 10, no. 1, pp. 48–54, Feb. 2023, doi: 10.18280/mmep.100106.

F. Gurcan and A. Soylu, “Learning from Imbalanced Data: Integration of Advanced Resampling Techniques and Machine Learning Models for Enhanced Cancer Diagnosis and Prognosis,” Cancers, vol. 16, no. 19, p. 3417, Oct. 2024, doi: 10.3390/cancers16193417.

L. D. Nurhayati and M. Rahardi, “Impact of SMOTE and ADASYN on Class Imbalance in Metabolic Syndrome Classification Using Random Forest Algorithm,” JAIC, vol. 9, no. 5, pp. 2807–2813, Oct. 2025, doi: 10.30871/jaic.v9i5.10657.

Z. Ye, Y. Xu, Q. He, M. Wang, W. Bai, and H. Xiao, “Feature Selection Based on Adaptive Particle Swarm Optimization with Leadership Learning,” Computational Intelligence and Neuroscience, vol. 2022, pp. 1–18, Aug. 2022, doi: 10.1155/2022/1825341.

Angga Maulana Akbar, R. Herteno, S. W. Saputro, M. R. Faisal, and R. A. Nugroho, “Optimizing Software Defect Prediction Models: Integrating Hybrid Grey Wolf and Particle Swarm Optimization for Enhanced Feature Selection with Popular Gradient Boosting Algorithm,” j.electron.electromedical.eng.med.inform, vol. 6, no. 2, pp. 169–181, Apr. 2024, doi: 10.35882/jeeemi.v6i2.388.

Y. Hidaka, T. Imai, K. Omae, T. Kagawa, S. Ishikawa, and T. Inaba, “Application and performance of tree model-based classifier and anomaly-detection approaches for medical imbalanced data,” Informatics in Medicine Unlocked, vol. 58, p. 101677, 2025, doi: 10.1016/j.imu.2025.101677.

I. Riadi, A. Yudhana, and G. C. Kurniawan, “Evaluating Synthetic Minority Oversampling Technique Strategies for Diabetes Mellitus Classification using K-Nearest Neighbors Algorithm,” J. Tek. Inform. (JUTIF), vol. 6, no. 5, pp. 3958–3970, Oct. 2025, doi: 10.52436/1.jutif.2025.6.5.5189.

U. J. Nzenwata, E. Edwin, E. A. Chukwu, D. Osilaja, J. O. Hinmikaiye, and C. Enyinnah, “Extra Trees Model for Heart Disease Prediction,” JDAIP, vol. 13, no. 02, pp. 125–139, 2025, doi: 10.4236/jdaip.2025.132008.

D. Melanson, “Extremely Randomized Trees with Multiparty Computation,” Senior Thesis, University of Washington Tacoma, 2020.

R. Vohra, A. Hussain, A. K. Dudyala, J. Pahareeya, and W. Khan, “Multi-class classification algorithms for the diagnosis of anemia in an outpatient clinical setting,” PLoS ONE, vol. 17, no. 7, p. e0269685, Jul. 2022, doi: 10.1371/journal.pone.0269685.

N. A. Putri, “A Comparative Analysis of Machine Learning Classifier of Anemia Diagnosis Based on Complete Blood Count (CBC) Data,” IJIIS Int. J. Informatics Inf. Syst., vol. 8, no. 4, pp. 188–200, Dec. 2025, doi: 10.47738/ijiis.v8i4.286.

R. Z. Haider, I. U. Ujjan, N. A. Khan, E. Urrechaga, and T. S. Shamsi, “Beyond the In-Practice CBC: The Research CBC Parameters-Driven Machine Learning Predictive Modeling for Early Differentiation among Leukemias,” Diagnostics, vol. 12, no. 1, p. 138, Jan. 2022, doi: 10.3390/diagnostics12010138.

D. C. E. Saputra, V. R. Oktavia, I. Futri, and A. M. Pertiwi, “An Extreme Gradient Boosting for Blood Disease Classification Using Hematological Parameters: A Comparative Evaluation with Ensemble and Non-Ensemble Models,” vol. 11, no. 4, 2025.

S. Ameen, R. Balachandran, and T. Theodoridis, “Deriving Hematological Disease Classes Using Fuzzy Logic and Expert Knowledge: A Comprehensive Machine Learning Approach with CBC Parameters,” 2024, arXiv. doi: 10.48550/ARXIV.2406.13015.

H. Sazak and M. Kotan, “Automated Blood Cell Detection and Classification in Microscopic Images Using YOLOv11 and Optimized Weights,” Diagnostics, vol. 15, no. 1, p. 22, Dec. 2024, doi: 10.3390/diagnostics15010022.

M. M. Alam and M. T. Islam, “Machine learning approach of automatic identification and counting of blood cells,” Healthcare Tech Letters, vol. 6, no. 4, pp. 103–108, Aug. 2019, doi: 10.1049/htl.2018.5098.

T. N. Annisa, J. Jasmir, and N. Nurhadi, “Comparison of ANOVA and Chi-Square Feature Selection Methods to Improve Machine Learning Performance in Anemia Classification,” J. Tek. Inform. (JUTIF), vol. 6, no. 4, pp. 1925–1940, Aug. 2025, doi: 10.52436/1.jutif.2025.6.4.5017.

S. Muharni, S. Andriyanto, and S. Supardi, “Optimization of the Naïve Bayes Algorithm Using Particle Swarm Optimization (PSO) for Predicting Heart Disease Symptoms,” J. Prisma. Sains, vol. 13, no. 3, pp. 582–601, Jun. 2025, doi: 10.33394/j-ps.v13i3.15573.

I. Wijayanto, S. Hadiyoso, A. S. Safitri, and T. D. Rahmaniar, “Application of Hybrid Metaheuristic Algorithms for Feature Selection in Event-Related Potential Classification in Problematic Gamers Using Electroencephalograph Signal,” j.electron.electromedical.eng.med.inform, vol. 7, no. 2, pp. 366–379, Mar. 2025, doi: 10.35882/jeeemi.v7i2.638.

M. B. Subkhi, C. Fatichah, and A. Zaenal Arifin, “Seleksi Fitur Menggunakan Hybrid Binary Grey Wolf Optimizer untuk Klasifikasi Hadist Teks Arab,” JTIIK, vol. 10, no. 5, pp. 1115–1122, Oct. 2023, doi: 10.25126/jtiik.2023106375.

A. A.-B. R. Saabia and M. Frikha, “Hybrid GTO-SA Metaheuristic for Feature Subset Selection in High-Dimensional Medical Datasets,” ISI, vol. 30, no. 10, Oct. 2025, doi: 10.18280/isi.301002.

N. Alamsyah, B. Budiman, V. R. Danestiara, T. P. Yoga, R. Nursyanti, and V. Kaunang, “A Metaheuristic wrapper approach to feature selection with genetic algorithm for enhancing XGBoost classification in diabetes prediction,” KINETIK, Oct. 2025, doi: 10.22219/kinetik.v10i4.2366.

D. Gürkan Kuntalp, N. Özcan, O. Düzyel, F. Y. Kababulut, and M. Kuntalp, “A Comparative Study of Metaheuristic Feature Selection Algorithms for Respiratory Disease Classification,” Diagnostics, vol. 14, no. 19, p. 2244, Oct. 2024, doi: 10.3390/diagnostics14192244.

Y. Pourardebil Khah, M. Hosseini Shirvani, and J. Taheri, “A survey study on meta-heuristic-based feature selection approaches of intrusion detection systems in distributed networks,” Computer Standards & Interfaces, vol. 96, p. 104074, Mar. 2026, doi: 10.1016/j.csi.2025.104074.

T. Stephan, P. S, C.-C. Lin, and S. Agarwal, “A Comprehensive Study of Grey Wolf Optimizer Variants for Optimizing Feature Selection in High-Dimensional Data,” Applied Artificial Intelligence, vol. 40, no. 1, p. 2601378, Dec. 2026, doi: 10.1080/08839514.2025.2601378.

X. Liu and H. Tian, “A novel hybrid feature selection method combining binary grey wolf optimization and cuckoo search,” Sci Rep, vol. 15, no. 1, p. 45190, Nov. 2025, doi: 10.1038/s41598-025-29018-x.

F. Yang, H. Wang, H. Mi, C. Lin, and W. Cai, “Using random forest for reliable classification and cost-sensitive learning for medical diagnosis,” BMC Bioinformatics, vol. 10, no. S1, p. S22, Jan. 2009, doi: 10.1186/1471-2105-10-S1-S22.

E. Erni and R. Sa’adah, “Comparison of Decision Trees, Naïve Bayes and Random Forest in Detecting Heart Disease,” SISTEMASI, vol. 13, no. 4, p. 1491, Jul. 2024, doi: 10.32520/stmsi.v13i4.4163.

M. A. Al Ghifari, I. Budiman, T. H. Saragih, M. I. Mazdadi, R. Herteno, and H. A. A. Rozaq, “Implementation of Extra Trees Classifier and Chi-Square Feature Selection for Early Detection of Liver Disease,” J. Tek. Inform. (JUTIF), vol. 6, no. 5, pp. 3925–3937, Oct. 2025, doi: 10.52436/1.jutif.2025.6.5.4261.

M. Naseriparsa, A. Al-Shammari, M. Sheng, Y. Zhang, and R. Zhou, “RSMOTE: improving classification performance over imbalanced medical datasets,” Health Inf Sci Syst, vol. 8, no. 1, p. 22, Dec. 2020, doi: 10.1007/s13755-020-00112-w.

A. R. B. Alamsyah, S. R. Anisa, N. S. Belinda, and A. Setiawan, “SMOTE and Nearmiss Methods for Disease Classification with Unbalanced Data: Case Study: IFLS 5,” icdsos, vol. 2021, no. 1, pp. 305–314, Jan. 2022, doi: 10.34123/icdsos.v2021i1.240.

Downloads

Published

2026-03-31

Issue

Section

Articles