Resource-Efficient Optimization for Multi-Class Hematological Diagnosis: A Hybrid BPSO-Extra Trees Approach with Data Imbalance Handling
DOI:
https://doi.org/10.12928/mf.v8i1.15703Keywords:
Complete Blood Count, Hematological Disorders, Feature Selection, Binary Particle Swarm Optimizatio, Computational EfficiencyAbstract
Complete Blood Count (CBC) remains the cornerstone for initial screening of hematological disorders, yet manual interpretation is often challenged by overlapping biological patterns and substantial inter-patient variability. Although machine learning approaches have demonstrated promise for automated diagnosis, many existing studies prioritize classification accuracy while neglecting computational efficiency and the persistent class imbalance inherent in medical datasets. This study develops a lightweight yet effective diagnostic framework for classifying nine hematological conditions using routine CBC parameters. Evaluated on a public dataset of 1,281 records from Kaggle, the proposed model is benchmarked against standard Random Forest, XGBoost, and Support Vector Machine (SVM) classifiers. The approach integrates the Synthetic Minority Oversampling Technique (SMOTE) to mitigate class imbalance, and Binary Particle Swarm Optimization (BPSO) to identify a compact and clinically informative feature subset of exactly 6 parameters, referred to as a clinical fingerprint, optimized for the Extra Trees classifier. Evaluated using ten-fold cross-validation, the BPSO-Extra Trees model achieved an average accuracy of 87.43 percent and an F1 score of 82.75 percent, while demonstrating superior resource efficiency with peak memory consumption of only 0.249 MB, corresponding to a 46.6 percent reduction compared with the standard Random Forest baseline. These findings confirm that swarm intelligence optimized models can effectively balance diagnostic performance with extreme computational frugality, enabling the potential deployment of accurate hematology-based decision support systems on portable devices and in resource-limited laboratory environments.
References
J. Yi, Z. Tang, L. Chen, and X. Meng, “The clinical value of complete blood count parameters in the prediction of postpartum hemorrhage severity,” Critical Public Health, vol. 35, no. 1, p. 2571873, 2025.
N. Mishra, A. Rana, P. Kumar, R. Gupta, and M. Kumar, “Complete Blood Count (CBC) Parameters as a Cost-Effective Tool for Early Diagnosis of Pediatric Sepsis: A Retrospective Cross-Sectional Study,” Cureus, vol. 17, no. 8, p. e89661, 2025.
J. Liu et al., “Development and application of machine learning models for hematological disease diagnosis using routine laboratory parameters: a user-friendly diagnostic platform,” Front. Med., vol. 12, p. 1605868, Oct. 2025, doi: 10.3389/fmed.2025.1605868.
G. Gulati, G. Uppal, and J. Gong, “Unreliable Automated Complete Blood Count Results: Causes, Recognition, and Resolution,” Ann Lab Med, vol. 42, no. 5, pp. 515–530, Sep. 2022, doi: 10.3343/alm.2022.42.5.515.
A. Kubiak, E. Ziółkowska, and A. Korycka-Wołowiec, “The diagnostic pitfalls and challenges associated with basic hematological tests,” Acta Haematol Pol., vol. 53, no. 2, pp. 104–111, Apr. 2022, doi: 10.5603/AHP.a2022.0014.
W. Hong et al., “A Comparison of XGBoost, Random Forest, and Nomograph for the Prediction of Disease Severity in Patients With COVID-19 Pneumonia: Implications of Cytokine and Immune Cell Profile,” Front. Cell. Infect. Microbiol., vol. 12, p. 819267, Apr. 2022, doi: 10.3389/fcimb.2022.819267.
H. Esmaily, M. Tayefi, H. Doosti, M. Ghayour-Mobarhan, H. Nezami, and A. Amirabadizadeh, “A Comparison between Decision Tree and Random Forest in Determining the Risk Factors Associated with Type 2 Diabetes”.
G. Airlangga, “Anemia Classification Using Hybrid Machine Learning Models: A Comparative Study of Ensemble Techniques on CBC Data,” JoSYC, vol. 5, no. 4, pp. 1108–1117, Aug. 2024, doi: 10.47065/josyc.v5i4.5848.
H. Amjad, Z. Hussain, M. Hasan, and M. Ul Hassan, “Machine learning-based models for screening of anemia and leukemia using features of complete blood count reports,” Sci Rep, vol. 15, no. 1, p. 33333, Sep. 2025, doi: 10.1038/s41598-025-21279-w.
G. Airlangga, “Leveraging Machine Learning for Accurate Anemia Diagnosis Using Complete Blood Count Data,” IJAIDM, vol. 7, no. 2, p. 318, May 2024, doi: 10.24014/ijaidm.v7i2.29869.
Y. Yang, H. A. Khorshidi, and U. Aickelin, “A review on over-sampling techniques in classification of multi-class imbalanced datasets: insights for medical problems,” Front. Digit. Health, vol. 6, p. 1430245, Jul. 2024, doi: 10.3389/fdgth.2024.1430245.
H. Hairani, T. Widiyaningtyas, and D. Dwi Prasetya, “Addressing Class Imbalance of Health Data: A Systematic Literature Review on Modified Synthetic Minority Oversampling Technique (SMOTE) Strategies,” JOIV : Int. J. Inform. Visualization, vol. 8, no. 3, p. 1310, Sep. 2024, doi: 10.62527/joiv.8.3.2283.
M. Zhu, “Solving Class Imbalance in Medical Image Classification Based on Federated Learning,” in Proceedings of the 2025 2nd International Conference on Generative Artificial Intelligence and Information Security, Hangzhou China: ACM, Feb. 2025, pp. 546–550. doi: 10.1145/3728725.3728811.
S. Dhibar, D. Gupta, G. E.A., and S. Vekkot, “An enhanced grey wolf optimization method for feature selection and explainable prediction in chronic disease analytics,” Decision Analytics Journal, vol. 17, p. 100655, Dec. 2025, doi: 10.1016/j.dajour.2025.100655.
L. K. Kumar, K. G. Suma, P. Udayaraju, V. Gundu, S. V. Mantena, and B. N. Jagadesh, “Clustering-based binary Grey Wolf Optimisation model with 6LDCNNet for prediction of heart disease using patient data,” Sci Rep, vol. 15, no. 1, p. 1270, Jan. 2025, doi: 10.1038/s41598-025-85561-7.
A. O. Salau, E. D. Markus, T. A. Assegie, C. O. Omeje, and J. N. Eneh, “Influence of Class Imbalance and Resampling on Classification Accuracy of Chronic Kidney Disease Detection,” MMEP, vol. 10, no. 1, pp. 48–54, Feb. 2023, doi: 10.18280/mmep.100106.
F. Gurcan and A. Soylu, “Learning from Imbalanced Data: Integration of Advanced Resampling Techniques and Machine Learning Models for Enhanced Cancer Diagnosis and Prognosis,” Cancers, vol. 16, no. 19, p. 3417, Oct. 2024, doi: 10.3390/cancers16193417.
L. D. Nurhayati and M. Rahardi, “Impact of SMOTE and ADASYN on Class Imbalance in Metabolic Syndrome Classification Using Random Forest Algorithm,” JAIC, vol. 9, no. 5, pp. 2807–2813, Oct. 2025, doi: 10.30871/jaic.v9i5.10657.
Z. Ye, Y. Xu, Q. He, M. Wang, W. Bai, and H. Xiao, “Feature Selection Based on Adaptive Particle Swarm Optimization with Leadership Learning,” Computational Intelligence and Neuroscience, vol. 2022, pp. 1–18, Aug. 2022, doi: 10.1155/2022/1825341.
Angga Maulana Akbar, R. Herteno, S. W. Saputro, M. R. Faisal, and R. A. Nugroho, “Optimizing Software Defect Prediction Models: Integrating Hybrid Grey Wolf and Particle Swarm Optimization for Enhanced Feature Selection with Popular Gradient Boosting Algorithm,” j.electron.electromedical.eng.med.inform, vol. 6, no. 2, pp. 169–181, Apr. 2024, doi: 10.35882/jeeemi.v6i2.388.
Y. Hidaka, T. Imai, K. Omae, T. Kagawa, S. Ishikawa, and T. Inaba, “Application and performance of tree model-based classifier and anomaly-detection approaches for medical imbalanced data,” Informatics in Medicine Unlocked, vol. 58, p. 101677, 2025, doi: 10.1016/j.imu.2025.101677.
I. Riadi, A. Yudhana, and G. C. Kurniawan, “Evaluating Synthetic Minority Oversampling Technique Strategies for Diabetes Mellitus Classification using K-Nearest Neighbors Algorithm,” J. Tek. Inform. (JUTIF), vol. 6, no. 5, pp. 3958–3970, Oct. 2025, doi: 10.52436/1.jutif.2025.6.5.5189.
U. J. Nzenwata, E. Edwin, E. A. Chukwu, D. Osilaja, J. O. Hinmikaiye, and C. Enyinnah, “Extra Trees Model for Heart Disease Prediction,” JDAIP, vol. 13, no. 02, pp. 125–139, 2025, doi: 10.4236/jdaip.2025.132008.
D. Melanson, “Extremely Randomized Trees with Multiparty Computation,” Senior Thesis, University of Washington Tacoma, 2020.
R. Vohra, A. Hussain, A. K. Dudyala, J. Pahareeya, and W. Khan, “Multi-class classification algorithms for the diagnosis of anemia in an outpatient clinical setting,” PLoS ONE, vol. 17, no. 7, p. e0269685, Jul. 2022, doi: 10.1371/journal.pone.0269685.
N. A. Putri, “A Comparative Analysis of Machine Learning Classifier of Anemia Diagnosis Based on Complete Blood Count (CBC) Data,” IJIIS Int. J. Informatics Inf. Syst., vol. 8, no. 4, pp. 188–200, Dec. 2025, doi: 10.47738/ijiis.v8i4.286.
R. Z. Haider, I. U. Ujjan, N. A. Khan, E. Urrechaga, and T. S. Shamsi, “Beyond the In-Practice CBC: The Research CBC Parameters-Driven Machine Learning Predictive Modeling for Early Differentiation among Leukemias,” Diagnostics, vol. 12, no. 1, p. 138, Jan. 2022, doi: 10.3390/diagnostics12010138.
D. C. E. Saputra, V. R. Oktavia, I. Futri, and A. M. Pertiwi, “An Extreme Gradient Boosting for Blood Disease Classification Using Hematological Parameters: A Comparative Evaluation with Ensemble and Non-Ensemble Models,” vol. 11, no. 4, 2025.
S. Ameen, R. Balachandran, and T. Theodoridis, “Deriving Hematological Disease Classes Using Fuzzy Logic and Expert Knowledge: A Comprehensive Machine Learning Approach with CBC Parameters,” 2024, arXiv. doi: 10.48550/ARXIV.2406.13015.
H. Sazak and M. Kotan, “Automated Blood Cell Detection and Classification in Microscopic Images Using YOLOv11 and Optimized Weights,” Diagnostics, vol. 15, no. 1, p. 22, Dec. 2024, doi: 10.3390/diagnostics15010022.
M. M. Alam and M. T. Islam, “Machine learning approach of automatic identification and counting of blood cells,” Healthcare Tech Letters, vol. 6, no. 4, pp. 103–108, Aug. 2019, doi: 10.1049/htl.2018.5098.
T. N. Annisa, J. Jasmir, and N. Nurhadi, “Comparison of ANOVA and Chi-Square Feature Selection Methods to Improve Machine Learning Performance in Anemia Classification,” J. Tek. Inform. (JUTIF), vol. 6, no. 4, pp. 1925–1940, Aug. 2025, doi: 10.52436/1.jutif.2025.6.4.5017.
S. Muharni, S. Andriyanto, and S. Supardi, “Optimization of the Naïve Bayes Algorithm Using Particle Swarm Optimization (PSO) for Predicting Heart Disease Symptoms,” J. Prisma. Sains, vol. 13, no. 3, pp. 582–601, Jun. 2025, doi: 10.33394/j-ps.v13i3.15573.
I. Wijayanto, S. Hadiyoso, A. S. Safitri, and T. D. Rahmaniar, “Application of Hybrid Metaheuristic Algorithms for Feature Selection in Event-Related Potential Classification in Problematic Gamers Using Electroencephalograph Signal,” j.electron.electromedical.eng.med.inform, vol. 7, no. 2, pp. 366–379, Mar. 2025, doi: 10.35882/jeeemi.v7i2.638.
M. B. Subkhi, C. Fatichah, and A. Zaenal Arifin, “Seleksi Fitur Menggunakan Hybrid Binary Grey Wolf Optimizer untuk Klasifikasi Hadist Teks Arab,” JTIIK, vol. 10, no. 5, pp. 1115–1122, Oct. 2023, doi: 10.25126/jtiik.2023106375.
A. A.-B. R. Saabia and M. Frikha, “Hybrid GTO-SA Metaheuristic for Feature Subset Selection in High-Dimensional Medical Datasets,” ISI, vol. 30, no. 10, Oct. 2025, doi: 10.18280/isi.301002.
N. Alamsyah, B. Budiman, V. R. Danestiara, T. P. Yoga, R. Nursyanti, and V. Kaunang, “A Metaheuristic wrapper approach to feature selection with genetic algorithm for enhancing XGBoost classification in diabetes prediction,” KINETIK, Oct. 2025, doi: 10.22219/kinetik.v10i4.2366.
D. Gürkan Kuntalp, N. Özcan, O. Düzyel, F. Y. Kababulut, and M. Kuntalp, “A Comparative Study of Metaheuristic Feature Selection Algorithms for Respiratory Disease Classification,” Diagnostics, vol. 14, no. 19, p. 2244, Oct. 2024, doi: 10.3390/diagnostics14192244.
Y. Pourardebil Khah, M. Hosseini Shirvani, and J. Taheri, “A survey study on meta-heuristic-based feature selection approaches of intrusion detection systems in distributed networks,” Computer Standards & Interfaces, vol. 96, p. 104074, Mar. 2026, doi: 10.1016/j.csi.2025.104074.
T. Stephan, P. S, C.-C. Lin, and S. Agarwal, “A Comprehensive Study of Grey Wolf Optimizer Variants for Optimizing Feature Selection in High-Dimensional Data,” Applied Artificial Intelligence, vol. 40, no. 1, p. 2601378, Dec. 2026, doi: 10.1080/08839514.2025.2601378.
X. Liu and H. Tian, “A novel hybrid feature selection method combining binary grey wolf optimization and cuckoo search,” Sci Rep, vol. 15, no. 1, p. 45190, Nov. 2025, doi: 10.1038/s41598-025-29018-x.
F. Yang, H. Wang, H. Mi, C. Lin, and W. Cai, “Using random forest for reliable classification and cost-sensitive learning for medical diagnosis,” BMC Bioinformatics, vol. 10, no. S1, p. S22, Jan. 2009, doi: 10.1186/1471-2105-10-S1-S22.
E. Erni and R. Sa’adah, “Comparison of Decision Trees, Naïve Bayes and Random Forest in Detecting Heart Disease,” SISTEMASI, vol. 13, no. 4, p. 1491, Jul. 2024, doi: 10.32520/stmsi.v13i4.4163.
M. A. Al Ghifari, I. Budiman, T. H. Saragih, M. I. Mazdadi, R. Herteno, and H. A. A. Rozaq, “Implementation of Extra Trees Classifier and Chi-Square Feature Selection for Early Detection of Liver Disease,” J. Tek. Inform. (JUTIF), vol. 6, no. 5, pp. 3925–3937, Oct. 2025, doi: 10.52436/1.jutif.2025.6.5.4261.
M. Naseriparsa, A. Al-Shammari, M. Sheng, Y. Zhang, and R. Zhou, “RSMOTE: improving classification performance over imbalanced medical datasets,” Health Inf Sci Syst, vol. 8, no. 1, p. 22, Dec. 2020, doi: 10.1007/s13755-020-00112-w.
A. R. B. Alamsyah, S. R. Anisa, N. S. Belinda, and A. Setiawan, “SMOTE and Nearmiss Methods for Disease Classification with Unbalanced Data: Case Study: IFLS 5,” icdsos, vol. 2021, no. 1, pp. 305–314, Jan. 2022, doi: 10.34123/icdsos.v2021i1.240.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 Dimas Chaerul Ekty Saputra, Zahid Abdullah Nur Mukhlishin, Affifah Mutiara Pertiwi, Mochammad Zulfikar Alfany, Irianna Futri, Raksmey Phann

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Start from 2019 issues, authors who publish with JURNAL MOBILE AND FORENSICS agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License (CC BY-SA 4.0) that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work.

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.





