Automated Multi-Scenario Classification of Food Allergens in Indonesian Culinary Datasets Using Bernoulli Naïve Bayes with Augmentation and Oversampling Methods

Authors

  • Aji Prasetya Wibawa Department of Electrical Engineering and Informatics, Faculty of Engineering, Universitas Negeri Malang, Malang, Indonesia
  • David Satria Alamsyah Department of Electrical Engineering and Informatics, Faculty of Engineering, Universitas Negeri Malang, Malang, Indonesia
  • Aldy Rahmat Yulianto Department of Electrical Engineering and Informatics, Faculty of Engineering, Universitas Negeri Malang, Malang, Indonesia
  • Ahmad 'Ammar Musyaffa' Department of Electrical Engineering and Informatics, Faculty of Engineering, Universitas Negeri Malang, Malang, Indonesia
  • Adil Zakaria Department of Electrical Engineering and Informatics, Faculty of Engineering, Universitas Negeri Malang, Malang, Indonesia
  • Agung Bella Putra Utama Department of Electrical Engineering and Informatics, Faculty of Engineering, Universitas Negeri Malang, Malang, Indonesia

DOI:

https://doi.org/10.12928/jafost.v7i3.14827

Keywords:

Augmentation, Food allergens, Indonesian culinary, Machine learning, Naive bayes

Abstract

Limited awareness of allergen content in traditional Indonesian foods poses a potential health risk because recipe platforms rarely provide explicit allergen information. This study develops an automated allergen classification system for Indonesian recipes using the Bernoulli Naïve Bayes algorithm. The contributions include the construction of a recipe-level allergen dataset, the integration of text preprocessing, augmentation, and resampling techniques, and the introduction of multi-scenario allergen classification to support safer food choices. Recipe data were collected from Cookpad, producing 9,921 cleaned recipes after preprocessing. To increase data diversity and address class imbalance, rule-based ingredient augmentation expanded the dataset to 15,030 recipes. The model was evaluated under three classification scenarios: binary allergen detection (0–1), allergen exposure levels (0–3), and detailed allergen counts (0–14). Each scenario was tested using four dataset configurations: raw data, raw + ROS (Random Oversampling), augmented data, and augmented + ROS. Class distributions become increasingly imbalanced as the number of classes grows, particularly in the 15-class scenario. The binary scenario achieved perfect performance, with accuracy, precision, recall, and F1-score all equal to 1.0 across all dataset variants. Scenario 2 achieved the best multi-class performance with an F1-score of 0.9293 using Augmented + ROS data. Performance decreased in the 15-class setting due to increased classification complexity. Although a perfect binary score may raise concerns about overfitting, the result likely reflects the clear separability of the dataset’s allergen features. It should be interpreted cautiously when applied to broader culinary data.

References

C. M. Warren, S. Sehgal, S. H. Sicherer, and R. S. Gupta, “Epidemiology and the growing epidemic of food allergy in children and adults across the globe,” Curr. Allergy Asthma Rep., vol. 24, no. 3, pp. 95–106, 2024, https://doi.org/10.1007/s11882-023-01120-y.

N. Roy, “Food allergy knowledge, attitudes, and practices among restaurant’s staff in Bangladesh: A structural modeling approach,” Food Sci. Nutr., vol. 13, no. 8, p. e70754, 2025, https://doi.org/10.1002/fsn3.70754.

A. A. Takrouni, “Knowledge gaps in food allergy among the general public in Jeddah, Saudi Arabia: Insights based on the Chicago food allergy research survey,” Front. Allergy, vol. 3, p. 1002694, 2022, https://doi.org/10.3389/falgy.2022.1002694.

F. Chang, L. Eng, and C. Chang, “Food allergy labeling laws: International guidelines for residents and travelers,” Clin. Rev. Allergy Immunol., vol. 65, no. 2, pp. 148–165, 2023, https://doi.org/10.1007/s12016-023-08960-6.

G. N. Konstantinou, O. Pampoukidou, D. Sergelidis, and M. Fotoulaki, “Managing food allergies in dining establishments: Challenges and innovative solutions,” Nutrients, vol. 17, no. 10, p. 1737, 2025, https://doi.org/10.3390/nu17101737.

S. La Vieille, J. O. Hourihane, and J. L. Baumert, “Precautionary allergen labeling: What advice is available for health care professionals, allergists, and allergic consumers?,” J. Allergy Clin. Immunol. Pract., vol. 11, no. 4, pp. 977–985, 2023, https://doi.org/10.1016/j.jaip.2022.12.042

S. Xiong, W. Tian, H. Si, G. Zhang, and L. Shi, “A survey of the applications of text mining for the food domain,” Algorithms, vol. 17, no. 5, p. 176, 2024, https://doi.org/10.3390/a17050176.

W. Liu, “Quantitative food allergen risk assessment: Evolving concepts, modern approaches, and industry implications,” Compr. Rev. Food Sci. Food Saf., vol. 24, no. 2, p. e70132, 2025, https://doi.org/10.1111/1541-4337.70132.

D. D. Purwanto, A. P. Wibawa, and S. Patmanthara, “Philosophy of culinary arts: Reading truth through language and taste,” Bull. Culin. Art Hosp., vol. 5, no. 2, 2025, https://doi.org/10.17977/um069v5i22025p62-72.

D. Dooley, “Food process ontology requirements,” Semant. Web, vol. 15, no. 4, pp. 1133–1164, 2024, https://doi.org/10.3233/SW-223096.

A. F. U. R. Khilji, “Multimodal recipe recommendation system using deep learning and rule-based approach,” SN Comput. Sci., vol. 4, no. 4, p. 421, 2023, https://doi.org/10.1007/s42979-023-01870-6.

H. Zhou, X. Chen, and X. Li, “Food allergenicity evaluation methods: Classification, principle, and applications,” Foods, vol. 14, no. 12, p. 2005, 2025, https://doi.org/10.3390/foods14122005.

F. K. K. Putra, M. K. Putra, and S. Novianti, “Taste of Asean: Traditional food images from Southeast Asian countries,” J. Ethn. Foods, vol. 10, no. 1, p. 20, 2023, https://doi.org/10.1186/s42779-023-00189-0.

Y. Zhang, “Deep learning in food category recognition,” Inf. Fusion, vol. 98, p. 101859, 2023, https://doi.org/10.1016/j.inffus.2023.101859.

A. Kumar and P. S. Rana, “A deep learning based ensemble approach for protein allergen classification,” PeerJ Comput. Sci., vol. 9, p. e1622, 2023, https://doi.org/10.7717/peerj-cs.1622.

A. Roither, M. Kurz, and E. Sonnleitner, “The chef’s choice: System for allergen and style classification in recipes,” Appl. Sci., vol. 12, no. 5, p. 2590, 2022, https://doi.org/10.3390/app12052590.

P. J. Turner, “Time to ACT-UP: Update on precautionary allergen labelling (PAL),” World Allergy Organ. J., vol. 17, no. 10, p. 100972, 2024, https://doi.org/10.1016/j.waojou.2024.100972.

Y. Zhong, W. Zhou, and Z. Wang, “A survey of data augmentation in domain generalization,” Neural Process. Lett., vol. 57, no. 2, p. 34, 2025, https://doi.org/10.1007/s11063-025-11747-9.

S. Jeong and J. Lee, “Effects of cultural background on consumer perception and acceptability of foods and drinks: A review of latest cross-cultural studies,” Curr. Opin. Food Sci., vol. 42, pp. 248–256, 2021, https://doi.org/10.1016/j.cofs.2021.07.004.

K. Maharana, S. Mondal, and B. Nemade, “A review: Data pre-processing and data augmentation techniques. Global Transitions Proceedings, vol. 3, no. 1, pp. 91–99, 2022, https://doi.org/10.1016/j.gltp.2022.04.020.

M. Lango and J. Stefanowski, “What makes multi-class imbalanced problems difficult? An experimental study,” Expert Syst. Appl., vol. 199, p. 116962, 2022, https://doi.org/10.1016/j.eswa.2022.116962.

J. F. Kurian and M. Allali, “Detecting drifts in data streams using Kullback-Leibler (KL) divergence measure for data engineering applications,” J. Data, Inf. Manag., vol. 6, no. 3, pp. 207–216, 2024, https://doi.org/10.1007/s42488-024-00119-y.

A. Figueira and B. Vaz, “Survey on synthetic data generation, evaluation methods and GANs,” Mathematics, vol. 10, no. 15, p. 2733, 2022, https://doi.org/10.3390/math10152733.

F. Gurcan and A. Soylu, “Synthetic boosted resampling using deep generative adversarial networks: A novel approach to improve cancer prediction from imbalanced datasets,” Cancers (Basel)., vol. 16, no. 23, p. 4046, 2024, https://doi.org/10.3390/cancers16234046.

E. Sarmas, E. Spiliotis, V. Marinakis, T. Koutselis, and H. Doukas, “A meta-learning classification model for supporting decisions on energy efficiency investments,” Energy Build., vol. 258, p. 111836, 2022, https://doi.org/10.1016/j.enbuild.2022.111836.

S. F. Taskiran, B. Turkoglu, E. Kaya, and T. Asuroglu, “A comprehensive evaluation of oversampling techniques for enhancing text classification performance,” Sci. Rep., vol. 15, no. 1, p. 21631, 2025, https://doi.org/10.1038/s41598-025-05791-7.

M. Scheschenja, M. B. Bastian, J. Wessendorf, A. D. Owczarek, A. M. König, S. Viniol, and A. H. Mahnken, “ChatGPT: Evaluating answers on contrast media related questions and finetuning by providing the model with the ESUR guideline on contrast agents,” Curr. Probl. Diagn. Radiol., vol. 53, no. 4, pp. 488–493, 2024, https://doi.org/10.1067/j.cpradiol.2024.04.005.

R. Blanquero, E. Carrizosa, P. Ramírez-Cobo, and M. R. Sillero-Denamiel, “Variable selection for naïve bayes classification,” Comput. Oper. Res., vol. 135, p. 105456, 2021, https://doi.org/10.1016/j.cor.2021.105456.

J. Li, A. Wu, L. Liu, A. Qu, C. Xu, H. Kuang, and L. Xu, “Analysis of food safety based on machine learning: A comprehensive review and future prospects,” Food Chem., vol. 490, p. 145170, 2025, https://doi.org/10.1016/j.foodchem.2025.145170.

L. A. Yates, Z. Aandahl, S. A. Richards, and B. W. Brook, “Cross validation for model selection: A review with examples from ecology,” Ecol. Monogr., vol. 93, no. 1, p. e1557, 2023, https://doi.org/10.1002/ecm.1557.

S. M. Malakouti, “Babysitting hyperparameter optimization and 10-fold-cross-validation to enhance the performance of ML methods in predicting wind speed and energy generation,” Intell. Syst. with Appl., vol. 19, p. 200248, 2023, https://doi.org/10.1016/j.iswa.2023.200248.

T. R. Mahesh, O. Geman, M. Margala, and M. Guduri, “The stratified k-folds cross-validation and class-balancing methods with high-performance ensemble classifiers for breast cancer classification,” Healthc. Anal., vol. 4, p. 100247, 2023, https://doi.org/10.1016/j.health.2023.100247.

O. Rainio, J. Teuho, and R. Klén, “Evaluation metrics and statistical tests for machine learning,” Sci. Rep., vol. 14, no. 1, p. 6086, 2024, https://doi.org/10.1038/s41598-024-56706-x.

E. Kına, “Real-time food allergen detection using OCR-enhanced machine learning techniques,” PeerJ Comput. Sci., vol. 11, p. e3338, 2025, https://doi.org/10.7717/peerj-cs.3338.

Zeeshan Ahmad, Anum Farooq, Muhammad Fuzail, Yasir Aziz, and Naeem Aslam, “Food allergy detection using machine learning approach,” Kashf J. Multidiscip. Res., vol. 2, no. 04, pp. 116–127, 2025, https://doi.org/10.71146/kjmr402.

S. Brahimi, “AI-powered dining: text information extraction and machine learning for personalized menu recommendations and food allergy management,” Int. J. Inf. Technol., vol. 17, no. 4, pp. 2107–2115, 2025, https://doi.org/10.1007/s41870-024-02154-9.

G. Charizanos, H. Demirhan, and D. İçen, “Binary classification with fuzzy logistic regression under class imbalance and complete separation in clinical studies,” BMC Med. Res. Methodol., vol. 24, no. 1, p. 145, 2024, https://doi.org/10.1186/s12874-024-02270-x.

T. Wongvorachan, S. He, and O. Bulut, “A Comparison of under sampling, oversampling, and SMOTE methods for dealing with imbalanced classification in educational data mining,” Information, vol. 14, no. 1, p. 54, 2023, https://doi.org/10.3390/info14010054.

Y. Li, Y. Yang, P. Song, L. Duan, and R. Ren, “An improved SMOTE algorithm for enhanced imbalanced data classification by expanding sample generation space,” Sci. Rep., vol. 15, no. 1, p. 23521, 2025, https://doi.org/10.1038/s41598-025-09506-w.

M. K. Vadiveloo, “Exploring facilitators and barriers for personalized dietary incentives among online shoppers at cardiovascular risk and key informants to inform an automated shopping platform,” J. Nutr. Educ. Behav., 2025, https://doi.org/10.1016/j.jneb.2025.07.010

K.-S. Lee and C.-W. W. Tao, “Culinary knowledge sharing on social media: Case of the 2019 Malaysian World Pastry Champion Wei Loon Tan,” J. Hosp. Tour. Manag., vol. 52, pp. 52–64, 2022, https://doi.org/10.1016/j.jhtm.2022.06.006.

W. M. Blom, “Allergen labelling: Current practice and improvement from a communication perspective,” Clin. Exp. allergy, vol. 51, no. 4, pp. 574–584, 2021, https://doi.org/10.1111/cea.13830.

S. Barker, “Allergy education and training for physicians,” World Allergy Organ. J., vol. 14, no. 10, p. 100589, 2021, https://doi.org/10.1016/j.waojou.2021.100589.

M. Shahab, J. Xiao, J. Wang, and Z. Huang, “Molecular basis of BACE1 modulation revealed by machine learning, molecular simulations, and experimental validation,” Int. J. Biol. Macromol., p. 152409, 2026, https://doi.org/10.1016/j.ijbiomac.2026.152409.

F. Gerz and M. Jelali, “Towards reliable data augmentation in machine learning: Practices to prevent data leakage,” Smart Agric. Technol., vol. 15, p. 102462, 2026, https://doi.org/10.1016/j.atech.2026.102462.

Y. Rimal and N. Sharma, “Ensemble machine learning prediction accuracy: local vs. global precision and recall for multiclass grade performance of engineering students,” Front. Educ., vol. 10, 2025, https://doi.org/10.3389/feduc.2025.1571133.

Graphical abstract 14827

Downloads

Published

2026-06-26

Issue

Section

Articles