A Hybrid Machine Learning and Digital Forensics Framework for Detecting Coordinated Suspicious Accounts on X
DOI:
https://doi.org/10.12928/mf.v8i2.16449Keywords:
Fake Accounts, Machine Learning, Stylometry, Digital Forensics, Coordinated Inauthentic BehaviorAbstract
Coordinated Inauthentic Behavior (CIB) on X threatens public discourse, yet single-method approaches fail to detect implicit coordination. This study proposes a conditional hybrid framework spanning four forensic stages: (1) automated classification using multilingual BERT text embeddings concatenated with 21 numerical features, (2) sockpuppet candidate generation via Sentence-BERT stylometric fingerprinting and cosine similarity, (3) graph-based coordination analysis using the Louvain algorithm, and (4) manual forensic confirmation with OSINT investigation. Trained on the Cresci-2017 benchmark which exhibits substantial temporal and linguistic domain mismatch from our 2025 Indonesian target domain, the BERT model achieved an accuracy, precision, recall, and F1-score of 1.00 on its held-out test partition prior to deployment. The framework was applied to 1,309 Indonesian Military Law (UU TNI) tweets from 720 accounts, which resulted in the identification of 114 suspicious accounts (15.8%) and the extraction of 28 candidate accounts from 6 groups. Manual validation by the authors through side-by-side timeline inspection confirmed one pair of accounts with synchronized and verbatim identical content. The absence of direct network interactions despite high stylometric similarity is interpreted as consistent with implicit CIB evasion strategies, but attribution to organized groups cannot be definitively established on the basis of platform data alone.
References
S. Gurajala, J. S. White, B. Hudson, B. R. Voter, and J. N. Matthews, “Profile characteristics of fake Twitter accounts,” Big Data & Society, vol. 3, no. 2, p. 2053951716674236, Dec. 2016, doi: 10.1177/2053951716674236.
E. Ferrara, O. Varol, C. Davis, F. Menczer, and A. Flammini, “The rise of social bots,” Commun. ACM, vol. 59, no. 7, pp. 96–104, Jun. 2016, doi: 10.1145/2818717.
O. Varol, E. Ferrara, C. Davis, F. Menczer, and A. Flammini, “Online Human-Bot Interactions: Detection, Estimation, and Characterization,” ICWSM, vol. 11, no. 1, pp. 280–289, May 2017, doi: 10.1609/icwsm.v11i1.14871.
Meta Transparency Center, “Inauthentic Behavior | Transparency Center.” Accessed: Jul. 18, 2026. [Online]. Available: https://transparency.meta.com/policies/community-standards/inauthentic-behavior
F. Giglietto, N. Righetti, L. Rossi, and G. Marino, “It takes a village to manipulate the media: coordinated link sharing behavior during 2018 and 2019 Italian elections,” Information, communication & society : ICS ; an international journal for the Information Age, vol. 23, no. 6, pp. 867–891, 2020, doi: 10.1080/1369118X.2020.1739732.
N. Reimers and I. Gurevych, “Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks,” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), Hong Kong, China: Association for Computational Linguistics, 2019, pp. 3980–3990. doi: 10.18653/v1/D19-1410.
L. Mannocci, M. Mazza, A. Monreale, M. Tesconi, and S. Cresci, “Detection and Characterization of Coordinated Online Behavior: A Survey,” 2024, arXiv. doi: 10.48550/ARXIV.2408.01257.
A. Homsi, J. Al Nemri, N. Naimat, H. Abdul Kareem, M. Al-Fayoumi, and M. Abu Snober, “Detecting Twitter Fake Accounts using Machine Learning and Data Reduction Techniques:,” in Proceedings of the 10th International Conference on Data Science, Technology and Applications, Online Streaming, --- Select a Country ---: SCITEPRESS - Science and Technology Publications, 2021, pp. 88–95. doi: 10.5220/0010604300880095.
B. Bharti, N. S. Gill, and P. Gulia, “Exploring machine learning techniques for fake profile detection in online social networks,” IJECE, vol. 13, no. 3, p. 2962, Jun. 2023, doi: 10.11591/ijece.v13i3.pp2962-2971.
M. Mohammadrezaei, M. E. Shiri, and A. M. Rahmani, “Identifying Fake Accounts on Social Networks Based on Graph Analysis and Classification Algorithms,” Security and Communication Networks, vol. 2018, pp. 1–8, Aug. 2018, doi: 10.1155/2018/5923156.
S. Cresci, R. Di Pietro, M. Petrocchi, A. Spognardi, and M. Tesconi, “The Paradigm-Shift of Social Spambots: Evidence, Theories, and Tools for the Arms Race,” in Proceedings of the 26th International Conference on World Wide Web Companion - WWW ’17 Companion, Perth, Australia: ACM Press, 2017, pp. 963–972. doi: 10.1145/3041021.3055135.
S. R. Sahoo and B. B. Gupta, “Hybrid approach for detection of malicious profiles in twitter,” Computers & Electrical Engineering, vol. 76, pp. 65–81, Jun. 2019, doi: 10.1016/j.compeleceng.2019.03.003.
M. Mazza, S. Cresci, M. Avvenuti, W. Quattrociocchi, and M. Tesconi, “RTbust: Exploiting Temporal Patterns for Botnet Detection on Twitter,” in Proceedings of the 10th ACM Conference on Web Science, Boston Massachusetts USA: ACM, Jun. 2019, pp. 183–192. doi: 10.1145/3292522.3326015.
L. Nizzoli, S. Tardelli, M. Avvenuti, S. Cresci, and M. Tesconi, “Coordinated Behavior on Social Media in 2019 UK General Election,” ICWSM, vol. 15, pp. 443–454, May 2021, doi: 10.1609/icwsm.v15i1.18074.
F. Cinus, M. Minici, L. Luceri, and E. Ferrara, “Exposing Cross-Platform Coordinated Inauthentic Activity in the Run-Up to the 2024 U.S. Election,” in Proceedings of the ACM on Web Conference 2025, Sydney NSW Australia: ACM, Apr. 2025, pp. 541–559. doi: 10.1145/3696410.3714698.
D. Pacheco, P.-M. Hui, C. Torres-Lugo, B. T. Truong, A. Flammini, and F. Menczer, “Uncovering Coordinated Networks on Social Media: Methods and Case Studies,” ICWSM, vol. 15, pp. 455–466, May 2021, doi: 10.1609/icwsm.v15i1.18075.
T. Solorio, R. Hasan, and M. Mizan, “A Case Study of Sockpuppet Detection in Wikipedia,” in Proc. Workshop on Language Analysis in Social Media (LASM), 2013, pp. 59–68.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” presented at the Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Jun. 2019, pp. 4171–4186. doi: 10.18653/v1/N19-1423.
V. D. Blondel, J.-L. Guillaume, R. Lambiotte, and E. Lefebvre, “Fast unfolding of communities in large networks,” J. Stat. Mech., vol. 2008, no. 10, p. P10008, Oct. 2008, doi: 10.1088/1742-5468/2008/10/P10008.
A. Alharbi, H. Dong, X. Yi, Z. Tari, and I. Khalil, “Social Media Identity Deception Detection: A Survey,” ACM Comput. Surv., vol. 54, no. 3, pp. 1–35, Apr. 2022, doi: 10.1145/3446372.
H. Satria, tweet-harvest: A Twitter crawler. (2025). GitHub repository. [GitHub repository]. Available: https://github.com/helmisatria/tweet-harvest
J. Echeverría, E. De Cristofaro, N. Kourtellis, I. Leontiadis, G. Stringhini, and S. Zhou, “LOBO: Evaluation of Generalization Deficiencies in Twitter Bot Classifiers,” in Proceedings of the 34th Annual Computer Security Applications Conference, San Juan PR USA: ACM, Dec. 2018, pp. 137–146. doi: 10.1145/3274694.3274738.
C. Hays, Z. Schutzman, M. Raghavan, E. Walk, and P. Zimmer, “Simplistic Collection and Labeling Practices Limit the Utility of Benchmark Datasets for Twitter Bot Detection,” in Proceedings of the ACM Web Conference 2023, Austin TX USA: ACM, Apr. 2023, pp. 3660–3669. doi: 10.1145/3543507.3583214.
F. Pedregosa et al., “Scikit-learn: Machine Learning in Python,” Journal of Machine Learning Research, vol. 12, no. 85, pp. 2825–2830, 2011.
T. Pires, E. Schlinger, and D. Garrette, “How Multilingual is Multilingual BERT?,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Florence, Italy: Association for Computational Linguistics, 2019, pp. 4996–5001. doi: 10.18653/v1/P19-1493.
E. Stamatatos, “A survey of modern authorship attribution methods,” J. Am. Soc. Inf. Sci., vol. 60, no. 3, pp. 538–556, Mar. 2009, doi: 10.1002/asi.21001.
A. A. Hagberg, D. A. Schult, and P. J. Swart, “Exploring Network Structure, Dynamics, and Function using NetworkX,” presented at the Python in Science Conference, Pasadena, California, Jun. 2008, pp. 11–15. doi: 10.25080/TCWV9851.
Techniques, I. I. E. C., “Iso/iec 27037: 2012 information technology security techniques guidelines for identification collection acquisition and preservation of digital evidence,” ISO/IEC-The standard was published in October. Accessed: Jul. 19, 2026. [Online]. Available: https://iecstandardstore.com/
K. Kent, S. Chevalier, T. Grance, and H. Dang, “Guide to integrating forensic techniques into incident response,” National Institute of Standards and Technology, Gaithersburg, MD, NIST SP 800-86, 2006. doi: 10.6028/NIST.SP.800-86.
S. Kudugunta and E. Ferrara, “Deep neural networks for bot detection,” Information Sciences, vol. 467, pp. 312–322, Oct. 2018, doi: 10.1016/j.ins.2018.08.019.
D. Dukic, D. Keca, and D. Stipic, “Are You Human? Detecting Bots on Twitter Using BERT,” in 2020 IEEE 7th International Conference on Data Science and Advanced Analytics (DSAA), sydney, Australia: IEEE, Oct. 2020, pp. 631–636. doi: 10.1109/DSAA49011.2020.00089.
L. Cima, L. Mannocci, M. Avvenuti, M. Tesconi, and S. Cresci, “Coordinated Behavior in Information Operations on Twitter,” IEEE Access, vol. 12, pp. 61568–61585, 2024, doi: 10.1109/ACCESS.2024.3393482.
A. S. Franzke, A. Bechmann, C. M. Ess, and M. Zimmer, “Internet Research: Ethical Guidelines 3.0,” AoIR (The International Association of Internet Researchers), Report, 2020.
C. Fiesler and N. Proferes, “‘Participant’ Perceptions of Twitter Research Ethics,” Social Media + Society, vol. 4, no. 1, p. 2056305118763366, Jan. 2018, doi: 10.1177/2056305118763366.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Riko Iman Decamarta, Yudi Prayudi

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Start from 2019 issues, authors who publish with JURNAL MOBILE AND FORENSICS agree to the following terms:
- Authors retain copyright and grant the journal right of first publication with the work simultaneously licensed under a Creative Commons Attribution License (CC BY-SA 4.0) that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work.

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.





