2608004823
  • Open Access
  • Article

Breaking and Defending LLM-Powered Social Media Bot Detection Systems

  • Nof Orenstein *,   
  • Yoni Birman

Received: 09 Jul 2026 | Revised: 15 Jul 2026 | Accepted: 05 Aug 2026 | Published: 11 Aug 2026

Abstract

The rise of social media bots poses a persistent threat, enabling misinformation, public opinion manipulation, and erosion of trust in online platforms. To combat this, machine learning systems have been developed to detect and limit bot activity. However, attackers continuously adapt through adversarial optimization, behavior imitation, and semantic manipulation strategies, creating an escalating arms race with detection tools. Recent advances in LLMs have significantly improved bot detection by enabling deeper semantic and contextual analysis. However, this shift also introduces new attack surfaces, allowing adversaries to craft exploits that directly target LLM reasoning and generation mechanisms. Industry tools like Anthropic’s Claude Code Security similarly leverage LLMs for security, motivating our study of their attack surfaces. In this work, we explore both offensive and defensive aspects of LLM-powered, threat-specific cybersecurity applications. While centered on the challenge of social media bot detection, our methodology and insights generalize to a broad class of LLM-powered cybersecurity systems, including phishing detection, email classification, fraud analysis, and more. We introduce two novel adversarial attack strategies that systematically exploit semantic and contextual weaknesses of LLM-based classifiers, degrading LLM performance in bot detection by up to 48%, and propose a robust multi-LLM defense architecture designed to preserve detection reliability under adaptive adversarial conditions. Our solution, LSABRE, is a multi-LLM framework that improves robustness across various attacks, maintaining 86% detection accuracy even under strong adaptive adversarial attacks.

References 

  • 1.

    Feng, S.; Wan, H.; Wang, N.; et al. What Does the Bot Say? Opportunities and Risks of Large Language Models in Social Media Bot Detection. arXiv 2024, arXiv:2402.00371.

  • 2.

    MITRE ATLAS. Available online: https://atlas.mitre.org/ (accessed on 15 July 2026).

  • 3.

    OWASP. OWASP Top 10 for LLM Applications. 2024. Available online: https://owasp.org/www-project-top-10-for-largelanguage-model-applications/ (accessed on 15 July 2026).

  • 4.

    Anthropic. Making Frontier Cybersecurity Capabilities Available to Defenders. Anthropic Blog, 20 February 2026. Available online : https://www.anthropic.com/news/claude-code-security (accessed on 24 February 2026).

  • 5.

    Feng, S.; Wan, H.; Wang, N.; et al. TwiBot-20: A Comprehensive Twitter Bot Detection Benchmark. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management, Gold Coast, QLD, Australia, 1–5 November 2021.

  • 6.

    Dubey, A.; Jauhri, A.; Pandey, A.; et al. The Llama 3 Herd of Models. arXiv 2024, arXiv:2407.21783.

  • 7.

    Jiang, A.Q.; Sablayrolles, A.; Mensch, A.; et al. Mistral 7B. arXiv 2023, arXiv:2310.06825.

  • 8.

    Gemma Team; Mesnard, T.; Hardin, C.; et al. Gemma: Open Models Based on Gemini Research and Technology. arXiv 2024, arXiv:2403.08295.

  • 9.

    Beskow, D.M.; Carley, K.M. Bot-hunter: A Tiered Approach to Detecting & Characterizing Automated Activity on Twitter. In Proceedings of the International Conference on Social Computing, Behavioral-Cultural Modeling and Prediction and Behavior Representation in Modeling and Simulation, Washington, DC, USA, 10–13 July 2018.

  • 10.

    Yang, K.C.; Ferrara, E.; Menczer, F. Botometer 101: Social bot practicum for computational social scientists. J. Comput. Soc. Sci. 2022, 5, 1511–1528.

  • 11.

    Hayawi, K.; Mathew, S.S.; Venugopal, N.; et al. DeeProBot: a hybrid deep neural network model for social bot detection based on user profile data. Soc. Netw. Anal. Min. 2022, 12, 43.

  • 12.

    Wu, J.; Ye, X.; Man, Y. BotTriNet: A Unified and Efficient Embedding for Social Bots Detection via Metric Learning. In Proceedings of the 2023 11th International Symposium on Digital Forensics and Security (ISDFS), Chattanooga, TN, USA, 11–12 May 2023; pp. 1–6.

  • 13.

    Huang, Z.; Lv, Z.; Han, X.; et al. Social Bot-Aware Graph Neural Network for Early Rumor Detection. In Proceedings of the 29th International Conference on Computational Linguistics, Gyeongju, Republic of Korea, 12–17 October 2022.

  • 14.

    Feng, S.; Tan, Z.; Li, R.; et al. Heterogeneity-aware Twitter Bot Detection with Relational Graph Transformers. In Proceedings of the 36th AAAI Conference on Artificial Intelligence, Virtual, 22 February–1 March 2022.

  • 15.

    Heidari, M.; Jones, J.H. Using BERT to Extract Topic-Independent Sentiment Features for Social Media Bot Detection. In Proceedings of the 2020 11th IEEE Annual Ubiquitous Computing, Electronics & Mobile Communication Conference (UEMCON), Virtual, 28–31 October 2020; pp. 0542–0547.

  • 16.

    Cai, Z.; Tan, Z.; Lei, Z.; et al. LMBot: Distilling Graph Knowledge into Language Model for Graph-less Deployment in Twitter Bot Detection. In Proceedings of the 17th ACM International Conference on Web Search and Data Mining, Merida, Mexico, 4–8 March 2024.

  • 17.

    Zou, A.; Wang, Z.; Kolter, J.Z.; et al. Universal and Transferable Adversarial Attacks on Aligned Language Models. arXiv 2023, arXiv:2307.15043.

  • 18.

    Perez, F.; Ribeiro, I. Ignore Previous Prompt: Attack Techniques for Language Models. arXiv 2022, arXiv:2211.09527.

  • 19.

    Feng, S.; Tan, Z.; Wan, H.; et al. TwiBot-22: Towards Graph-Based Twitter Bot Detection. arXiv 2022, arXiv:2206.04564.

  • 20.

    Cresci, S.; Pietro, R.D.; Petrocchi, M.; et al. The Paradigm-Shift of Social Spambots: Evidence, Theories, and Tools for the Arms Race. In Proceedings of the 26th International Conference on World Wide Web Companion, Perth, WA, Australia, 3–7 April 2017.

  • 21.

    Wan, A.;Wallace, E.; Shen, S.; et al. Poisoning Language Models During Instruction Tuning. arXiv 2023, arXiv:2305.00944.

  • 22.

    Dong, Z.; Zhou, Z.; Yang, C.; et al. Attacks, Defenses and Evaluations for LLM Conversation Safety: A Survey. arXiv 2024, arXiv:2402.09283.

  • 23.

    Ganguli, D.; Lovitt, L.; Kernion, J.; et al. Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned. arXiv 2022, arXiv:2209.07858.

  • 24.

    Wallace, E.; Rodriguez, P.; Feng, S.; et al. Trick Me If You Can: Human-in-the-Loop Generation of Adversarial Examples for Question Answering. Trans. Assoc. Comput. Linguist. 2018, 7, 387–401.

  • 25.

    Shen, X.; Chen, Z.J.; Backes, M.; et al. “Do Anything Now”: Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models. arXiv 2023, arXiv:2308.03825.

  • 26.

    Tian, Y.; Yang, X.; Zhang, J.; et al. Evil Geniuses: Delving into the Safety of LLM-based Agents. arXiv 2023, arXiv:2311.11855.

  • 27.

    Shah, R.; Feuillade-Montixi, Q.; Pour, S.; et al. Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation. arXiv 2023, arXiv:2311.03348.

  • 28.

    Liu, X.; Xu, N.; Chen, M.; et al. AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models. arXiv 2023, arXiv:2310.04451.

  • 29.

    Liu, Y.; Jia, Y.; Geng, R.; et al. Formalizing and Benchmarking Prompt Injection Attacks and Defenses. In Proceedings of the 33rd USENIX Security Symposium (USENIX Security 24), Philadelphia, PA, USA, 14–16 August 2024.

  • 30.

    Lin, G.; Tanaka, T.; Zhao, Q. Large Language Model Sentinel: LLM Agent for Adversarial Purification. arXiv 2024, arXiv:2405.20770.

Share this article:
How to Cite
Orenstein, N.; Birman, Y. Breaking and Defending LLM-Powered Social Media Bot Detection Systems . Pragmatic Cybersecurity 2026, 1 (2), 10. https://doi.org/10.53941/pc.2026.100010.
RIS
BibTex
Copyright & License
article copyright Image
Copyright (c) 2026 by the authors.