Large language models (LLMs) have demonstrated remarkable capabilities across languages, yet their safety alignment remains predominantly evaluated in monolingual, especially English, settings. Code-switching (also known as code-mixing)—the alternation between two or more languages within a single utterance or conversation—is a pervasive phenomenon in multilingual societies and digital communication. Recent evidence reveals that code-switched inputs can substantially degrade the safety robustness of LLMs, enabling jailbreaks and harmful outputs that monolingual safety mechanisms fail to intercept. This survey provides the first comprehensive review of research on code-switched LLM safety and robustness evaluation. We systematize existing attack methodologies (including Code-Switching Red-Teaming, Multilingual Blending, and attributional analyses), evaluation datasets and metrics, empirical findings on attack success rates and linguistic factors, and emerging mitigation strategies. We further situate these works within the broader multilingual safety literature, highlight critical gaps in linguistic coverage, cultural contextualization, and mechanistic understanding, and outline a research agenda toward equitable, linguistically inclusive LLM safety. Our synthesis aims to guide researchers and practitioners in developing more robust evaluation frameworks and alignment techniques for real-world multilingual deployments.
Introduction
The text is a survey of safety and robustness problems in Large Language Models (LLMs) when users employ code-switching or code-mixing—the practice of combining two or more languages, scripts, dialects, or transliterated forms within the same conversation or sentence.
The main concern is that most existing LLM safety benchmarks and alignment methods are English-centric and monolingual. As a result, models that safely reject harmful prompts in English may become significantly more vulnerable when the same intent is expressed using mixed languages. Research shows that code-switched prompts can substantially increase Attack Success Rates (ASR), sometimes reaching very high levels, particularly when low-resource languages or different language families are combined.
Key points
Code-switching is common in real-world communication, particularly on social media, messaging platforms, and online communities across multilingual regions.
LLM safety does not consistently transfer across languages. High-resource languages generally receive stronger safety alignment, while low-resource languages, transliteration, and mixed-language inputs can expose weaknesses.
Code-Switching Red-Teaming (CSRT) provides an automated method for generating multilingual adversarial prompts. Studies report substantially higher attack success than conventional English-only attacks, with vulnerability increasing as more languages are mixed.
Multilingual Blending demonstrates that simply interleaving multiple languages can bypass safety mechanisms. Reported bypass rates were substantially higher for some mixed-language prompts than for monolingual baselines.
Mechanistic studies suggest that code-mixing can cause saliency drift, where the model's attention or attribution moves away from safety-critical words, making harmful intent harder for safety mechanisms to detect.
Romanized and transliterated languages, such as Roman Urdu, as well as phonetic variations and code-switched speech, can also create safety vulnerabilities.
Evaluation itself is difficult because many automatic safety classifiers are trained primarily on English. Human evaluation requires bilingual or multilingual expertise and is expensive.
Mitigation strategies
The survey identifies several possible defenses:
Translating code-switched input into a high-resource language before applying safety filters.
Adding realistic code-switched examples to safety fine-tuning and preference-training datasets.
Developing multilingual and code-switch-aware preference optimization.
Detecting mixed-language harmful intent at runtime.
However, no current mitigation provides comprehensive protection across languages, scripts, and cultural contexts, and some approaches introduce additional computational cost or latency.
Research gaps
The paper identifies six major areas requiring further research:
Greater linguistic and cultural coverage, especially low-resource and indigenous languages.
Better mechanistic understanding of why safety mechanisms fail under code-switching.
Multi-turn evaluation that reflects realistic conversations rather than isolated prompts.
Multimodal safety testing, including code-switched speech, images, and video.
More reliable evaluation metrics and multilingual safety classifiers.
Better alignment datasets and training methods specifically designed for code-switched communication.
Conclusion
Code-switching is a fundamental feature of human multilingual communication and an increasingly recognized attack surface for large language models. Empirical evidence from CSRT, Multilingual Blending, attributional analyses, and related studies demonstrates that safety guardrails optimized for monolingual English (and, to a lesser extent, other high-resource languages) are systematically fragile under mixed-language inputs. Attack success rates rise substantially, sometimes dramatically, and the failures appear rooted in attributional and representational mismatches rather than mere incomprehension.
Addressing this vulnerability is both a technical and an equity imperative. Billions of users communicate in code-switched forms; safety systems that fail them are incomplete. By systematizing the current literature, clarifying evaluation practices, and charting open problems, this survey aims to accelerate the development of linguistically inclusive, robust LLM safety evaluation and alignment. The path forward requires collaboration across NLP, linguistics, safety research, and the communities whose languages and communicative practices have been historically underrepresented in AI development
References
[1] Banerjee, S., Chatterjee, P., Kumar, S., Layek, S., Agrawal, P., Hazra, R., & Mukherjee, A. (2025). Attributional Safety Failures in Large Language Models under Code-Mixed Perturbations. arXiv preprint arXiv:2505.14469.
[2] Gupta, H., Jayarao, P., Varshney, N., & Dwivedi, C. (2026). Code-Mixing and Code-Switching for Text in the LLM Era: A Playbook of Models, Data, Evaluation, and Open Problems. arXiv preprint arXiv:2602.11181.
[3] Sheth, R., Sinha, S. R., Patil, M., Beniwal, H., & Singh, M. (2025). Beyond Monolingual Assumptions: A Survey of Code-Switched NLP in the Era of Large Language Models. arXiv preprint arXiv:2510.07037.
[4] Song, J., Huang, Y., Zhou, Z., & Ma, L. (2025). Multilingual Blending: Large Language Model Safety Alignment Evaluation with Language Mixture. In Findings of the Association for Computational Linguistics: NAACL 2025, pages 3433–3449.
[5] Winata, G. I., Aji, A. F., Yong, Z.-X., & Solorio, T. (2023). The Decades Progress on Code-Switching Research in NLP: A Systematic Survey on Trends and Challenges. In Findings of the Association for Computational Linguistics: ACL 2023, pages 2936–2978.
[6] Yong, Z.-X., Ermis, B., Fadaee, M., Bach, S. H., & Kreutzer, J. (2025). The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 15845–15860.
[7] Yoo, H., Yang, Y., & Lee, H. (2025). Code-Switching Red-Teaming: LLM Evaluation for Safety and Multilingual Understanding. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 13392–13413.