LLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review
Summary
This systematic literature review analyzes Large Language Model (LLM) safety alignment in low-resource languages, synthesizing 50 studies selected from approximately 1,500 candidates using PRISMA 2020 guidelines. The authors identify a persistent safety gap where models trained on high-resource languages, primarily English, fail to generalize safety constraints to underrepresented languages. Key vulnerabilities include significantly higher rates of cross-lingual jailbreaks, susceptibility to code-switching attacks, and the generation of culturally inappropriate or toxic content, such as unsafe product recommendations in Hausa. The review categorizes alignment approaches into data adaptation, objective optimization, and mechanistic alignment. While methods like culturally grounded dataset creation and reward gap optimization show promise, the authors note that translated English benchmarks often fail to capture locally specific harms. Furthermore, uneven pre-training coverage and insufficient native-language preference data hinder effective cross-lingual transfer. African languages are particularly underrepresented in existing safety benchmarks compared to other regions. The study concludes that future progress requires moving beyond translation-based datasets toward participatory data collection, balanced multilingual pre-training, and culturally aware evaluation frameworks to ensure equitable safety protections for all language communities.
PDF viewer
Chunks(31)
Chunk 0 · 1,998 chars
LLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review Valdini Douglace Lemofouet1 Blessing Ngozi Uzor1 Paula Chikaodinaka Anyanwu1 Danielle Blanche Kapsa1 Sukairaj Hafiz Imam2 P Sam Sahil7 Abigail Oppong3 Tassallah Abdullahi4 Clemencia Siro5 Idris Abdulmumin6 Seid Muhie Yimam7 Shamsuddeen Hassan Muhammad8,2 1African Institute for Mathematical Sciences (AIMS), Cameroon 2Bayero University Kano 3 Independent Researcher 4Brown University 5Centrum Wiskunde & Informatica 6University of Pretoria 7University of Hamburg 8Imperial College London Abstract Large Language Models (LLMs) have achieved substantial progress in safety alignment, yet their safety guarantees remain significantly weaker in low-resource and multilingual set- tings than in high-resource languages. In this paper, we conduct a Systematic Literature Re- view (SLR) of LLM safety alignment in low- resource languages by adopting the PRISMA 2020 methodology. Out of roughly 1,500 pa- pers identified from Semantic Scholar, arXiv, and OpenAlex, 50 relevant studies have been selected and analyzed. Our review is organized around four themes: safety alignment meth- ods, multilingual safety risks, evaluation bench- marks, and cross-lingual transferability. We fur- ther propose a taxonomy of safety alignment approaches based on three adaptation mech- anisms: data adaptation, objective optimiza- tion, and mechanistic alignment. Across lit- erature, translated English benchmarks fail to sufficiently represent culturally rooted harms, and multilingual models are more vulnerable to cross-lingual jailbreaks, code-switching at- tacks, and safety degradation in underrepre- sented languages. These failures are driven by several key factors, including uneven multilin- gual pre-training coverage, insufficient native- language preference data, poor transfer of safety representations, and a lack of culturally aware evaluation frameworks. The review also notes that many low-resource languages, es- pecially
Chunk 1 · 1,989 chars
es. These failures are driven by several key factors, including uneven multilin- gual pre-training coverage, insufficient native- language preference data, poor transfer of safety representations, and a lack of culturally aware evaluation frameworks. The review also notes that many low-resource languages, es- pecially African languages, have fewer safety benchmarks available than other multilingual regions. Overall, the results reveal a persistent multilingual safety gap, and suggest that fu- ture progress will require culturally grounded benchmarks, participatory data collection, bal- anced multilingual pre-training, and scalable multilingual alignment methods. 1 Introduction In recent years, Large Language Models (LLMs) have been increasingly deployed in domains such as education, healthcare, governance, and digital communication, raising concerns about their safety, reliability, and alignment with human values. A language model is considered safe if it avoids gen- erating harmful, illegal, toxic, dangerous, or biased content, while remaining useful to users. Recent advances in safety alignment have introduced sev- eral techniques for reducing harmful, toxic, biased, or otherwise unsafe model behavior. These tech- niques include reinforcement learning from human feedback (RLHF), constitutional AI, safety fine- tuning, and adversarial red teaming (Ouyang et al., 2022; Bai et al., 2022; Perez et al., 2022). As these systems continue to grow in capability and adop- tion, ensuring their safe and responsible behavior has become a major research priority. Most existing alignment techniques and eval- uation frameworks have been designed primar- ily for high-resource languages, particularly En- glish (Yong et al., 2025a). These approaches often depend on large-scale annotated datasets, bench- marks, and an extensive human feedback pipeline, resources that remain limited or unavailable for many underrepresented languages. Safety alignment learned in English does
Chunk 2 · 1,999 chars
y for high-resource languages, particularly En- glish (Yong et al., 2025a). These approaches often depend on large-scale annotated datasets, bench- marks, and an extensive human feedback pipeline, resources that remain limited or unavailable for many underrepresented languages. Safety alignment learned in English does not transfer reliably to other languages. Recent stud- ies show that LLMs trained to reject harmful in- structions in English may comply with equivalent prompts written in low-resource languages (Yong et al., 2023). Previous studies also suggest that these failures are highly related to the imbalanced coverage of multilingual pre-training (Shen et al., 2024). In practice, these kinds of gaps can trans- late into weaker safety protections for speakers of underrepresented languages. For example, Inuwa- Dutse (2025) presented examples of an LLM de- ployed in Hausa producing false and unsafe rec- ommendations for toxic local products that were presented as safe for humans. In order to better understand this emerging area, this paper presents a Systematic Literature Re- view (SLR) of safety alignment in low-resource arXiv:2608.14626v1 [cs.CL] 20 Jul 2026 -- 1 of 13 -- languages using the PRISMA 2020 framework. We synthesize existing research across four di- mensions: safety alignment methods, multilin- gual safety risks, evaluation benchmarks and cross- lingual transferability. We further propose a taxon- omy of safety alignment approaches by their main adaptation mechanisms and identify key challenges limiting effective safety transfer across languages. In this review, we highlight current research trends, methodological limitations, and future directions for building safer and more inclusive multilingual language models. This review is guided by the following research questions: • RQ1: How are LLM safety alignment meth- ods adapted for low-resource languages? • RQ2: What safety risks, adversarial behav- iors, and cultural harms emerge in multilin- gual
Chunk 3 · 1,985 chars
ture directions for building safer and more inclusive multilingual language models. This review is guided by the following research questions: • RQ1: How are LLM safety alignment meth- ods adapted for low-resource languages? • RQ2: What safety risks, adversarial behav- iors, and cultural harms emerge in multilin- gual LLM settings? • RQ3: What datasets, benchmarks, and eval- uation frameworks exist for safety alignment in low-resource languages? • RQ4: What factors affect cross-lingual trans- fer of safety alignment to low-resource lan- guages? 2 Related Work Several studies have been conducted on the safety of multilingual LLMs. Some of these works aimed to analyze and review the field. Among these, Yong et al. (2025b) conducted a descrip- tive systematic review of nearly 300 publications, revealing a persistent and growing language gap in LLM safety research. Also, Banerjee et al. (2026) synthesized findings for Global South lan- guages: "safety guardrails weaken sharply on low- resource and code-mixed inputs". They also pro- posed parameter- efficient steering and participa- tory workflows that enable communities to define and mitigate harm. Shen et al. (2024) examines the safety challenges of LLMs in multilingual set- tings. They observed that LLMs tend to generate unsafe responses much more often when a mali- cious prompt is written in a lower-resource lan- guage. These works share four limitations: (1) none fol- lows a formal SLR methodology (explicit research questions, inclusion/exclusion criteria, PRISMA flow), (2) African languages are subsumed under generic “low - resource” categories, (3) findings are not structured into actionable dimensions (methods, risks, benchmarks), and (4) they do not synthesise mechanistic explanations for cross- lingual transfer failures. Our work addresses these gaps by presenting the first systematic literature review (PRISMA) dedicated exclusively to LLM safety alignment in African and low-resource languages. 3
Chunk 4 · 1,997 chars
nable dimensions (methods, risks, benchmarks), and (4) they do not synthesise mechanistic explanations for cross- lingual transfer failures. Our work addresses these gaps by presenting the first systematic literature review (PRISMA) dedicated exclusively to LLM safety alignment in African and low-resource languages. 3 Methodology We performed a systematic literature review based on the PRISMA 2020 guidelines (Page et al., 2021). The PRISMA flow diagram is shown in Fig. 1. Relevant articles were collected from Semantic Scholar, arXiv, and OpenAlex using searches in- volving queries belonging to the following three keyword sets: (i) technology (LLM, instruction tuning, foundation models), (ii) safety (align- ment, jailbreak, adversarial attack, toxicity, RLHF, DPO), and (iii) linguistic scope (low-resource lan- guages, multilingual, African languages, cross- lingual, code-switching, Hausa, Swahili, Yoruba, Amharic). Detailed search strings are available in Appendix A.1.This process generated approxi- mately 1,500 papers, which were pared down to 1,300 after de-duplication. A Python-based filter system was used with criteria such as year cut- offs (>2020) and multi-faceted keyword relevancy, generating 242 potential papers to review. The inclusion criteria included the focus on safety or adversarial robustness for LLMs in multilingual and low-resource contexts, while purely English and non-empirical works were excluded. We then applied a two-step screening process. The first step involved LLM-based evaluation (Claude Sonnet) of each potential paper based on research questions and inclusion/exclusion criteria: 70 papers were selected, and after verification, none of the papers excluded by Claude were useful. We retained 50 papers in the final review set. 4 Safety Alignment Methods for Low-Resource Languages RQ1 — Adaptation of safety methods How are LLM safety alignment methods adapted for low-resource languages? Safety alignment for low-resource languages has become
Chunk 5 · 1,986 chars
ion, none of the papers excluded by Claude were useful. We retained 50 papers in the final review set. 4 Safety Alignment Methods for Low-Resource Languages RQ1 — Adaptation of safety methods How are LLM safety alignment methods adapted for low-resource languages? Safety alignment for low-resource languages has become increasingly important as large language -- 2 of 13 -- Figure 1: PRISMA flow diagram. 1,500 records were narrowed through deduplication, automated scoring, LLM-based eligibility screening, and two rounds of manual revision producing 50 papers for synthesis. models are deployed across multilingual settings. Most safety alignment methods were designed for English and rely on large annotated datasets that do not exist for the majority of the world’s lan- guages. Applying them to low-resource settings requires adapting one of four components: (i) the data adaptation, (ii) the optimization objective, (iii) the model’s internal parameters, or (iv) data. 4.1 Culturally Grounded Data Adaptation The most direct response to low-resource safety failures is to build better multilingual training data. Methods in this family share a common premise: safety quality is primarily a function of how well the training data covers local language patterns and cultural harm definitions. Joshi et al. (2025) addressed this issue through CultureGuard, a four-stage pipeline that generates a culturally grounded safety dataset in target lan- guages using LLMs, filters it with a safety classifier, validates it with native speakers, and fine-tunes a guard model with LoRA. Trained on 386k samples across nine languages, including Hindi and Thai, the resulting model outperforms models trained on translated English data on non-English bench- marks and achieves zero-shot transfer to unseen languages. Shan et al. (2025) introduce SEALGuard, a multilingual guardrail designed to improve the safety alignment across for Southeast Asian lan- guages. SEALGuard is fine-tuned(LoRA)
Chunk 6 · 1,998 chars
ng model outperforms models trained on translated English data on non-English bench- marks and achieves zero-shot transfer to unseen languages. Shan et al. (2025) introduce SEALGuard, a multilingual guardrail designed to improve the safety alignment across for Southeast Asian lan- guages. SEALGuard is fine-tuned(LoRA) on SEALSBench, (266k-prompt benchmark of locally sourced and translated harmful prompts across nine languages). The model improves Defence Success Rate by 48% over LlamaGuard, confirming that English-centric guardrails fail on region-specific jailbreak attempts. Another angle is taken by Tan et al. (2025) with LionGuard 2,a lightweight, data- efficient, localized multilingual content modera- tion classifier for Singapore’s linguistic context. By incorporating Singapore-specific harm cate- gories and code-mixed Singlish data, the model outperforms commercial moderation APIs across 17 benchmarks. The authors also find that naively translated training data reduces performance, rein- forcing the case for localized data construction. Paul et al. (2025) demonstrate that data qual- ity matters more than scale for alignment(Hindi). They translate preference data using Llama-3.1- 405B, which preserves code, formulas, and URLs that standard machine translation corrupts, then filter outputs with FAITH-based quality metrics be- fore running a two-stage alignment process consist- ing of Supervised Fine-Tuning (SFT) and Direct -- 3 of 13 -- Figure 2: Taxonomy of safety alignment methods for low-resource languages, grouped by primary transfer mechanism: data adaptation, objective optimization, and mechanistic alignment. Preference Optimization (DPO). Forty thousand filtered samples match or exceed the alignment per- formance of 200k unfiltered ones. In parallel, Bell et al. (2025) investigate translation-based multi- lingual toxicity detection pipelines, where the ini- tial step is to translate potentially offensive con- tent to English before classifying it. The
Chunk 7 · 1,995 chars
Forty thousand filtered samples match or exceed the alignment per- formance of 200k unfiltered ones. In parallel, Bell et al. (2025) investigate translation-based multi- lingual toxicity detection pipelines, where the ini- tial step is to translate potentially offensive con- tent to English before classifying it. The authors demonstrate that translate-and-classify frameworks always perform better than multilingual classifiers when dealing with most low-resource languages. Apart from multilingual moderation datasets, other researchers have also developed resources for foundational safety alignment datasets in low- resource languages.Huang et al. (2026) present the Tibetan Foundation Dataset (TFD), which is a massive dataset in the Tibetan language that can be used in all stages of LLM training processes, from pre-training to safety alignment and reasoning. For safety alignment purposes, some methods for translating existing datasets suffer from a problem of preserving intent (harmful, toxic). It is in this context that Ge et al. (2025) proposes a framework for toxicity-preserving translation, demonstrated on a code-mixed Singlish safety corpus. 4.2 Cross-Lingual Transfer and Objective Optimization When multilingual safety data is scarce, a second strategy is to modify the training objective so that safety behaviour learned in one language trans- fers to others. This family splits into two com- plementary directions: reward-based optimization and reasoning-based consistency. 4.2.1 Preference and Reward Optimization Zhao et al. (2025) address the paired preference data bottleneck with MPO (Multilingual reward gaP Optimization). Rather than requiring (chosen, rejected) pairs in each target language, MPO min- imizes the reward gap between safe and unsafe outputs directly across languages. It consistently outperforms RLHF and DPO baselines on cross- lingual safety benchmarks under noisy conditions. Lim et al. (2025) compare SFT, DPO, and Kahneman-Tversky Optimization
Chunk 8 · 1,991 chars
chosen, rejected) pairs in each target language, MPO min- imizes the reward gap between safe and unsafe outputs directly across languages. It consistently outperforms RLHF and DPO baselines on cross- lingual safety benchmarks under noisy conditions. Lim et al. (2025) compare SFT, DPO, and Kahneman-Tversky Optimization (KTO) for Singlish, where preference pairs are difficult to collect. Their combined SFT+KTO pipeline re- duces toxicity by 99% and introduces KTO-S, an improved regularization strategy that stabilizes fine- tuning under data scarcity. Zhang et al. (2026a) use knowledge distillation to transfer refusal behaviour -- 4 of 13 -- from OpenAI o1-mini into open-source multilin- gual models via LoRA. Jailbreak resistance im- proves in both high and low-resource languages, but the authors find a critical caveat: unfiltered distillation increases jailbreak success rates by up to 16.6 points because ambiguous “boundary” re- fusals from the teacher confuse the student. Filter- ing these responses mitigates the degradation. 4.2.2 Reasoning and Consistency Transfer Yang et al. (2025) introduced MrGuard, which combines synthetic multilingual data generation, SFT( with QLoRA), and GRPO-based reinforce- ment learning. A distinctive feature is that the model produces explicit reasoning chains rather than binary labels, which improves generalization across eight languages(including LRL) by more than 15% in average F1 and maintains robustness against code-switching attacks. Chen et al. (2025) presented ConsistentGuard, a guardrail model trained using only 1,000 samples. They combines supervised fine-tuning (to provide the model with initial task-specific knowledge), fol- lowed by GRPO (to promote reasoning diversity and length), and finally a CAO (constrained Align- ment Optimization) to align the model’s reasoning process across different languages. The model ob- tained outperforms larger classifiers while provid- ing interpretable reasoning chains, suggesting
Chunk 9 · 1,998 chars
ecific knowledge), fol- lowed by GRPO (to promote reasoning diversity and length), and finally a CAO (constrained Align- ment Optimization) to align the model’s reasoning process across different languages. The model ob- tained outperforms larger classifiers while provid- ing interpretable reasoning chains, suggesting that reasoning supervision is an efficient substitute for scale. Bu et al. (2026) proposes a resource-efficient method for improving multilingual safety align- ment. Their Multi-Lingual Consistency (MLC) loss enforces directional consistency between mul- tilingual representation vectors in a single train- ing update, enabling simultaneous safety align- ment across languages without response-level su- pervision in target languages. Further evaluation across languages and tasks indicates improved cross-lingual generalization, suggesting the ap- proach as a practical solution for multilingual con- sistency alignment under limited supervision. 4.3 Mechanistic and Parameter-Level Alignment A third line of research investigates whether safety alignment can be achieved through direct modifica- tion of model internals, thereby bypassing data col- lection and full retraining entirely. The hypothesis is that safety behaviour is concentrated in identifi- able subsets of parameters. In this line of research , Liang et al. (2026) pro- pose the Sparse Weight Editing (SWE), based on the finding that safety capabilities are concentrated in a sparse subset of model weights. They identify this subset in a safety-aligned high-resource model and compute a closed-form linear transformation to project harmful low-resource representations into the corresponding safety subspaces without gradi- ent computation. SWE was evaluated across eight languages and two model families (Llama-3, Qwen- 2.5). It reduces attack success rates in low-resource languages, with negligible impact on general rea- soning performance. Shin and Hwang (2026) takes a layer-level ap- proach: Activation
Chunk 10 · 1,990 chars
g safety subspaces without gradi- ent computation. SWE was evaluated across eight languages and two model families (Llama-3, Qwen- 2.5). It reduces attack success rates in low-resource languages, with negligible impact on general rea- soning performance. Shin and Hwang (2026) takes a layer-level ap- proach: Activation analysis identifies which trans- former layers carry safety-critical features in a safety-aligned high-resource model. Those lay- ers are then transplanted into a low-resource expert model. The resulting training-free method achieves safety gains on MultiJail while preserving perfor- mance on general benchmarks (MMMLU, BELE- BELE, MGSM). Banerjee et al. (2025) localize the intervention further with Soteria, which identifies language- specific attention heads responsible for harmful outputs via gradient-based attribution and steers only those components. Modifying a small fraction of parameters drastically reduces policy violations while preserving language quality, even in low- resource settings. 5 Safety Risks in Multilingual and Low resource Language Contexts RQ2 - Risks and cultural harms What safety risks, adversarial behaviors, and cultural harms emerge in multilingual LLM settings? Although much research is focused on reducing the safety gap between HRLs and LRLs, under- standing the safety risks faced by low-resource lan- guages is crucial. Here, We present these LLM vulnerabilities in LRL settings. 5.1 Cross-Lingual Jailbreaks A consistent body of work shows that safety vul- nerabilities increase significantly in low-resource and underrepresented languages. Early evidence from Yong et al. (2023) demon- strates that simply translating harmful English -- 5 of 13 -- prompts into low-resource languages can achieve a 79% jailbreak success rate on GPT-4, while high- resource languages remain below 15%. This dis- parity suggests that safety alignment does not gen- eralize evenly across linguistic space. Building on this, Deng et al. (2024)
Chunk 11 · 1,998 chars
translating harmful English -- 5 of 13 -- prompts into low-resource languages can achieve a 79% jailbreak success rate on GPT-4, while high- resource languages remain below 15%. This dis- parity suggests that safety alignment does not gen- eralize evenly across linguistic space. Building on this, Deng et al. (2024) identifies multilingual jailbreak challenges within LLMs and examines two potential risk scenarios: uninten- tional and intentional. They found that LRLs have about 3 times the likelihood of encountering harm- ful content as HRLs. Shen et al. (2024) also sup- ports these findings, noting that in low-resource languages, models not only become less safe but also follow instructions less reliably as linguistic resources decrease. It is important to note that these weaknesses are not limited to translation attacks. In fact, Chrab ˛aszcz et al. (2025) show that for Polish, a small proxy model can generate transferable ad- versarial perturbations at low cost, producing at- tack success rates significantly higher than those observed in English. Pattnayak and Chowdhuri (2026a) reach a related finding across 12 Indic lan- guages through IndicJR: contract-bound prompts in JSON format inflate refusal counts without ac- tually preventing jailbreaks, and safety alignment degrades consistently as language resource level decreases. Atil et al. (2025) extend this picture to ten languages, finding that both logical-expression- based and adversarial-prompt-based jailbreak meth- ods generalize poorly across languages, with low- resource languages remaining more vulnerable. 5.2 Fine-Tuning on New Languages as an Attack Vector Beyond prompting-based attacks, a more subtle risk arises from model adaptation itself. Upadhayay and Behzadan (2025) show that fine- tuning aligned models on new or synthetic lan- guages, even using only benign data, can still de- grade safety alignment. This suggests that learning new linguistic mappings can interfere with previ- ously established
Chunk 12 · 1,993 chars
a more subtle risk arises from model adaptation itself. Upadhayay and Behzadan (2025) show that fine- tuning aligned models on new or synthetic lan- guages, even using only benign data, can still de- grade safety alignment. This suggests that learning new linguistic mappings can interfere with previ- ously established safety constraints. For African language adaptation pipelines, this implies that in- troducing new linguistic domains without safety recalibration may reduce robustness. 5.3 Code-Switching and Multilingual Blending Multilingual interaction further expands the attack surface through code-switching. Song et al. (2025) show that mixing languages within a single prompt significantly reduces safety robustness, with bypass rates reaching 67.23% on GPT-3.5 and 40.34% on GPT-4. These effects are not uniform and depend on linguistic distance and prompt structure. This is particularly relevant in African contexts, where code-switching is a common communicative norm rather than an adversarial construct. 5.4 Culturally Specific Harm A growing body of work also highlights that safety failures are not only linguistic but cultural in nature. For example, Inuwa-Dutse (2025) shows that GPT- OSS-20B produces unsafe or misleading outputs in Hausa, including toxic product recommenda- tions and culturally inappropriate content that can amplify hate speech. Shukla et al. (2026) further demonstrates that composite harms often change meaning during translation, leading multilingual safety systems to miss culturally contextualized unsafe content. Beyond explicit toxicity (Saeed et al., 2026), their multilingual evaluations also reveal subtle stereotype propagation and culturally dependent bias patterns (mostly in low-resource settings) that are often missed by conventional safety bench- marks. 6 Safety Evaluation Benchmarks RQ3 - Datasets and evaluation What datasets, benchmarks, and evaluation frameworks exist for safety alignment in low- resource languages? 6.1 Global
Chunk 13 · 1,998 chars
agation and culturally dependent bias patterns (mostly in low-resource settings) that are often missed by conventional safety bench- marks. 6 Safety Evaluation Benchmarks RQ3 - Datasets and evaluation What datasets, benchmarks, and evaluation frameworks exist for safety alignment in low- resource languages? 6.1 Global Multilingual Safety Benchmarks Recent multilingual safety benchmarks have pro- gressively moved beyond translation toward more culturally diverse evaluation settings. Early efforts such as XSafety (Wang et al., 2024) introduced a benchmark of 28,000 annotated instances across 14 safety issues and 10 languages spanning multiple language families(including Hindi). Building on this direction, LinguaSafe (Ning et al., 2025) com- bines translated, trans-created, and natively writ- ten prompts across 12 languages(including Malay and Bengali). On his side (Kumar et al., 2025) in- troduce POLYGUARDPROMPTS a multilingual benchmark for the evaluation of safety guardrails, Created by combining naturally occurring multilin- gual human-LLM interactions and human-verified -- 6 of 13 -- machine translations of an English-only safety dataset. Other benchmarks focus more specifically on toxicity and adversarial robustness. (Jain et al., 2024) introduce PolygloToxicityPrompts a large- scale multilingual toxicity evaluation benchmark of 425K naturally occurring prompts spanning 17 lan- guages. whereas ML-Bench (Zhao et al., 2026) in- troduces policy-grounded multilingual safety eval- uation covering 14 languages, based on regional regulations. 6.2 Regional and Culturally Localized Benchmarks In other side, a growing body of work focuses on culturally localized safety evaluation rather than direct translation from English benchmarks. In Southeast Asia, SEALSBench (Shan et al., 2025) and SEA-SafeguardBench (Tasawong et al., 2025) introduce multilingual safety datasets built around regional socio-cultural risks and native-language prompts. Several benchmarks also target
Chunk 14 · 1,995 chars
alized safety evaluation rather than direct translation from English benchmarks. In Southeast Asia, SEALSBench (Shan et al., 2025) and SEA-SafeguardBench (Tasawong et al., 2025) introduce multilingual safety datasets built around regional socio-cultural risks and native-language prompts. Several benchmarks also target multilingual ad- versarial behavior and online toxicity. SGTox- icGuard (Hu et al., 2025) evaluates conversa- tional toxicity and red-teaming scenarios across Singapore’s multilingual setting, while Qorgau (Goloburda et al., 2025) was designed for safety evaluation in Kazakh(a low-resource language) and Russian. In South Asia, IndicSafe (Pattnayak and Chowdhuri, 2026b) and IndicJR (Pattnayak and Chowdhuri, 2026a) extend evaluation to culturally grounded harms and jailbreak robustness across Indic languages. Complementing these efforts, SEAHateCheck (Ng et al., 2026) and IndoSafety (Azmi et al., 2025) introduce hate speech and safety evaluation datasets for Southeast Asian languages and regional language varieties. 6.3 African Language Safety Benchmarks Compared with other multilingual regions, bench- mark resources for African languages remain lim- ited. LSR (Faruna, 2026) evaluates the cross- lingual refusal degradation in West African lan- guages such as Yoruba, Hausa, Igbo, and Igala, while (Abdullahi et al., 2026) introduces Ubuntu- Guard, the first African policy-based safety bench- mark. Beyond toxicity evaluation, Uhura (Bayes et al., 2024) studies truthfulness and safety con- straints in African and low-resource settings. Overall, current benchmark development shows a gradual shift toward native-language evaluation, culturally grounded harms, and multilingual ad- versarial testing. However, African languages remain substantially under-represented compared with other multilingual benchmark ecosystems. 7 Cross-Lingual Transferability of Safety Alignment RQ4 — Cross-lingual transfer factors What factors affect cross-lingual transfer of safety
Chunk 15 · 1,997 chars
grounded harms, and multilingual ad- versarial testing. However, African languages remain substantially under-represented compared with other multilingual benchmark ecosystems. 7 Cross-Lingual Transferability of Safety Alignment RQ4 — Cross-lingual transfer factors What factors affect cross-lingual transfer of safety alignment to low-resource languages? A recurring finding across multilingual safety research is that safety alignment is not language- agnostic: alignment behaviors learned in high- resource languages often degrade when transferred to low-resource languages. Most studies attribute this limitation to uneven multilingual representa- tions acquired during pre-training, where safety- relevant knowledge remains concentrated in high- resource languages. Several prior works identify pre-training cov- erage as the primary bottleneck behind transfer degradation. Shen et al. (2024) show that align- ment methods mainly improve safety in languages that are well represented during pre-training, while gains remain limited in low-resource settings. Similarly, Verma and Bharadwaj (2025) show that safety-relevant features cluster around high- resource regions of the latent space, resulting in weaker safety enforcement in underrepresented lan- guages. Upadhayay and Behzadan (2025) further show that introducing new languages during fine- tuning can disrupt previously learned safety behav- ior, suggesting that multilingual adaptation itself may destabilize alignment. Transfer is also affected by linguistic and struc- tural differences across languages. Morphologi- cally rich languages and underrepresented scripts often suffer from fragmented tokenization and weaker semantic representations, reducing the re- liability of safety reasoning and refusal behavior. These issues become particularly visible in African and other low-resource languages, where both pre- training data and alignment supervision remain scarce. More recent work explores whether safety trans- fer can be
Chunk 16 · 1,999 chars
d weaker semantic representations, reducing the re- liability of safety reasoning and refusal behavior. These issues become particularly visible in African and other low-resource languages, where both pre- training data and alignment supervision remain scarce. More recent work explores whether safety trans- fer can be improved directly at the representation level. Liang et al. (2026) propose Sparse Weight Editing (SWE), which transfers safety represen- tations through lightweight parameter transforma- -- 7 of 13 -- tions, while Shin and Hwang (2026) improves mul- tilingual robustness by identifying and replacing safety-critical transformer layers without full re- training. Mechanistic studies such as Zhang et al. (2026b) further suggest that cross-lingual safety behavior may depend on a sparse set of shared multilingual safety neurons that can be selectively targeted during alignment.Wang et al. (2026) find a shared directional structure in refusal behavior across safety-aligned languages, implying that mul- tilingual safety alignment could depend on partially universal latent safety representations. Recent approaches also investigate how safety transfer can scale efficiently across many languages. Bansal and Mishra (2026) show that alignment learned from a carefully selected subset of lan- guages can generalize broadly through multilingual representation clustering. Complementing this di- rection, Bu et al. (2026) introduces a multilingual consistency objective that explicitly enforces align- ment agreement across languages, improving trans- fer without requiring extensive target-language su- pervision. 8 Discussion The common denominator for these four research questions is that the prevailing methodologies used for developing safety alignment are mostly based on English-language data and have not been prop- erly extended to accommodate the other languages of the world. This imbalance spans from the data used in pre-training to alignment procedures them- selves
Chunk 17 · 1,996 chars
four research questions is that the prevailing methodologies used for developing safety alignment are mostly based on English-language data and have not been prop- erly extended to accommodate the other languages of the world. This imbalance spans from the data used in pre-training to alignment procedures them- selves as well as evaluation techniques. Methodologically speaking(RQ1 and RQ4), techniques like parameter-efficient fine-tuning, data synthesis, and cross-lingual transfer might be helpful, yet they are poorly validated in the context of the African languages. As long as the pre-trained models lack coverage of these languages, the gains brought by such techniques are rather marginal. Partly because, at the pre-training stage, these mod- els already under-represent the languages, which makes recovery of any kind of multilingual repre- sentation impossible. A similar discrepancy appears in the safety risk landscape (RQ2). Issues such as cross-lingual jail- break transfer, code-switching based vulnerabili- ties, and unintended safety degradation through benign fine-tuning into new languages have been reported, but seldom evaluated. Furthermore, cul- turally specific risks, such as those reported in Hausa (Inuwa-Dutse, 2025), are almost completely neglected in benchmarks designed with English- centric assumptions. According to Vajjala (2025), however, such limitations do not apply exclusively to models; annotation discrepancies and culturally biased benchmarks contribute to the inadequacy of evaluation measures. The evaluation landscape (RQ3) reveals a sim- ilar pattern. While benchmarks for multiple lan- guages continue to emerge, there is still limited representation of African languages. Translation is still the primary strategy used in the creation of datasets, although this methodology has repeatedly demonstrated its ability to manipulate meanings and mislead safety annotations. Combined, these limitations create a reinforc- ing cycle whereby inadequate
Chunk 18 · 1,999 chars
l limited representation of African languages. Translation is still the primary strategy used in the creation of datasets, although this methodology has repeatedly demonstrated its ability to manipulate meanings and mislead safety annotations. Combined, these limitations create a reinforc- ing cycle whereby inadequate pre-training cover- age translates into poor representation, which then limits the capabilities of the alignment methods, whereas poor benchmarks limit the ability to assess the deficiencies. To overcome this problem, we need to see progress on all fronts simultaneously, from moving away from translation-based dataset building, to better native language benchmark cre- ation, to creating alignment methods that take this problem into consideration. 9 Conclusion This systematic literature review synthesized find- ings from 50 studies on LLM safety alignment in low-resource language settings. The review showed that current safety alignment approaches re- main largely designed for high-resource languages and do not generalize effectively across diverse lin- guistic and cultural contexts. Existing benchmarks rely heavily on translated datasets that often fail to capture culturally grounded harms. The find- ings further indicated that multilingual adaptation and fine-tuning may reduce safety performance in low-resource settings. However, this review is lim- ited to English-language studies from major aca- demic databases and may not fully represent locally published or recent research. Additionally, group- ing languages into broad regional categories may obscure important linguistic differences. Future research should focus on native-language safety benchmarks, balanced multilingual pre-training, culturally aware evaluation methods, and partic- ipatory frameworks that actively involve African language communities in defining AI safety stan- dards and practices. -- 8 of 13 -- References Tassallah Abdullahi, Macton Mgonzo, Mardiyyah Odu- wole, Paul Okewunmi,
Chunk 19 · 1,991 chars
fety benchmarks, balanced multilingual pre-training, culturally aware evaluation methods, and partic- ipatory frameworks that actively involve African language communities in defining AI safety stan- dards and practices. -- 8 of 13 -- References Tassallah Abdullahi, Macton Mgonzo, Mardiyyah Odu- wole, Paul Okewunmi, Abraham Owodunni, Rita- mbhara Singh, and Carsten Eickhoff. 2026. Ubun- tuguard: A culturally-grounded policy benchmark for equitable ai safety in african languages. Preprint, arXiv:2601.12696. Berk Atil, Rebecca J. Passonneau, and Fred Morstatter. 2025. Do methods to jailbreak and defend llms gener- alize across languages? Preprint, arXiv:2511.00689. Muhammad Falensi Azmi, Muhammad Dehan Al Kaut- sar, Alfan Farizki Wicaksono, and Fajri Koto. 2025. IndoSafety: Culturally grounded safety for LLMs in Indonesian languages. In Proceedings of the 2025 Conference on Empirical Methods in Natural Lan- guage Processing, pages 9135–9166, Suzhou, China. Association for Computational Linguistics. Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christo- pher Olah, Danny Hernandez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, and 32 others. 2022. Constitutional ai: Harmlessness from ai feedback. Preprint, arXiv:2212.08073. Somnath Banerjee, Rima Hazra, and Animesh Mukher- jee. 2026. Bridging the multilingual safety divide: Efficient, culturally-aware alignment for global south languages. Preprint, arXiv:2602.13867. Somnath Banerjee, Sayan Layek, Pratyush Chatterjee, Animesh Mukherjee, and Rima Hazra. 2025. Soteria: Language-specific functional parameter steering for multilingual safety alignment. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 9347–9364, Suzhou, China. Association for Computational Linguistics. Lavish Bansal and Naman Mishra. 2026. Crest: Uni- versal safety
Chunk 20 · 1,995 chars
a Hazra. 2025. Soteria: Language-specific functional parameter steering for multilingual safety alignment. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 9347–9364, Suzhou, China. Association for Computational Linguistics. Lavish Bansal and Naman Mishra. 2026. Crest: Uni- versal safety guardrails through cluster-guided cross- lingual transfer. Preprint, arXiv:2512.02711. Edward Bayes, Israel Abebe Azime, Jesujoba O. Al- abi, Jonas Kgomo, Tyna Eloundou, Elizabeth Proehl, Kai Chen, Imaan Khadir, Naome A. Etori, Sham- suddeen Hassan Muhammad, Choice Mpanza, Igne- ciah Pocia Thete, Dietrich Klakow, and David Ife- oluwa Adelani. 2024. Uhura: A benchmark for evaluating scientific question answering and truth- fulness in low-resource african languages. Preprint, arXiv:2412.00948. Samuel Bell, Eduardo Sánchez, David Dale, Pontus Stenetorp, Mikel Artetxe, and Marta R. Costa-Jussà. 2025. Translate, then detect: Leveraging machine translation for cross-lingual toxicity classification. In Proceedings of the Tenth Conference on Machine Translation, pages 253–268, Suzhou, China. Associ- ation for Computational Linguistics. Yuyan Bu, Xiaohao Liu, ZhaoXing Ren, Yaodong Yang, and Juntao Dai. 2026. Align once, benefit multi- lingually: Enforcing multilingual consistency for LLM safety alignment. In The Fourteenth Interna- tional Conference on Learning Representations, Rio de Janeiro, Brazil. Zhuowei Chen, Bowei Zhang, Nankai Lin, Tian Hou, and Lianxi Wang. 2025. Unlocking LLM safeguards for low-resource languages via reasoning and align- ment with minimal training data. In Proceedings of the 5th Workshop on Multilingual Representation Learning (MRL 2025), pages 96–105, Suzhuo, China. Association for Computational Linguistics. Maciej Chrab ˛aszcz, Katarzyna Lorenc, and Karolina Seweryn. 2025. Evaluating llms robustness in less resourced languages with proxy models. Preprint, arXiv:2506.07645. Yue Deng, Wenxuan Zhang, Sinno Jialin Pan, and Li- dong
Chunk 21 · 1,998 chars
n Learning (MRL 2025), pages 96–105, Suzhuo, China. Association for Computational Linguistics. Maciej Chrab ˛aszcz, Katarzyna Lorenc, and Karolina Seweryn. 2025. Evaluating llms robustness in less resourced languages with proxy models. Preprint, arXiv:2506.07645. Yue Deng, Wenxuan Zhang, Sinno Jialin Pan, and Li- dong Bing. 2024. Multilingual jailbreak challenges in large language models. In The Twelfth Interna- tional Conference on Learning Representations, Vi- enna, Austria. Godwin Abuh Faruna. 2026. Lsr: Linguistic safety robustness benchmark for low-resource west african languages. Preprint, arXiv:2603.19273. Ziyu Ge, Gabriel Chua, Leanne Tan, and Roy Ka- Wei Lee. 2025. Toxicity-aware few-shot prompt- ing for low-resource singlish translation. Preprint, arXiv:2507.11966. Maiya Goloburda, Nurkhan Laiyk, Diana Turmakhan, Yuxia Wang, Mukhammed Togmanov, Jonibek Mansurov, Askhat Sametov, Nurdaulet Mukhituly, Minghan Wang, Daniil Orel, Zain Muhammad Mu- jahid, Fajri Koto, Timothy Baldwin, and Preslav Nakov. 2025. Qor ´gau: Evaluating safety in Kazakh- Russian bilingual contexts. In Findings of the As- sociation for Computational Linguistics: ACL 2025, pages 9765–9784, Vienna, Austria. Association for Computational Linguistics. Yujia Hu, Ming Shan Hee, Preslav Nakov, and Roy Ka- Wei Lee. 2025. Toxicity red-teaming: Benchmarking LLM safety in Singapore’s low-resource languages. In Proceedings of the 2025 Conference on Empiri- cal Methods in Natural Language Processing, pages 12183–12201, Suzhou, China. Association for Com- putational Linguistics. Cheng Huang, Fan Gao, Nyima Tashi, Yutong Liu, Xi- angxiang Wang, Thupten Tsering, Ban Ma-bao, Xiao Feng, Renzeg Duojie, Gadeng Luosang, Rinchen Dongrub, Dorje Tashi, Hao Wang, and Yongbin Yu. 2026. Tfd: A comprehensive structured ti- betan foundation dataset for low-resource language processing and large-scale modeling. Preprint, arXiv:2503.18288. Isa Inuwa-Dutse. 2025. Openai’s gpt-oss-20b model and safety alignment issues
Chunk 22 · 1,999 chars
zeg Duojie, Gadeng Luosang, Rinchen Dongrub, Dorje Tashi, Hao Wang, and Yongbin Yu. 2026. Tfd: A comprehensive structured ti- betan foundation dataset for low-resource language processing and large-scale modeling. Preprint, arXiv:2503.18288. Isa Inuwa-Dutse. 2025. Openai’s gpt-oss-20b model and safety alignment issues in a low-resource lan- guage. Preprint, arXiv:2510.01266. -- 9 of 13 -- Devansh Jain, Priyanshu Kumar, Samuel Gehman, Xuhui Zhou, Thomas Hartvigsen, and Maarten Sap. 2024. Polyglotoxicityprompts: Multilingual evalua- tion of neural toxic degeneration in large language models. Preprint, arXiv:2405.09373. Raviraj Bhuminand Joshi, Rakesh Paul, Kanishk Singla, Anusha Kamath, Michael Evans, Katherine Luna, Shaona Ghosh, Utkarsh Vaidya, Eileen Margaret Pe- ters Long, Sanjay Singh Chauhan, and Niranjan Wartikar. 2025. CultureGuard: Towards culturally- aware dataset and guard model for multilingual safety applications. In Proceedings of the 14th Interna- tional Joint Conference on Natural Language Pro- cessing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Lin- guistics, pages 2666–2685, Mumbai, India. The Asian Federation of Natural Language Processing and The Association for Computational Linguistics. Priyanshu Kumar, Devansh Jain, Akhila Yerukola, Li- wei Jiang, Himanshu Beniwal, Thomas Hartvigsen, and Maarten Sap. 2025. Polyguard: A multilingual safety moderation tool for 17 languages. In Sec- ond Conference on Language Modeling, Montreal, Canada. Jiaming Liang, Zhaoxin Wang, and Handing Wang. 2026. Multilingual safety alignment via sparse weight editing. Preprint, arXiv:2602.22554. Isaac Lim, Shaun Khoo, Watson Wei Khong Chua, Jessica Foo, Jia Yi Goh, and Roy Ka-Wei Lee. 2025. Safe at the margins: A general approach to safety alignment in low-resource english languages – a singlish case study. In Second Workshop on Language Models for Underserved Communities (LM4UC). Ri Chi Ng, Aditi Kumaresan, Yujia Hu, and Roy
Chunk 23 · 1,990 chars
Khoo, Watson Wei Khong Chua, Jessica Foo, Jia Yi Goh, and Roy Ka-Wei Lee. 2025. Safe at the margins: A general approach to safety alignment in low-resource english languages – a singlish case study. In Second Workshop on Language Models for Underserved Communities (LM4UC). Ri Chi Ng, Aditi Kumaresan, Yujia Hu, and Roy Ka- Wei Lee. 2026. Seahatecheck: Functional tests for detecting hate speech in low-resource languages of southeast asia. Preprint, arXiv:2603.16070. Zhiyuan Ning, Tianle Gu, Jiaxin Song, Shixin Hong, Lingyu Li, Huacan Liu, Jie Li, Yixu Wang, Meng Lingyu, Yan Teng, and Yingchun Wang. 2025. Linguasafe: A comprehensive multilingual safety benchmark for large language models. Preprint, arXiv:2508.12733. Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Gray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, New Orleans, Louisiana, United States of America. Matthew J Page, David Moher, Patrick M Bossuyt, Is- abelle Boutron, Tammy C Hoffmann, Cynthia D Mulrow, Larissa Shamseer, Jennifer M Tetzlaff, Elie A Akl, Sue E Brennan, Roger Chou, Julie Glanville, Jeremy M Grimshaw, Asbjørn Hróbjarts- son, Manoj M Lalu, Tianjing Li, Elizabeth W Loder, Evan Mayo-Wilson, Steve McDonald, and 7 others. 2021. Prisma 2020 explanation and elaboration: up- dated guidance and exemplars for reporting system- atic reviews. BMJ, 372. Priyaranjan Pattnayak and Sanchari Chowdhuri. 2026a. IndicJR: A judge-free benchmark of jailbreak robust- ness in South Asian languages. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 5: Industry Track), pages 649–668, Rabat, Morocco. Association for Computational
Chunk 24 · 1,996 chars
ayak and Sanchari Chowdhuri. 2026a. IndicJR: A judge-free benchmark of jailbreak robust- ness in South Asian languages. In Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 5: Industry Track), pages 649–668, Rabat, Morocco. Association for Computational Linguistics. Priyaranjan Pattnayak and Sanchari Chowdhuri. 2026b. Indicsafe: A benchmark for evaluating multilingual llm safety in south asia. Preprint, arXiv:2603.17915. Rakesh Paul, Anusha Kamath, Kanishk Singla, Ravi- raj Joshi, Utkarsh Vaidya, Sanjay Singh Chauhan, and Niranjan Wartikar. 2025. Aligning large lan- guage models to low-resource languages through LLM-based selective translation: A systematic study. In Proceedings of the 1st Workshop on Benchmarks, Harmonization, Annotation, and Standardization for Human-Centric AI in Indian Languages (BHASHA 2025), pages 69–82, Mumbai, India. Association for Computational Linguistics. Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, Amelia Glaese, Nat McAleese, and Geoffrey Irving. 2022. Red teaming language models with language models. In Proceed- ings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 3419–3448, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics. Muhammed Saeed, Muhammad Abdul-mageed, and Shady Shehata. 2026. Surfacing subtle stereotypes: A multilingual, debate-oriented evaluation of modern llms. Preprint, arXiv:2511.01187. Wenliang Shan, Michael Fu, Rui Yang, and Chakkrit Tantithamthavorn. 2025. SEALGuard: Safeguarding the Multilingual Conversations in Southeast Asian Languages for AI-Powered Software . In 2025 2nd IEEE/ACM International Conference on AI-powered Software (AIware), pages 197–206, Los Alamitos, CA, USA. IEEE Computer Society. Lingfeng Shen, Weiting Tan, Sihao Chen, Yunmo Chen, Jingyu Zhang, Haoran Xu, Boyuan Zheng, Philipp Koehn, and Daniel Khashabi. 2024. The language barrier:
Chunk 25 · 1,997 chars
or AI-Powered Software . In 2025 2nd IEEE/ACM International Conference on AI-powered Software (AIware), pages 197–206, Los Alamitos, CA, USA. IEEE Computer Society. Lingfeng Shen, Weiting Tan, Sihao Chen, Yunmo Chen, Jingyu Zhang, Haoran Xu, Boyuan Zheng, Philipp Koehn, and Daniel Khashabi. 2024. The language barrier: Dissecting safety challenges of LLMs in mul- tilingual contexts. In Findings of the Association for Computational Linguistics: ACL 2024, pages 2668– 2680, Bangkok, Thailand. Association for Computa- tional Linguistics. Hyunseo Shin and Wonseok Hwang. 2026. Layer-wise swapping for generalizable multilingual safety. In Proceedings of the 19th Conference of the European -- 10 of 13 -- Chapter of the Association for Computational Lin- guistics (Volume 1: Long Papers), pages 2223–2238, Rabat, Morocco. Association for Computational Lin- guistics. Vaibhav Shukla, Hardik Sharma, Adith N Reganti, So- ham Wasmatkar, Bagesh Kumar, and Vrijendra Singh. 2026. Lost in translation? a comparative study on the cross-lingual transfer of composite harms. Preprint, arXiv:2602.07963. Jiayang Song, Yuheng Huang, Zhehua Zhou, and Lei Ma. 2025. Multilingual blending: Large language model safety alignment evaluation with language mixture. In Findings of the Association for Com- putational Linguistics: NAACL 2025, pages 3433– 3449, Albuquerque, New Mexico. Association for Computational Linguistics. Leanne Tan, Gabriel Chua, Ziyu Ge, and Roy Ka-Wei Lee. 2025. LionGuard 2: Building lightweight, data- efficient & localised multilingual content moderators. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 264–285, Suzhou, China. As- sociation for Computational Linguistics. Panuthep Tasawong, Jian Gang Ngui, Alham Fikri Aji, Trevor Cohn, and Peerat Limkonchotiwat. 2025. SEA-safeguardbench: Evaluating AI safety in SEA languages and cultures. Bibek Upadhayay and Vahid Behzadan. 2025. Tongue- tied: Breaking LLMs
Chunk 26 · 1,996 chars
ons, pages 264–285, Suzhou, China. As- sociation for Computational Linguistics. Panuthep Tasawong, Jian Gang Ngui, Alham Fikri Aji, Trevor Cohn, and Peerat Limkonchotiwat. 2025. SEA-safeguardbench: Evaluating AI safety in SEA languages and cultures. Bibek Upadhayay and Vahid Behzadan. 2025. Tongue- tied: Breaking LLMs safety through new language learning. In Proceedings of the 7th Workshop on Computational Approaches to Linguistic Code- Switching, pages 32–47, Albuquerque, New Mexico, USA. Association for Computational Linguistics. Sowmya Vajjala. 2025. The problem with safety classification is not just the models. Preprint, arXiv:2507.21782. Nikhil Verma and Manasa Bharadwaj. 2025. The hidden space of safety: Understanding preference- tuned llms in multilingual context. Preprint, arXiv:2504.02708. Wenxuan Wang, Zhaopeng Tu, Chang Chen, Youliang Yuan, Jen-tse Huang, Wenxiang Jiao, and Michael Lyu. 2024. All languages matter: On the multilingual safety of LLMs. In Findings of the Association for Computational Linguistics: ACL 2024, pages 5865– 5877, Bangkok, Thailand. Association for Computa- tional Linguistics. Xinpeng Wang, Mingyang Wang, Yihong Liu, Hin- rich Schütze, and Barbara Plank. 2026. Refusal di- rection is universal across safety-aligned languages. Preprint, arXiv:2505.17306. Yahan Yang, Soham Dan, Shuo Li, Dan Roth, and In- sup Lee. 2025. MrGuard: A multilingual reasoning guardrail for universal LLM safety. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 27377–27396, Suzhou, China. Association for Computational Lin- guistics. Zheng Xin Yong, Beyza Ermis, Marzieh Fadaee, Stephen Bach, and Julia Kreutzer. 2025a. The state of multilingual LLM safety research: From measur- ing the language gap to mitigating it. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 15845–15860, Suzhou, China. Association for Computational Lin- guistics. Zheng Xin Yong, Beyza Ermis,
Chunk 27 · 1,998 chars
ulia Kreutzer. 2025a. The state of multilingual LLM safety research: From measur- ing the language gap to mitigating it. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 15845–15860, Suzhou, China. Association for Computational Lin- guistics. Zheng Xin Yong, Beyza Ermis, Marzieh Fadaee, Stephen Bach, and Julia Kreutzer. 2025b. The state of multilingual LLM safety research: From measur- ing the language gap to mitigating it. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 15845–15860, Suzhou, China. Association for Computational Lin- guistics. Zheng Xin Yong, Cristina Menghini, and Stephen Bach. 2023. Low-resource languages jailbreak GPT-4. In Socially Responsible Language Modelling Research, New Orleans, United States. Max Zhang, Derek Liu, Kai Zhang, Joshua Franco, Hai- hao Liu, and Kevin Zhu. 2026a. Response-based knowledge distillation for multilingual jailbreak pre- vention unwittingly compromises safety. In 4th De- ployable AI Workshop, Singapore Expo, Singapore. Xianhui Zhang, Chengyu Xie, Linxia Zhu, Yonghui Yang, Weixiang Zhao, Zifeng Cheng, Cong Wang, Fei Shen, and Tat-Seng Chua. 2026b. Who transfers safety? identifying and targeting cross-lingual shared safety neurons. Preprint, arXiv:2602.01283. Weixiang Zhao, Yulin Hu, Yang Deng, Tongtong Wu, Wenxuan Zhang, Jiahe Guo, An Zhang, Yanyan Zhao, Bing Qin, Tat-Seng Chua, and Ting Liu. 2025. MPO: Multilingual safety alignment via reward gap opti- mization. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Vol- ume 1: Long Papers), pages 23564–23587, Vienna, Austria. Association for Computational Linguistics. Yunhan Zhao, Zhaorun Chen, Xingjun Ma, Yu- Gang Jiang, and Bo Li. 2026. Ml-bench&guard: Policy-grounded multilingual safety benchmark and guardrail for large language models. Preprint, arXiv:2605.00689. A Appendix A.1 Search Strings and Criteria To extract the articles
Chunk 28 · 1,995 chars
stria. Association for Computational Linguistics.
Yunhan Zhao, Zhaorun Chen, Xingjun Ma, Yu-
Gang Jiang, and Bo Li. 2026. Ml-bench&guard:
Policy-grounded multilingual safety benchmark and
guardrail for large language models. Preprint,
arXiv:2605.00689.
A Appendix
A.1 Search Strings and Criteria
To extract the articles from these different
databases, we used the following queries:
A.2 LLM Usage: LLM Screening Prompt
The following prompt was submitted once per can-
didate record to Claude Sonnet (claude-sonnet-4.5).
No cross-record context was provided. The JSON
-- 11 of 13 --
Figure 3: Safety Risks in Multilingual and low-resource settings
Table 1: Search strings per database.
Database Query (simplified)
Semantic
Scholar
(LLM OR large language model) AND (safety
OR jailbreak OR toxicity OR alignment) AND
(multilingual OR low-resource OR African lan-
guage)
arXiv Same term clusters submitted via arXiv search
API;
OpenAlex Same term clusters submitted via OpenAlex
search API
output was parsed programmatically; decision
and justification were appended as columns to
the screening spreadsheet.
You are a systematic literature review
assistant. You will be given the title,
abstract, and metadata of one academic
paper. Assess whether it is relevant
to a review on LLM safety alignment for
low-resource and African languages.
Research questions: RQ1 (alignment
methods for low-resource languages), RQ2
(safety risks and cultural harms in
multilingual settings), RQ3 (datasets
and benchmarks for low-resource safety),
RQ4 (cross-lingual transfer of safety
alignment).
Inclusion: I1 post-2020, I2 English, I3
primary LLM safety focus, I4 covers or
is applicable to low-resource languages,
I5 empirical content. Exclusion: E1
English-only, E2 non-LLM models, E3
peripheral safety, E4 no empirical
results, E5 inaccessible.
Paper—Title: {title}. Year: {year}.
Venue: {venue}. Abstract: {abstract}.
Return ONLY valid JSON:
{"rqs_addressed": [...],
Table 2: Inclusion (I) and exclusion (E)Chunk 29 · 1,989 chars
ow-resource languages,
I5 empirical content. Exclusion: E1
English-only, E2 non-LLM models, E3
peripheral safety, E4 no empirical
results, E5 inaccessible.
Paper—Title: {title}. Year: {year}.
Venue: {venue}. Abstract: {abstract}.
Return ONLY valid JSON:
{"rqs_addressed": [...],
Table 2: Inclusion (I) and exclusion (E) criteria.
Inclusion
I1 Published or preprinted after January 2020.
I2 Written in English.
I3 Primary focus on LLM safety, adversarial robustness,
toxicity detection, or content moderation.
I4 Covers ≥1 non-English, low-resource, or African lan-
guage, or proposes methods applicable to such settings.
I5 Peer-reviewed or substantive preprint with methods
and empirical results.
Exclusion
E1 Safety studied exclusively in English; no multilingual
component or discussion.
E2 Focused on classical NLP models (LSTM, CNN) with
no connection to LLMs.
E3 Safety is peripheral, not a primary contribution.
E4 Editorial or opinion piece without empirical results.
E5 Full text inaccessible.
"inclusion_met": [...],
"exclusion_triggered": [...],
"decision": "keep" | "remove",
"confidence": "high" | "medium" |
"low", "justification": "one sentence"}
A.3 Further Analysis
The contrast in research coverage illustrated in Fig-
ure 4 underpins the safety disparities quantified in
Figure 3. As shown in Figure 4, English dominates
across all study categories (including benchmarks,
alignment methods, and adversarial attacks), while
low-resource languages remain confined to a long
tail of minimal representation. This imbalance
in research attention is reflected in downstream
safety performance. Figure 3(A) shows that low-
resource languages exhibit cross-lingual jailbreak
failure rates more than seven times higher than
-- 12 of 13 --
those observed in high-resource baselines. Further-
more, Figure 3(B) highlights a clear separation in
the risk–utility landscape: high-resource languages
cluster in a region characterized by low harmful out-
put rates and limited utilityChunk 30 · 698 chars
ngual jailbreak failure rates more than seven times higher than -- 12 of 13 -- those observed in high-resource baselines. Further- more, Figure 3(B) highlights a clear separation in the risk–utility landscape: high-resource languages cluster in a region characterized by low harmful out- put rates and limited utility degradation, whereas low-resource and code-switched settings occupy a high-risk region marked by elevated harmful out- puts and substantial utility loss. Together, these results highlight a systematic disparity in safety alignment coverage and underscore the need for more equitable multilingual safety strategies. Figure 4: Languages Distribution Across Studies -- 13 of 13 --