Cross-Lingual Sentiment Misalignment: Auditing Multilingual Language Models for Inversion Risk, Dialectal Representation, and Affective Stability
Summary
This study audits cross-lingual sentiment alignment between Bengali and English, revealing critical failures in multilingual transformer models. Using a benchmark of 7,350 parallel sentence pairs stratified by dialect, the authors evaluate four architectures: XLM-T, IndicBERT, Tabularis, and mDistilBERT. The research identifies three primary issues. First, compressed models like mDistilBERT exhibit a 28.7% "Sentiment Inversion Rate," fundamentally misinterpreting positive semantics as negative or vice versa, posing severe risks for safety classifiers. Second, the authors define "Asymmetric Empathy," where models systematically dampen or amplify the affective weight of Bengali text relative to English, creating directional bias. Third, a "Modern Bias" is observed, where models perform significantly worse on formal Sadhu Bengali compared to colloquial Cholito, with IndicBERT showing a 57% increase in alignment error for formal registers. The findings suggest that model compression and limited training data diversity degrade representational stability. While large-scale models like XLM-T show better resilience, the distilled Tabularis model demonstrates that synthetic data augmentation can mitigate these issues. Consequently, the authors propose the "Affective Stability Index" as a new metric to penalize polarity inversions and dialectal divergence. They argue that current benchmarks overlook affective fidelity, urging the integration of these stability metrics to ensure equitable and reliable multilingual NLP systems, particularly for low-resource languages like Bengali.
PDF viewer
Chunks(25)
Chunk 0 · 1,993 chars
Cross-Lingual Sentiment Misalignment: Auditing Multilingual Language Models for Inversion Risk, Dialectal Representation, and Affective Stability Nusrat Jahan Lia Institute of Information Technology University of Dhaka, Dhaka, Bangladesh bsse1306@iit.du.ac.bd Shubhashis Roy Dipta University of Maryland, Baltimore County Baltimore, Maryland, USA sroydip1@umbc.edu Abstract Recent advances in multilingual representa- tion learning aim to bridge the performance gap between high- and low-resource languages, yet their ability to preserve affective meaning across languages remains underexplored, par- ticularly for underrepresented languages like Bengali. This research addresses cross-lingual sentiment misalignment between Bengali and English by introducing a controlled benchmark- ing framework evaluating four multilingual transformer models on parallel Bengali-English sentence pairs, stratified by dialect, to assess their representational stability. We demonstrate that a compressed model architecture exhibits a 28.7% “Sentiment Inversion Rate,” funda- mentally misinterpreting positive semantics as negative (or vice versa). Consequently, we identify a cross-lingual sentiment skew that we call “Asymmetric Empathy”, where models systematically dampen or artificially amplify the affective weight of Bengali text relative to its exact English counterpart. Finally, we ex- pose a key vulnerability regarding dialectal rep- resentation: a “Modern Bias” in the regional model, which exhibits a 57% increase in align- ment error when processing the formal Bengali register compared to modern colloquial text. As foundational encoders continue to serve as safety classifiers and reward models for LLM pipelines, cross-lingual reliability becomes a critical concern. We therefore advocate for the integration of “Affective Stability” metrics into future cross-lingual benchmarks to detect and penalize polarity inversions, particularly in low- resource settings. 1 Introduction Multilingual
Chunk 1 · 1,994 chars
iers and reward models for LLM pipelines, cross-lingual reliability becomes a critical concern. We therefore advocate for the integration of “Affective Stability” metrics into future cross-lingual benchmarks to detect and penalize polarity inversions, particularly in low- resource settings. 1 Introduction Multilingual models have rapidly become the backbone of language-agnostic information ac- cess, spanning sentiment analysis, content modera- tion, retrieval-augmented systems, and downstream knowledge-intensive applications. However, when Negative Neutral Positive Negative Neutral Positive Alignment Corridor (Y=X) All that day, his insulted and angry heart began to rule his instincts সিদন সমস্ত িদন ধিরয়া তাহার অপ মানাহত ুব্ধ িচত্ত তাহার প্র বৃিত্ত েক শাসন কিরেত লািগল Figure 1: The plot maps the predicted English sentiment vector (X-axis) against the Bengali vector (Y -axis) for a test sentence. The green corridor represents ideal cross-lingual sentiment alignment (Y ≈ X). While the large-scale XLM-T model successfully preserves polarity, the compressed (mDistilBERT) and regionally specialized (IndicBERT) architectures are projecting a negative English statement into a positive Bengali latent space (Red Zone). a multilingual model correctly identifies that an En- glish sentence carries negative sentiment but simul- taneously classifies its Bengali semantic equivalent as positive (Figure 1), the representational promise of multilingualism collapses in practice. Multilingual sentiment encoders, including those for Bengali, are deployed as core components in content moderation systems, where biases can prop- agate to real-world safety failures; for instance, an audit of Bengali sentiment tools revealed identity- based inconsistencies that undermine moderation accuracy (Das et al., 2024). Similarly, Tan et al. (2025) use multilingual classifiers for government- scale moderation in diverse languages with reliance on foundational encoders for LLM safety
Chunk 2 · 1,998 chars
ilures; for instance, an audit of Bengali sentiment tools revealed identity- based inconsistencies that undermine moderation accuracy (Das et al., 2024). Similarly, Tan et al. (2025) use multilingual classifiers for government- scale moderation in diverse languages with reliance on foundational encoders for LLM safety pipelines. Such instances illustrate the risks of sentiment mis- alignment in operational safety classifiers. 1 arXiv:2602.17469v2 [cs.CL] 30 Apr 2026 -- 1 of 10 -- To mitigate the curse of multilinguality, the phe- nomenon where adding more languages to a fixed- capacity model degrades per-language performance (Conneau et al., 2020), we investigate the “Senti- ment Inversion” crisis and its broader representa- tional implications. We demonstrate that the multi- lingual curse manifests not only as accuracy degra- dation but also as affective inversion. We employ a standardized cross-lingual sentiment alignment framework to evaluate the semantic consistency of four multilingual sentiment classifier models on parallel Bengali–English text (see Table 4). We show that alignment is not uniform across languages or dialects, as summarized in the follow- ing key findings: Capacity Constraints & Alignment Instabil- ity: We demonstrate that the capacity-constrained mDistilBERT architecture exhibits a 28.7% “Senti- ment Inversion Rate,” accompanied by a heavy- tailed distribution of errors (Figure 3), suggest- ing that model compression may compromise the safety margins required for cross-lingual alignment. We further find that regional specialization alone does not resolve this issue; however, a distilled ar- chitecture (Tabularis) that leverages synthetic data can reduce such representational misalignments. Asymmetric Empathy: We reveal that mod- els exhibit consistent cross-lingual directional bias, systematically dampening or artificially inflating the emotional intensity of non-English text. The Dialectal Sentiment Gap: We find a “Mod- ern Bias” in which
Chunk 3 · 1,998 chars
s synthetic data can reduce such representational misalignments. Asymmetric Empathy: We reveal that mod- els exhibit consistent cross-lingual directional bias, systematically dampening or artificially inflating the emotional intensity of non-English text. The Dialectal Sentiment Gap: We find a “Mod- ern Bias” in which models align well with col- loquial text but underperform on formal dialects, which are the backbone of Bengali literature. 2 Related Works This section reviews prior work in multilingual rep- resentation, cross-lingual sentiment analysis, and dialect-aware NLP, with a focus on gaps in affective alignment and evaluation. 2.1 Bengali in the Multilingual NLP Landscape Bengali remains one of the most underrepresented major languages in NLP despite its global speaker base. Kabir et al. (2024) provide a comprehen- sive audit of LLM performance on Bengali NLP tasks, identifying key failure modes including lan- guage generation errors, verbose mismatching with evaluation metrics, and task-specific weaknesses. Bhowmik et al. (2025) document consistent perfor- mance gaps for Bengali relative to English across recent LLMs, tracing these to tokenization ineffi- ciency, where models fragment Bengali script into excessive subword units, degrading semantic coher- ence. The BnMMLU benchmark further demon- strates that even large-scale frontier models show sublinear returns in Bengali reasoning as model size increases, an empirical fingerprint of the mul- tilingual curse at scale (Joy and Shatabda, 2025; Conneau et al., 2020). For sentiment analysis specifically, BanglaBERT (Bhattacharjee et al., 2022) established a strong monolingual baseline, and ensemble transformer systems (Hoque et al., 2024) have achieved high ag- gregate accuracy. However, in a landmark audit of Bengali sentiment analysis tools, Das et al. (2024) reveal that aggregate accuracy conceals systematic identity-based biases, with tools exhibiting differ- ential performance across gender, religious,
Chunk 4 · 1,993 chars
mble transformer systems (Hoque et al., 2024) have achieved high ag- gregate accuracy. However, in a landmark audit of Bengali sentiment analysis tools, Das et al. (2024) reveal that aggregate accuracy conceals systematic identity-based biases, with tools exhibiting differ- ential performance across gender, religious, and national identity signals. This colonial impulse in tool design highlights how reductionist representa- tions reanimate historical hierarchies and motivates the external auditing approach we adopt. Our work extends this critical perspective by focusing specif- ically on cross-lingual affective misalignment and dialectal representational harm. 2.2 The Multilingual Curse and Capacity Constraints The curse of multilinguality, originally formalized by Conneau et al. (2020), describes the empiri- cal observation that, under fixed model capacity, adding more languages to pretraining initially ben- efits low-resource languages through positive trans- fer but eventually degrades per-language perfor- mance due to inter-language parameter competi- tion. Recent work has substantially refined this understanding. Blevins et al. (2024) demonstrate that the curse can be partially lifted through Cross- lingual Expert Language Models (X-ELM), which decouple per-language capacity via modular train- ing and outperform jointly trained multilingual models across 16 languages. Foroutan et al. (2025) further argue that the curse arises not from language count per se but from finite model capacity ampli- fying the impact of noisy, low-quality data in low- resource languages. This has direct implications for Bengali, where pretraining data is both scarce and noisier than for high-resource languages. 2 -- 2 of 10 -- 2.3 Cross-Lingual Sentiment Analysis and Representational Failures The adoption of transformer architectures has substantially advanced sentiment analysis in low- resource languages (Bhowmick and Jana, 2021). Cross-lingual transfer from high-resource to
Chunk 5 · 1,996 chars
e and noisier than for high-resource languages. 2 -- 2 of 10 -- 2.3 Cross-Lingual Sentiment Analysis and Representational Failures The adoption of transformer architectures has substantially advanced sentiment analysis in low- resource languages (Bhowmick and Jana, 2021). Cross-lingual transfer from high-resource to low- resource languages has been enabled both through shared representations in pretrained multilingual models (Conneau et al., 2020) and through machine translation strategies (Poncelas et al., 2020). Chen et al. (2025) document that while GPT-4 achieves approximately 84.4% F1 in English sentiment, this drops to around 67% for low-resource languages, and propose adaptive self-alignment strategies with data augmentation to partially close this gap. Re- cent literature shows that hybrid approaches retain- ing lexicon features maintain stability advantages over purely neural representations (Mahmud et al., 2024), and ensemble methods can achieve high aggregate accuracy (Hoque et al., 2024). A growing body of literature documents that high aggregate accuracy routinely masks represen- tational failures (Das et al., 2024). Wasi et al. (2024) show that LLMs acquire social biases through surface linguistic cues, and that Bengali di- alectal variation, particularly religious dialect vari- ation, induces systematic performance divergence in large models. Ochieng et al. (2025) extend this critique, demonstrating that reasoning-based LLM sentiment evaluation in low-resource, culturally nuanced contexts reveals failures invisible to label- prediction benchmarks. The CuLEmo benchmark (Belay et al., 2025) further shows that multilin- gual LLMs systematically fail to capture culturally grounded variations in emotional expression across languages. These cultural layers of failure are dis- tinct from, but related to, the cross-lingual affective misalignment we document. 2.4 Bengali Diglossia and Dialectal NLP Dialectal variation represents a significant chal- lenge
Chunk 6 · 1,992 chars
matically fail to capture culturally grounded variations in emotional expression across languages. These cultural layers of failure are dis- tinct from, but related to, the cross-lingual affective misalignment we document. 2.4 Bengali Diglossia and Dialectal NLP Dialectal variation represents a significant chal- lenge for multilingual NLP, as Wasi et al. (2024) demonstrate through empirical evaluation of LLMs on Bengali religious dialects. Bengali exhibits a well-documented diglossic structure comprising Sadhu Bhasha (formal/literary, Sanskrit-derived vo- cabulary, archaic conjugation) and Cholito Bhasha (colloquial/standard, simplified morphology, con- temporary vocabulary), a stylistic split that presents significant challenges for multilingual NLP due to the frequent blending of these forms in everyday communication (Ayman et al., 2025). The critical insight motivating our dialect stratification is that training data for multilingual models is overwhelm- ingly drawn from contemporary digital sources such as large-scale web crawls, social media, and news corpora (Conneau et al., 2020; Kakwani et al., 2020), creating a training distribution that is inher- ently skewed toward the more common, modern Cholito Bengali form (Ayman et al., 2025). 2.5 Benchmarking Gaps and the Need for Affective Stability Metrics Current multilingual benchmarks, including XTREME, XNLI, and their derivatives, evaluate cross-lingual performance on semantic tasks such as natural language inference, question answering (e.g., MLQA, TyDiQA), and named entity recog- nition (e.g., WikiAnn), with XNLI focusing on en- tailment classification (Hu et al., 2020). While ef- fective at measuring semantic transfer, these bench- marks largely overlook affective fidelity. Recent efforts such as MMAFFBen (Liu et al., 2025) begin to address this gap by introducing affective eval- uation, but comprehensive measurement of senti- ment preservation and cross-lingual affective align- ment remains limited.
Chunk 7 · 1,992 chars
at measuring semantic transfer, these bench- marks largely overlook affective fidelity. Recent efforts such as MMAFFBen (Liu et al., 2025) begin to address this gap by introducing affective eval- uation, but comprehensive measurement of senti- ment preservation and cross-lingual affective align- ment remains limited. Ochieng et al. (2025) ex- plicitly call for benchmarks that measure LLM sentiment in low-resource, culturally nuanced con- texts beyond label accuracy. Miah et al. (2024) note that translation-based cross-lingual sentiment approaches can achieve high aggregate accuracy while failing to capture culturally grounded vari- ations in emotional expression, often introducing translation biases in affective intensity. We interpret these findings as evidence of di- rectional distortions in how sentiment is mapped across languages, which we formalize as asym- metric empathy. Recent work on multilingual bias evaluation (Wasi et al., 2024; Sadhu et al., 2025) has established that social bias in Bengali LLMs op- erates across gender and religious lines, but has not examined the cross-lingual affective alignment di- mension we investigate. Our proposal for affective stability metrics, which explicitly penalize polar- ity inversions and dialectal divergence, responds to this benchmarking gap, extending recent work on stability-focused evaluation metrics (Atil et al., 2024) to the multilingual affective domain. 3 -- 3 of 10 -- 3 Methodology We employ a controlled experimental framework to quantify cross-lingual sentiment alignment in multilingual transformer architectures. We adopt a within-model comparative design in which each transformer processes parallel Bengali-English text pairs independently, enabling direct measurement of semantic divergence without inter-model archi- tectural confounds. 3.1 Dataset Specification We utilize a parallel corpus comprising n = 7,350 Bengali-English sentence pairs, sourced from the publicly available “BanglaBlend” dataset
Chunk 8 · 1,995 chars
sses parallel Bengali-English text
pairs independently, enabling direct measurement
of semantic divergence without inter-model archi-
tectural confounds.
3.1 Dataset Specification
We utilize a parallel corpus comprising n = 7,350
Bengali-English sentence pairs, sourced from the
publicly available “BanglaBlend” dataset (Ayman
et al., 2025). Formally, the dataset D is defined as
a set of tuples:
D = {(Bi, Ei, Di) | i ∈ [1, n]} (1)
where:
• Bi ∈ Σ∗
Bengali represents the i-th Bengali sen-
tence (original text),
• Ei ∈ Σ∗
English represents the corresponding
English translation,
• Di ∈ {Sadhu, Cholito} denotes the Bengali
dialect classification.
Bengali exhibits diglossia with two primary writ-
ten forms:
1. Sadhu Bhasha: Formal/literary register, char-
acterized by Sanskrit-derived vocabulary and
archaic verb conjugations.
2. Cholito Bhasha: Colloquial/standard regis-
ter with relatively simplified morphology and
contemporary vocabulary.
3.2 Models Evaluated
We benchmark four multilingual transformer ar-
chitectures representing distinct design paradigms:
XLM-T, a large-scale model fine-tuned on high-
volume multilingual social media data; IndicBERT,
a regionally specialized encoder for Indian lan-
guages; Tabularis, a distilled multilingual model en-
hanced with synthetic data for broad cross-lingual
coverage; and mDistilBERT, a compressed multi-
lingual model designed for efficient zero-shot sen-
timent transfer. Full repository mappings are pro-
vided in Table 4. This selection enables systematic
analysis of how scale, regional specialization, syn-
thetic augmentation, and compression interact with
cross-lingual affective alignment.
These specific models were selected based on
three core inclusion criteria: they are publicly ac-
cessible, they are capable of zero-shot Bengali
sentiment inference without requiring further task-
specific fine-tuning, and together they cover a prin-
cipled spectrum of architectural capacity (large-
scale vs. compressed) and designChunk 9 · 1,989 chars
models were selected based on
three core inclusion criteria: they are publicly ac-
cessible, they are capable of zero-shot Bengali
sentiment inference without requiring further task-
specific fine-tuning, and together they cover a prin-
cipled spectrum of architectural capacity (large-
scale vs. compressed) and design intent (global vs.
regional). Furthermore, because our experimental
design isolates cross-lingual representation and af-
fective transfer as the primary variables of interest,
we focus exclusively on multilingual architectures,
purposefully excluding monolingual models.
3.3 Experimental Design
For each model M and sentence pair (Bi, Ei), we
perform independent inference on both languages
using the same model weights θM :
ˆyB,i = Mθ(Bi) [Bengali stream] (2)
ˆyE,i = Mθ(Ei) [English stream] (3)
Any divergence between ˆyB,i and ˆyE,i is at-
tributable to cross-lingual representation, calibra-
tion, or decision boundary alignment within the
same parameter space.
3.4 Score Normalization and Metric
Formulation
3.4.1 Universal Score Normalizer
To enable direct comparison, we define a universal
normalization function φ : Predictions → [−1, 1]:
φ(ˆy) =
ˆy.score if ˆy.label ∈ {positive}
−ˆy.score if ˆy.label ∈ {negative}
0 if ˆy.label ∈ {neutral}
(4)
*Note: Standard “Positive” and “Negative” labels
map to ±s for 2-class and 3-class models, while
intermediate positive/negative classes are scaled by
0.5 in the 5-class Tabularis model to accommodate
the “Very Positive” and “Very Negative” extremes.
This maintains a uniform linear spacing across
sentiment intensity levels, ensuring that the five
classes are equidistant on the sentiment continuum.
4
-- 4 of 10 --
After normalization, we obtain continuous senti-
ment scores:
SB,i = φ(ˆyB,i) ∈ [−1, 1] (5)
SE,i = φ(ˆyE,i) ∈ [−1, 1] (6)
where SB,i and SE,i denote the Bengali and En-
glish sentiment scores, respectively.
3.4.2 Sentence-Level Alignment Metrics
For each sentence pair i, we compute fourChunk 10 · 1,997 chars
nt continuum. 4 -- 4 of 10 -- After normalization, we obtain continuous senti- ment scores: SB,i = φ(ˆyB,i) ∈ [−1, 1] (5) SE,i = φ(ˆyE,i) ∈ [−1, 1] (6) where SB,i and SE,i denote the Bengali and En- glish sentiment scores, respectively. 3.4.2 Sentence-Level Alignment Metrics For each sentence pair i, we compute four align- ment metrics: M1. Alignment Divergence Di = |SB,i − SE,i| ∈ [0, 2] (7) Interpretation: Di = 0.0 indicates perfect align- ment (identical sentiment), Di = 0.5 indicates moderate divergence, Di = 2.0 indicates maximal divergence (opposite extremes). M2. Directional Bias Bi = SE,i − SB,i ∈ [−2, 2] (8) Interpretation: Bi > 0 indicates that the English text is predicted as more positive (or less negative) than its Bengali counterpart; Bi < 0 indicates the reverse; Bi ≈ 0 indicates minimal cross-lingual divergence for that pair. M3. Polarity Inversion (Safety Metric) Ii = 1 "(SB,i > τ ∧ SE,i < −τ ) ∨ (SB,i < −τ ∧ SE,i > τ ) # (9) where 1[·] is the indicator function and τ = 0.1 is a noise threshold to avoid false positives from near-zero scores. Interpretation: Ii = 1 indicates sentiment in- version (e.g., Bengali=Positive, English=Negative); Ii = 0 indicates polarity preserved. Inversion is the most severe failure mode, as it indicates mis- alignment of sentiment direction. 3.5 Population-Level Aggregation and Statistical Computation To characterize model-level performance and eval- uate representational equity across Bengali diglos- sia, we aggregate the sentence-level metrics into population-level statistics. We compute these across the entire dataset D as well as its stratified dialectal subsets (DSadhu and DCholito). Metric Tabularis XLM-T IndicBERT mDistilBERT Mean Div. 0.200 0.276 0.375 0.417 Std Dev. 0.214 0.298 0.607 0.429 Sadhu Div. 0.239 0.286 0.459 0.456 Cholito Div. 0.161 0.266 0.292 0.379 Dialect Gap 0.078 0.020 0.167 0.077 Sadhu Err. Inc. (%) 48.4 7.6 57.1 20.5 Robustness (%) 43.1 42.1 58.3 34.2 Inversions 635
Chunk 11 · 1,989 chars
aris XLM-T IndicBERT mDistilBERT Mean Div. 0.200 0.276 0.375 0.417 Std Dev. 0.214 0.298 0.607 0.429 Sadhu Div. 0.239 0.286 0.459 0.456 Cholito Div. 0.161 0.266 0.292 0.379 Dialect Gap 0.078 0.020 0.167 0.077 Sadhu Err. Inc. (%) 48.4 7.6 57.1 20.5 Robustness (%) 43.1 42.1 58.3 34.2 Inversions 635 267 1471 2107 Inv. Rate (%) 8.6 3.6 20.0 28.7 Dir. Bias (En-Bn) 0.002 0.057 0.106 -0.066 Table 1: Metric scores per model Table 2 details the mathematical formulations and interpretations for all population-level evalua- tion criteria, including overall alignment statistics, safety indicators, and specialized metrics designed to quantify the dialectal gap. 4 Results Beyond average alignment error, our evaluation identifies three critical failure modes in multilin- gual model behavior. Table 1 presents the corre- sponding quantitative results. 4.1 Finding 1: Sentiment Inversion and Alignment Instability We define a sentiment inversion as a case where a Bengali-English translation pair (similar mean- ing) receives opposite polarity classifications (pos- itive vs. negative). Such inversions represent ma- jor alignment failures, as propositional meaning is preserved but affective interpretation is reversed. Across model architectures, inversion rates vary dramatically, revealing a potential relationship be- tween compression, divergence magnitude, and op- timization (see Figure 2). Figure 2: Sentiment Inversion Rate Across Models • Compression and Elevated Inversion Risk. The distilled multilingual architecture (mDis- tilBERT) exhibits the highest mean alignment divergence and lowest robustness (see Ta- ble 1). Nearly one in three sentence pairs 5 -- 5 of 10 -- Metric Formulation Interpretation Mean Alignment Error (Divergence) μD (M ) = 1 n Pn i=1 Di Lower μD indicates better average cross- lingual consistency. Std. Deviation (Divergence) σD (M ) = q 1 n−1 Pn i=1(Di − μD )2 Quantifies the variability in alignment qual- ity across the
Chunk 12 · 1,989 chars
n three sentence pairs 5 -- 5 of 10 -- Metric Formulation Interpretation Mean Alignment Error (Divergence) μD (M ) = 1 n Pn i=1 Di Lower μD indicates better average cross- lingual consistency. Std. Deviation (Divergence) σD (M ) = q 1 n−1 Pn i=1(Di − μD )2 Quantifies the variability in alignment qual- ity across the dataset. Robustness Index R(M ) = 1 n Pn i=1 1[Di < 0.1] × 100% % of pairs with negligible divergence; mea- sures the “safe operating zone.” Inversion Rate Irate(M ) = 1 n Pn i=1 Ii × 100% % of sentence pairs exhibiting sentiment polarity flips. Mean Directional Bias μB (M ) = 1 n Pn i=1 Bi > 0: English favored; < 0: Bengali fa- vored; ≈ 0: No systematic language skew. Formal Penalty (Dialect Gap) ∆dialect(M ) = μD (M, Sadhu) − μD (M, Cholito) > 0 implies a “Modern Bias” where the model struggles with formal registers. Relative Dialect Error ∆% dialect(M ) = ∆dialect(M ) μD (M,Cholito) × 100% Normalizes the formal penalty by the base- line colloquial error rate for fair compari- son. Table 2: Summary of population-level alignment and dialectal metrics. processed by the compressed model receives directly contradictory affective classifications across Bengali and English. This indicates that while compression improves efficiency, it may disproportionately reduce the representa- tional capacity required for reliable affective calibration. As a result, sentiment polarity re- versal emerges as a critical failure mode that distorts core cross-lingual meaning. • Heavy-Tailed Failure Distribution. Align- ment error density analysis (Figure 3) shows that divergence is not normally distributed. In- stead, the compressed architecture exhibits a long right tail, corresponding to extreme po- larity flips. While IndicBERT achieves the highest robustness metric, it simultaneously exhibits a massive standard deviation in diver- gence and a high inversion rate with alignment errors that are not normally distributed. This non-normal distribution has practical
Chunk 13 · 1,996 chars
its a long right tail, corresponding to extreme po- larity flips. While IndicBERT achieves the highest robustness metric, it simultaneously exhibits a massive standard deviation in diver- gence and a high inversion rate with alignment errors that are not normally distributed. This non-normal distribution has practical implica- tions: mean divergence understates actual risk, and models cannot be reliably characterized by their average behavior for deployment in critical downstream applications. • Scale-Driven Inversion Resilience. The large-scale multilingual model (XLM-T) records the lowest polarity inversion rate (Ta- ble 1) across all evaluation pairs. This sug- gests that massive parameter scale and diverse pre-training may preserve coherent affect map- pings more effectively than regional special- ization, buffering against semantic instability. Crucially, however, the robust distilled Tabu- laris model (8.6% inversion rate) implies that compression with data-centric optimization can substantially close the gap between com- pressed and full-scale architectures. Hence, scale is not the only path to alignment stability. As a DistilBERT-based model fine-tuned with diverse synthetic multilingual data, Tabularis shows that targeted training strategies, rather than scale alone, can drive alignment stability. Figure 3: Distribution of Alignment Error Density 4.2 Finding 2: Representational Harm and the Dialectal Gap To evaluate robustness under Bengali diglossia, we compare alignment divergence across colloquial (Cholito) and formal (Sadhu) variants. A multilin- gual system should maintain stable cross-lingual 6 -- 6 of 10 -- calibration regardless of lexical register. However, we observed a dialectal sensitivity pattern. • Modern-Register Overfitting. Both In- dicBERT and Tabularis exhibit sharp in- creases in divergence when processing for- mal Sadhu text (see Table 1). This indicates that alignment quality is concentrated in mod- ern and high-frequency lexical
Chunk 14 · 1,996 chars
f lexical register. However, we observed a dialectal sensitivity pattern. • Modern-Register Overfitting. Both In- dicBERT and Tabularis exhibit sharp in- creases in divergence when processing for- mal Sadhu text (see Table 1). This indicates that alignment quality is concentrated in mod- ern and high-frequency lexical distributions, while archaic or formal constructions fall out- side the model’s calibrated semantic mani- fold. We term this phenomenon Modern Bias: strong alignment in contemporary usage, but underperformance in formal registers. The Modern Bias finding has direct consequences for linguistic equity: users who communi- cate in formal registers including academic, literary, and administrative Bengali, receive poorer cross-lingual sentiment alignment than users of colloquial varieties. As a result, cur- rent multilingual systems risk institutionaliz- ing a structural inequity in which access to reliable inference is contingent on conforming to simplified or non-native linguistic norms. Conversely, XLM-T demonstrates strong di- alectal resilience, likely due to its large-scale multilingual training over diverse text distri- butions, which provides broader lexical and syntactic coverage and supports stable affec- tive mappings across regional variation. 4.3 Finding 3: Asymmetric Empathy and Directional Bias In a multilingual architecture, the system must pre- serve not just the polarity but the intensity of user intent. We evaluate this using Directional Bias (En- glish score − Bengali score). An ideally aligned system should yield a distribution centered at zero with low variance. Instead, we observed two dis- tinct alignment regimes (as seen in Table 1, Fig- ure 4 and Figure 5) • Compression-Induced Bengali Positivity Skew. The distilled architecture (mDistil- BERT) exhibits a negative directional bias, in- dicating that it systematically scores Bengali text as more positive (or less negative) than its exact English translation. For a human user in a
Chunk 15 · 1,996 chars
seen in Table 1, Fig- ure 4 and Figure 5) • Compression-Induced Bengali Positivity Skew. The distilled architecture (mDistil- BERT) exhibits a negative directional bias, in- dicating that it systematically scores Bengali text as more positive (or less negative) than its exact English translation. For a human user in a safety-critical context: the model may artificially dampen the severity of a neg- ative Bengali sentiment, underweighting the severity of negative Bengali content relative to equivalent English content. As demonstrated in the case study (Figure 5), mDistilBERT ex- hibits a safety failure by correctly scoring an English statement as deeply negative (-0.981) while assigning a positive score (+0.533) for its exact Bengali equivalent. • Regional English Optimism Bias. Con- versely, IndicBERT demonstrates a positive mean directional bias, assigning higher posi- tivity (or lower negativity) to English inputs relative to their Bengali counterparts. This imposes an equity penalty on Bengali users, as their neutral or moderately negative state- ments are penalized with harsher negative clas- sifications compared to English speakers. These findings indicate that cross-lingual affec- tive misalignment is consistently and directionally structured rather than stochastic, which can lead to downstream distortions in applications relying on Bengali sentiment signals. 4.4 Proposed Metric: Affective Stability Index Our empirical evaluation reveals that aggregate alignment error (μD) alone is insufficient to capture the critical safety failures inherent in cross-lingual representation. Specifically, models with relatively moderate average divergence may still exhibit high rates of sentiment inversion (Irate) (Table 1). To address this benchmarking gap and formally quan- tify cross-lingual reliability, we introduce the Af- fective Stability Index (AS). We define Affective Stability as a composite metric that rewards tight se- mantic alignment while strictly penalizing
Chunk 16 · 1,989 chars
still exhibit high rates of sentiment inversion (Irate) (Table 1). To address this benchmarking gap and formally quan- tify cross-lingual reliability, we introduce the Af- fective Stability Index (AS). We define Affective Stability as a composite metric that rewards tight se- mantic alignment while strictly penalizing polarity inversions. Utilizing the population-level metrics defined in Section 3.5, it is computed as: AS(M ) = 1 − μD(M ) 2 × (1 − Irate(M )) (10) The first term normalizes the Mean Alignment Divergence (μD) into a similarity score bounded by [0, 1], as the maximum theoretical divergence in our normalized space is 2.0. The second term acts as a strict penalty mask based on the Inversion Rate (Irate), expressed as a probability. An AS score of 1.0 indicates perfect cross-lingual affective fidelity, whereas lower scores reflect compounding representational misalignments. 7 -- 7 of 10 -- Table 3 presents the Affective Stability scores for all evaluated architectures. The results validate our observations regarding scale and compression. While the large-scale XLM-T maintains high affec- tive stability, mDistilBERT suffers a degradation in overall reliability. Notably, the distilled Tabu- laris model achieves an AS score highly compet- itive with XLM-T. This empirically demonstrates that targeted training strategies and synthetic data utilization can preserve affective calibration even under capacity constraints. Model Mean Div. (μD) Inv. Rate (Irate) Affective Stability (AS) XLM-T 0.276 0.036 0.831 Tabularis 0.200 0.086 0.823 IndicBERT 0.375 0.200 0.650 mDistilBERT 0.417 0.287 0.564 Table 3: Affective Stability (AS) of evaluated models. 5 Discussion: Alignment Under Linguistic Pluralism Our findings instantiate the multilingual curse at an affective, rather than purely accuracy-based, level. The inversion gap we observe is consistent with Blevins et al. (2024)’s theoretical framing: fixed model capacity induces inter-language
Chunk 17 · 1,996 chars
AS) of evaluated models. 5 Discussion: Alignment Under Linguistic Pluralism Our findings instantiate the multilingual curse at an affective, rather than purely accuracy-based, level. The inversion gap we observe is consistent with Blevins et al. (2024)’s theoretical framing: fixed model capacity induces inter-language parameter competition that degrades per-language represen- tation. We hypothesize that compression amplifies this competition in ways that may disrupt the fine- grained representational signals required for stable sentiment polarity mapping. Our results suggest two distinct intervention points for multilingual model development. First, at the pretraining data level: the Modern Bias find- ing points to a gap in training corpus composi- tion for Bengali, even when utilizing synthetic data. Formal Sadhu text is underrepresented in web-crawled multilingual corpora (dominated by social media and news in contemporary registers), and correcting this imbalance would directly ad- dress dialectal representational harm. Second, at the post-distillation fine-tuning level: our finding that Tabularis substantially outperforms mDistil- BERT despite both being distilled architectures demonstrates that targeted fine-tuning with syn- thetic multilingual corpora can preserve affective calibration through compression. Addressing both stages is particularly vital in low-resource settings like Bengali, where computationally efficient, com- pressed multilingual models must still deliver eq- uitable and reliable affective understanding across all linguistic communities. The multilingual NLP evaluation landscape cur- rently lacks standardized metrics for cross-lingual affective consistency. Existing benchmarks (Hu et al., 2020; Ruder et al., 2021; Han et al., 2025; Goldman et al., 2025) primarily measure semantic and syntactic transfer but do not penalize polarity inversions or account for affective fidelity. We pro- pose that multilingual benchmarks incorporate the metric
Chunk 18 · 1,993 chars
-lingual affective consistency. Existing benchmarks (Hu et al., 2020; Ruder et al., 2021; Han et al., 2025; Goldman et al., 2025) primarily measure semantic and syntactic transfer but do not penalize polarity inversions or account for affective fidelity. We pro- pose that multilingual benchmarks incorporate the metric suite introduced in this work (Inversion Rate, Affective Stability) as standard evaluation dimen- sions. This is especially critical for downstream ap- plications that depend on affective signals: content moderation, mental health monitoring, customer feedback analysis, and social listening systems op- erating across languages all face systematic failure risks due to sentiment inversion. We further argue that the choice is not binary between scale and alignment; targeted distillation strategies, data augmentation, and curated metrics that explicitly prioritize affective calibration can simultaneously address efficiency constraints and ensure representational equity, particularly in low- resource deployment settings. 6 Conclusion We present a cross-lingual sentiment alignment au- dit comparing four transformer architectures on Bengali-English parallel data stratified by dialect. Our findings reveal that current multilingual mod- els exhibit structured affective representational fail- ures: sentiment inversion under compression, di- alectal bias against formal registers, and direc- tional asymmetry in emotional intensity calibra- tion. These failures are distinct from accuracy degradation measured by standard benchmarks and constitute specific threats to equitable language technology access for Bengali users. We argue that building better multilingual representations requires evaluative interventions that make affec- tive alignment failures visible. Future multilingual benchmarks should incorporate Inversion Rate and Affective Stability as standard dimensions. Future training and alignment research should address the formal register gap in Bengali
Chunk 19 · 1,992 chars
tter multilingual representations requires evaluative interventions that make affec- tive alignment failures visible. Future multilingual benchmarks should incorporate Inversion Rate and Affective Stability as standard dimensions. Future training and alignment research should address the formal register gap in Bengali pretraining data and explore dialect-stratified post-training alignment as a path to equitable compression. We believe these directions are generalizable beyond Bengali to the broader landscape of low-resource and dialectally diverse languages underrepresented in current mul- tilingual NLP infrastructure. 8 -- 8 of 10 -- References Berk Atil, Alexa Chittams, Liseng Fu, Ferhan Ture, Lixinyu Xu, and Breck Baldwin. 2024. Llm stabil- ity: A detailed analysis with some surprises. arXiv preprint arXiv:2408.04667, 1. Umme Ayman, Chayti Saha, Azmain Mahtab Rahat, and Sharun Akter Khushbu. 2025. Banglablend: A large- scale nobel dataset of bangla sentences categorized by saint and common form of bangla language. Data in Brief, 58:111240. Francesco Barbieri, Luis Espinosa Anke, and Jose Camacho-Collados. 2022. Xlm-t: Multilingual lan- guage models in twitter for sentiment analysis and beyond. In Proceedings of the thirteenth language resources and evaluation conference, pages 258–266. Tadesse Destaw Belay, Ahmed Haj Ahmed, Alvin C Grissom Ii, Iqra Ameer, Grigori Sidorov, Olga Kolesnikova, and Seid Muhie Yimam. 2025. Culemo: Cultural lenses on emotion-benchmarking llms for cross-cultural emotion understanding. In Proceed- ings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Pa- pers), pages 18894–18909. Abhik Bhattacharjee, Tahmid Hasan, Wasi Ahmad, Kazi Samin Mubasshir, Md Saiful Islam, Anindya Iqbal, M Sohel Rahman, and Rifat Shahriyar. 2022. Banglabert: Language model pretraining and bench- marks for low-resource language understanding eval- uation in bangla. In Findings of the Association for Computational
Chunk 20 · 1,996 chars
18909. Abhik Bhattacharjee, Tahmid Hasan, Wasi Ahmad, Kazi Samin Mubasshir, Md Saiful Islam, Anindya Iqbal, M Sohel Rahman, and Rifat Shahriyar. 2022. Banglabert: Language model pretraining and bench- marks for low-resource language understanding eval- uation in bangla. In Findings of the Association for Computational Linguistics: NAACL 2022, pages 1318–1327. Anirban Bhowmick and Abhik Jana. 2021. Sentiment analysis for bengali using transformer based models. In Proceedings of the 18th International Conference on Natural Language Processing (ICON), pages 481– 486. Shimanto Bhowmik, Tawsif Tashwar Dipto, Md Saz- zad Islam, Sheryl Hsu, and Tahsin Reasat. 2025. Evaluating llms’ multilingual capabilities for ben- gali: Benchmark creation and performance analysis. arXiv preprint arXiv:2507.23248. Terra Blevins, Tomasz Limisiewicz, Suchin Gururan- gan, Margaret Li, Hila Gonen, Noah A Smith, and Luke Zettlemoyer. 2024. Breaking the curse of multi- linguality with cross-lingual expert language models. In Proceedings of the 2024 conference on empiri- cal methods in natural language processing, pages 10822–10837. Vadim Borisov, Samuel Gyamfi, and Richard H. Schreiber. 2025. Multilingual sentiment analysis. Revision 69afb83. Li Chen, Shifeng Shang, and Yawen Wang. 2025. Bridg- ing resource gaps in cross-lingual sentiment anal- ysis: adaptive self-alignment with data augmenta- tion and transfer learning. PeerJ Computer Science, 11:e2851. Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettle- moyer, and Veselin Stoyanov. 2020. Unsupervised cross-lingual representation learning at scale. In Pro- ceedings of the 58th annual meeting of the associa- tion for computational linguistics, pages 8440–8451. Dipto Das, Shion Guha, Jed R Brubaker, and Bryan Semaan. 2024. The“colonial impulse" of natural language processing: An audit of bengali sentiment analysis tools and their identity-based biases.
Chunk 21 · 1,989 chars
g at scale. In Pro- ceedings of the 58th annual meeting of the associa- tion for computational linguistics, pages 8440–8451. Dipto Das, Shion Guha, Jed R Brubaker, and Bryan Semaan. 2024. The“colonial impulse" of natural language processing: An audit of bengali sentiment analysis tools and their identity-based biases. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, pages 1–18. Negar Foroutan, Paul Teiletche, Ayush Kumar Tarun, and Antoine Bosselut. 2025. Revisiting multilingual data mixtures in language model pretraining. arXiv preprint arXiv:2510.25947. Omer Goldman, Uri Shaham, Dan Malkin, Sivan Eiger, Avinatan Hassidim, Yossi Matias, Joshua Maynez, Adi Mayrav Gilady, Jason Riesa, Shruti Rijhwani, and 1 others. 2025. Eclektic: a novel challenge set for evaluation of cross-lingual knowledge transfer. arXiv preprint arXiv:2502.21228. Wenhan Han, Yifan Zhang, Zhixun Chen, Binbin Liu, Haobin Lin, Bingni Zhang, Taifeng Wang, Mykola Pechenizkiy, Meng Fang, and Yin Zheng. 2025. Mubench: Assessment of multilingual capabilities of large language models across 61 languages. arXiv preprint arXiv:2506.19468. Md Nesarul Hoque, Umme Salma, Md Jamal Uddin, Md Martuza Ahamad, and Sakifa Aktar. 2024. Ex- ploring transformer models in the sentiment analysis task for the under-resource bengali language. Natu- ral Language Processing Journal, 8:100091. Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020. Xtreme: A massively multilingual multi-task bench- mark for evaluating cross-lingual generalisation. In International conference on machine learning, pages 4411–4421. PMLR. Saman Sarker Joy and Swakkhar Shatabda. 2025. Bnmmlu: Measuring massive multitask lan- guage understanding in bengali. arXiv preprint arXiv:2505.18951. Mohsinul Kabir, Mohammed Saidul Islam, Md Tah- mid Rahman Laskar, Mir Tafseer Nayeem, M Saiful Bari, and Enamul Hoque. 2024. Benllm-eval: A com- prehensive evaluation into the
Chunk 22 · 1,996 chars
rker Joy and Swakkhar Shatabda. 2025. Bnmmlu: Measuring massive multitask lan- guage understanding in bengali. arXiv preprint arXiv:2505.18951. Mohsinul Kabir, Mohammed Saidul Islam, Md Tah- mid Rahman Laskar, Mir Tafseer Nayeem, M Saiful Bari, and Enamul Hoque. 2024. Benllm-eval: A com- prehensive evaluation into the potentials and pitfalls of large language models on bengali nlp. In Pro- ceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 2238– 2252. Divyanshu Kakwani, Anoop Kunchukuttan, Satish Golla, Gokul NC, Avik Bhattacharyya, Mitesh M Khapra, and Pratyush Kumar. 2020. Indicnlpsuite: Monolingual corpora, evaluation benchmarks and 9 -- 9 of 10 -- pre-trained multilingual language models for indian languages. In Findings of the association for compu- tational linguistics: EMNLP 2020, pages 4948–4961. Zhiwei Liu, Lingfei Qian, Qianqian Xie, Jimin Huang, Kailai Yang, and Sophia Ananiadou. 2025. Mmaff- ben: a multilingual and multimodal affective analy- sis benchmark for evaluating llms and vlms. arXiv preprint arXiv:2505.24423. Hemal Mahmud, Hasan Mahmud, and Mohammad Rifat Ahmmad Rashid. 2024. Enhancing sen- timent analysis in bengali texts: A hybrid ap- proach using lexicon-based algorithm and pre- trained language model bangla-bert. arXiv preprint arXiv:2411.19584. Md Saef Ullah Miah, Md Mohsin Kabir, Talha Bin Sarwar, Mejdl Safran, Sultan Alfarhood, and Md F Mridha. 2024. A multimodal approach to cross- lingual sentiment analysis with ensemble of trans- former and llm. Scientific Reports, 14(1):9603. Millicent Ochieng, Anja Thieme, Ignatius Ezeani, Risa Ueno, Samuel Maina, Keshet Ronen, Javier Gonzalez, and Jacki O’Neill. 2025. Reasoning beyond labels: Measuring llm sentiment in low- resource, culturally nuanced contexts. arXiv preprint arXiv:2508.04199. Alberto Poncelas, Pintu Lohar, James Hadley, and Andy Way. 2020. The impact of indirect machine transla- tion on
Chunk 23 · 1,989 chars
a Ueno, Samuel Maina, Keshet Ronen, Javier Gonzalez, and Jacki O’Neill. 2025. Reasoning beyond labels: Measuring llm sentiment in low- resource, culturally nuanced contexts. arXiv preprint arXiv:2508.04199. Alberto Poncelas, Pintu Lohar, James Hadley, and Andy Way. 2020. The impact of indirect machine transla- tion on sentiment classification. In Proceedings of the 14th Conference of the Association for Machine Translation in the Americas (Volume 1: Research Track), pages 78–88. Sebastian Ruder, Noah Constant, Jan Botha, Aditya Sid- dhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, and 1 others. 2021. Xtreme-r: Towards more challenging and nuanced multilingual evaluation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Lan- guage Processing, pages 10215–10245. Jayanta Sadhu, Maneesha Rani Saha, and Rifat Shahri- yar. 2025. Social bias in large language models for bangla: An empirical study on gender and religious bias. In Proceedings of the First Workshop on Lan- guage Models for Low-Resource Languages, pages 204–218. Leanne Tan, Gabriel Chua, Ziyu Ge, and Roy Ka-Wei Lee. 2025. Lionguard 2: Building lightweight, data- efficient & localised multilingual content moderators. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 264–285. Azmine Toushik Wasi, Raima Islam, Mst Rafia Islam, Taki Hasan Rafi, and Dong-Kyu Chae. 2024. Explor- ing bengali religious dialect biases in large language models with evaluation perspectives. arXiv preprint arXiv:2407.18376. A Supplementary Figures and Tables This appendix contains the visualizations and table referenced in the main findings of the paper. Model Name Repository (HuggingFace) XLM-T cardiffnlp/XLM-Toberta-sentiment (Barbieri et al., 2022) IndicBERT ai4bharat/IndicBERTv2-sentiment Tabularis tabularisai/multilingual-sentiment mDistilBERT lxyuan/distilbert-multilingual Table 4: Model Repository
Chunk 24 · 873 chars
izations and table referenced in the main findings of the paper. Model Name Repository (HuggingFace) XLM-T cardiffnlp/XLM-Toberta-sentiment (Barbieri et al., 2022) IndicBERT ai4bharat/IndicBERTv2-sentiment Tabularis tabularisai/multilingual-sentiment mDistilBERT lxyuan/distilbert-multilingual Table 4: Model Repository Mapping Note on Tabularis: This model is a fine-tuned version of the model distilbert/distilbert-base-multilingual-cased for multilingual sentiment analysis. It utilizes synthetic data from multiple sources to achieve robust performance across different languages and cultural contexts (Borisov et al., 2025). Figure 4: Directional Bias in Sentiment Scores (English − Bengali) Figure 5: Illustrative case study validating Asymmetric Empathy in cross-lingual sentiment alignment, demon- strating severe instance-level directional bias. 10 -- 10 of 10 --