Camellia : Benchmarking Cultural Biases in LLMs for Asian Languages
Summary
This paper introduces Camellia, a benchmark designed to evaluate entity-centric cultural biases in Large Language Models (LLMs) across nine Asian languages representing six distinct cultures. The dataset comprises 19,530 manually annotated entities contrasting Asian and Western cultures, alongside 2,173 masked contexts derived from social media posts. The authors assess four multilingual LLM families—Llama, Qwen, Aya, and Gemma—across three tasks: cultural context adaptation, sentiment association, and entity extractive question answering. Results indicate that LLMs frequently struggle to distinguish between Asian and Western entities, often assigning higher likelihood to Western entities even in culturally grounded Asian contexts. Performance varies significantly by model family; for instance, Qwen demonstrates superior adaptation in Chinese, Japanese, and Korean, likely due to better tokenizer coverage of those scripts. Sentiment analysis reveals distinct biases, with some models associating Western entities with negative sentiments while others link Asian entities to positive ones. Furthermore, extractive QA tasks show substantial accuracy gaps in Asian languages when comparing native versus Western entities, a disparity that largely disappears when testing in English. These findings highlight persistent challenges in cultural fairness and context understanding for non-Western languages, suggesting that current models lack sufficient sensitivity to cultural nuances outside of English-centric training data.
PDF viewer
Chunks(54)
Chunk 0 · 1,998 chars
Camellia : Benchmarking Cultural Biases in LLMs for Asian Languages Tarek Naous1, Anagha Savit1, Carlos Rafael Catalan2, Geyang Guo1, Jaehyeok Lee3, Kyungdon Lee3, Lheane Marie Dizon2, Mengyu Ye4, Neel Kothari1, Sahajpreet Singh5, Sarah Masud6, Tanish Patwa1, Trung Tanh Tran7, Zohaib Khan8, Alan Ritter1, Tanmoy Chakraborty9, Yuki Arase10, Keisuke Sakaguchi4, JinYeong Bak3, Wei Xu1 1Georgia Institute of Technology, 2Samsung R&D Institute Philippines, 3Sungkyunkwan University, 4Tohoku University, 5National University of Singapore, 6University of Copenhagen, 7Takenote.ai, 8University of Michigan, 9Indian Institute of Technology Delhi, 10Institute of Science Tokyo tareknaous@gatech.edu tareknaous/camellia Abstract As Large Language Models (LLMs) develop stronger multilingual capabilities, their sensi- tivity to culturally diverse entities becomes increasingly important. Prior work by Naous et al. (2024) has shown that LLMs often favor Western-associated entities in Arabic. Due to the lack of entity-centric multilingual bench- marks, it remains unclear if such biases also manifest in various non-Western languages. In this paper, we introduce Camellia, a bench- mark for evaluating entity-centric cultural bi- ases in nine Asian languages, spanning six Asian cultures. Camellia includes 19,530 man- ually annotated entities associated with the cov- ered Asian or Western cultures, as well as 2,173 masked contexts for these entities derived from social media posts. Using Camellia, we eval- uate cultural biases in four recent multilingual LLMs across three tasks: cultural context adap- tation, sentiment association, and entity extrac- tive QA. Our analyses show that LLMs struggle with cultural adaptation across these languages, with performance differing across models de- veloped in different regions. We further observe that different LLM families can hold distinct biases, reflected in the ways they link cultures to particular sentiments. Lastly, we find that LLMs can struggle
Chunk 1 · 1,989 chars
at LLMs struggle with cultural adaptation across these languages, with performance differing across models de- veloped in different regions. We further observe that different LLM families can hold distinct biases, reflected in the ways they link cultures to particular sentiments. Lastly, we find that LLMs can struggle with context understanding in some Asian languages, creating performance gaps between cultures in entity extraction. 1 Introduction LLMs have rapidly integrated into modern tech- nology, serving users from diverse cultures (Adi- lazuarda et al., 2024). In the wide range of text they process, LLMs frequently encounter entities such as people’s names, locations, or food dishes, which are pervasive in text corpora (Wolfe and Caliskan, 2021; Pawar et al., 2025b) and often appear in user prompts (Li et al., 2025; Wang et al., 2025). More than just words, entities carry cultural significance. For example, “Oxford” does not merely refer to a city; it also evokes associations with Western academia. Similar connotations exist for the enti- ties in every culture, shaped by historical, linguistic, or religious factors (Whorf, 2012). In light of these cultural associations, it is impor- tant for LLMs to handle culturally diverse entities fairly. Prior work has shown that such associations can influence model behavior, leading to implicit stereotyping (Wan et al., 2023) and discriminatory decision-making (An et al., 2024). For example, Naous et al. (2024) found that, in Arabic, LLMs perform better on entities associated with Western culture than on those linked to Arab culture. This raises the question of whether similar cultural bi- ases also emerge in other non-Western languages. To this end, we introduce Camellia (Cultural Appropriateness Measure Set for LLMs in Asian Languages), a benchmark for measuring entity- centric cultural biases in 9 non-Western languages spoken in the Asian continent: Chinese (zh), Japanese (ja), Korean (ko), Vietnamese (vi),
Chunk 2 · 1,995 chars
so emerge in other non-Western languages. To this end, we introduce Camellia (Cultural Appropriateness Measure Set for LLMs in Asian Languages), a benchmark for measuring entity- centric cultural biases in 9 non-Western languages spoken in the Asian continent: Chinese (zh), Japanese (ja), Korean (ko), Vietnamese (vi), Urdu (ur), Hindi (hi), Malayalam (ml), Marathi (mr), and Gujarati (gu), covering 6 distinct cultures (see Figure 1). We undertook a year-long collaboration with native speakers to collect and annotate 19,530 cultural entities across six entity types that contrast Asian and Western cultures (§3.1). We also curate 2,173 masked contexts for entities derived from natural discussions on social media (§3.2). The en- tities and masked contexts in Camellia enable the evaluation of cultural biases in LLMs via diverse testing setups. Moreover, Camellia includes an English translation for each entity and masked con- arXiv:2510.05291v2 [cs.CL] 27 May 2026 -- 1 of 24 -- Pakistani Culture Urdu (ur) Indian Culture Hindi (hi) Gujarati (gu) Marathi (mr) Malayalam (ml) Chinese Culture Chinese (zh) Japanese Culture Japanese (ja) Korean Culture Korean (ko) Vietnamese Culture Vietnamese (vi) Camellia Multilingual benchmark for entity-centric biases 6 Asian cultures, 9 Asian languages Cultural Entities 15,821 Asian-centric Entities Pakistani food dishes: زردہ , مٹھائی , … Korean person names: 시우, 태현, … Indian authors: अंबिका आनंद, ബിന്ദു ഭട്ട് , … 19,530 manually annotated cultural entities for 6 entity types 3,709 Western-centric Entities (parallel in all languages) Western beverage: Irish cream - ഐറിഷ് ക്രീം - … Western locations: Flagstaff - 弗拉格斯塔夫 –फ्लै गस्टाफ - … Western sports clubs: LA Galaxy - 로스앤젤레스 갤럭시, … Cultural Context Adaptation Sentiment Association Entity Extractive QA 𝑖,𝑗,𝑘 𝕀[𝑃 𝑀𝐴𝑆𝐾 𝑏𝑗 𝑡𝑘 > 𝑃 𝑀𝐴𝑆𝐾 𝑎𝑖 𝑡𝑘 ] Measure LLM preference of Western entities vs Asian entities Asian Asian Western Western Measure LLM association of
Chunk 3 · 1,996 chars
–फ्लै गस्टाफ - … Western sports clubs: LA Galaxy - 로스앤젤레스 갤럭시, … Cultural Context Adaptation Sentiment Association Entity Extractive QA 𝑖,𝑗,𝑘 𝕀[𝑃 𝑀𝐴𝑆𝐾 𝑏𝑗 𝑡𝑘 > 𝑃 𝑀𝐴𝑆𝐾 𝑎𝑖 𝑡𝑘 ] Measure LLM preference of Western entities vs Asian entities Asian Asian Western Western Measure LLM association of entities with negative or positive sentiments ------------------------- ---- tiramisu -------- ------------------------ Measure LLM accuracy at extracting entities from long implicit contexts ------------------------- ------ cá kho tộ ----- ------------------------ Western Entity Vietnamese Entity 88% 74% 73% 76% 73% 89% Evaluation Tasks Masked Contexts 2,173 masked contexts Culturally-Grounded contexts 한국 도시 [MASK]에서 제일 긴 케이블카야 (It's the longest cable car in the Korean city of [MASK]) Culturally-Neutral contexts 这里的这道菜为什么比隔壁卖的 [MASK] 贵? (Why is this dish here more expensive than the [MASK] sold next door?) Extractive QA contexts Lúc sáng vừa mới đăng bài …. hơn 3 tỷ tiền USDT chút xíu a giao về [MASK] giúp em được không ? … sâu vùng xa nào cũng được (Just this morning, I posted ... a little over 3 billion USDT. Could you help me transfer it to [MASK]? ... Anywhere in a remote area is fine) … … 93% 88% Figure 1: We construct Camellia, a benchmark to measure cultural biases for six Asian cultures, covering nine languages. Camellia provides 2,173 masked contexts categorized into: culturally-grounded, culturally-neutral, and extractive question-answering (QA). Camellia also provides 19,530 culturally relevant entities that contrast the respective Asian vs. Western culture across six entity types that exhibit cultural variation. The masked contexts and entities in Camellia enable the evaluation of biases in LLMs via versatile task setups. text, enabling cross-lingual comparisons for testing LLMs in English vs the respective Asian language. Using Camellia, we examine biases in four multilingual LLM families (Llama, Qwen, Aya, Gemma) (§4). We first
Chunk 4 · 1,997 chars
masked contexts and entities in Camellia enable the evaluation of biases in LLMs via versatile task setups. text, enabling cross-lingual comparisons for testing LLMs in English vs the respective Asian language. Using Camellia, we examine biases in four multilingual LLM families (Llama, Qwen, Aya, Gemma) (§4). We first evaluate the ability of LLMs at generating culturally appropriate entities in each language and culture we study. We find a strug- gle by models at distinguishing between Asian and Western cultural entities, assigning higher likeli- hood for Western entities in 30-40% of cases, even when inappropriate to the cultural context (§4.1). We then analyze whether LLMs hold specific senti- ment associations towards Asian and Western cul- tures, revealing distinct biases in different model families (§4.2). This raises concerns about how models developed in different regions may reflect conflicting biases, impacting their downstream per- formance that rely on fair and neutral language un- derstanding. Lastly, we evaluate models at extract- ing entities from paragraph-long contexts where, for several languages, we observe large accuracy gaps when entities in the same text were associated with different cultures suggesting a lack of ability to efficiently grasp context (§4.3). 2 Related Work Multilingual Cultural Evaluation of LLMs. The rapid deployment of LLMs has sparked re- cent interest from the research community in their cultural awareness (Liu et al., 2025; Qadri et al., 2025a,b; Singh et al., 2025), resulting in the release of various cultural evaluation benchmarks (Pawar et al., 2025a). Past work has introduced several QA-style datasets that evaluate models on open- ended culture-specific questions (Myung et al., 2024; Chiu et al., 2025). Other works have focused on constructing knowledge bases to evaluate spe- cific cultural domains such as culinary practices (Palta and Rudinger, 2023; Zhou et al., 2025) or social norms (Fung et al., 2024; Rao et al.,
Chunk 5 · 1,990 chars
at evaluate models on open-
ended culture-specific questions (Myung et al.,
2024; Chiu et al., 2025). Other works have focused
on constructing knowledge bases to evaluate spe-
cific cultural domains such as culinary practices
(Palta and Rudinger, 2023; Zhou et al., 2025) or
social norms (Fung et al., 2024; Rao et al., 2025).
Multilingual resources have also been introduced
to evaluate LLMs on geo-diverse facts (Yin et al.,
2022; Keleg and Magdy, 2023; Dammu et al., 2024;
Tanwar et al., 2025; Maji et al., 2025), regional
exam questions (Romanou et al., 2025; Singh et al.,
2025), and questions on local norms sourced from
native speakers (Guo et al., 2025; Alwajih et al.,
2025). A few studies have also introduced bench-
marks for multilingual multi-modal cultural evalu-
ations, such as the recognition of culture-specific
traditions (Mogrovejo et al., 2024) or food dishes
-- 2 of 24 --
(Li et al., 2024; Winata et al., 2025; Lavrouk et al.,
2025). Less work has evaluated the sensitivity of
LLMs to entities that exhibit cultural variation (An
et al., 2024; Nghiem et al., 2024; Nikandrou et al.,
2025; Zhao et al., 2025; Arora et al., 2025), es-
pecially in non-English languages (Naous et al.,
2024). Our work addresses this gap by introducing
Camellia, a benchmark to measure entity-centric
cultural biases in 6 non-Western cultures in Asia,
covering 9 Asian languages. Camellia includes
2,173 masked contexts constructed from social me-
dia posts and 19,530 entities extracted from Wiki-
data and mC4 web-crawls with manual annotation.
LLM Biases in Asian Languages. Various stud-
ies have introduces multilingual resources for mea-
suring biases in LLMs, which cover Asian lan-
guages. Much of the prior resources probe LLMs
for demographic biases using manually written
templates (e.g.; Everyone hates {attribute}) (Levy
et al., 2023), focusing on attributes such as gender
(Kaneko et al., 2022; Vashishtha et al., 2023; Ding
et al., 2025), race (Costa-jussà et al., 2023),Chunk 6 · 1,996 chars
which cover Asian lan-
guages. Much of the prior resources probe LLMs
for demographic biases using manually written
templates (e.g.; Everyone hates {attribute}) (Levy
et al., 2023), focusing on attributes such as gender
(Kaneko et al., 2022; Vashishtha et al., 2023; Ding
et al., 2025), race (Costa-jussà et al., 2023), reli-
gion (Rinki et al., 2025), age (Zhao et al., 2023),
and more (Hsieh et al., 2024; Lan et al., 2025). An-
other line of research measures the reflection of
culture-specific stereotypes (Sahoo et al., 2024) by
introducing resources of stereotype pairs (Bhutani
et al., 2024) or natural language statements that
reflect stereotypes (Mitchell et al., 2025).
Other works have adapted existing English
benchmarks (Parrish et al., 2022) for measuring
stereotypes in QA model outputs into Chinese
(Huang and Xiong, 2024), Japanese (Yanaka et al.,
2025), and Korean (Jin et al., 2024). Monolingual
resources have been introduced to measure moral
bias in Chinese (Hämmerl et al., 2023), and polit-
ical bias in Urdu (Nadeem et al., 2025). Different
from existing research, our work focuses on evalu-
ating biases in LLMs when handling native Asian
vs Western-centric entities.
3 Constructing Camellia
This section describes our process for construct-
ing Camellia, which extends the Arabic CAMeL
benchmark (Naous et al., 2024) into 9 new Asian
languages. First, we describe the procedure for
collecting culturally-relevant entities across nine
Asian languages (§3.1). We then describe how we
collect masked contexts for entities, which enable
testing for entity-centric cultural biases in LLMs
across versatile setups (§3.2). Additional details
about data collection are outlined in Appendix A.
3.1 Collecting Cultural Entities
Our objective is to collect a comprehensive list of
culturally-relevant entities in each language. This
includes entities tied to Asian cultures where the
language is spoken (e.g., entities associated with
Pakistani culture in Urdu, Chinese culture forChunk 7 · 1,996 chars
ction are outlined in Appendix A. 3.1 Collecting Cultural Entities Our objective is to collect a comprehensive list of culturally-relevant entities in each language. This includes entities tied to Asian cultures where the language is spoken (e.g., entities associated with Pakistani culture in Urdu, Chinese culture for Chi- nese, etc.) and entities written in those Asian lan- guages but associated with Western culture (North America and Europe). We consider 6 entity types that exhibit variation across cultures: authors, food dishes, beverages, first names, locations, and sports clubs. To collect entities, we use the multilingual Wikidata knowledge base and perform pattern- based extraction on web-crawled data. Figure 2 shows the statistics of Asian-centric and Western entities collected in Camellia. Defining Asian vs. Western cultures. For the Asian cultures that we study, there exists a clear distinction between entities that are associated with their respective Asian culture and entities that are typically viewed as Western by those cultures. For example, native Chinese associate the first name “Weili" with Chinese culture and the first name “Valentina” with Western culture. Similarly, native Pakistanis associate the dish “Nihari" with Pak- istani culture and the dish “Lasagna" with Western culture. We follow this phenomenon to distinguish between entities that are native to each Asian cul- ture in Camellia from Western entities. Western culture encompasses many countries across different continents. We limited our West- ern entities to countries in North America (i.e., the United States, Canada, and Mexico) and Europe, as these regions account for the vast majority of West- ern entities that appeared in the Asian languages we study. We grouped all entities from these countries under a broad Western culture, rather than analyz- ing each country separately, which simplifies the design of the benchmark. We note that this excludes some regions such as Australia, a
Chunk 8 · 1,993 chars
the vast majority of West- ern entities that appeared in the Asian languages we study. We grouped all entities from these countries under a broad Western culture, rather than analyz- ing each country separately, which simplifies the design of the benchmark. We note that this excludes some regions such as Australia, a limitation that we discuss at the end of our paper. Extracting Entities from Wikidata. We started by collecting entities from Wikidata by querying the corresponding Wikidata classes for our target entity categories in each language and extracting all registered entities under each class. We found the coverage in Wikidata to be generally suffi- cient for authors, locations, and sports clubs in all languages. However, the coverage for the other -- 3 of 24 -- Authors 王林 Wang Lin 坂木司 Tsukasa Sakaki 공지영 Gong Jiyeong Trần Tiêu Tran Tieu إشفاق أحمد Ashfaq Ahmed अजीत कौर Ajit Kaur आयन रैंड Ayn Rand Beverage 乌龙茶 Oolong ほうじ茶 Hōjicha 식혜 Sikhye Cà phê trứng Egg coffee بورہانی Burhani चास Chaas 버번 Bourbon Food 光饼 Kompyang バラ焼き Barayaki 돔배기 Dombaegi Cơm hến Baby Clam Rice برفی Barfi സ ിംഗ്ജു Singju ティラミス Tiramisu Locations 红桥镇 Hongqiao 臼杵 Usuki 신암 Sinam Nam Dinh Nam Dinh سمندری Samundri વારાણસી Varanasi 코체부 Kotzebue Names (M) 涛鑫 Taoxin 均 Hitoshi 윤호 Yunho Thiết Thiet مظفر Muzaffar अभिषेक Abhishek કોનર Connor Names (F) 瑶宇 Yaoyu 祐希 Yuki 민설 Minseol Liên Lien نوشین Nosheen સંગીતા Sangita ഇസബെല്ല Isabella Sports 天津立飞 Tianjin Lifei FC刈谷 FC Kariya 포항 스틸러스 Pohang Steelers Thanh Hóa Thanh Hoa کراچی زیبراز Karachi Zebras ગુજરાત ટાઇટન્ સ Gujarat Titans AC プラート A.C. Prato Figure 2: Example per entity type and statistics of respective Asian entities per culture and Western entities in Camellia. Western entities are parallel for all 9 languages while Indian entities are parallel in all Indian languages (§3.1). Camellia also provides an English translation for each entity. entity types (food dishes, beverages, names) was much less extensive and varied by language. We
Chunk 9 · 1,989 chars
r culture and Western entities in Camellia. Western entities are parallel for all 9 languages while Indian entities are parallel in all Indian languages (§3.1). Camellia also provides an English translation for each entity. entity types (food dishes, beverages, names) was much less extensive and varied by language. We ob- served that higher-resource languages had a sizable amount of entities in Wikidata (e.g., 253 Indian food dishes written in Hindi) while lower-resource languages had much less representation (e.g., only 24 Indian food dishes in Malayalam, 37 Pakistani names in Urdu, etc.). Pattern-based Extraction from Web-Crawls. To expand on the initial lists obtained from Wiki- data for entity types that had little coverage, we performed pattern-based extraction of entities from web-crawled corpora in each language. We man- ually defined patterns in each language that typi- cally precede entities (e.g., brother/sister named for first names, recipe of for food dishes, etc.). Using the patterns, we scanned through each language’s partition in the open-source mC4 web- crawl corpus (Xue et al., 2021) and extracted uni- grams and bigrams that appeared after a detected pattern. We also accounted for gender inflections if required. This resulted in 5k-10k extractions per type and language, which were then manually fil- tered to remove irrelevant ones. Since Chinese and Japanese do not use word-separating spaces, we re- trieved both the detected pattern (e.g., “喝”, which means “to drink”) and up to ten surrounding char- acters, and then prompted GPT-4o-mini to extract the entity from the captured characters, if any were mentioned. This was followed by manual filtering to remove irrelevant characters. Annotation by Native Speakers. The annota- tion was conducted by a total of nine different au- thors of this paper, each a native speaker of one of the 9 Asian languages in Camellia. This involved manual filtering of the Wikidata and mC4 extrac- tions to identify
Chunk 10 · 1,997 chars
ed by manual filtering to remove irrelevant characters. Annotation by Native Speakers. The annota- tion was conducted by a total of nine different au- thors of this paper, each a native speaker of one of the 9 Asian languages in Camellia. This involved manual filtering of the Wikidata and mC4 extrac- tions to identify culturally relevant entities. The collected entities were then annotated for being as- sociated with the respective Asian culture of the language or associated with Western culture. To ensure quality, we performed double annotation of the entities in each language. The second set of annotators consisted of undergraduate or mas- ter’s students hired for zh, ja, ko, hi, ml, mr, and gu; and native speaker volunteers for vi and ur. We achieved high inter-annotator agreements as measured by Cohen’s Kappa (zh: 0.85, ja: 0.78, ko: 0.92, vi: 0.80, ur: 0.88, hi: 0.94, ml: 0.83, mr: 0.93, gu: 0.97). The disagreements were resolved in an adjudication step to decide the final label. Translating Entities to English. To support comparative analyses of LLM performance when tested in both the native language and English, we mapped each entity in Camellia to its English translation. When possible, we retrieved the En- glish label directly from Wikidata (available for 86.58% of Wikidata-sourced entities). For entities without an English label and ones extracted from mC4, we manually searched for their most com- monly used English transliterated form found on- line, ensuring that the translations reflect how enti- ties appear in real-world usage. -- 4 of 24 -- Parallelizing Western Entities. To enable lan- guage comparisons in our experiments, we paral- lelized the Western entities across all languages (i.e., each Western entity has a written version in every language). For authors, locations, and sports clubs, we constructed their parallel Western sets directly from Wikidata by extracting the entities of each Western country in North America and Eu- rope that had a
Chunk 11 · 1,980 chars
lelized the Western entities across all languages (i.e., each Western entity has a written version in every language). For authors, locations, and sports clubs, we constructed their parallel Western sets directly from Wikidata by extracting the entities of each Western country in North America and Eu- rope that had a written form in at least 6 of the languages. A lot of these Western entities did not have written versions in Wikidata in lower-resource languages (ur, ml, gu, and mr). For those cases, we manually filled in their missing translations. For food, beverage, and names, Western en- tities were collected independently in each lan- guage via pattern-based extractions mC4. We uni- fied these language-specific sets by first using their English translations as the common key. Specifi- cally, when the same English translation appeared for multiple languages, we treated it as the com- mon “parallel” entity. This revealed large over- laps for high-resource languages (hi, zh, ja, ko), which shared many common Western entities, but also showed substantial gaps for low-resource lan- guages in which data was already scarce (e.g., 1k–1.5k food entities needed to be translated to ur). To balance translation efforts and ensuring high quality, we randomly sampled 500 unified entities per type and, with the help of annotators, manually completed the missing entries by translating them from English into their languages. Parallelizing Entities in Indian Languages. To enable direct comparisons between Indian lan- guages, we also parallelized the Indian entities across the four Indian languages (hi, ml, mr, gu). Since Indian entities were independently collected and annotated for each language, we used their En- glish translations as an intermediate representation to map equivalent entities across languages. An- notators then translated the missing gaps from En- glish. The majority of Indian cultural entities were initially collected in Hindi (hi), being the
Chunk 12 · 1,996 chars
ndently collected and annotated for each language, we used their En- glish translations as an intermediate representation to map equivalent entities across languages. An- notators then translated the missing gaps from En- glish. The majority of Indian cultural entities were initially collected in Hindi (hi), being the most resource-rich Indian language. In contrast, transla- tion efforts were mostly required to map entities into lower-resource languages (ml, mr, gu). 3.2 Constructing Masked Contexts To evaluate whether LLMs can distinguish between entities associated with each Asian culture vs. those associated with Western cultures, Camellia provides 2,173 masked contexts for entities de- Camellia-Grounded Camellia-Neutral Camellia-QA #Contexts Avg Len #Contexts Avg Len #Contexts Avg Len zh 131 37.95±8.99 126 32.02±8.44 64 81.59±18.20 ja 137 58.47±16.39 140 43.41±12.20 60 115.08±26.27 ko 150 9.13±4.62 208 7.52±3.27 70 32.30±6.57 vi 165 38.44±13.44 192 40.05±14.50 78 64.17±6.92 ur 70 15.63±6.27 70 14.20±4.67 58 47.02±13.42 hi 215 21.21±12.02 192 14.36±7.85 47 45.40±13.05 ml 215 13.78±7.48 192 9.21±4.94 47 29.06±9.31 mr 215 16.59±9.30 192 11.41±6.32 47 33.66±9.82 gu 215 18.21±9.94 192 12.62±7.04 47 36.91±11.55 Table 1: Statistics of masked contexts in Camellia. In- dian contexts are parallel across all Indian languages. We report word length for all languages except Chi- nese and Japanese, for which we report character length. Each context is also available as an English translation. rived from naturally-occurring discussions by na- tive speakers on X (formerly Twitter). We source our contexts from X for all languages except Chi- nese, for which we use Weibo and Xiaohongshu instead, since X is officially blocked in China. The masked contexts in Camellia are split into the fol- lowing three types as summarized in Table 1: • Camellia-Grounded: contexts that are uniquely suited for entities associated with each
Chunk 13 · 1,998 chars
rom X for all languages except Chi- nese, for which we use Weibo and Xiaohongshu instead, since X is officially blocked in China. The masked contexts in Camellia are split into the fol- lowing three types as summarized in Table 1: • Camellia-Grounded: contexts that are uniquely suited for entities associated with each Asian cul- ture, enabling us to assess cultural adaptation (e.g., ᄒ ᅡ ᆫᄀ ᅮ ᆨᄀ ᅡᄆ ᅧ ᆫ ᄆ ᅮᄌ ᅩᄀ ᅥ ᆫ ᄆ ᅡᄉ ᅧᄋ ᅣᄀ ᅦ ᆻᄃ ᅡ [MASK] - When I go to Korea, I absolutely have to try [MASK]). • Camellia-Neutral: neutral contexts where enti- ties from any culture are appropriate, helping de- termine the default inclinations of models in the absence of cultural cues (e.g., ᄋ ᅩᄂ ᅳ ᆯ ᄆ ᅮᄌ ᅩᄀ ᅥ ᆫ ᄆ ᅥ ᆨᄂ ᅳ ᆫ ᄃ ᅡ [MASK] - I’m definitely eating [MASK] today). • Camellia-QA: paragraph-long contexts that ref- erence entities implicitly, presenting a challenging setup for testing models at entity identification in an extractive QA format. Contexts for Evaluating Cultural Adaptation. To construct Camellia-Grounded, we searched us- ing two types of search queries: randomly sampled Asian entities (e.g., [Korean entity], [Indian entity], etc.), and manually designed patterns that mention a culturally-relevant entity (e.g., the [Korean] city of, the [Indian] dish, etc.). We then manually in- spected the retrieved tweets to identify ones that provide contexts where only an entity associated with the respective Asian culture can be placed. From these, we constructed our masked contexts by replacing the entity mentioned in the tweet with a [MASK] token. Similarly, to construct the culturally- neutral contexts of Camellia-Neutral, we identi- fied tweets where entities from any culture would be appropriate as [MASK]. -- 5 of 24 -- Sentiment Annotation. We annotated each con- text with one of three sentiment labels: positive, negative, or neutral. This helps evaluate whether substituting the [MASK] token with the respective Asian or Western entities changes the sentiment predicted by
Chunk 14 · 1,992 chars
om any culture would be appropriate as [MASK]. -- 5 of 24 -- Sentiment Annotation. We annotated each con- text with one of three sentiment labels: positive, negative, or neutral. This helps evaluate whether substituting the [MASK] token with the respective Asian or Western entities changes the sentiment predicted by LLMs (§4.2). Contexts for Extractive QA. In addition to the contexts used to evaluate cultural adaptation in LLMs, we constructed longer, paragraph-level con- texts (Table 1) in which entities are mentioned im- plicitly. These longer contexts enable a challenging evaluation setup for entity extraction, as they re- quire understanding of the underlying context to identify the entity. We follow the same keyword search strategy to identify such contexts, and re- place the mentioned entity with the [MASK] token. Camellia-QA provides ∼8-10 of such contexts for each entity type in each language. Parallelizing Indian Contexts. The contexts in hi, ml, mr, and gu were originally collected in- dependently for each language. To enable compar- isons across these Indian languages, we parallelized them by first translating the contexts into English and then into the other Indian languages. 4 Are Cultural Biases Consistent Across Languages and LLMs? We leverage the cultural entities and masked con- texts in Camellia to investigate whether cultural bi- ases are persistent across languages and LLMs. We experiment with four open-weight LLMs with mul- tilingual abilities: Llama3.3-70b (Grattafiori et al., 2024), Qwen2.5-72b (Yang et al., 2025), Aya- expanse-32b (Dang et al., 2024), and Gemma3- 27b (Team et al., 2025). We test the LLMs under three setups: cultural adaptation (§4.1), sentiment association (§4.2), and extractive QA (§4.3). Addi- tional experimental details and results are provided in Appendices B and C, respectively. 4.1 Cultural Context Adaptation We first assess the ability of LLMs to adapt to dif- ferent Asian cultural contexts by analyzing their
Chunk 15 · 1,995 chars
ups: cultural adaptation (§4.1), sentiment
association (§4.2), and extractive QA (§4.3). Addi-
tional experimental details and results are provided
in Appendices B and C, respectively.
4.1 Cultural Context Adaptation
We first assess the ability of LLMs to adapt to dif-
ferent Asian cultural contexts by analyzing their as-
signed likelihood for the respective Asian vs West-
ern entities as [MASK] token fillings.
Cultural Bias Score (CBS). We use the CBS
(Naous et al., 2024) to measure the level of West-
ern bias in an LLMθ. CBS is a likelihood-based
measure that computes the percentage of an LLM’s
preference for Western entities over Asian ones
within the same cultural context. Given an en-
tity type D, two type-specific sets of respective
Asian entities A = {ai}N
i=1 and Western entities
B = {bj }M
j=1, and a masked context ck, we com-
pute CBSD(LLMθ, A, B, ck) per language as:
1
N × M
N X
i=1
M X
j=1
1[P[MASK](bj |ck) > P[MASK](ai|ck)]
(1)
where P[MASK] is the LLM’s probability of an en-
tity filling the [MASK] token. For entities tokenized
into multiple tokens, we take the product of the
conditional probabilities of each token. For a set of
prompts C = {ck}K
k=1, the CBS per entity type for
an LLM is computed by averaging over all ck ∈ C.
An LLM is considered more Western-biased as its
CBS gets close to 100%.
Results. Figures 3 and 4 show the average CBS
across entity types when tested in each language.
We observe the following key insights:
LLMs can struggle to distinguish Asian vs. West-
ern entities. Since the contexts we test on are
grounded in each Asian culture (only entities asso-
ciated with the specific Asian culture are appropri-
ate for filling the [MASK]), models should always
assign higher likelihood to the native Asian enti-
ties in those contexts, and the CBS is expected
to be low (closer to the 0-5% range). However,
as Figure 3 shows, in most cases the CBS is in
the 30-40% range. This highlights a large num-
ber of cases where LLMs struggle toChunk 16 · 1,998 chars
r filling the [MASK]), models should always assign higher likelihood to the native Asian enti- ties in those contexts, and the CBS is expected to be low (closer to the 0-5% range). However, as Figure 3 shows, in most cases the CBS is in the 30-40% range. This highlights a large num- ber of cases where LLMs struggle to differentiate between Asian and Western entities, assigning a higher likelihood to Western ones despite them be- ing inappropriate to the context. Are models sensitive to cultural grounding? We further analyze if performance changes when testing on the contexts that are culturally neutral (i.e., any entity is an appropriate [MASK] filling in the context). The results in Figure 4 show that CBS scores are slightly higher when contexts are neutral. In the majority of cases, the scores remain very close to when contexts are culturally grounded. This suggests a lack of sensitivity to cultural con- texts in LLMs, whereby their ability to select the appropriate entities at generation time is not greatly impacted by cultural grounding. Adaptation ability varies by LLM family. No- ticeable differences can be seen in the performance of LLM families developed in different regions. -- 6 of 24 -- Chinese Japanese Korean Vietnamese Pakistani Indian Indian Indian Indian 20 30 40 50 CBS zh ja ko vi ur hi ml mr gu Llama3.3-70b Chinese Japanese Korean Vietnamese Pakistani Indian Indian Indian Indian 20 30 40 50 CBS zh ja ko vi ur hi ml mr gu Qwen2.5-72b Chinese Japanese Korean Vietnamese Pakistani Indian Indian Indian Indian 20 30 40 50 CBS zh ja ko vi ur hi ml mr gu Aya-expanse-32b Chinese Japanese Korean Vietnamese Pakistani Indian Indian Indian Indian 20 30 40 50 CBS zh ja ko vi ur hi ml mr gu Gemma3-27b Figure 3: Average Cultural Bias Score (CBS) (↓) across entity types achieved by LLMs on culturally-grounded contexts (Camellia-Grounded). LLMs can struggle to generate the appropriate Asian entities in each culture, assigning better likelihood to
Chunk 17 · 1,997 chars
Indian Indian 20 30 40 50 CBS zh ja ko vi ur hi ml mr gu Gemma3-27b Figure 3: Average Cultural Bias Score (CBS) (↓) across entity types achieved by LLMs on culturally-grounded contexts (Camellia-Grounded). LLMs can struggle to generate the appropriate Asian entities in each culture, assigning better likelihood to Western entities 30-40% of the time. See results per entity type in Appendix C.1. zh ja ko vi ur hi ml mr gu 20 30 40 50 CBS Llama3.3-70b zh ja ko vi ur hi ml mr gu 20 30 40 50 CBS Qwen2.5-72b zh ja ko vi ur hi ml mr gu 20 30 40 50 CBS Aya-expanse-32b zh ja ko vi ur hi ml mr gu 20 30 40 50 CBS Gemma3-27b Culturally Grounded Contexts Culturally Neutral Contexts Figure 4: Average CBS across entity types on culturally-grounded contexts (Camellia-Grounded) vs culturally- neutral contexts (Camellia-Neutral) . LLMs show more preference towards Western entities in culturally-neutral contexts (higher CBS). CBS scores are lower in culturally-grounded contexts, yet remain close to the neutral case. Specifically, the Qwen2.5-72b model that is devel- oped by China-based Alibaba performs the best on Chinese, Japanese, and Korean, compared to the rest of the models. This corroborates the results of past work that shows a better ability of Qwen models at answering questions specific to Chinese culture (Guo et al., 2025). One likely reason for such a gap could be varying access to culturally relevant pre-training data. Because the pre-training datasets of these open-source models are not pub- licly disclosed, directly analyzing differences in their training data remains challenging. However, we present an additional analysis in Appendix C.1, where we compare model tokenizers, which of- fer a lens into their pre-training data characteris- tics (Hayase et al., 2024). Our results indicate that greater representation of a language’s script in a model’s tokenizer, which suggests stronger data coverage during frequency-based vocabulary con- struction, is associated with
Chunk 18 · 1,996 chars
compare model tokenizers, which of- fer a lens into their pre-training data characteris- tics (Hayase et al., 2024). Our results indicate that greater representation of a language’s script in a model’s tokenizer, which suggests stronger data coverage during frequency-based vocabulary con- struction, is associated with improved performance. 4.2 Sentiment Association We now examine whether LLMs subtly associate entities from Asian or Western cultures with spe- cific sentiment labels. Setup. We leverage the masked contexts in Camellia-Grounded and Camellia-Neutral that were manually annotated for sentiment to create a test set in each language. For each context, we replace the [MASK] token with 50 randomly sam- pled culture-specific Asian and Western entities. This results in two separate evaluation sets of ∼20k sentences per language: one with culture-specific Asian entities and the other with Western entities. Importantly, the contexts remain the same across both sets, allowing us to isolate the effect of entity cultural association on changes in the LLMs’ pre- dictions. We prompt LLMs to predict the sentiment of each sample and compare their false negative sentiment and false positive sentiment predictions between sentences containing Asian entities vs. Western entities. Fair LLMs should have near-zero false negative or false positive differences since their sentiment prediction should be based on the sentence’s context and not the swap of entities. Results. Figure 5 shows the average differences in false negative and false positive predictions by LLMs for each language. We find that sentiment associations vary greatly across different LLMs. Llama and Gemma generally show a stronger ten- dency to associate Western entities with negative sentiment, whereas Qwen often associates Asian entities with positive sentiment, particularly in In- dian languages. These results highlight how current LLMs can be sensitive to cultural associations of entities leading to biased
Chunk 19 · 1,958 chars
nd Gemma generally show a stronger ten- dency to associate Western entities with negative sentiment, whereas Qwen often associates Asian entities with positive sentiment, particularly in In- dian languages. These results highlight how current LLMs can be sensitive to cultural associations of entities leading to biased misclassifications. -- 7 of 24 -- Llama3.3-70b Qwen2.5-72b Aya-expanse-32b Gemma3-27b zh ja ko vi ur hi ml mr gu -45.7 ±7.5 43.7 ±6.1 -19.0 ±12.0 -72.3 ±8.4 -28.3 ±4.3 -2.7 ±3.7 -49.0 ±10.3 -74.3 ±4.3 -85.3 ±14.8 26.0 ±12.2 11.7 ±19.9 -226.3 ±16.4 -22.7 ±10.9 -30.3 ±4.5 -34.7 ±9.6 -115.0 ±4.4 2.3 ±2.6 -1.0 ±1.2 18.0 ±4.8 -12.7 ±2.7 -88.7 ±13.2 -65.3 ±11.9 10.0 ±14.1 -131.0 ±12.3 -115.3 ±12.5 22.7 ±14.3 -36.3 ±14.8 34.0 ±11.0 -43.7 ±16.3 41.0 ±16.0 -16.3 ±20.0 -61.7 ±15.6 -77.7 ±17.3 -69.3 ±15.9 -16.7 ±14.5 -103.7 ±14.7 Llama3.3-70b Qwen2.5-72b Aya-expanse-32b Gemma3-27b zh ja ko vi ur hi ml mr gu -73.3 ±10.6 -55.7 ±17.9 58.7 ±15.0 -33.7 ±10.4 49.3 ±8.9 19.7 ±11.9 81.7 ±17.0 -11.3 ±11.5 -20.3 ±14.5 79.3 ±21.5 170.0 ±26.1 255.7 ±15.3 -90.7 ±12.8 32.7 ±12.7 72.3 ±13.2 11.7 ±6.4 -6.0 ±4.8 10.0 ±3.9 -25.3 ±7.6 2.3 ±2.8 -33.0 ±16.1 134.0 ±17.4 3.7 ±14.1 -10.7 ±9.0 32.3 ±18.5 250.3 ±22.0 4.0 ±17.8 110.7 ±9.4 -84.7 ±18.4 187.0 ±22.8 131.0 ±19.1 27.0 ±13.9 -34.3 ±21.1 189.0 ±20.4 108.0 ±23.6 -6.7 ±9.9 100 50 0 50 100 False Negatives Asian Negativity Western Negativity 200 100 0 100 200 False Positives Asian Positivity Western Positivity Figure 5: Differences in False Negative (FN) and False Positive (FP) sentiment predictions by LLMs on Camellia contexts filled with Asian vs Western entities. Results are averaged across 3 runs of 50 randomly sampled Asian vs Western entities in each language. zh ko ja vi ur hi ml mr gu 50 75 100 QA Accuracy Llama3.3-70b zh ko ja vi ur hi ml mr gu 50 75 100 QA Accuracy Qwen2.5-72b zh ko ja vi ur hi ml mr gu 50 75 100 QA Accuracy
Chunk 20 · 1,996 chars
s filled with Asian vs Western entities. Results are averaged across 3 runs of 50 randomly sampled Asian vs Western entities in each language. zh ko ja vi ur hi ml mr gu 50 75 100 QA Accuracy Llama3.3-70b zh ko ja vi ur hi ml mr gu 50 75 100 QA Accuracy Qwen2.5-72b zh ko ja vi ur hi ml mr gu 50 75 100 QA Accuracy Aya-expanse-32b zh ko ja vi ur hi ml mr gu 50 75 100 QA Accuracy Gemma3-27b Entity Cultural Association Respective Asian Culture Western Culture Figure 6: Extractive QA accuracy by LLMs on Camellia-QA contexts containing Asian vs Western entities when tested in each Asian language. LLMs generally achieve higher accuracy on extracting entities associated with each Asian culture rather than Western-associated entities. 4.3 Entity Extractive QA Finally, we analyze the ability of LLMs to extract entities from paragraph-long contexts. We compare their performance when these entities are associ- ated with Asian vs. Western cultures. Setup. Using the contexts from Camellia-QA, we construct Asian and Western test sets in each language. For each context, we replace the [MASK] with 50 randomly sampled entities, in a similar manner to our earlier experiment for sentiment as- sociation (§4.2). We then prompt LLMs to extract the entity from each context and compute their ac- curacy on the Asian vs Western test sets. Results. Figure 6 shows the accuracies achieved by LLMs. We find many cases where LLMs gen- erally achieve higher accuracy in extracting the native Asian entities rather than Western ones. A few languages show the opposite behavior, specif- ically in Vietnamese and Urdu, where Llama and Qwen achieve higher accuracy on Western entities. To compare whether similar gaps also appear in English, we evaluate all models on parallel English data for each culture. The results are shown in Ta- ble 2. In English, cultural gaps are generally small, Llama3.3-70b Qwen2.5-70b Aya-expanse-32b Gemma3-27b Culture Asian English Asian English Asian English Asian
Chunk 21 · 1,995 chars
es. To compare whether similar gaps also appear in English, we evaluate all models on parallel English data for each culture. The results are shown in Ta- ble 2. In English, cultural gaps are generally small, Llama3.3-70b Qwen2.5-70b Aya-expanse-32b Gemma3-27b Culture Asian English Asian English Asian English Asian English Chinese -1.32 0.30 0.43 -2.84 2.84 -5.83 -1.36 -5.63 Japanese 7.55 2.72 18.87 4.53 8.84 -0.73 16.40 -3.22 Korean 9.69 0.66 16.47 -2.49 13.94 1.43 7.94 2.54 Vietnamese -13.53 1.95 -14.33 -3.61 2.83 -1.88 4.15 1.65 Pakistani -4.71 10.54 -4.99 12.16 0.12 4.54 21.11 4.54 Indian (hi) 10.05 6.71 3.63 10.67 11.54 1.07 6.81 3.25 Indian (ml) 13.15 — 4.22 — 10.93 — 9.01 — Indian (mr) 11.07 — 1.68 — 12.64 — 3.50 — Indian (gu) 14.44 — 6.02 — 12.89 — 6.54 — Table 2: ∆Accuracy on extractive QA between Western and Asian entities when testing models on parallel data in the respective Asian language of each culture vs. in English. Gaps between cultures are generally much smaller in English, while gaps in Asian languages are larger, falling mostly in the range of 10-20%. mostly between 1% and 5%, with no consistent ad- vantage for either culture. In contrast, gaps in Asian languages are often larger, reaching 12%-20%, ex- cept in Chinese where the differences remain mini- mal. These results suggest that LLMs still struggle to capture implicit cultural context in many non- English languages, leading to substantially larger performance disparities across cultures. -- 8 of 24 -- 5 Conclusion We introduced Camellia, a benchmark for evalu- ating entity-centric cultural biases in 9 Asian lan- guages across 6 Asian cultures. Through various evaluations, we showed that current multilingual LLMs exhibit various types of cultural biases in these non-Western languages which can manifest in poor cultural adaptation, biased sentiment as- sociations, and accuracy gaps in entity extraction tasks. We hope
Chunk 22 · 1,993 chars
Asian lan- guages across 6 Asian cultures. Through various evaluations, we showed that current multilingual LLMs exhibit various types of cultural biases in these non-Western languages which can manifest in poor cultural adaptation, biased sentiment as- sociations, and accuracy gaps in entity extraction tasks. We hope that Camellia will serve as a valu- able resource to support future research aimed at developing more culturally fair multilingual LLMs. Limitations In Camellia, we defined the broad Western culture as countries that are exclusively in North America and Europe. However, there are several countries in other geographical regions where Western cul- ture dominates such as Australia, New Zealand, and South American countries, that were excluded from our definition. Our focus was to explore cul- tural biases in LLMs when contrasting entities as- sociated each Asian culture we study against those associated with the broad Western culture. We thus followed the view of North America and Europe as representing Western culture, and for which data could be more easily collected in the Asian lan- guages we study. Future work can expand on our set of Western entities to include more representa- tion from these regions to enable more fine-grained comparisons. Ethics Statement While collecting data from social media posts to construct the masked contexts in Camellia, we dis- carded any tweets that included offensive or toxic language, hate speech, stereotypes, or included any personally identifiable information. Data collection was done through a manual process by searching on the X, Weibo, and Xiaohongshu platforms, with- out the use of automated scraping. We do not share the raw social media posts but modified, slightly re-written versions where cultural entities are re- placed by a [MASK], which can be used for research purposes. Camellia is constructed for the purpose of testing cultural biases in LLMs and enabling future research on the development of LLMs
Chunk 23 · 1,989 chars
scraping. We do not share the raw social media posts but modified, slightly re-written versions where cultural entities are re- placed by a [MASK], which can be used for research purposes. Camellia is constructed for the purpose of testing cultural biases in LLMs and enabling future research on the development of LLMs that work efficiently and fairly for all entities regardless of the cultural associations they carry. Reproducibility Statement The Camellia benchmark will be made publicly available to the community, which includes the col- lected entities with their annotations for cultural association and the masked contexts for all lan- guages. We provide in Appendix A the annotation guideline we used to annotate entities, and addi- tional experimental details in Appendix B, such as the prompts and decoding configurations that can be used to replicate our experiments for all languages. References Muhammad Farid Adilazuarda, Sagnik Mukherjee, Pradhyumna Lavania, Siddhant Shivdutt Singh, Al- ham Fikri Aji, Jacki O’Neill, Ashutosh Modi, and Monojit Choudhury. 2024. Towards measuring and modeling “culture” in LLMs: A survey. In Proceed- ings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 15763–15784, Miami, Florida, USA. Association for Computational Linguistics. Fakhraddin Alwajih, Abdellah El Mekki, Samar Mo- hamed Magdy, Abdelrahim A Elmadany, Omer Nacar, El-Moatez-Billah Nagoudi, Reem Abdel- Salam, Hanin Atwany, Youssef Nafea, Abdulfat- tah Mohammed Yahya, and 1 others. 2025. Palm: A culturally inclusive and linguistically diverse dataset for Arabic LLMs. In Proceedings of the 63rd An- nual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 32871– 32894. Haozhe An, Christabel Acquaye, Colin Wang, Zongxia Li, and Rachel Rudinger. 2024. Do large language models discriminate in hiring decisions on the basis of race, ethnicity, and gender? In Proceedings of the 62nd Annual Meeting of the
Chunk 24 · 1,988 chars
ssociation for Computational Linguistics (Volume 1: Long Papers), pages 32871– 32894. Haozhe An, Christabel Acquaye, Colin Wang, Zongxia Li, and Rachel Rudinger. 2024. Do large language models discriminate in hiring decisions on the basis of race, ethnicity, and gender? In Proceedings of the 62nd Annual Meeting of the Association for Compu- tational Linguistics (Volume 2: Short Papers), pages 386–397. Shane Arora, Marzena Karpinska, Hung-Ting Chen, Ip- sita Bhattacharjee, Mohit Iyyer, and Eunsol Choi. 2025. CaLMQA: Exploring culturally specific long- form question answering across 23 languages. In Proceedings of the 63rd Annual Meeting of the As- sociation for Computational Linguistics (Volume 1: Long Papers), pages 11772–11817, Vienna, Austria. Association for Computational Linguistics. Ivona Barešová and Petr Janda. 2023. Tradition and change: naming practices in contemporary Japan and Taiwan. Continuity and change in Asia, pages 393– 411. Mukul Bhutani, Kevin Robinson, Vinodkumar Prab- hakaran, Shachi Dave, and Sunipa Dev. 2024. SeeG- ULL multilingual: A dataset of geo-culturally situ- ated stereotypes. In Proceedings of the 62nd Annual -- 9 of 24 -- Meeting of the Association for Computational Lin- guistics (Volume 2: Short Papers), pages 842–854. Yu Ying Chiu, Liwei Jiang, Bill Yuchen Lin, Chan Young Park, Shuyue Stella Li, Sahithya Ravi, Mehar Bhatia, Maria Antoniak, Yulia Tsvetkov, Vered Shwartz, and 1 others. 2025. CulturalBench: A robust, diverse and challenging benchmark for mea- suring lms’ cultural knowledge through human-ai red- teaming. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Vol- ume 1: Long Papers), pages 25663–25701. John Connell. 2018. Globalisation, soft power, and the rise of football in China. Geographical research, 56(1):5–15. Marta R Costa-jussà, Pierre Andrews, Eric Smith, Prangthip Hansanti, Christophe Ropers, Elahe Kalbassi, Cynthia Gao, Daniel Licht, and Carleigh Wood. 2023.
Chunk 25 · 1,991 chars
(Vol- ume 1: Long Papers), pages 25663–25701. John Connell. 2018. Globalisation, soft power, and the rise of football in China. Geographical research, 56(1):5–15. Marta R Costa-jussà, Pierre Andrews, Eric Smith, Prangthip Hansanti, Christophe Ropers, Elahe Kalbassi, Cynthia Gao, Daniel Licht, and Carleigh Wood. 2023. Multilingual holistic bias: Extending descriptors and patterns to unveil demographic bi- ases in languages at scale. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 14141–14156. Preetam Prabhu Srikar Dammu, Hayoung Jung, Anjali Singh, Monojit Choudhury, and Tanu Mitra. 2024. “They are uncultured”: Unveiling covert harms and so- cial threats in LLM generated conversations. In Pro- ceedings of the 2024 Conference on Empirical Meth- ods in Natural Language Processing, pages 20339– 20369, Miami, Florida, USA. Association for Com- putational Linguistics. John Dang, Shivalika Singh, Daniel D’souza, Arash Ahmadian, Alejandro Salamanca, Madeline Smith, Aidan Peppin, Sungjin Hong, Manoj Govindassamy, Terrence Zhao, and 1 others. 2024. Aya expanse: Combining research breakthroughs for a new multi- lingual frontier. arXiv preprint arXiv:2412.04261. YiTian Ding, Jinman Zhao, Chen Jia, Yining Wang, Zifan Qian, Weizhe Chen, and Xingyu Yue. 2025. Gender bias in large language models across multiple languages: A case study of ChatGPT. In Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025), pages 552–579. Negar Foroutan, Clara Meister, Debjit Paul, Joel Niklaus, Sina Ahmadi, Antoine Bosselut, and Rico Sennrich. 2025. Parity-aware byte-pair encoding: Im- proving cross-lingual fairness in tokenization. arXiv preprint arXiv:2508.04796. Yi Fung, Ruining Zhao, Jae Doo, Chenkai Sun, and Heng Ji. 2024. Massively multi-cultural knowledge acquisition & LM benchmarking. arXiv preprint arXiv:2402.09369. Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha
Chunk 26 · 1,986 chars
n tokenization. arXiv preprint arXiv:2508.04796. Yi Fung, Ruining Zhao, Jae Doo, Chenkai Sun, and Heng Ji. 2024. Massively multi-cultural knowledge acquisition & LM benchmarking. arXiv preprint arXiv:2402.09369. Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, and 1 others. 2024. The Llama 3 herd of models. arXiv preprint arXiv:2407.21783. Geyang Guo, Tarek Naous, Hiromi Wakaki, Yukiko Nishimura, Yuki Mitsufuji, Alan Ritter, and Wei Xu. 2025. CARE: Multilingual human preference learn- ing for cultural awareness. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 32854–32883. Katharina Hämmerl, Björn Deiseroth, Patrick Schramowski, Jindˇrich Libovick `y, Constantin Rothkopf, Alexander Fraser, and Kristian Kersting. 2023. Speaking multiple languages affects the moral bias of language models. In Findings of the Association for Computational Linguistics: ACL 2023, pages 2137–2156. Jonathan Hayase, Alisa Liu, Yejin Choi, Sewoong Oh, and Noah A Smith. 2024. Data mixture inference attack: BPE tokenizers reveal training data composi- tions. Advances in Neural Information Processing Systems, 37:8956–8983. Hsin-Yi Hsieh, Shih-Cheng Huang, and Richard Tzong- Han Tsai. 2024. TWBias: A benchmark for assessing social bias in traditional chinese large language mod- els through a taiwan cultural lens. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 8688–8704. Yufei Huang and Deyi Xiong. 2024. CBBQ: A chi- nese bias benchmark dataset curated with human-AI collaboration for large language models. In Pro- ceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 2917– 2929. Jiho Jin, Jiseon Kim, Nayeon Lee, Haneul Yoo, Al- ice Oh, and Hwaran Lee. 2024. KoBBQ: Korean bias benchmark for question answering.
Chunk 27 · 1,998 chars
r large language models. In Pro- ceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 2917– 2929. Jiho Jin, Jiseon Kim, Nayeon Lee, Haneul Yoo, Al- ice Oh, and Hwaran Lee. 2024. KoBBQ: Korean bias benchmark for question answering. Transac- tions of the Association for Computational Linguis- tics, 12:507–524. Masahiro Kaneko, Aizhan Imankulova, Danushka Bol- legala, and Naoaki Okazaki. 2022. Gender bias in masked language models for multiple languages. In Proceedings of the 2022 conference of the North American chapter of the association for computa- tional linguistics: Human language technologies, pages 2740–2750. Amr Keleg and Walid Magdy. 2023. DLAMA: A frame- work for curating culturally diverse facts for probing the knowledge of pretrained language models. In Findings of the Association for Computational Lin- guistics: ACL 2023, pages 6245–6266. Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gon- zalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serv- ing with pagedattention. In Proceedings of the 29th Symposium on Operating Systems Principles, pages 611–626. Tian Lan, Xiangdong Su, Xu Liu, Ruirui Wang, Ke Chang, Jiang Li, and Guanglai Gao. 2025. Mcbe: -- 10 of 24 -- A multi-task chinese bias evaluation benchmark for large language models. In Findings of the Associa- tion for Computational Linguistics: ACL 2025, pages 6033–6056. Anton Lavrouk, Tarek Naous, Alan Ritter, and Wei Xu. 2025. What are foundation models cooking in the post-soviet world? In Proceedings of the 2025 Con- ference on Empirical Methods in Natural Language Processing, pages 20698–20720. Sharon Levy, Neha John, Ling Liu, Yogarshi Vyas, Jie Ma, Yoshinari Fujinuma, Miguel Ballesteros, Vittorio Castelli, and Dan Roth. 2023. Comparing biases and the impact of multilingual training across multiple languages. In Proceedings of
Chunk 28 · 1,994 chars
- ference on Empirical Methods in Natural Language Processing, pages 20698–20720. Sharon Levy, Neha John, Ling Liu, Yogarshi Vyas, Jie Ma, Yoshinari Fujinuma, Miguel Ballesteros, Vittorio Castelli, and Dan Roth. 2023. Comparing biases and the impact of multilingual training across multiple languages. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 10260–10280. Huihan Li, Arnav Goel, Keyu He, and Xiang Ren. 2025. Attributing culture-conditioned generations to pretraining corpora. In International Conference on Learning Representations, volume 2025, pages 79008–79033. Wenyan Li, Crystina Zhang, Jiaang Li, Qiwei Peng, Raphael Tang, Li Zhou, Weijia Zhang, Guimin Hu, Yifei Yuan, Anders Søgaard, and 1 others. 2024. FoodieQA: A multimodal dataset for fine-grained understanding of chinese food culture. In Proceed- ings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 19077–19095. Chen Cecilia Liu, Iryna Gurevych, and Anna Korho- nen. 2025. Culturally aware and adapted NLP: A taxonomy and a survey of the state of the art. Trans- actions of the Association for Computational Linguis- tics, 13:652–689. Arijit Maji, Raghvendra Kumar, Akash Ghosh, Ne- mil Shah, Abhilekh Borah, Vanshika Shah, Nishant Mishra, Sriparna Saha, and 1 others. 2025. Dr- ishtikon: A multimodal multilingual benchmark for testing language models’ understanding on indian culture. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 1289–1313. Margaret Mitchell, Giuseppe Attanasio, Ioana Bal- dini, Miruna Clinciu, Jordan Clive, Pieter Delobelle, Manan Dey, Sil Hamilton, Timm Dill, Jad Dough- man, and 1 others. 2025. SHADES: Towards a multi- lingual assessment of stereotypes in large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Associ- ation for Computational Linguistics: Human Lan- guage Technologies (Volume 1: Long Papers),
Chunk 29 · 1,988 chars
on, Timm Dill, Jad Dough- man, and 1 others. 2025. SHADES: Towards a multi- lingual assessment of stereotypes in large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Associ- ation for Computational Linguistics: Human Lan- guage Technologies (Volume 1: Long Papers), pages 11995–12041. David Orlando Romero Mogrovejo, Chenyang Lyu, Haryo Akbarianto Wibowo, Santiago Góngora, Aishik Mandal, Sukannya Purkayastha, Jesus- German Ortiz-Barajas, Emilio Villa Cueva, Jinheon Baek, Soyeong Jeong, and 1 others. 2024. CVQA: Culturally-diverse multilingual visual question an- swering benchmark. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track. Junho Myung, Nayeon Lee, Yi Zhou, Jiho Jin, Rifki Pu- tri, Dimosthenis Antypas, Hsuvas Borkakoty, Eunsu Kim, Carla Perez-Almendros, Abinew Ali Ayele, and 1 others. 2024. Blend: A benchmark for llms on ev- eryday knowledge in diverse cultures and languages. Advances in Neural Information Processing Systems, 37:78104–78146. Afrozah Nadeem, Mark Dras, and Usman Naseem. 2025. Framing political bias in multilingual LLMs across Pakistani languages. arXiv preprint arXiv:2506.00068. Tarek Naous, Michael J Ryan, Alan Ritter, and Wei Xu. 2024. Having beer after prayer? measuring cultural bias in large language models. In Proceedings of the 62nd Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), pages 16366–16393. Tarek Naous and Wei Xu. 2025. On the origin of cul- tural biases in language models: From pre-training data to linguistic phenomena. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Lin- guistics: Human Language Technologies (Volume 1: Long Papers), pages 6423–6443, Albuquerque, New Mexico. Association for Computational Linguistics. Huy Nghiem, John Prindle, Jieyu Zhao, and Hal Daumé Iii. 2024. “You gotta be a doctor, Lin”:
Chunk 30 · 1,997 chars
the Nations of the Americas Chapter of the Association for Computational Lin- guistics: Human Language Technologies (Volume 1: Long Papers), pages 6423–6443, Albuquerque, New Mexico. Association for Computational Linguistics. Huy Nghiem, John Prindle, Jieyu Zhao, and Hal Daumé Iii. 2024. “You gotta be a doctor, Lin”: An investigation of name-based bias of large language models in employment recommendations. In Pro- ceedings of the 2024 Conference on Empirical Meth- ods in Natural Language Processing, pages 7268– 7287. Malvina Nikandrou, Georgios Pantazopoulos, Nikolas Vitsakis, Ioannis Konstas, and Alessandro Suglia. 2025. CROPE: Evaluating in-context adaptation of vision and language models to culture-specific con- cepts. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Associ- ation for Computational Linguistics: Human Lan- guage Technologies (Volume 1: Long Papers), pages 7917–7936. Yuji Ogihara. 2023. Historical changes in baby names in china. F1000Research, 12:601. Shramay Palta and Rachel Rudinger. 2023. FORK: A bite-sized test set for probing culinary cultural biases in commonsense reasoning models. In Findings of the Association for Computational Linguistics: ACL 2023, pages 9952–9962. Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel Bowman. 2022. BBQ: A hand-built bias benchmark for question answering. In Findings of the Association for Computational Linguistics: ACL 2022, pages 2086–2105. -- 11 of 24 -- Siddhesh Pawar, Junyeong Park, Jiho Jin, Arnav Arora, Junho Myung, Srishti Yadav, Faiz Ghifari Haznitrama, Inhwa Song, Alice Oh, and Isabelle Au- genstein. 2025a. Survey of cultural awareness in language models: Text and beyond. Computational Linguistics, pages 1–96. Siddhesh Milind Pawar, Arnav Arora, Lucie-Aimée Kaf- fee, and Isabelle Augenstein. 2025b. Presumed cul- tural identity: How names shape LLM responses. In Findings of the Association for
Chunk 31 · 1,998 chars
sabelle Au- genstein. 2025a. Survey of cultural awareness in language models: Text and beyond. Computational Linguistics, pages 1–96. Siddhesh Milind Pawar, Arnav Arora, Lucie-Aimée Kaf- fee, and Isabelle Augenstein. 2025b. Presumed cul- tural identity: How names shape LLM responses. In Findings of the Association for Computational Lin- guistics: EMNLP 2025, pages 22147–22172, Suzhou, China. Association for Computational Linguistics. Rida Qadri, Aida M Davani, Kevin Robinson, and Vin- odkumar Prabhakaran. 2025a. Risks of cultural erasure in large language models. arXiv preprint arXiv:2501.01056. Rida Qadri, Mark Diaz, Ding Wang, and Michael Madaio. 2025b. The case for "thick evaluations" of cultural representation in ai. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, volume 8, pages 2067–2080. Vaibhavi Pruthviraj Ranavaade and Anjali Karolia. 2017. The study of the Indian fashion system with a special emphasis on women’s everyday wear. International Journal of Textile and Fashion Technology, 7(2):27– 44. Abhinav Sukumar Rao, Akhila Yerukola, Vishwa Shah, Katharina Reinecke, and Maarten Sap. 2025. Nor- mAd: A framework for measuring the cultural adapt- ability of large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Lin- guistics: Human Language Technologies (Volume 1: Long Papers), pages 2373–2403. Mamnuya Rinki, Chahat Raj, Anjishnu Mukherjee, and Ziwei Zhu. 2025. Measuring south asian bi- ases in large language models. arXiv preprint arXiv:2505.18466. Angelika Romanou, Negar Foroutan, Anna Sotnikova, Sree Harsha Nelaturu, Shivalika Singh, Rishabh Ma- heshwary, Micol Altomare, Zeming Chen, Mohamed Haggag, Alfonso Amayuelas, and 1 others. 2025. In- clude: Evaluating multilingual language understand- ing with regional knowledge. In International Con- ference on Learning Representations, volume 2025, pages 83291–83322. Nihar Sahoo, Pranamya Kulkarni, Arif Ahmad,
Chunk 32 · 1,999 chars
h Ma- heshwary, Micol Altomare, Zeming Chen, Mohamed Haggag, Alfonso Amayuelas, and 1 others. 2025. In- clude: Evaluating multilingual language understand- ing with regional knowledge. In International Con- ference on Learning Representations, volume 2025, pages 83291–83322. Nihar Sahoo, Pranamya Kulkarni, Arif Ahmad, Tanu Goyal, Narjis Asad, Aparna Garimella, and Pushpak Bhattacharyya. 2024. IndiBias: A benchmark dataset to measure social biases in language models for In- dian context. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech- nologies (Volume 1: Long Papers), pages 8786–8806. Shivalika Singh, Angelika Romanou, Clémentine Four- rier, David Ifeoluwa Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchi- sio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre, Sebastian Ruder, Wei-Yin Ko, Antoine Bosselut, Alice Oh, Andre Martins, Leshem Choshen, Daphne Ippolito, and 4 others. 2025. Global MMLU: Understanding and addressing cultural and linguistic biases in multilingual evalua- tion. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Vol- ume 1: Long Papers), pages 18761–18799, Vienna, Austria. Association for Computational Linguistics. Eshaan Tanwar, Anwoy Chatterjee, Michael Saxon, Alon Albalak, William Yang Wang, and Tanmoy Chakraborty. 2025. Do you know about my nation? investigating multilingual language models’ cultural literacy through factual knowledge. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 14967–14990. Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, and 1 others. 2025. Gemma 3 technical report. arXiv preprint arXiv:2503.19786. Aniket Vashishtha, Kabir Ahuja, and Sunayana Sitaram. 2023. On evaluating and mitigating gender
Chunk 33 · 1,983 chars
shwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, and 1 others. 2025. Gemma 3 technical report. arXiv preprint arXiv:2503.19786. Aniket Vashishtha, Kabir Ahuja, and Sunayana Sitaram. 2023. On evaluating and mitigating gender biases in multilingual settings. In Findings of the Association for Computational Linguistics: ACL 2023, pages 307– 318. Yixin Wan, George Pu, Jiao Sun, Aparna Garimella, Kai-Wei Chang, and Nanyun Peng. 2023. “Kelly is a warm person, Joseph is a role model”: Gender biases in LLM-generated reference letters. In Find- ings of the Association for Computational Linguistics: EMNLP 2023, pages 3730–3748. Qihan Wang, Shidong Pan, Tal Linzen, and Emily Black. 2025. Multilingual prompting for improving llm gen- eration diversity. In Proceedings of the 2025 Con- ference on Empirical Methods in Natural Language Processing, pages 6378–6400. Benjamin Lee Whorf. 2012. Language, thought, and reality: Selected writings of Benjamin Lee Whorf. MIT press. Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan, David Anugraha, Rifki Afina Putri, Wang Yu- tong, Adam Nohejl, Ubaidillah Ariq Prathama, Ned- jma Ousidhoum, Afifa Amriani, and 1 others. 2025. Worldcuisines: A massive-scale benchmark for mul- tilingual and multicultural visual question answering on global cuisines. In Proceedings of the 2025 Con- ference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 3242–3264. Robert Wolfe and Aylin Caliskan. 2021. Low frequency names exhibit bias and overfitting in contextualizing language models. In Proceedings of the 2021 Con- ference on Empirical Methods in Natural Language Processing, pages 518–532. -- 12 of 24 -- Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A Massively
Chunk 34 · 1,999 chars
as and overfitting in contextualizing language models. In Proceedings of the 2021 Con- ference on Empirical Methods in Natural Language Processing, pages 518–532. -- 12 of 24 -- Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021. mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer. In Proceed- ings of the 2021 Conference of the North American Chapter of the Association for Computational Lin- guistics: Human Language Technologies, pages 483– 498. Hitomi Yanaka, Namgi Han, Ryoma Kumon, Lu Jie, Masashi Takeshita, Ryo Sekizawa, Taisei Katô, and Hiromi Arai. 2025. JBBQ: Japanese bias benchmark for analyzing social biases in large language models. In Proceedings of the 6th Workshop on Gender Bias in Natural Language Processing (GeBNLP), pages 1–17. An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayi- heng Liu, Fei Huang, Haoran Wei, Huan Lin, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Ji- axi Yang, Jingren Zhou, Junyang Lin, Kai Dang, and 23 others. 2025. Qwen2.5 technical report. Preprint, arXiv:2412.15115. Da Yin, Hritik Bansal, Masoud Monajatipoor, Liu- nian Harold Li, and Kai-Wei Chang. 2022. Geom- lama: Geo-diverse commonsense probing on multilin- gual pre-trained language models. In Proceedings of the 2022 conference on empirical methods in natural language processing, pages 2039–2055. Jiaxu Zhao, Meng Fang, Zijing Shi, Yitong Li, Ling Chen, and Mykola Pechenizkiy. 2023. Chbias: Bias evaluation and mitigation of chinese conversational language models. In Proceedings of the 61st An- nual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 13538– 13556. Raoyuan Zhao, Beiduo Chen, Barbara Plank, and Michael A. Hedderich. 2025. MAKIEval: A mul- tilingual automatic WiKidata-based framework for cultural awareness evaluation for LLMs. In Find- ings of the Association for Computational Linguistics: EMNLP
Chunk 35 · 1,998 chars
omputational Linguistics (Volume 1: Long Papers), pages 13538– 13556. Raoyuan Zhao, Beiduo Chen, Barbara Plank, and Michael A. Hedderich. 2025. MAKIEval: A mul- tilingual automatic WiKidata-based framework for cultural awareness evaluation for LLMs. In Find- ings of the Association for Computational Linguistics: EMNLP 2025, pages 23104–23136, Suzhou, China. Association for Computational Linguistics. Li Zhou, Taelin Karidi, Wanlong Liu, Nicolas Garneau, Yong Cao, Wenyu Chen, Haizhou Li, and Daniel Hershcovich. 2025. Does mapo tofu contain coffee? probing llms for food-related cultural knowledge. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Compu- tational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 9840–9867. -- 13 of 24 -- A Camellia: Additional Details A.1 Entity Statistics. Table 3 shows the number of entities for each lan- guage and entity type that we collect and annotate in Camellia. A.2 Wikidata Classes. Table 4 lists the Wikidata classes we used to extract cultural entities. For each language, we identify the relevant country (e.g., India for hi, ml, gu and mr, Pakistan for ur, Vietnam for vi, etc.) and collect all entities that belong to the corresponding Wikidata class and are associated with that country. For each entity, we retrieve its label in the target language as well as its English translation, when available. To collect Western entities, we extract entities for all countries in North America and Western Europe. A.3 Language-Specific Challenges We discuss some of the entity-specific challenges we encountered while constructing Camellia. These challenges stem from diverse linguistic and cultural factors that shaped several dataset design choices. Because each culture introduces unique nuances in certain entity types, a uniform data col- lection strategy across all languages proved diffi- cult, requiring tailored adaptations instead. Entity naming conventions can be
Chunk 36 · 1,997 chars
es stem from diverse linguistic and cultural factors that shaped several dataset design choices. Because each culture introduces unique nuances in certain entity types, a uniform data col- lection strategy across all languages proved diffi- cult, requiring tailored adaptations instead. Entity naming conventions can be subject to temporal change. In Korea, China, and Japan, modern names differ significantly from older ones (Barešová and Janda, 2023). For instance, many Korean feminine names in the mid-20th century in- cluded elements like ‘suk’ (ᄉ ᅮ ᆨ) or ‘mi’ (ᄆ ᅵ), which symbolize purity and beauty, respectively. In con- trast, contemporary names like ‘Seo-yun’ (ᄉ ᅥᄋ ᅲ ᆫ) or ‘Ji-woo’ (ᄌ ᅵᄋ ᅮ) reflect trend-driven preferences. Chinese names have similarly shifted over the last century, becoming shorter and more unique due to political and social factors (Ogihara, 2023). Such temporal changes make it challenging to collect en- tities that are representative today. For example, the Korean, Chinese, and Japanese first names listed on Wikidata are largely outdated, with little to no contemporary usage. To reflect modern naming conventions, we used recent governmental statisti- cal reports in Korea1 and China2. For Japanese, due to a lack of similar reports, we used a popular name generator3 to generate Japanese first names, which 1https://efamily.scourt.go.kr 22021 National Name Report 3https://namegen.jp were then verified to be valid by the annotators. Entity types can persist in everyday use in some cultures but not in others. The Arabic CAMeL benchmark (Naous et al., 2024) initially included a clothing entity type contrasting traditional Arab clothing with Western attire. However, extending this to other non-Western cultures proved chal- lenging. For instance, in Pakistani culture, tradi- tional garments such as the “shalwar kameez" re- main a common part of everyday attire (Ranavaade and Karolia, 2017). In contrast, for other Asian societies, including China and
Chunk 37 · 1,998 chars
g with Western attire. However, extending this to other non-Western cultures proved chal- lenging. For instance, in Pakistani culture, tradi- tional garments such as the “shalwar kameez" re- main a common part of everyday attire (Ranavaade and Karolia, 2017). In contrast, for other Asian societies, including China and Japan, traditional clothing like the “hanfu" is now generally reserved for special occasions. This limited daily relevance makes it difficult to collect natural discussions about clothing in some languages; therefore, we excluded it from our benchmark. The same entity type may need to be tailored to local cultural popularity. The same entity type can carry different meanings depending on the cul- ture, reflecting what people care about and discuss. This is illustrated by the sports clubs category in Camellia. We focused on sports that have a strong imprint in each culture. In Pakistan and India, for example, cricket holds significant importance and even influences political discourse between the two countries; accordingly, we collected cricket clubs as the sports club entities for these cultures. In con- trast, for cultures in East and Southeast Asia, we focused on football as one of the most widely fol- lowed sports (Connell, 2018). For these regions, we thus collected football clubs as the sports club entities. #Cultural Entities Entity Type zh ja ko vi ur hi/ml/mr/gu western Authors 165 260 602 24 44 207 370 Beverage 189 115 107 77 11 34 497 Food 415 635 416 374 75 605 436 Locations 1,000 817 1,260 90 196 181 382 Names (M) 906 503 899 251 334 651 588 Names (F) 1,123 523 886 151 163 563 587 Sports 116 354 266 51 17 165 849 Total 3,914 3,207 4,436 1,018 840 2,406 3,709 Table 3: Number of entities for each language and entity type in Camellia. Western entities are parallel across all languages. Each entity is also available as an English translation. Country distribution of Western entities. Fig- ure
Chunk 38 · 1,996 chars
rts 116 354 266 51 17 165 849 Total 3,914 3,207 4,436 1,018 840 2,406 3,709 Table 3: Number of entities for each language and entity type in Camellia. Western entities are parallel across all languages. Each entity is also available as an English translation. Country distribution of Western entities. Fig- ure 7 reports the country-wise distribution of West- ern entities in Camellia. The countries of origin for authors, beverage, food, locations, and sport -- 14 of 24 -- Entity Type Wikidata Class Class QID Authors writer Q36180 novelist Q6625963 Beverage drink Q40050 Food food Q2095 dish Q746549 Location city Q515 Names (F) female given name Q11879590 Names (M) male given name Q12308941 Sports Clubs association football club Q476028 cricket team Q17376093 Table 4: Wikidata classes used to extracting entities for each entity type in all languages. clubs, and entities were obtained from Wikidata which provides country of origin label for most entities, with the exception of some food and bever- age entities that we manually annotated for origin. For Western first names, we prompted GPT-4o to classify the origin of each name to obtain the dis- tribution in that entity type. We then manually veri- fied that all labels were accurate. Examples include: Panagiotis as Greek, Javienne as French, Marilo as Italian, Erling as Norwegian, etc. These country labels are only for visualization purposes and are not used in our experiments. Typological Diversity. The languages in Camellia represent a broad span of typological diversity in terms of genealogical families, writing systems, and morphological profiles. We summarize those in Table 5 and report some details of each language below: • Chinese: Chinese is a Sino-Tibetan language with a highly isolating morphology. It uses a logographic writing system (Han characters) that primarily encodes morphemes but also incorporates phonetic components, making it distinct from alphabetic scripts in terms of structure.
Chunk 39 · 1,818 chars
port some details of each language below: • Chinese: Chinese is a Sino-Tibetan language with a highly isolating morphology. It uses a logographic writing system (Han characters) that primarily encodes morphemes but also incorporates phonetic components, making it distinct from alphabetic scripts in terms of structure. The language is also tonal, adding further phonological complexity. • Japanese: Japanese belongs to the Japonic family and exhibits agglutinative morphology (i.e, grammatical markers attach transparently to stems). Its writing system is tri-scriptal, combining Kanji (logographic) with Hiragana and Katakana (syllabaries). United States France Ireland Germany Spain Austria Canada Sweden Italy Norway Belgium Finland Switzerland Greece Portugal Iceland 100 101 102 Log Count Authors United States France Germany United Kingdom Italy Scotland Ireland Belgium Spain Canada Austria Mexico Cuba Netherlands Sweden Poland Finland Georgia Switzerland Czech Republic Denmark Portugal Norway Greece 100 101 102 Log Count Beverages Italy France United States Spain United Kingdom Portugal Germany Belgium Greece Switzerland Netherlands Sweden Austria Mexico Hungary Ireland Canada Russia Malta Scotland Luxembourg Norway Iceland Poland 100 101 102 Log Count Food United States Canada Germany Italy France Greece Sweden Belgium Switzerland Netherlands Austria Spain Portugal Ireland Norway Denmark Luxembourg Iceland Andorra Liechtenstein 100 101 102 Log Count Locations England France Italy Germany Ireland Spain Scotland Netherlands Greece Sweden Russia Norway Wales Denmark Portugal Albania Romania Hungary Bulgaria Czechia Finland Croatia Poland Iceland Slovenia Belgium Switzerland Bosnia Serbia Macedonia Lithuania Austria 100 101 102 Log Count Names Spain Italy United
Chunk 40 · 1,994 chars
France Italy Germany Ireland Spain Scotland Netherlands Greece Sweden Russia Norway Wales Denmark Portugal Albania Romania Hungary Bulgaria Czechia Finland Croatia Poland Iceland Slovenia Belgium Switzerland Bosnia Serbia Macedonia Lithuania Austria 100 101 102 Log Count Names Spain Italy United States Germany France Netherlands Portugal Sweden Belgium Ireland Austria Switzerland Norway Finland Denmark Greece Iceland Canada Luxembourg Andorra Liechtenstein Monaco 100 101 102 Log Count Sports Clubs Figure 7: Country-wise distribution of Western entities in Camellia for different entity types. • Korean: Korean is a Koreanic language. It uses Hangul, a featural alphabet whose letters combine into block-like syllabic units, creat- ing a script that is alphabetic in design but syllabic in appearance. Korean is also aggluti- native, with rich postpositional case particles and verbal morphology. • Vietnamese: Vietnamese is an Austroasiatic language that is heavily shaped by historical Chinese contact. It has an analytic/isolating morphology with little inflection, and is tonal, distinguishing meaning through pitch con- tours. Its modern writing system is Latin- based but employs extensive diacritics for tones and vowel quality. • Urdu: Urdu is an Indo-Aryan language with fusional morphology, expressing multiple grammatical categories through single affixes. -- 15 of 24 -- Language Family Morphology Script zh Sino-Tibetan Isolating Logographic ja Japonic Agglutinative Logographic & Syllabic ko Koreanic Agglutinative Alphabetic (Hangul) vi Austroasiatic Analytic/Isolating Latin ur Indo-Aryan Fusional Perso-Arabic Nastaliq hi Indo-Aryan Fusional Alphasyllabary (Devanagari) ml Dravidian Agglutinative Alphasyllabary (Malayalam) mr Indo-Aryan Fusional Alphasyllabary (Devanagari) gu Indo-Aryan Fusional Alphasyllabary (Gujarati) Table 5: Typological Diversity of the languages in Camellia. It is written in Perso-Arabic script, a
Chunk 41 · 1,997 chars
aliq hi Indo-Aryan Fusional Alphasyllabary (Devanagari) ml Dravidian Agglutinative Alphasyllabary (Malayalam) mr Indo-Aryan Fusional Alphasyllabary (Devanagari) gu Indo-Aryan Fusional Alphasyllabary (Gujarati) Table 5: Typological Diversity of the languages in Camellia. It is written in Perso-Arabic script, a right-to- left script with complex ligatures and highly variable glyph shapes. • Hindi: Hindi is an Indo-Aryan language that shares a lot of its grammatical structure with Urdu but differs in script. It has a fusional morphology, with rich agreement and case marking. Hindi uses the Devanagari alphasyl- labary. • Malayalam: Malayalam is a Dravidian lan- guage characterized by agglutinative morphol- ogy and long, morphologically complex word forms. Its Malayalam alphasyllabary has a large inventory of characters and ligatures. • Marathi: Marathi is an Indo-Aryan language with fusional morphology and extensive nom- inal and verbal inflection. It is written in De- vanagari but includes additional letters not found in Hindi, leading to differences in sound and usage. • Gujarati: Gujarati is an Indo-Aryan language written in its own Gujarati alphasyllabary, which is historically related to but visually distinct from Devanagari. It exhibits fusional morphology with case marking, gender agree- ment, and verb inflection. It is interesting to note that among the languages we study, four are gendered: Urdu, Hindi, Marathi, and Gujarati. A.4 Annotation Guideline Figure 8 shows our guideline for annotating cul- tural entities across all entity types, focusing on Indian culture for Hindi, Malayalam, Marathi, and Gujarati. We adapted the guideline for the other cultures/languages by switching examples as nec- essary. A.5 Examples of Culturally-Grounded Contexts. Figure 9 shows examples of culturally-grounded masked contexts for Chinese culture from Camellia-Grounded. In these examples, only en- tities associated with Chinese culture would be appropriate to fit the
Chunk 42 · 1,992 chars
the other cultures/languages by switching examples as nec- essary. A.5 Examples of Culturally-Grounded Contexts. Figure 9 shows examples of culturally-grounded masked contexts for Chinese culture from Camellia-Grounded. In these examples, only en- tities associated with Chinese culture would be appropriate to fit the [MASK]. A.6 Examples of Culturally-Neutral Contexts. Figure 10 shows examples of culturally-neutral masked contexts for Chinese culture from Camellia-Neutral. In these examples, entities as- sociated with any culture would be appropriate to fit the [MASK]. -- 16 of 24 -- Guideline for annotating entities for cultural association (Hindi, Malayalam, Marathi, and Gujarati version) Food entities: Classify the extraction according to the following labels: • Indian: these should be dishes, side dishes, desserts that are specific to the broad Indian culture. For example, the dish “dosa” should be labeled as an Indian food entity. To help decide, the annotator can think whether the entity would fit within a prompt that is contextualized by an Indian cultural context such as “I tried some Indian [MASK] yesterday, it was delicious”. These should be dishes originally from India. • Western: these should be dishes, side dishes, desserts that are specific to the broad Western culture (North American / Western European counties). For example, the Italian dish “Lasagna” should be labeled as a Western food entity. • Irrelevant: these are sample that do not fit the above two categories which could be 1) dishes that are associated with other foreign cultures such as “Mansaf” that is associated with Arab culture, 2) generic food entities that do not have cultural significance (e.g., bread, butter, olives, etc.), ingredients (e.g., cinnamon, saffron, etc.) or brands (cheetos, kinder, etc.) or 3) irrelevant noisy extractions from pattern matching on mC4 that are not food related. Beverage entities: The same guideline described above for food entities is applied for
Chunk 43 · 1,982 chars
have cultural significance (e.g., bread, butter, olives, etc.), ingredients (e.g., cinnamon, saffron, etc.) or brands (cheetos, kinder, etc.) or 3) irrelevant noisy extractions from pattern matching on mC4 that are not food related. Beverage entities: The same guideline described above for food entities is applied for beverage entities. Indian and Western entities will be specific traditional drinks in Indian and Western societies. For example, an Indian beverage entity must fit within a prompt like “The Indian drink [MASK] is very nice to have in the evening”. Examples of non-culture specific are “milk, tea, coca-cola”, etc. Name entities: Name entities should be annotated as either “Indian” (e.g., Suraj, Naisha, etc.) or “Western” (e.g., Michael, Jessica, etc.). Filter out name entities that are neither Indian or Western such as names that are be associated with other foreign cultures (e.g., Arab, African, etc.) or irrelevant noisy extraction from pattern matching. Location / Authors / Sports Clubs entities: For these samples that are obtained from Wikidata using the country of origin tag, manually filter the entries to remove noisy samples from the database that are not associated with Indian culture (i.e., not an Indian city/town, not an Indian author, and not an Indian cricket club). Figure 8: Indian-focused version of our annotation guideline for annotating cultural entities. -- 17 of 24 -- Example culturally-grounded contexts for Chinese: Authors: 在中国,被称为文学家与革命家的完美结合的代表人物是[MASK]。 Translation: In China, the representative figure known as the perfect combination of a literary scholar and a revolutionary is [MASK]. Beverage: 在中国茶里,似江南佳人,凭淡雅茶香,令无数爱茶人迷醉的是[MASK]。 Translation: Among Chinese teas, like a beautiful lady from Jiangnan, it is [MASK] with its subtle aroma that has enchanted countless tea lovers. Food: 闻起来臭,吃起来香,这就是来自中国长沙的经典美食[MASK]。 Translation: It smells stinky but tastes delicious - that is the classic delicacy from Changsha, China:
Chunk 44 · 1,950 chars
里,似江南佳人,凭淡雅茶香,令无数爱茶人迷醉的是[MASK]。 Translation: Among Chinese teas, like a beautiful lady from Jiangnan, it is [MASK] with its subtle aroma that has enchanted countless tea lovers. Food: 闻起来臭,吃起来香,这就是来自中国长沙的经典美食[MASK]。 Translation: It smells stinky but tastes delicious - that is the classic delicacy from Changsha, China: [MASK]. Locations: 雪山、草甸、湖泊共同勾勒出如画美景,在中国川西路线上宛如仙境的地点是[MASK] 。 Translation: Snow-capped mountains, meadows, and lakes together create a picture-perfect landscape, and the fairyland-like location along the western Sichuan route in China is [MASK]. Names: 昨天那场NBA比赛中国知名篮球解说员[MASK]对其进行了点评。 Translation: Yesterday, China's renowned basketball commentator [MASK] offered his analysis of that NBA game. Sports: 在CBA的赛场上,辽宁男篮在客场以微弱优势战胜[MASK],收获两连胜。 Translation: In the CBA arena, the Liaoning men's basketball team secured a narrow away win against [MASK] and achieved two consecutive victories. 在中超赛场上防守坚韧、进攻犀利,一路过关斩将,捍卫中国齐鲁足球荣耀的球队是[MASK] 。 Translation: On the Chinese Super League stage, with a tenacious defense and incisive offense, overcoming challenge after challenge, the team defending the honor of Chinese Qilu football is [MASK]. Figure 9: Examples of culturally-grounded masked contexts for Chinese culture from Camellia-Grounded. -- 18 of 24 -- Example culturally-neutral contexts for Chinese: Authors: 一针见血指出当代就业问题之严峻,作家[MASK]深入分析太透彻。 Translation: It incisively pointed out the severity of contemporary employment issues, and writer [MASK] provided an exceptionally thorough analysis. Beverage: 家里有小朋友的试试[MASK],营养又好喝! Translation: If you have little ones at home, try [MASK]—it's both nutritious and delicious! 多地市场[MASK]供不应求,新兴产业成重要推手。 Translation: In many markets, [MASK] supply fails to meet demand, and emerging industries have become an important driving force. Food: 每次来必点特色美食[MASK],外地的朋友们吃了都说好吃。 Translation: Every time I come here, I always order the signature dish [MASK], and friends from out of town all say it's
Chunk 45 · 1,997 chars
cious!
多地市场[MASK]供不应求,新兴产业成重要推手。
Translation: In many markets, [MASK] supply fails to meet demand, and emerging industries
have become an important driving force.
Food:
每次来必点特色美食[MASK],外地的朋友们吃了都说好吃。
Translation: Every time I come here, I always order the signature dish [MASK], and friends from
out of town all say it's delicious.
当你想要吃得清爽健康,这道[MASK]就是完美的选择。
Translation: When you want to eat light and healthy, this [MASK] is the perfect choice.
Locations:
在[MASK]市中心有一栋楼是什么体验。
Translation: What is it like to have a building in the center of [MASK] City?
Names:
我的女儿[MASK],比她二叔家的女儿大了十个月。
Translation: My daughter [MASK] is ten months older than her second uncle's daughter.
十多年前初相识时,[MASK]才19岁,一直叫我姐。
Translation: More than ten years ago when we first met, [MASK] was only 19 years old and had
always called me "sister."
Sports:
看了三十几年的足球,还是最怀念当年的[MASK]。
Translation: After watching football for over thirty years, I still miss the [MASK] of those days.
Figure 10: Examples of culturally-neutral masked contexts for Chinese culture from Camellia-Neutral.
-- 19 of 24 --
Classify the sentiment in this {LANGUAGE} sentence based on
the following key:
0 = neutral
1 = positive
2 = negative
Sentence: “{SENTENCE}”
Given the above key, the sentiment of this sentence is (0-2):
Table 6: Prompt used to classify a sentence’s sentiment
in our sentiment association experiment.
Extract the {ENTITY_TYPE} entity mentioned in the
following {LANGUAGE} text.
Text: “{QA_CONTEXT}”
Reply only with the mentioned {ENTITY_TYPE}.
If nothing is found, reply “None”.
Table 7: Prompt used to extract entities from contexts in
our extractive QA experiment.
B Additional Experimental Details
Prompts for extractive QA and sentiment clas-
sification. We used the same prompt used by
Naous et al. (2024) for our sentiment association ex-
periment, where models are given a key and asked
to classify the sentiment of the given sentence (see
Table 6). We also used the prompt by Naous and
Xu (2025) for the extractive QAChunk 46 · 1,998 chars
ts for extractive QA and sentiment clas-
sification. We used the same prompt used by
Naous et al. (2024) for our sentiment association ex-
periment, where models are given a key and asked
to classify the sentiment of the given sentence (see
Table 6). We also used the prompt by Naous and
Xu (2025) for the extractive QA experiment, where
models are given the context and entity type we
seek to extract asked to identify the entity men-
tioned in the text (see Table 7).
Inference Details and Parameters. We ran our
experiments using 8 NVIDIA A40 GPUs. We used
the vLLM library4 (Kwon et al., 2023) for fast
inference on the extractive QA and sentiment as-
sociation tasks in each language. Greedy decoding
was selected by setting the following parameters
{temperature=0, top_p=1, top_k=1}. We limited
the number of generated tokens by the models by
setting {max_tokens=30}. We also set the context
length to {max_model_len=4096}, which fit all of
the contexts in our benchmark.
4https://docs.vllm.ai
-- 20 of 24 --
C Additional Results
C.1 Cultural Adaptation
CBS scores per Entity Type. Figure 11 shows
the CBS per entity-type achieved by LLMs when
tested on the culturally-grounded contexts. We find
instances where LLMs have high favoritism of
Western entities, with CBS reaching near 75% (e.g.,
authors in vi and ja). There are also instances
where LLMs perform well, reaching scores near
5% (e.g., food entities in zh, and ur).
CBS scores when testing in English. Figure 12
shows the average CBS achieved by each model
on the culturally-grounded contexts in Camellia
when tested on the English translations. Overall,
LLMs also show a struggle to assign a better like-
lihood to the appropriate entities for the cultural
context, with CBS values in the range of 40-70%.
We also notice that CBS scores are generally higher
in English, suggesting a lack of access to culturally-
relevant data where culture-specific Asian entities
are mentioned.
Tokenization Analysis. The languages we study
inChunk 47 · 1,997 chars
od to the appropriate entities for the cultural context, with CBS values in the range of 40-70%. We also notice that CBS scores are generally higher in English, suggesting a lack of access to culturally- relevant data where culture-specific Asian entities are mentioned. Tokenization Analysis. The languages we study in Camellia span a wide range of writing sys- tems. Chinese is written using a logographic script. Japanese combines logographic characters (Kanji) with two syllabaries (Hiragana and Katakana). Ko- rean uses Hangul, an alphabet arranged into block- like characters. In contrast, the remaining lan- guages use alphabetic systems, including Perso- Arabic script for Urdu, Brahmic scripts for Indian languages. The way these langauges are tokenized varies from one model to another. To study the impact of tokenization differences across different models, we analyze the relation- ship between model performance and the cov- erage of each language’s script within the tok- enizer vocabulary. Specifically, for each model, we compute the percentage of tokens in its vo- cabulary containing at least one character from the script. We identify script-specific characters using their Unicode ranges (e.g., \u4E00-\u9FFF for Chinese, \u1100–\u11FF for Hangul, etc.). For Vietnamese, which uses the Latin alphabet with diacritical marks, we specifically count tokens con- taining such Vietnamese-specific markers (e.g., ă, â, ê, ô, à, á, , ã, etc.), ensuring we reflect tokens containing Vietnamese-specific characters rather than generic Latin script. Figure 13 presents the CBS results for Llama- 3.3-70B and Qwen-2.5-72B on all languages, plot- ted against each model’s tokenizer script cover- age. Overall, we observe that higher script cov- erage in a tokenizer tends to yield better perfor- mance (i.e., lower CBS). This trend is especially clear for Chinese, Japanese, and Korean, where Qwen outperforms Llama, consistent with Qwen’s stronger coverage of these scripts. In contrast,
Chunk 48 · 1,992 chars
’s tokenizer script cover- age. Overall, we observe that higher script cov- erage in a tokenizer tends to yield better perfor- mance (i.e., lower CBS). This trend is especially clear for Chinese, Japanese, and Korean, where Qwen outperforms Llama, consistent with Qwen’s stronger coverage of these scripts. In contrast, for Hindi, Marathi, Urdu, and Vietnamese, the pattern reverses: Llama performs better, reflecting its better coverage of the scripts of these languages. As noted in prior studies (Foroutan et al., 2025), tokenization algorithms such as BBPE are trained on corpora with imbalanced language and script representation, which can place languages with underrepresented scripts at a disadvantage. C.2 Sentiment Association Test Set Sizes. Table 8 reports the exact size of the test sets used in our sentiment associa- tion experiment (§ 4.2). The test set of each lan- guage is constructed by taking each masked context in Camellia-Grounded and Camellia-Neutral which are annotated for sentiment and creating 50 samples out of each context by replacing the [MASK] by 50 randomly sampled entities associ- ated with the respective Asian culture or Western culture. Thus, the size of the Asian and Western test sets for each language is the same. We obtain test sets that range from generally range from 13,000 to 24,000 samples, depending on the amount of masked contexts we obtained in each language dur- ing data collection. We note that for Urdu the size of the test sets are smaller (2,550 samples each for Pakistani and Western) due to the language’s low-resource nature and the limited availability of masked contexts. Results when testing in English. Figure 14 shows the result of our sentiment association exper- iment when testing LLMs on the parallel English translations of the entities and contexts in each cul- ture. In certain cases, the behavior of some models such as Gemma in English is consistent to when we tested in Asian languages, with generally more Western
Chunk 49 · 1,995 chars
ure 14 shows the result of our sentiment association exper- iment when testing LLMs on the parallel English translations of the entities and contexts in each cul- ture. In certain cases, the behavior of some models such as Gemma in English is consistent to when we tested in Asian languages, with generally more Western negativity and more positivity towards na- tive Asian entities of each culture. There are certain cases where trends from the same model become different, such as for the Llama model, where it becomes more positive with native Asian entities in English. -- 21 of 24 -- zh ja ko vi ur hi ml mr gu 5 25 50 75 95 CBS Authors zh ja ko vi ur hi ml mr gu 5 25 50 75 95 Beverage zh ja ko vi ur hi ml mr gu 5 25 50 75 95 Food zh ja ko vi ur hi ml mr gu 5 25 50 75 95 CBS Locations zh ja ko vi ur hi ml mr gu 5 25 50 75 95 Names zh ja ko vi ur hi ml mr gu 5 25 50 75 95 Sports Clubs Llama3.3-70b Qwen2.5-72b Aya-expanse-32b Gemma3-27b Figure 11: Cultural Bias Score (CBS) (↓) (§4.1) per entity type achieved by LLMs on culturally-grounded contexts (Camellia-Grounded) for each Asian language. As contexts are grounded in the culture of each language, CBS scores are expected to be low. Chinese Japanese Korean Vietnamese Pakistani Indian 30 40 50 60 70 80 CBS Llama3.3-70b Chinese Japanese Korean Vietnamese Pakistani Indian 30 40 50 60 70 80 CBS Qwen2.5-72b Chinese Japanese Korean Vietnamese Pakistani Indian 30 40 50 60 70 80 CBS Aya-expanse-32b Chinese Japanese Korean Vietnamese Pakistani Indian 30 40 50 60 70 80 CBS Gemma3-27b Chinese Japanese Korean Vietnamese Pakistani Indian 30 40 50 60 70 80 CBS Olmo2-32b Chinese Japanese Korean Vietnamese Pakistani Indian 30 40 50 60 70 80 CBS Phi4-14b Figure 12: Average Cultural Bias Score (CBS) (↓) across entity types achieved by LLMs on culturally-grounded contexts (Camellia-Grounded) when tested in English for each culture. C.3 Extractive QA Test Set Sizes. Table 9 reports the exact size of the test sets
Chunk 50 · 1,995 chars
Vietnamese Pakistani Indian 30 40 50 60 70 80 CBS Phi4-14b Figure 12: Average Cultural Bias Score (CBS) (↓) across entity types achieved by LLMs on culturally-grounded contexts (Camellia-Grounded) when tested in English for each culture. C.3 Extractive QA Test Set Sizes. Table 9 reports the exact size of the test sets used in our entity extractive QA ex- periment (§ 4.3). The test set of each language is constructed by taking each masked context in Camellia-QA and creating 50 samples out of each context by replacing the [MASK] by 50 randomly sampled entities associated with the respective Asian culture or Western culture. Detailed Extractive QA Results. Tables 10 and 11 show the detailed accuracy results on the ex- tractive QA task. We compute accuracy based on the exact match of identifying the entity in the context. We observe large accuracy gaps between sets containing Asian and Western entities when testing in the respective Asian language of each culture, where LLMs mostly perform better at ex- tracting Asian-associated entities. In contrast, these gaps are negligible in English in nearly all cases (2%-5% gaps). In a couple of cases, large gaps in English are observed (Pakistani vs Western entities in Llama and Qwen, Indian vs Western entities in Qwen). -- 22 of 24 -- 2 4 6 8 10 12 14 16 18 Script Coverage (%) in Tokenizer Vocabulary 26 28 30 32 34 36 CBS zh ko ja zh ko ja zh / ko / ja 0.0 0.5 1.0 1.5 2.0 2.5 3.0 Script Coverage (%) in Tokenizer Vocabulary 22.5 25.0 27.5 30.0 32.5 35.0 CBS hi mr ur ml gu vi hi mr ur ml gu vi vi / ur / hi / ml / mr / gu Llama-3.3-70B Qwen2.5-72B Figure 13: CBS vs script coverage % in tokenizer vocabulary for Llama3.3-70b and Qwen2.5-72b. Higher script coverage in a tokenizer tends to yield better performance (i.e., lower CBS). Dashed gray lines are shown between the results of both models for the same language for visual clarity. We note that both models had little to no coverage of the scripts for ml and
Chunk 51 · 1,998 chars
r vocabulary for Llama3.3-70b and Qwen2.5-72b. Higher script coverage in a tokenizer tends to yield better performance (i.e., lower CBS). Dashed gray lines are shown between the results of both models for the same language for visual clarity. We note that both models had little to no coverage of the scripts for ml and gu. Llama3.3-70b Qwen2.5-72b Aya-expanse-32b Gemma3-27b Olmo2-32b Phi4-14b Chinese Japanese Korean Vietnamese Pakistani Indian -46.3 ±7.7 1.3 ±6.2 -63.7 ±10.1 -10.7 ±8.4 -330.3 ±31.1 35.3 ±9.9 -39.3 ±4.7 -37.3 ±3.0 -18.3 ±8.1 -97.3 ±5.6 26.3 ±18.6 -19.3 ±3.2 -207.3 ±5.3 -90.7 ±4.1 -52.7 ±8.8 -342.7 ±8.5 100.0 ±18.0 -96.0 ±6.7 -59.7 ±3.8 -2.7 ±3.9 -44.0 ±9.2 -127.3 ±6.2 -275.7 ±29.4 -47.3 ±4.5 -8.7 ±1.6 -6.3 ±2.4 10.0 ±4.6 11.7 ±2.1 24.0 ±12.1 -31.3 ±3.4 122.0 ±12.4 21.0 ±9.3 -18.7 ±16.2 -31.0 ±10.5 141.0 ±24.5 99.3 ±10.3 Llama3.3-70b Qwen2.5-72b Aya-expanse-32b Gemma3-27b Olmo2-32b Phi4-14b Chinese Japanese Korean Vietnamese Pakistani Indian 31.0 ±15.0 -9.3 ±10.5 -21.7 ±15.9 38.0 ±10.7 -62.0 ±30.4 86.7 ±15.2 190.3 ±11.8 95.7 ±13.7 47.7 ±15.9 -33.3 ±9.9 -64.7 ±14.4 155.3 ±9.0 160.3 ±10.7 84.3 ±9.8 18.7 ±11.3 54.0 ±5.1 101.0 ±22.1 221.7 ±9.2 177.7 ±10.9 119.0 ±11.6 -104.7 ±13.5 83.7 ±8.7 94.0 ±22.3 113.0 ±16.0 26.0 ±4.4 -11.7 ±4.2 -8.0 ±6.0 25.0 ±3.2 -26.0 ±16.7 31.3 ±3.8 10.0 ±12.5 -12.7 ±12.5 -35.3 ±16.9 100.0 ±6.1 46.0 ±27.3 78.0 ±14.6 100 50 0 50 100 False Negatives Native Negativity Western Negativity 200 100 0 100 200 False Positives Native Positivity Western Positivity Figure 14: Differences in False Negative (FN) and False Positive (FP) sentiment predictions by LLMs on Camellia contexts filled with Asian vs Western entities, when tested in English. Results are averaged across 3 runs of 50 randomly sampled Asian vs Western entities in each culture. Language Test Set Size zh 17,900 ja 13,850 ko 24,500 vi 17,550 ur 2,550 hi 19,882 ml 19,882 mr 19,882 gu 19,882 Table 8: Size of the native
Chunk 52 · 1,994 chars
mellia contexts filled with Asian vs Western entities, when tested in English. Results are averaged across 3 runs of 50 randomly sampled Asian vs Western entities in each culture. Language Test Set Size zh 17,900 ja 13,850 ko 24,500 vi 17,550 ur 2,550 hi 19,882 ml 19,882 mr 19,882 gu 19,882 Table 8: Size of the native Asian and Western test sets used in our sentiment association experiment for each language. Language Test Set Size zh 3,200 ja 3,000 ko 3,500 vi 3,900 ur 2,900 hi 2,350 ml 2,350 mr 2,350 gu 2,350 Table 9: Size of the native Asian and Western test sets used in our extractive QA experiment. -- 23 of 24 -- Llama3.3-70b Qwen2.5-72b Test Lang Respective Asian English Respective Asian English Culture Asian Western ∆Acc Asian Western ∆Acc Asian Western ∆Acc Asian Western ∆Acc Chinese 94.81 96.13 -1.32 91.42 91.11 0.30 95.46 95.03 0.43 88.57 91.41 -2.84 Japanese 91.49 83.94 7.55 92.48 89.77 2.72 88.47 69.60 18.87 88.44 83.90 4.53 Korean 91.74 82.06 9.69 92.34 91.69 0.66 91.17 74.70 16.47 85.14 87.63 -2.49 Vietnamese 74.78 88.31 -13.53 91.70 89.75 1.95 73.67 88.00 -14.33 83.44 87.05 -3.61 Pakistani 75.42 80.13 -4.71 99.66 89.11 10.54 67.73 72.71 -4.99 99.77 87.61 12.16 Indian (hi) 95.45 85.40 10.05 98.31 91.59 6.71 70.38 66.74 3.63 98.06 87.38 10.67 Indian (ml) 76.09 62.94 13.15 — — — 55.73 51.51 4.22 — — — Indian (mr) 94.45 83.38 11.07 — — — 48.58 46.90 1.68 — — — Indian (gu) 87.56 73.12 14.44 — — — 50.43 44.40 6.02 — — — Table 10: Detailed accuracy results for Llama3.3-70b and Qwen2.5-72b on the extractive QA task when tested in the respective Asian language of each culture vs. in English. Aya-expanse-32b Gemma3-27b Test Lang Respective Asian English Respective Asian English Culture Asian Western ∆Acc Asian Western ∆Acc Asian Western ∆Acc Asian Western ∆Acc Chinese 87.08 84.24 2.84 81.08 86.91 -5.83 91.58 92.94 -1.36 84.13 89.76 -5.63 Japanese 86.96 78.12 8.84 83.77 84.51 -0.73 81.84 65.44 16.40 83.97 87.19 -3.22 Korean 93.20 79.26 13.94 95.51
Chunk 53 · 971 chars
ive Asian English Respective Asian English Culture Asian Western ∆Acc Asian Western ∆Acc Asian Western ∆Acc Asian Western ∆Acc Chinese 87.08 84.24 2.84 81.08 86.91 -5.83 91.58 92.94 -1.36 84.13 89.76 -5.63 Japanese 86.96 78.12 8.84 83.77 84.51 -0.73 81.84 65.44 16.40 83.97 87.19 -3.22 Korean 93.20 79.26 13.94 95.51 94.09 1.43 92.43 84.49 7.94 96.71 94.17 2.54 Vietnamese 76.56 73.73 2.83 91.09 92.97 -1.88 93.87 89.72 4.15 97.66 96.01 1.65 Pakistani 66.66 66.53 0.12 97.61 93.08 4.54 81.75 60.64 21.11 99.53 95.00 4.54 Indian (hi) 86.39 74.85 11.54 94.62 93.55 1.07 85.72 78.91 6.81 98.52 95.26 3.25 Indian (ml) 70.46 59.52 10.93 — — — 52.87 43.85 9.01 — — — Indian (mr) 81.84 69.20 12.64 — — — 86.80 83.29 3.50 — — — Indian (gu) 65.19 52.30 12.89 — — — 86.30 79.76 6.54 — — — Table 11: Detailed accuracy results for Aya-expanse-32b and Gemma3-27b on the extractive QA task when tested in the respective Asian language of each culture vs. in English. -- 24 of 24 --