[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85178-en":3,"doc-seo-85178-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85178,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Minionese Comprehensive Benchmark and Mechanistic Study of Multilingual LLM Safety","Safety alignment in large language models can break across languages: requests refused in English may receive harmful compliance in non-English and low-resource contexts. Minionese introduces a multilingual jailbreak benchmark covering 18 languages, four resource tiers, and four perturbation types, paired with geometric mechanistic analysis. Attack types yield distinct vulnerability profiles, and multilingual low-resource jailbreaks succeed via misaligned subspaces that insufficiently project onto refusal directions, leaving refusal mechanisms untriggered. English-only evaluations are therefore inadequate, requiring script, perturbation, and per-language alignment coverage.","Minionese: Comprehensive Benchmark and Mechanistic Study of Multilingual LLM Safety  \nChigozirim Ifebi * 1 Brent Kong * 1 Ayushi Mehrotra 1  \nContent Warning: This paper contains harmful data and model-generated content that can be offensive in  \nnature.  \narXiv :2607 . 10 1 12v 1 [ cs .CR] 11 Jul 2026  \nAbstract  \nSafety alignment in large language models remains brittle across languages: prompts reliably refused in English can elicit harmful compliance in non-English and low-resource settings. We introduce Minionese, a multilingual jailbreak benchmark spanning 18 languages, 4 resource tiers, and 4 perturbation types (standard translation, code-switching, transliteration, and translationese), paired with a geometric mechanistic analysis of refusal failure across language tiers. We show that each attack type produces a distinct vulnerability profile: transliteration vulnerability is mediated by script identity, codeswitching maintains effectiveness through the lowest-resource tier, and a sharp safety regime transition between Tiers 2 and 3 is consistent across all models. Mechanistically, low-resource jailbreaks succeed by routing harmful content through a geometrically misaligned subspace that projects insuﬀicientlyonto the refusal directions, leaving the refusal mechanism intact but untriggered. These findings show that English-only safety evaluations are insuﬀicient; they require accounting for script family, perturbation type, and per-language alignment coverage. The benchmark and analysis code is at [https://gith](https://gith)[ub.com/Brentkong/Minionese-Comprehen](ub.com/Brentkong/Minionese-Comprehen)[sive-Benchmark-and-Mechanistic-Study](sive-Benchmark-and-Mechanistic-Study)[-of-Multilingual-LLM-Safety.git](-of-Multilingual-LLM-Safety.git.)[.](-of-Multilingual-LLM-Safety.git.)  \n* Equal contribution 1 California Institute of Technology. Correspondence to: Chigozirim Ifebi  \n\u003C[cifebi@caltech.edu](cifebi@caltech.edu) >, Brent Kong \u003C[bkong@caltech.edu](bkong@caltech.edu) >.  \nSecond Workshop on Technical AI Governance Research (TAIGR) @ ICML 2026, Seoul, South Korea. 2026. Copyright 2026 by the author(s) .  \n1. Introduction  \nLarge language models (LLMs) are increasingly deployed in multilingual settings, where they are expected to remain both helpful and safe across languages, scripts, and user populations. In practice, however, LLMs systematically behave less reliably in non-English and low-resource language contexts due to disparities in training data, evaluation coverage, and representational fidelity (Pava et al., 2025) . The consequence is a language-gap: users in lower-resource language communities receive weaker safety guarantees than users in English, and current evaluation frameworks largely fail to surface this disparity. We argue that closing this gap requires moving beyond aggregate jailbreak rates and toward a mechanistic account of where and why safety enforcement degrades under multilingual distribution shift.  \nMultilingual prompting can substantially amplify jailbreak success rates (Deng et al., 2024; Wang et al. , 2025), yet the failure modes are poorly characterized at the level of internal model behavior. Mechanistic work in English-centric settings has established that refusal can be mediated by a low-dimensional direction in activation space: ablating this direction suppresses refusal on harmful requests, and adding it can induce refusal on benign ones (Arditi et al., 2024) . This finding reframes safety as a geometric property of internal encodings, and opens the possibility of mechanistically auditing and repairing safety failures rather than simply measuring them at the output level.  \nRecent work has also shown that refusal among safetyaligned languages is mediated by a single direction Wang et al. (2025) . They further show that multilingual jailbreaks can persist even given this shared mechanism, because harmful and harmless prompts are often less cleanly separated in non-English representation","cbCair6J7b0IZ6nl","https://ap.wps.com/l/cbCair6J7b0IZ6nl","pdf",2038095,3,1,16,"English","en",105,"# Introduction\n## Multilingual safety gap and motivation\n## Limits of existing evaluations\n## Mechanistic framing and prior refusal geometry\n## Minionese benchmark scope and perturbations\n## Mechanistic findings and vulnerability profiles\n## Implications and contributions","[{\"question\":\"What problem does the Minionese study address in multilingual LLM safety?\",\"answer\":\"It targets brittleness in safety alignment across languages, where English refusals may not hold for non-English and low-resource settings, allowing harmful compliance through multilingual jailbreaks.\"},{\"question\":\"What does the Minionese benchmark include?\",\"answer\":\"It spans 18 languages, four resource tiers, and four perturbation types: standard translation, code-switching, transliteration, and translationese.\"},{\"question\":\"How does the mechanistic analysis explain why jailbreaks succeed?\",\"answer\":\"Low-resource jailbreaks route harmful content through geometrically misaligned, low-rank subspaces that project insufficiently onto refusal directions, yielding a subthreshold activation failure even when harmful representations exist.\"}]",1784201565,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"minionese-comprehensive-benchmark-and-mechanistic-study-of-multilingual-llm-safety","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/minionese-comprehensive-benchmark-and-mechanistic-study-of-multilingual-llm-safety/85178/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the Minionese study address in multilingual LLM safety?","Question",{"text":75,"@type":76},"It targets brittleness in safety alignment across languages, where English refusals may not hold for non-English and low-resource settings, allowing harmful compliance through multilingual jailbreaks.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the Minionese benchmark include?",{"text":80,"@type":76},"It spans 18 languages, four resource tiers, and four perturbation types: standard translation, code-switching, transliteration, and translationese.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the mechanistic analysis explain why jailbreaks succeed?",{"text":84,"@type":76},"Low-resource jailbreaks route harmful content through geometrically misaligned, low-rank subspaces that project insufficiently onto refusal directions, yielding a subthreshold activation failure even when harmful representations exist.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]