[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82476-en":3,"doc-seo-82476-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82476,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","ALEE Any Language Evaluation of Embeddings via English Centric Minimal Pairs","Text embeddings support semantic similarity and alignment tasks, but evaluation remains limited by static benchmarks, narrow language coverage, domain dependence, overfitting to fixed pairs, and weak representation of low-resource languages. ALEE introduces a dynamic crosslingual framework extending Sentence Smith to sentence and paragraph level. It uses English Abstract Meaning Representations to generate controlled minimal pairs, pairs them with target-language translations, and evaluates 275+ languages across multiple embedding datasets, revealing persistent gaps tied to training prevalence and tokenization fragmentation.","ALEE: Any-Language Evaluation of Embeddings via English-Centric  \nMinimal Pairs  \nAndrianos Michail Stylianos Psychias Michelle Wastl Simon Clematide Rico Sennrich Juri Opitz  \nDepartment of Computational Linguistics  \nUniversity of Zurich  \nandrianos.michail@cl.[uzh.ch](uzh.ch)  \narXiv :2607 .00171v1 [ cs .CL] 30 Jun 2026  \nAbstract  \nText embeddings are standard for semantic similarity tasks, yet their evaluation remains an open challenge. Current benchmarks are static, cover only a limited set of languages, are often domain-specific, susceptible to overfitting, and poorly representative of low-resource languages. To address these limitations, we introduce ALEE, a framework that extends Sentence Smith (Li et al., 2025) to the crosslingual and paragraph level. ALEE uses Abstract Meaning Representations (AMR) to generate English minimal pairs with controlled, fine-grained semantic shifts, which are paired with translations in target languages. This approach enables targeted diagnostics for models in any language with English parallel data. We conduct a large-scale empirical study across a diverse set of embedding models and 275+ languages spanning three parallel datasets. On ALEE, performance varies substantially across languages, text lengths, and linguistic phenomena, exposing persistent gaps in cross-lingual semantic representation that track language prevalence in training resources and subword tokenization. We release ALEE at [https://github.com/Andrian0s/any](https://github.com/Andrian0s/any)lang-embed-eval.  \n1 Introduction  \nSemantic text embeddings are central to modern information retrieval, clustering, and cross-lingual alignment (Reimers and Gurevych, 2019 ; Gao et al., 2021) . However, their evaluation is largely based on static benchmarks with coarse-grained similarity judgments (Muennighoff et al., 2023) . These datasets have three major limitations: they are biased toward high-resource languages, they are vulnerable to data leakage and overfitting because they are fixed, and they are too coarse-grained to distinguish semantic equivalence from lexical overlap.  \nTo address these limitations, we propose a dynamic minimal-pair evaluation setting leverag-  \nRomansh –(Sursilvan variant)  \nIl medi crei che la terapia funcziuna.  \n\n| Parallel English\u003Cbr>The doctor believes the treatment works. |  |  |  |\n| --- | --- | --- | --- |\n|  |  | PARSE |  |\n|  | \u003Cbr>believe-01\u003Cbr>\u003Cbr>-\u003Cbr> :polarity\u003Cbr> |  |  |\n|  |  | GENERATE |  |\n| Edited English\u003Cbr>The doctor doesn’t believe the treatment works. |  |  |  |\n\nFigure 1: Controlled cross-lingual polarity flip in ALEE.  \ning Abstract Meaning Representation (AMR, Banarescu et al., 2013) . AMR is a formal semantic representation of a text’s meaning, explicating semantic phenomena such as entities and their roles, negation, cause, and more—making it suitable for controlled representation and generation tasks (Wein and Opitz, 2024 ; Sadeddine et al., 2024) . Instead of relying on fixed sentence pairs as in previous embedding evaluations, we use AMR to generate challenging sentence examples that are lexically and syntactically close to a source sentence but differ by one controlled semantic operation. Figure 1 illustrates the minimal pair generation with a sentence in Sursilvan, a Romansh variety spoken in the Swiss canton of Graubünden (Moseley, 2010) . The English parallel sentence,“The doctor believes the treatment works” is parsed into an AMR graph, and a :polarity -edge is added to the predicate believe-01 . This yields the sentence “The doctor  \ndoesn’t believe the treatment works,” while the Sursilvan sentence remains unchanged: “Il medi creiche la terapia funcziuna”. Henceforth, we denote the generated sentence also by foil or confounder, as this indicates its functional purpose within our embedding evaluation framework.  \nThis construction allows us to test whether a  \nmodel assigns higher similarity to the original English sentence than to its minimally perturbed foil when","cbCaiinH0w43Hq4l","https://ap.wps.com/l/cbCaiinH0w43Hq4l","pdf",9403991,1,25,"English","en",105,"# Introduction\n## Motivation and limitations of existing benchmarks\n## Dynamic minimal-pair evaluation with AMR\n## Contributions and key findings","[{\"question\":\"What problem does ALEE address in embedding evaluation?\",\"answer\":\"Existing benchmarks are static, cover limited languages, are often domain-specific, prone to data leakage and overfitting, and too coarse to separate true semantic equivalence from lexical overlap. ALEE targets these limitations with a dynamic, crosslingual diagnostic setup.\"},{\"question\":\"How does ALEE generate minimal pairs for different languages?\",\"answer\":\"The framework parses English sentences into AMR graphs, applies controlled semantic operations (e.g., polarity flip) to create English minimal pairs, and pairs the original and foil with translations in target languages.\"},{\"question\":\"What factors most strongly influence ALEE performance across languages?\",\"answer\":\"Reported results show that performance varies by language, text length, and linguistic phenomena. Performance correlates with how prevalent each language is in the training corpus, and with subword tokenization fragmentation.\"}]",1784180756,63,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"alee-any-language-evaluation-of-embeddings-via-english-centric-minimal-pairs","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/alee-any-language-evaluation-of-embeddings-via-english-centric-minimal-pairs/82476/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does ALEE address in embedding evaluation?","Question",{"text":75,"@type":76},"Existing benchmarks are static, cover limited languages, are often domain-specific, prone to data leakage and overfitting, and too coarse to separate true semantic equivalence from lexical overlap. ALEE targets these limitations with a dynamic, crosslingual diagnostic setup.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does ALEE generate minimal pairs for different languages?",{"text":80,"@type":76},"The framework parses English sentences into AMR graphs, applies controlled semantic operations (e.g., polarity flip) to create English minimal pairs, and pairs the original and foil with translations in target languages.",{"name":82,"@type":73,"acceptedAnswer":83},"What factors most strongly influence ALEE performance across languages?",{"text":84,"@type":76},"Reported results show that performance varies by language, text length, and linguistic phenomena. Performance correlates with how prevalent each language is in the training corpus, and with subword tokenization fragmentation.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]