[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84752-en":3,"doc-seo-84752-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84752,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","MTEB-BR: A Text Embedding Benchmark for Brazilian Portuguese","Text embeddings for Portuguese lack a dedicated, native evaluation benchmark, leaving practitioners to rely on translated corpora or scattered multilingual coverage. MTEBBR introduces 22 Brazilian-Portuguese tasks across seven categories—classification, multilabel, pair classification, semantic similarity, clustering, retrieval, and reranking—constructed exclusively from Portuguese data and excluding translations by design. Results evaluate 93 models (73 open-weight, 20 closed APIs) with bootstrap intervals, significance tests, Item Response Theory-based discrimination, and cross-leaderboard correlations.","MTEB-BR: A Text Embedding Benchmark for  \nBrazilian Portuguese  \nTardelli Ronan Coelho Stekel  \nFederal Institute of So Paulo (IFSP)  \nSo Paulo, Brazil  \n[stekel@ifsp.edu.br](stekel@ifsp.edu.br)  \narXiv :2607 .0458 1v2 [ cs .CL] 7 Jul 2026  \nAbstract—Text embeddings for Portuguese have no dedicated benchmark: evaluation rests on translated corpora such as English MS MARCO or on thin multilingual coverage, with native tasks scattered and unconsolidated. We introduce MTEBBR, a benchmark of 22 native Brazilian-Portuguese tasks across seven categories (classification, multilabel classification, pair classification, semantic textual similarity, clustering, retrieval, andreranking), admitting only data created or found in Portuguese and excluding translations by construction. We evaluate 93 models spanning 23M to 27B parameters: 73 open-weight and 20 closed commercial APIs. Alongside the leaderboard we report a statistical layer for every headline comparison: per-task bootstrap confidence intervals, paired-bootstrap significance, a task- and instance-level discrimination analysis (how sharply each task separates models) adapted from Item Response Theory, anda cross-leaderboard correlation. Three findings stand out. The benchmark cleanly separates about a dozen tiers of models, though the top six are statistically too close to order. An openly licensed, self-hostable model reaches that leading tier, so strong Portuguese embedding quality does not require a commercial API. And a model’s rank on the global multilingual leaderboard predicts its Portuguese rank only moderately (Spearman ρ = 0 .75 over 55 shared models; one model ranks 3rd there and 49th here), so a native benchmark measures something the multilingual boards do not. We release every task, our code, and a public leaderboard, so practitioners can choose Portuguese embedding models on native evidence.  \nI. INTRODUCTION  \nBrazilian Portuguese is spoken natively by over 200 million people, yet a practitioner who needs to deploy a sentenceembedding model for it, whether for semantic search, classification, or retrieval-augmented generation, has no comprehensive native benchmark to guide that choice. The dominant evaluation frameworks, MTEB [1] and its multilingual successor MMTEB [2], do include Portuguese tasks, but they area small fraction of MMTEB’s 500-task suite, and the most heavily reported Portuguese retrieval task is mMARCO-PT [3], a machine translation of English MS MARCO. Translation introduces systematic artifacts (translation noise, domain drift, idiomatic flattening) that can mask real differences between models on native text [4] .  \nLanguage-specific MTEB extensions have closed this gap for Chinese [5], Scandinavian [6], French [7], Polish [8], Russian [9], Persian [10], Dutch [11], German [12], Vietnamese [4], and Japanese [13], but no comprehensive native one existed for Portuguese. We close that gap, and add a layer of statistical rigor that complements the established MTEB methodology. We make four contributions:  \n1) A native, non-translated task suite. We curate 22 tasks across seven MTEB categories from existing BrazilianPortuguese resources spanning legal, medical, tax, scientific, encyclopedic, and social-media domains, excluding machine-translated benchmarks (notably mMARCO-PT and mkqa-PT) by construction.  \n2) A large and diverse model panel. We evaluate 93 models: 73 open-weight (23M–27B parameters) and 20 closed commercial APIs, covering the proprietary embedding services alongside the open ecosystem.  \n3) A statistical-rigor layer. For every headline comparison we report per-task bootstrap confidence intervals, pairedbootstrap p-values, Item Response Theory discrimination at the task and instance levels, and a Borda-count robustness check. This layer shows the benchmark cleanly orders 78.7% of all model pairs into about a dozen distinguishable tiers and places the leader above 87 of 92 models, while the top six converge into one unresolved frontier","cbCaih0ZXPyp7AyB","https://ap.wps.com/l/cbCaih0ZXPyp7AyB","pdf",2292768,2,1,16,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"Does performance on multilingual leaderboards reliably predict Portuguese performance?\",\"answer\":\"Agreement is only moderate: the correlation between MTEB-BR and the HuggingFace multilingual leaderboard is Spearman ρ = 0.75 over 55 shared models. One example shows a model ranked 3rd on the multilingual board but 49th on Portuguese.\"}]",1784198044,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"mteb-br-a-text-embedding-benchmark-for-brazilian-portuguese","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/mteb-br-a-text-embedding-benchmark-for-brazilian-portuguese/84752/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"Does performance on multilingual leaderboards reliably predict Portuguese performance?","Question",{"text":75,"@type":76},"Agreement is only moderate: the correlation between MTEB-BR and the HuggingFace multilingual leaderboard is Spearman ρ = 0.75 over 55 shared models. One example shows a model ranked 3rd on the multilingual board but 49th on Portuguese.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,111,114,119,122,126],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":29,"slug":110},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":112,"slug":113},30,"research-report",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":98,"slug":129},19,"General","general"]