[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-210577-en":3,"doc-seo-210577-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},210577,5909890329169,"McQueen","https://ap-avatar.wpscdn.com/davatar_9964176cb1d06d4a9deccf72a44ae3dc",8,"Research & Report","ALVORADA-BENCH - CAN LANGUAGE MODELS SOLVE BRAZILIAN UNIVERSITY ENTRANCE EXAMS? - Abstract","Language models are increasingly used in Brazil, but evaluation remains largely English-centric. This paper introduces Alvorada-Bench 1, a 4,515-question text-only benchmark drawn from five Brazilian university entrance examinations. Twenty models are tested under zero-shot, role-playing, and chain-of-thought prompting, producing 270,900 responses with structured confidence, perceived difficulty, and Bloom-level self-reports. Top systems achieve 94%+ accuracy overall, with notable drops in Mathematics and IME/ITA engineering exams, indicating multi-step reasoning gaps.","arXiv :2508 . 15835v1 [ cs .CL] 19 Aug 2025  \nALVORADA-BENCH: CAN LANGUAGE MODELS SOLVE BRAZILIAN UNIVERSITY ENTRANCE EXAMS?  \nHenrique Godoy  \nInteli  \nSão Paulo, Brazil  \n[henrique.godoy@sou.inteli.edu.br](henrique.godoy@sou.inteli.edu.br)  \nABSTRACT  \nLanguage models are increasingly used in Brazil, but most evaluation remains English-centric. This paper presents Alvorada-Bench 1 , a 4,515-question, text-only benchmark drawn from five Brazilian university entrance examinations. Evaluating twenty models under zero-shot, role-playing, and chain-of-thought prompting, producing 270,900 responses with structured self-reports of confidence, perceived difficulty, and Bloom level. The top models exceed 94% accuracy overall, but accuracy declines on Mathematics and on the engineering oriented IME and ITA exams, indicating persistent weaknesses in multi-step reasoning. Confidence is well calibrated and correlates with perceived difficulty, revealing that models can accurately assess their own certainty capabilities. A cost-accuracy analysis shows that high accuracy is achievable at under $2 per 1K tokens. On ENEM 2024 the topmodel (O3) achieved perfect scores in Languages subject questions while even the weakest system (GPT-4.1 Nano) only underperforms humans in Mathematics. Through exams that distill decades of Brazilian educational priorities and assess millions of students yearly, Alvorada-Bench establishes whether language models can navigate the intersection of language, culture, and reasoning that defines academic readiness in Brazil.  \n1 Introduction  \nLanguage models increasingly mediate critical decisions across diverse applications, from educational assessment to medical diagnosis, yet their evaluation remains predominantly English-centric. As these models expand into global markets serving linguistically and culturally diverse populations, this evaluation gap poses significant risks.  \nCurrent evaluations demonstrate remarkable performance on standardized tests. GPT-4 scores at the 90th percentile on the SAT, passes the Bar Exam in the top 10%, and outperforms 85% of participants in coding contests [2, 5] . However, these benchmarks embed cultural assumptions that limit global applicability. Translation cannot address implicit cultural frameworks, as SAT questions about financial aid assume familiarity with concepts irrelevant in countries with free universities. This cultural specificity produces measurable degradation: performance drops from 70.9% on English educational tasks to 49.7% in Telugu [3], while Chinese outputs exhibit 41% lexical divergence from native usage despite English-like syntactic patterns [4] .  \nThese issues are evident in non-English contexts such as Brazil, where Portuguese serves a population exceeding 220 million, making it the sixth most spoken language worldwide. However, Portuguese remains underrepresented in the benchmarks. Brazilian university entrance exams offer a compelling solution that combines cultural specificity with rigorous standardization. Refined over decades through expert review, statistical validation, and millions of student responses, these exams serve as a natural experiment in knowledge assessment, capturing both cognitive demands and the cultural knowledge expected of educated Brazilians.  \nTo address this gap, this paper introduce Alvorada-Bench: a benchmark comprising 4,515 questions drawn from five Brazilian university entrance examinations—ENEM (Exame Nacional do Ensino Médio), FUVEST (São Paulo),  \n1Data and code available at [https://huggingface.co/datasets/HenriqueGodoy/Alvorada-bench](https://huggingface.co/datasets/HenriqueGodoy/Alvorada-bench) and [https://](https://)[ ](https://)[github.com/herniqeu/Alvorada-bench](github.com/herniqeu/Alvorada-bench)  \nUNICAMP (Campinas), IME (Instituto Militar de Engenharia), and ITA (Instituto Tecnológico de Aeronáutica) spanning from 1981 to 2025 . Using Alvorada-Bench, we conduct a controlled evaluation of 20 models from Op","cbCaiihI8vBM0VC5","https://ap.wps.com/l/cbCaiihI8vBM0VC5","pdf",2488031,1,14,"English","en",105,"# Introduction\n## Motivation and evaluation gap\n## Cultural specificity and benchmark limitations\n## Alvorada-Bench overview\n# Dataset and Methodology\n## The Alvorada-Bench Dataset","[{\"question\":\"What is Alvorada-Bench and what does it contain?\",\"answer\":\"Alvorada-Bench is a text-only benchmark with 4,515 multiple-choice questions drawn from five Brazilian university entrance examinations. The dataset covers years spanning 1981 to 2025 and aligns with disciplinary areas linked to the BNCC.\"},{\"question\":\"How were language models evaluated on Alvorada-Bench?\",\"answer\":\"Twenty models were evaluated using zero-shot, role-playing, and chain-of-thought prompting. Each interaction includes structured self-reports such as confidence, perceived difficulty, and Bloom level.\"},{\"question\":\"What performance patterns appear across subjects and exams?\",\"answer\":\"Overall accuracy can exceed 94%, but accuracy declines on Mathematics and on engineering-oriented IME and ITA exams. This suggests persistent weaknesses in multi-step reasoning despite strong performance elsewhere.\"}]","ALVORADA-BENCH - CAN LANGUAGE MODELS SOLVE BRAZILIAN UNIVERSITY ENTRANCE EXAMS? - Abstract | PDF",1788666355,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"alvorada-bench-can-language-models-solve-brazilian-university-entrance-exams-abstract","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/alvorada-bench-can-language-models-solve-brazilian-university-entrance-exams-abstract/210577/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-06",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is Alvorada-Bench and what does it contain?","Question",{"text":75,"@type":76},"Alvorada-Bench is a text-only benchmark with 4,515 multiple-choice questions drawn from five Brazilian university entrance examinations. The dataset covers years spanning 1981 to 2025 and aligns with disciplinary areas linked to the BNCC.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How were language models evaluated on Alvorada-Bench?",{"text":80,"@type":76},"Twenty models were evaluated using zero-shot, role-playing, and chain-of-thought prompting. Each interaction includes structured self-reports such as confidence, perceived difficulty, and Bloom level.",{"name":82,"@type":73,"acceptedAnswer":83},"What performance patterns appear across subjects and exams?",{"text":84,"@type":76},"Overall accuracy can exceed 94%, but accuracy declines on Mathematics and on engineering-oriented IME and ITA exams. This suggests persistent weaknesses in multi-step reasoning despite strong performance elsewhere.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]