[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-216539-en":3,"doc-seo-216539-105":30,"detail-sidebar-cat-0-en-105":97},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},216539,962088006270,"eBook King","https://ap-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","BLUEX v2 - Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams","Large Language Models (LLMs) excel across many tasks, yet rigorous evaluation in Portuguese remains limited, especially for open-ended, discursive settings that require deeper reasoning and generation. BLUEX v2 extends the original BLUEX benchmark to Brazil’s second-phase university entrance exams, covering UNICAMP (Comvest) and USP (Fuvest) for 2022–2025. The dataset includes 395 questions (919 graded subquestions) with image-linked captions, annotated answers, LLM-generated rubric criteria, and six cognitive tags. Twenty-one state-of-the-art LLMs are assessed via an LLM-as-a-judge protocol, revealing a 4.92-point performance spread and identifying mathematical reasoning and image understanding as hardest.","arXiv :2606 .22723v2 [ cs .CL] 30 Jun 2026  \nBLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance  \nExams  \nJoão Guilherme Alves Santos 1 ,2[0000 −0001 −5307 −5338](􀀌), Giovana Kerche Bonás 1 ,2 ,3[0009 −0001 −9460 −8353], Thiago Laitz 1 ,2 ,3[0000 −0001 −7205 −2094], Thales Sales Almeida 1 ,2 ,3[0009 −0006 −9568 −9331], and Helio Pedrini 1[0000 −0003 −0125 −630X]  \n1 University of Campinas (UNICAMP), Campinas-SP, Brazil  \n2 Tropic AI, Campinas-SP, Brazil  \n3 Maritaca AI, Campinas-SP, Brazil  \n[j199624@dac.unicamp.br](j199624@dac.unicamp.br)  \nAbstract. Although Large Language Models (LLMs) excel in many tasks, their assessment in Portuguese has received less attention, particularly for open-ended, discursive tasks that demand deeper reasoning and generation capabilities. While the original BLUEX benchmark addressed the scarcity of Portuguese evaluation datasets through multiplechoice questions from Brazilian university entrance exams, it did not cover the more challenging second-phase examinations, which require free-form written responses. In this work, we introduce BLUEX v2, a benchmark derived from the second-phase entrance exams of Brazil’s two leading universities: UNICAMP (Comvest) and USP (Fuvest), spanning exam years 2022–2025 . Our dataset comprises 395 questions unfolding into 919 graded subquestions, with 55 .7% of questions containing associated images (represented as context-aware captions during inference to enable evaluation across both vision-capable and text-only models) . Each question is annotated with subject area, official reference answers, LLM-generated rubric criteria, and six cognitive capability tags. We evaluate 21 state-of-the-art LLMs using an LLM-as-ajudge protocol. Results reveal a 4.92-point performance spread across models (4.18–9.10 on a 0–10 scale), with Mathematical Reasoning and Image Understanding emerging as the hardest capability dimensions.  \nThe evaluation code, model outputs, and dataset are publicly available  \nat [https://github.com/TropicAI-Research/BLUEXv2](https://github.com/TropicAI-Research/BLUEXv2) and on Hugging  \nFace at [https://huggingface.co/datasets/Tropic-AI/BLUEX-v2](https://huggingface.co/datasets/Tropic-AI/BLUEX-v2) .  \nKeywords: LLMs benchmark · Portuguese · open-ended evaluation · large language models · university entrance exams · discursive questions  \n1 Introduction  \nThe evaluation of Large Language Models (LLMs) has predominantly relied on benchmarks designed for English, leaving a significant gap in the assessment of  \n2 Santos et al.  \nmodel capabilities for other widely spoken languages. Portuguese, despite being the fifth most spoken language in the world with over 250 million native speakers, remains underrepresented in rigorous LLM evaluation [1] .  \nThe original BLUEX benchmark [1] took an important step in addressing this gap by introducing multiple-choice questions from the first-phase entrance exams of UNICAMP and USP, Brazil’s two most prestigious universities. However, the first phase tests primarily recognition and selection abilities. The second phase, in contrast, requires candidates to produce free-form, discursive answers demonstrating deeper understanding, multi-step reasoning, and the ability to articulate complex ideas in written Portuguese.  \nSecond-phase exams at UNICAMP and USP present characteristics that make them particularly valuable for LLM evaluation:  \n– Open-ended responses: Candidates must generate coherent structured answers, simultaneously testing understanding and generation capabilities.  \n– Multi-step reasoning: Questions frequently require integrating knowledge across domains, performing mathematical derivations, or constructing logical arguments over several steps.  \n– Subject-specific depth: Nine academic subjects are covered at the depth demanded by some of Brazil’s most competitive selection processes.  \n– Structured rubrics: Grading criteria enable nuanced evaluation beyond binary co","cbCaiuG0qYQLOMoo","https://ap.wps.com/l/cbCaiuG0qYQLOMoo","pdf",928632,1,15,"English","en",105,"# Introduction\n## Evaluation gap in Portuguese LLM assessment\n## Why second-phase exams matter for LLM evaluation\n## LLM-as-a-judge protocol and rubric grounding\n## Contributions and dataset characteristics","[{\"question\":\"What makes BLUEX v2 different from the original BLUEX benchmark?\",\"answer\":\"BLUEX v2 targets Brazil’s second-phase entrance exams with free-form, discursive written responses, while the original BLUEX primarily covered first-phase multiple-choice questions.\"},{\"question\":\"What does the BLUEX v2 dataset contain?\",\"answer\":\"It comprises 395 discursive questions that unfold into 919 graded subquestions. Each question includes subject areas, official reference answers, LLM-generated rubric criteria, and six cognitive capability tags, and 55.7% link to images represented as context-aware captions.\"},{\"question\":\"How are the LLMs evaluated in BLUEX v2 and what key results are reported?\",\"answer\":\"Models are evaluated using an LLM-as-a-judge protocol grounded in rubric criteria derived from official expected answers. Results show a 4.92-point performance spread across 21 models, with Mathematical Reasoning and Image Understanding emerging as the hardest dimensions.\"}]","BLUEX v2 - Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams | PDF",1788817873,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":92,"head_meta":94,"extra_data":96,"updated_unix":28},"bluex-v2-benchmarking-llms-on-open-ended-questions-from-brazilian-university-entrance-exams","",{"@graph":36,"@context":91},[37,54,74],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/bluex-v2-benchmarking-llms-on-open-ended-questions-from-brazilian-university-entrance-exams/216539/",4,{"url":52,"name":13,"@type":55,"image":56,"author":61,"headline":13,"publisher":63,"fileFormat":66,"inLanguage":23,"description":14,"dateModified":67,"datePublished":68,"encodingFormat":66,"isAccessibleForFree":69,"interactionStatistic":70},"DigitalDocument",{"url":57,"@type":58,"width":59,"height":60},"https://docshare.wps.com/thumbnails/bluex-v2-benchmarking-llms-on-open-ended-questions-from-brazilian-university-entrance-exams/216539.png","ImageObject",300,407,{"name":9,"@type":62},"Person",{"url":41,"name":64,"@type":65},"DocShare","Organization","application/pdf","2026-09-11","2026-09-07",true,{"@type":71,"interactionType":72,"userInteractionCount":20},"InteractionCounter",{"@type":73},"ViewAction",{"@type":75,"mainEntity":76},"FAQPage",[77,83,87],{"name":78,"@type":79,"acceptedAnswer":80},"What makes BLUEX v2 different from the original BLUEX benchmark?","Question",{"text":81,"@type":82},"BLUEX v2 targets Brazil’s second-phase entrance exams with free-form, discursive written responses, while the original BLUEX primarily covered first-phase multiple-choice questions.","Answer",{"name":84,"@type":79,"acceptedAnswer":85},"What does the BLUEX v2 dataset contain?",{"text":86,"@type":82},"It comprises 395 discursive questions that unfold into 919 graded subquestions. Each question includes subject areas, official reference answers, LLM-generated rubric criteria, and six cognitive capability tags, and 55.7% link to images represented as context-aware captions.",{"name":88,"@type":79,"acceptedAnswer":89},"How are the LLMs evaluated in BLUEX v2 and what key results are reported?",{"text":90,"@type":82},"Models are evaluated using an LLM-as-a-judge protocol grounded in rubric criteria derived from official expected answers. Results show a 4.92-point performance spread across 21 models, with Mathematical Reasoning and Image Understanding emerging as the hardest dimensions.","https://schema.org",{"og:url":52,"og:type":93,"og:title":13,"og:site_name":64,"og:description":14},"article",{"robots":95,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":98},[99,103,107,111,116,121,126,129,134,137,141],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":100,"show_sort_weight":101,"slug":102},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},"Exam",70,"exam",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},5,"Comic",60,"comic",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},6,"Technology",50,"technology",{"id":122,"doc_module":4,"doc_module_name":46,"category_name":123,"show_sort_weight":124,"slug":125},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":127,"slug":128},30,"research-report",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":132,"slug":133},9,"Religion & Spirituality",20,"religion-spirituality",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":132,"slug":136},"World Cup","world-cup",{"id":138,"doc_module":4,"doc_module_name":46,"category_name":139,"show_sort_weight":138,"slug":140},10,"Lifestyle","lifestyle",{"id":142,"doc_module":4,"doc_module_name":46,"category_name":143,"show_sort_weight":112,"slug":144},19,"General","general"]