[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-210656-en":3,"doc-seo-210656-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},210656,137451211410,"\tCallum ","https://ap-avatar.wpscdn.com/avatar/2000bb0a9246f588df?x-image-process=image/resize,m_fixed,w_180,h_180&k=1786362646172706240",8,"Research & Report","Evaluating Large Language Models through Multidimensional Item Response Theory - A Comprehensive Case Study on ENEM","LLM evaluations on tasks such as high-stakes multidisciplinary tests often rely on raw accuracy, treating easy and hard items equally and overlooking the chance of correct guessing. The study repurposes the official Brazilian three-parameter logistic Item Response Theory (IRT) calibration used by INEP to score ENEM, and applies it to LLM outputs. It fits a four-dimensional 3-PL model aligned with ENEM knowledge domains, revealing domain-specific proficiency gaps hidden by similar accuracies.","Evaluating Large Language Models through Multidimensional Item Response Theory:  \nA Comprehensive Case Study on ENEM  \nLeonardo Taschetto 1 , Renato Fileto 1  \n1 Dept. of Computer Science, Federal Univ. of Santa Catarina, Florianpolis-SC, Brazil  \nAbstract. LLM evaluations on tasks like high-stakes multidisciplinary tests still rely on raw accuracy, a metric that weights easy and difficult questions equally and ignores guessing. To help bridge this methodological gap, we repurpose the official three-parameter logistic Item Response Theory (IRT) calibration that the Brazilian education authority (INEP) uses to score humans on the Exame Nacional do Ensino Me´dio (ENEM), and apply it to LLM responses. We then fit a four-dimensional 3-PL model aligned with ENEM’s knowledge domains.  \nResults show that similar accuracies can mask proficiency gaps exceeding one standard deviation across domains. Mathematics remains the toughest domain for both humans and models, whereas questions on Human Sciences are system atically easier for both.  \n1. Introduction  \nLarge Language Models (LLMs) now reach state-of-the-art (SOTA) performance on highstakes multidisciplinary tests, notably national university admission exams such as the SAT in the United States [OpenAI 2024a], the Gaokao in China [Zong and Qiu 2024], and Brazil’s ENEM [Abonizio et al. 2024] . These tests have been used as benchmarks to compare LLM capabilities. However, LLM performance studies on these exams mostly report a single metric – accuracy – which offers only a coarse view of test-taker proficiency. It weighs easy and hard questions equally, ignores how well each question differentiates examinees who are stronger or weaker in particular abilities, and fails to adjust for the non-zero chance of correctly guessing answers.  \nMeanwhile, in the realm of human testing, where reliable ranking of candidates is critical, exam authorities apply Item Response Theory (IRT) [Baker 2001] to calibrate each question weight on the candidate score according to how sharply it discriminates high-and low-performing examinees. Thereby, it can produce scaled proficiency scores that distinguish candidates who achieve the same number of correct answers, according to the relevance of each question answered correctly and incorrectly for scoring proficiency in particular abilities. IRT also takes into account question difficulty, based on the percentage of a population that answers it correctly, and the probability of guessing.  \nA few recent studies applying IRT to evaluate LLMs on university-admission exams [Zhang et al. 2023, Zong and Qiu 2024] show that IRT provides finer-grained rankings than accuracy alone. However, they do not exploit IRT distinct dimensions to assess specific abilities as we propose in this work. In addition, to the best of our knowledge, no work has applied IRT to evaluate LLMs on the Portuguese-language ENEM. Motivated by this gap, we ask two linked questions. First, using uni-dimensional IRT scoring, how do SOTA LLMs of varying sizes compare with human candidates across the ENEM’s four  \nknowledge domains? Second, once baselines are established, how do multidimensional IRT (MIRT) shift each model’s performance within those domains?  \nThe major contributions of this paper are: (i) replication of ENEM’s official unidimensional IRT models to evaluate LLM performance on this exam on the human scale;  \n(ii) an extended analysis using a multidimensional IRT model aligned with the exam’s official four knowledge domains. Our experimental results show that identical accuracy does not imply identical proficiency across knowledge domains, and illustrate how these differences depend on model architecture.  \n2. The Multidimensional Item Response Theory  \nThe Item Response Theory (IRT) is a family of mathematical models used to measure individuals’ latent abilities, i.e., unobservable characteristics or proficiencies (e.g., numerical reasoning, reading comprehension, scientific knowledg","cbCaitTckRs5Lrno","https://ap.wps.com/l/cbCaitTckRs5Lrno","pdf",510332,1,12,"English","en",105,"# 1. Introduction\n# 2. The Multidimensional Item Response Theory","[{\"question\":\"Why is accuracy an insufficient metric for evaluating LLMs on exams like ENEM?\",\"answer\":\"Accuracy weights easy and difficult questions equally, ignores how well items discriminate between stronger and weaker examinees, and does not adjust for guessing. This can mask proficiency differences across abilities.\"},{\"question\":\"What methodological approach does the paper use to evaluate LLMs?\",\"answer\":\"It repurposes the official three-parameter logistic IRT calibration used by INEP for ENEM human scoring, then fits a four-dimensional 3-PL model aligned with ENEM’s knowledge domains.\"},{\"question\":\"What do the results show about similar accuracies across ENEM domains?\",\"answer\":\"The study finds that identical accuracy can hide proficiency gaps exceeding one standard deviation across domains. It also reports that Mathematics is the toughest domain for both humans and models, while Human Sciences items are systemically easier.\"}]","Evaluating Large Language Models through Multidimensional Item Response Theory - A Comprehensive Case Study on ENEM | PDF",1788667431,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"evaluating-large-language-models-through-multidimensional-item-response-theory-a-comprehensive-case-study-on-enem","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/evaluating-large-language-models-through-multidimensional-item-response-theory-a-comprehensive-case-study-on-enem/210656/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-06",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is accuracy an insufficient metric for evaluating LLMs on exams like ENEM?","Question",{"text":75,"@type":76},"Accuracy weights easy and difficult questions equally, ignores how well items discriminate between stronger and weaker examinees, and does not adjust for guessing. This can mask proficiency differences across abilities.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What methodological approach does the paper use to evaluate LLMs?",{"text":80,"@type":76},"It repurposes the official three-parameter logistic IRT calibration used by INEP for ENEM human scoring, then fits a four-dimensional 3-PL model aligned with ENEM’s knowledge domains.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the results show about similar accuracies across ENEM domains?",{"text":84,"@type":76},"The study finds that identical accuracy can hide proficiency gaps exceeding one standard deviation across domains. It also reports that Mathematics is the toughest domain for both humans and models, while Human Sciences items are systemically easier.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]