[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-210539-en":3,"doc-seo-210539-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},210539,687207024478,"Mia  ","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","From Zero-shot to Self-generated References - Leveraging LLMs for Scoring ENEM Essays","This study investigates the application of Large Language Models (LLMs) to Automated Essay Scoring (AES) in the context of Brazil’s Exame Nacional do Ensino Médio (ENEM). Five state-of-the-art LLMs are evaluated across three prompting scenarios: zero-shot, one-shot using high-score references, and a self-generated reference approach where models create their own ideal reference before scoring. Using the Essay-BR corpus, performance is measured with classification and regression metrics. Results indicate one-shot prompting delivers the best overall metrics, while self-generated references remain a viable option when real references are unavailable.","From Zero-shot to Self-generated References: Leveraging LLMs for Scoring ENEM Essays  \nMatheus Yasuo Ribeiro Utino 1 , Paulo Mann2  \n1 Institute of Mathematics and Computer Science, University of So Paulo  \n2Institute of Computing, Federal University of Rio de Janeiro  \n[matheusutino@usp.br](matheusutino@usp.br) , [paulomannjr@gmail.com](paulomannjr@gmail.com)  \nAbstract. This study investigates the application of Large Language Models (LLMs) to Automated Essay Scoring (AES) in the context of Brazil’s Exame Na cional do Ensino Me´dio (ENEM). We evaluate five state-of-the-art LLMs across three prompting scenarios: zero-shot, one-shot (with high-score references), and a novel self-generated reference approach, where the model generates its own ideal reference before evaluation. Using the Essay-BR corpus, we assess performance using both classification and regression metrics. Results show that one-shot prompting consistently yields the best metrics, while the self-generated reference method presents a viable alternative when no real references are avail able. Our findings highlight the promise of LLMs for educational scoring.  \n1. Introduction  \nThe Exame Nacional do Ensino M e´dio (ENEM) is Brazil’s primary educational assessment, playing a central role in public university admissions and access to government programs such as Sisu, Prouni, and FIES [Inep 2020] . With millions of participants annually, ENEM evaluates a broad range of competencies, including written production, which is graded by human evaluators based on objective criteria. The essay component, in particular, represents a critical stage, as it synthesizes students’ linguistic, argumentative, and socio-cognitive skills.  \nMore than a mere evaluation tool, ENEM functions as a mechanism for social mobility and the democratization of higher education access [Pires 2023] . Consequently, accurately predicting student performance, especially in the essay section, can inform the design of public education policies, the personalization of pedagogical interventions, and the early identification of risks related to dropout or poor academic achievement.  \nWith recent advances in Large Language Models (LLMs), it has become possible to explore novel approaches to complex tasks requiring sophisticated textual comprehension and production, such as Automated Essay Scoring (AES)[Atkinson and Palma 2025] . These models have demonstrated capabilities not only in classifying texts with high accuracy but also in providing structured textual justifications, making them particularly promising for educational applications where interpretability is a fundamental requirement. These outcomes are closely tied to the emergent abilities of LLMs, which enable them to perform complex reasoning and explanation tasks beyond their training objectives [Berti et al. 2025] .  \nWithin this context, the present study investigates the use of LLMs to predict ENEM essay scores, evaluating different usage configurations (zero-shot, one-shot, and  \na novel approach herein termed self-generated reference) . In addition to analyzing the predictive performance of the models, we explore their capacity to simulate human evaluative behavior, generate patterns of textual excellence, and provide formative feedback. Through this work, we aim to contribute to the advancement of AES in Portuguese, emphasizing scalable, transparent, and pedagogically valuable solutions.  \nThe main contributions of this work are:  \n• The proposal of a novel self-generated reference strategy, in which the model creates its own ideal essay based solely on the theme and uses it as a benchmark for evaluation.  \n• A comparative evaluation of five open-source LLMs across three usage configurations (zero-shot, one-shot with real reference, and self-generated reference) in the context of ENEM essay scoring.  \n• The creation and public release of a new synthetic dataset containing LLMgenerated essays, predicted scores, and detailed justifications","cbCaiaCxf1wOy2uw","https://ap.wps.com/l/cbCaiaCxf1wOy2uw","pdf",191890,1,11,"English","en",105,"# Introduction\n## Automated Essay Scoring for ENEM\n## Study Objectives and Contributions\n# Related Work","[{\"question\":\"What prompting strategies are evaluated for ENEM essay scoring?\",\"answer\":\"The study evaluates three strategies: zero-shot, one-shot with high-score references, and a self-generated reference method where the model generates an ideal reference before evaluation.\"},{\"question\":\"How is model performance measured in the study?\",\"answer\":\"Performance is assessed on the Essay-BR corpus using both classification and regression metrics, aligned to the scoring task.\"},{\"question\":\"What does the study conclude about one-shot vs. self-generated references?\",\"answer\":\"One-shot prompting consistently produces the best metrics, while the self-generated reference approach is a viable alternative when real references are not available.\"}]","From Zero-shot to Self-generated References - Leveraging LLMs for Scoring ENEM Essays | PDF",1788665563,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"from-zero-shot-to-self-generated-references-leveraging-llms-for-scoring-enem-essays","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/from-zero-shot-to-self-generated-references-leveraging-llms-for-scoring-enem-essays/210539/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-06",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What prompting strategies are evaluated for ENEM essay scoring?","Question",{"text":75,"@type":76},"The study evaluates three strategies: zero-shot, one-shot with high-score references, and a self-generated reference method where the model generates an ideal reference before evaluation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is model performance measured in the study?",{"text":80,"@type":76},"Performance is assessed on the Essay-BR corpus using both classification and regression metrics, aligned to the scoring task.",{"name":82,"@type":73,"acceptedAnswer":83},"What does the study conclude about one-shot vs. self-generated references?",{"text":84,"@type":76},"One-shot prompting consistently produces the best metrics, while the self-generated reference approach is a viable alternative when real references are not available.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]