[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-210549-en":3,"doc-seo-210549-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},210549,3985741905716,"Kyle","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Evaluating Automated Scoring Models on Official ENEM Essays - Research paper","Automated Essay Scoring systems can reduce the heavy workload of teachers and enable students to practice more through faster feedback cycles. In Brazilian Portuguese, interest is growing in automatic scoring for the standardized ENEM exam, yet available datasets largely contain essays from practice mock exams. This work evaluates official ENEM essays using a new labeled dataset of 157 essays and official scores, analyzes similarity with mock datasets, and tests encoder models and pretraining effects using QWK and F1 metrics, showing an average 0.27 QWK gain from LLM pretraining.","Evaluating Automated Scoring Models on Official ENEM Essays  \nLaís Nuto Rossman and Igor Cataneo Silveira and Denis Deratani Mauá  \nInstitute of Mathematics and Statistics, University of São Paulo, São Paulo, Brazil [laisnuto@gmail.com](laisnuto@gmail.com) , {igorcs,[ddm}@ime.usp.br](ddm}@ime.usp.br)  \nAbstract  \nAutomated Essay Scoring systems can relieve teachers of this laborious task and allow students to practice more frequently due to faster feedback cycles. In Brazilian Portuguese, thereis growing interest in automatic scoring systems for the standardized ENEM exam. However, the only available datasets consist of essays written as practice for the official exam.  \nIn the literature, to the best of our knowledge, there is no work that evaluates official ENEM essays using mock-exam datasets. This work fills that gap by presenting a new labeled dataset composed of 157 essays written for the official ENEM exam. The analysis shows that this dataset shares characteristics similar to existing datasets of mock exam essays. The results also indicate that, for small datasets such as this one, the use of LLMspretrained on mock exams significantly improves the performance of automatic scorers for official ENEM essays, yielding an average gain of 0.27 points in the Quadratic Weighted Kappa metric compared to training solely on official data.  \n1 Introduction  \nAutomated Essay Scoring (AES) was proposed as a method to unburden teachers from the laborintensive task of scoring essays (Page, 1966) . By making the scoring process faster, it also allows students to practice more frequently. Being able to practice more is especially important in contexts where students need to take a standardized exam that requires writing a long text—usually an essay.  \nIn Brazil, the standardized exam Exame Nacional do Ensino Médio (ENEM) is the main entrance evaluation for universities. This exam is divided into two parts: answering four multiplechoice sections and writing an argumentative essay. Thus, writing a proper essay is crucial for securing a place in higher education.  \nThe vast majority of public datasets for AES in Brazilian Portuguese consist of essays submitted  \nto mock exams administered by websites that simulate the ENEM (Amorim and Veloso, 2017 ; Marinho et al., 2021 ; Silveira et al., 2024) . Previous work (Silveira et al., 2024) has noted that experienced graders found that the prompts proposed by these websites do not exactly match the characteristics of the official exam. Moreover, test-takers of mock exams have different incentives than testtakers of the official ENEM exam, and essays are written under different (and unknown) conditions. Mock exams are also subject to selection biases, asthe process of collecting, grading, and publishing essays on such websites is not disclosed. Thus, the validity and usefulness of existing AES systems as practice tools for the ENEM exams have not yet been established.  \nThis validation gap is complicated by the lack of proper representative datasets of real ENEM essays. Although the prompts of the official exam are made public, only a few perfect-scoring essays are disclosed, making it impossible for third parties to train models using official data. Furthermore, it is also impossible to validate that the models trained on the mock essays are actually assigning scores that would be assigned in the official exam.  \nThe present work fills this gap by presenting adataset composed of 157 essays written for the official exam, along with their official scores. In order to create this dataset, three steps were taken: first, creating an online form so that students could voluntarily submit their data; second, transforming the essays from images to text; and finally, verifying the OCR output for each essay.  \nUsing this dataset, four research questions are investigated: (1) How similar are the mock essays to the official ENEM essays? (2) Are models trained on mock test datasets able to grade official essay","cbCaivQYIStLwjSm","https://ap.wps.com/l/cbCaivQYIStLwjSm","pdf",1143878,1,11,"English","en",105,"# Abstract\n# 1 Introduction\n## Automated Essay Scoring and ENEM context\n## Dataset construction and research questions\n## Feature extraction and model evaluation\n## Fine-tuning results and ablation study","[{\"question\":\"Why are automated essay scoring systems important for ENEM practice?\",\"answer\":\"They reduce the labor-intensive task of scoring long argumentative essays and provide faster feedback cycles, which helps students practice more effectively for standardized exams.\"},{\"question\":\"What dataset is introduced in this work and how was it created?\",\"answer\":\"A labeled dataset of 157 essays written for the official ENEM exam is provided with their official scores. Creation involved collecting submissions via an online form, converting essays from images to text, and verifying OCR output for each essay.\"},{\"question\":\"How does pretraining on mock exams affect scoring official ENEM essays?\",\"answer\":\"For small official-dataset settings, models pretrained on simulated/mock essays significantly improve performance. The reported average gain is 0.27 points in Quadratic Weighted Kappa compared with training only on official data.\"}]","Evaluating Automated Scoring Models on Official ENEM Essays - Research paper | PDF",1788665665,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"evaluating-automated-scoring-models-on-official-enem-essays-research-paper","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/evaluating-automated-scoring-models-on-official-enem-essays-research-paper/210549/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-11","2026-09-06",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why are automated essay scoring systems important for ENEM practice?","Question",{"text":76,"@type":77},"They reduce the labor-intensive task of scoring long argumentative essays and provide faster feedback cycles, which helps students practice more effectively for standardized exams.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What dataset is introduced in this work and how was it created?",{"text":81,"@type":77},"A labeled dataset of 157 essays written for the official ENEM exam is provided with their official scores. Creation involved collecting submissions via an online form, converting essays from images to text, and verifying OCR output for each essay.",{"name":83,"@type":74,"acceptedAnswer":84},"How does pretraining on mock exams affect scoring official ENEM essays?",{"text":85,"@type":77},"For small official-dataset settings, models pretrained on simulated/mock essays significantly improve performance. The reported average gain is 0.27 points in Quadratic Weighted Kappa compared with training only on official data.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]