[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-209289-en":3,"doc-seo-209289-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},209289,24189269381491,"Bill Black","https://ap-avatar.wpscdn.com/avatar/160000cf11732dd8392?x-image-process=image/resize,m_fixed,w_180,h_180&k=1788146458752108895",8,"Research & Report","A New Benchmark for Automatic Essay Scoring in Portuguese","Automatic Essay Scoring enables scalable, timely feedback on student writing, yet high-quality resources for Portuguese remain limited, often inaccessible, incomplete in provenance, or inaccurate in ways that reduce system performance. This work introduces a new benchmark for Brazilian Portuguese by collecting publicly available ENEM-simulating essays, releasing both raw and processed data, and assigning expert grades to estimate annotation quality and task difficulty. Comprehensive experiments evaluate state-of-the-art predictors using multiple official-style criteria.","A New Benchmark for Automatic Essay Scoring in Portuguese  \nIgor Cataneo Silveira and André Barbosa and Denis Deratani Mauá Institute of Mathematics and Statistics, University of São Paulo, São Paulo, Brazil {igorcs, aborbosa, [ddm}@ime.usp.br](ddm}@ime.usp.br)  \nAbstract  \nAutomatic Essay Scoring promises to scale up student feedback of written input, considerably improving learning. Resources for Automatic Essay Scoring in Portuguese are however scarce, not publicly available or contain inaccuracies that degrade performance. Moreover, they lack data provenance and a richer annotation and analysis. In this work we mitigate those issues by presenting a new benchmark for the task in Brazilian Portuguese. We accomplish that by downloading a collection of publicly available essays from websites that simulate University Entrance Exams, making both processed and raw data available, having a subset of the essays graded by expert annotators to assess the quality and difficulty of the task, and carrying out an extensive empirical analysis of state-of-the-art predictors considering multiple evaluation criteria.  \n1 Introduction  \nGrading essays is a ubiquitous and crucial task in Education. For the instructor, the task consumes valuable time and effort in both the grading process per se and in training and preparation (especially for junior teachers and assistants or in standardized exams) . For the student, having adequate and timely feedback is essential to correct misunderstandings, encourage reflection, support engagement and maintain trust in the evaluation process.  \nWhile the importance of both scoring and commenting (i.e., providing feedback in written form) has been stressed since Page (1966)’s seminal work, most research and technological developments have focused on the scoring aspect, known as Automatic Essay Scoring (AES) .  \nAES systems are now widespread (Beigman Klebanov and Madnani, 2021); popular standardized exams such as TOEFL, GMAT, GRE and PTE all rely on some form of AES (Attali and Burstein, 2006 ; Beigman Klebanov and Madnani, 2020) . In  \naddition to English, there are AES systems for a large variety of languages such as French (Lemaire and Dessus, 2003), Danish, Finnish (Beigman Klebanov and Madnani, 2020), Chinese (Song et al., 2016), Arabic (Mezher and Omar, 2016) and Japanese (Ishioka and Kameda, 2006), to name a few.  \nAES systems for (Brazilian) Portuguese have been developed by Amorim and Veloso (2017); Fonseca et al. (2018); Marinho et al. (2021) . They are variously based on training Machine Learning models from corpora of human-annotated essays. The data sources are web sites and platforms used by high-school students for practicing for University Admission Exams, where students submit essays in exchange of feedback in the form of scoresand comments. While important, those systems fall short of providing a good benchmark for AES in Portuguese, for the following reasons.  \nThe annotated essays in the work of Amorim and Veloso (2017) were graded using a scale different from the the standardized exam it attempts to simulate, and contains no information about the scoring guidelines used by annotators. This makes it difficulty to enlarge the dataset with new essays and to validate or assess annotations. The very large data used by Fonseca et al. (2018) are proprietary and were not made publicly available. The Essay-Br corpus, used by Marinho et al. (2021), despite being relatively large and accessible, has many shortcomings. First, the HTML sources were not properly parsed to strip out unwanted content, which resulted in having annotator comments appearing in the middle of the text, ill-formed sentences, and artificial artifacts such as blank spacesand noticeable marks where comments appeared in the HTML source. That can artificially boost a machine-learning approach performance by data leakage as well as hurt the system’s performance due to noisy input. Second, there was no analysis of the quality of the","cbCaimYk7XVajn0h","https://ap.wps.com/l/cbCaimYk7XVajn0h","pdf",333907,1,10,"English","en",105,"# Introduction\n## ENEM Essays\n# Related Work\n## Corpus Construction\n# Benchmarking Methods\n## Evaluation Results\n# Final Remarks","[{\"question\":\"What problem does the work address in Portuguese automatic essay scoring resources?\",\"answer\":\"Existing Portuguese AES resources are scarce, not publicly available, or inaccurate, with limited annotation detail and missing data provenance that harms training and evaluation reliability.\"},{\"question\":\"How is the new benchmark for Brazilian Portuguese constructed?\",\"answer\":\"The benchmark is built by downloading essays from websites simulating University Entrance Exams, providing both processed and raw HTML sources, and including a subset graded by expert annotators.\"},{\"question\":\"What evaluation is performed to benchmark automatic essay scoring methods?\",\"answer\":\"The study benchmarks state-of-the-art predictors using standard machine learning methodology and multiple evaluation criteria aligned with official standardized exam practices, also estimating task difficulty via inter-annotator agreement.\"}]","A New Benchmark for Automatic Essay Scoring in Portuguese | PDF",1788612170,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"a-new-benchmark-for-automatic-essay-scoring-in-portuguese","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-new-benchmark-for-automatic-essay-scoring-in-portuguese/209289/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-11","2026-09-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the work address in Portuguese automatic essay scoring resources?","Question",{"text":76,"@type":77},"Existing Portuguese AES resources are scarce, not publicly available, or inaccurate, with limited annotation detail and missing data provenance that harms training and evaluation reliability.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How is the new benchmark for Brazilian Portuguese constructed?",{"text":81,"@type":77},"The benchmark is built by downloading essays from websites simulating University Entrance Exams, providing both processed and raw HTML sources, and including a subset graded by expert annotators.",{"name":83,"@type":74,"acceptedAnswer":84},"What evaluation is performed to benchmark automatic essay scoring methods?",{"text":85,"@type":77},"The study benchmarks state-of-the-art predictors using standard machine learning methodology and multiple evaluation criteria aligned with official standardized exam practices, also estimating task difficulty via inter-annotator agreement.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":21,"slug":134},"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]