[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-209264-en":3,"doc-seo-209264-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},209264,687207024478,"Mia  ","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Neuro-symbolic Approaches for Rubric-Based Automatic Evaluation of ENEM Essays - Part 1","Trait-specific automated scoring of essays in the Brazilian National Entrance Exam (ENEM) supports both timely classroom feedback and scalable, consistent official evaluation. Current systems mainly use statistical prediction and largely ignore rubric knowledge that guides human graders. This work leverages the official ENEM grader guidelines to build two neuro-symbolic methods: one uses an LLM to generate subcriteria-based explanations that train a predictive model, and the other converts rubric rules into logical inferences supported by gradient-based learning with weak supervision. Experiments on 63 annotated essays show performance comparable to purely statistical baselines with finer-grained, more actionable feedback.","Neuro-symbolic Approaches for Rubric-Based Automatic Evaluation of ENEM Essays  \nIgor Cataneo Silveira and Denis Deratani Mauá  \nInstitute of Mathematics and Statistics, University of São Paulo, São Paulo, Brazil  \n{igorcs,[ddm}@ime.usp.br](ddm}@ime.usp.br)  \nAbstract  \nTrait-specific automated scoring of essays written for the standardized Brazilian National Entrance Exam (ENEM) has received significant attention in recent years. The task is both im  \nportant in a classroom setting, to provide timely and personalized learning feedback, and in the official exam, to make the scoring process more scalable and consistent. The state-of-the-art systems approach the task as a purely statistical predictive task, ignoring the knowledge provided to human graders and test takers in the form of rubrics and guidelines. Aiming to produce more interpretable and informative formative feedback in this work, we leverage the official ENEM Grader’s handbook and develop two neuro-symbolic approaches to trait-specific essay scoring. The first approach uses a Large Language Model (GPT4o) to write an evaluative explanation of the essay score according to the subcriteria described in the guidelines; the explanation is then fed into a statistical model to effectively predict the score; the good performance of the scoring validates the quality of the explanations. The second approach formalizes the Guideline grading rubrics as logical rules that derive the essay score as a function of subcriteria, mimicking the recommended human grader’s scoring approach. In order to provide weak supervision in training and to evaluate the quality of the model, we build a dataset of 63 essays annotated with their subcriteria by two expert human graders. Our empirical results suggest that both approaches perform on par with purely statistical methods while providing more helpful and fine-grained feedback.  \n1 Introduction  \nAutomated Essay Evaluation (AEE) seeks to ease the burden on teachers and scale up personalized feedback for students (Page, 1966 ; Shermis and Burstein, 2013) . Most existing AEE systems are developed for and evaluated purely by their ability to produce essay scores that are aligned with  \nhuman-assigned scores.  \nIn Brazil, the Exame Nacional do Ensino Médio (ENEM) is a standardized exam used by the majority of higher education institutions as part of the admission process. The exam is divided into two parts: the first evaluates subject-matter knowledge through a multiple-choice questionnaire, and the second evaluates critical thinking and writing skills by means of an argumentative essay. Thus, mastery of essay writing is crucial for students who desire to enter higher education. Consequently, being able to simulate its evaluation is of major importance in High School. Being a high-stakes standardized exam implies that the ENEM essay exam has welldefined grading criteria and important societal and economic importance.  \nIt is not surprising that the majority of the literature on AEE for Brazilian Portuguese consists of training machine learning models on essays submitted for mock ENEM exams hosted on online web portals (Amorim and Veloso, 2017 ; Marinho et al., 2022a ; Silveira et al., 2024) . Those works cast essay scoring as a simple prediction task. More recently, Barbosa et al. (2025) have investigated the use of instruction-guided LLM-based approaches that utilize additional materials such as official exam rubrics. Still, their end task consisted in providing a trait-specific score for each essay.  \nIn this work, we leverage the released official ENEM Grader Guidelines 1 to provide more finegrained feedback in the form of explanations and rubric-based scoring. More precisely, the rubrics are stated in the form of high-level logical rules over sub-score criteria for each trait. We develop two neuro-symbolic approaches to score the essays that make use of the official rubrics.  \nThe first approach builds upon the work by Bar-  \n1Available at: [","cbCairfqD7WNJvXT","https://ap.wps.com/l/cbCairfqD7WNJvXT","pdf",229209,1,10,"English","en",105,"# Introduction\n## Automated Essay Evaluation in ENEM\n## Neuro-symbolic approaches using official rubrics\n## Contributions and experimental setup","[{\"question\":\"Why is automated evaluation of ENEM essays important?\",\"answer\":\"ENEM relies on well-defined grading criteria for high-stakes admissions, and students need feedback on argumentative writing. Automated scoring helps scale consistent, timely evaluation and supports personalized learning in high school contexts.\"},{\"question\":\"What is the first neuro-symbolic approach described?\",\"answer\":\"It uses a large language model to produce subcriteria-based evaluative explanations for a given essay score, then trains a statistical model using these explanations to predict the trait-specific score.\"},{\"question\":\"How does the second approach incorporate rubric information?\",\"answer\":\"It formalizes the grader rubrics as logical rules over subcriteria, using statistical models to infer truth values of rubric propositions, trained with a semantic loss and supported by a small subcriteria-labeled dataset.\"}]","Neuro-symbolic Approaches for Rubric-Based Automatic Evaluation of ENEM Essays - Part 1 | PDF",1788611599,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"neuro-symbolic-approaches-for-rubric-based-automatic-evaluation-of-enem-essays-part-1","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/neuro-symbolic-approaches-for-rubric-based-automatic-evaluation-of-enem-essays-part-1/209264/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-11","2026-09-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is automated evaluation of ENEM essays important?","Question",{"text":76,"@type":77},"ENEM relies on well-defined grading criteria for high-stakes admissions, and students need feedback on argumentative writing. Automated scoring helps scale consistent, timely evaluation and supports personalized learning in high school contexts.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is the first neuro-symbolic approach described?",{"text":81,"@type":77},"It uses a large language model to produce subcriteria-based evaluative explanations for a given essay score, then trains a statistical model using these explanations to predict the trait-specific score.",{"name":83,"@type":74,"acceptedAnswer":84},"How does the second approach incorporate rubric information?",{"text":85,"@type":77},"It formalizes the grader rubrics as logical rules over subcriteria, using statistical models to infer truth values of rubric propositions, trained with a semantic loss and supported by a small subcriteria-labeled dataset.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":21,"slug":134},"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]