[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83465-en":3,"doc-seo-83465-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83465,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","SEFORA Student Essays with Feedback Corpus and LLM Feedback Evaluation Framework","Effective writing feedback accelerates student learning, but producing it at scale is costly and labor-intensive. Large language models enable scalable writing support, yet two barriers remain: limited public corpora reflecting how instructors deliver inline feedback in real classrooms, and the lack of dependable ways to measure whether generated feedback matches instructor intent. SEFORA provides a public corpus of 564 drafts with prompt-, rubric-, and revision-linked feedback and 8,240 instructor annotations, while UNIMATCH evaluates LLM feedback via unit-level semantic alignment, yielding interpretable precision, recall, and F1.","SEFORA: Student Essays with Feedback Corpus and LLM Feedback Evaluation Framework  \nShayan Peyghambari Oskoui, Norah Almousa, Zhaoyi Joey Hou, Carolina Gustafson, Gayle Rogers, Raquel Coelho, Diane Litman, Xiang Lorraine Li  \nUniversity of Pittsburgh {shayan.p, [xianglli](xianglli}@pitt.edu)[}](xianglli}@pitt.edu)[@pitt.edu](xianglli}@pitt.edu)  \narXiv :2607 .00274v1 [ cs .CL] 30 Jun 2026  \nAbstract  \nEffective writing feedback is among the strongest drivers of student learning, yet producing it at scale is labor-intensive. LLMs offer a natural path to scaling writing support, but two gaps stand in the way: few public corpora capture how instructors actually deliver feedback in real classrooms, and no reliable method measures whether generated feedback aligns with what an instructor would write. We address both. SEFORA is a public corpus 1 pairing instructor inline feedback with assignment prompts, rubrics, scores, and multi-draft revisions across various college writing genres, comprising 564 drafts and 8,240 instructor annotations. UNIMATCH is a referencebased evaluation framework for open-ended generation: it segments feedback into feedback units, scores their semantic correspondence under instructor-derived criteria, and aligns them via optimal matching to yield interpretable precision, recall, and F 1 . Across 74 experimental configurations spanning multiple LLMs, no setting exceeds 0.4 F 1 . UNIMATCH reveals that models struggle to identify the feedback instructors would prioritize, and performance degrades as models generate more.  \n1 Introduction  \nFeedback plays a vital role in student learning in writing. It helps students correct misunderstandings and refine how they apply knowledge (Ahea et al., 2016 ; Banihashem et al., 2024), and it is consistently identified as one of the strongest influences on learning and achievement (Hattie and Timperley, 2007) . But effective feedback is not onesize-fits-all. Students report that the most useful feedback is specific to their own writing (Lipnevichand Smith, 2009), and effective feedback accounts for the draft’s position within the revision process (Carless and Boud, 2018); poorly targeted feedback can even be detrimental (Kluger and DeNisi, 1996) .  \n1 [https://github.com/ShayanPey/SEFORA](https://github.com/ShayanPey/SEFORA)  \nGenerate feedback UNIMATCH  \n Segmentation  \nFigure 1: Overview of the UNIMATCH evaluation pipeline. For each paragraph, LLM-generated feedback is segmented into units and compared with instructor feedback units. The resulting semantic similarity scores are used to compute an optimal matching between instructor and model feedback units, producing the final evaluation metrics.  \nIn writing instruction, producing such feedback is labor-intensive, and its cost at scale discourages the sustained practice effective instruction requires (Applebee and Langer, 2011 ; Graham, 2019) . This opens a natural opportunity for NLP: systems that generate useful feedback on drafts could help scale writing support.  \nProgress on this problem is constrained by two bottlenecks. First, few public datasets preserve how instructor feedback is actually delivered in coursework (Table 1): multifaceted comments (often addressing several points at once) anchored to specific spans (a paragraph, sentence, or word) and interpretable alongside the assignment prompt, rubric, and revision history. Some resources substitute structured labels (error tags or analytic scores) for formative commentary (Crossley et al., 2024 ; Mathias and Bhattacharyya, 2018 ; Dahlmeier et al., 2013 ; Lee et al., 2015), restrict coverage to a single prompt or narrow genre (Kashefi et al., 2022 ; Zyska et al., 2026), or forgo expert annotation for crowd-or model-generated feedback (Behzad et al., 2024) . Without datasets that capture feedback as instructors deliver it, evaluating whether LLMs pro-  \nduce what an instructor would write remains out of reach.  \nSecond, feedback evaluation is difficult for three","cbCaioMsclRwKiLB","https://ap.wps.com/l/cbCaioMsclRwKiLB","pdf",1531479,2,1,28,"English","en",105,"# Introduction\n## SEFORA: Student Essays with Feedback Corpus\n## UNIMATCH: LLM Feedback Evaluation Framework\n## Feedback Bottlenecks and Evaluation Challenges","[{\"question\":\"What are the main problems addressed by SEFORA and UNIMATCH?\",\"answer\":\"The work targets two gaps: few public datasets capture how instructors provide span-anchored inline feedback in real classroom settings, and there is no reliable reference-based method to judge whether LLM-generated feedback matches what an instructor would write.\"},{\"question\":\"What does the SEFORA corpus include?\",\"answer\":\"SEFORA pairs instructor inline feedback with assignment prompts, rubrics, scores, and multi-draft revisions across multiple college writing genres, containing 564 drafts and 8,240 instructor annotations.\"},{\"question\":\"How does UNIMATCH evaluate LLM feedback?\",\"answer\":\"UNIMATCH segments feedback into feedback units, scores semantic correspondence under instructor-derived criteria, and aligns instructor and model feedback units via optimal matching to compute interpretable precision, recall, and F1.\"}]",1784188153,71,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"sefora-student-essays-with-feedback-corpus-and-llm-feedback-evaluation-framework","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/sefora-student-essays-with-feedback-corpus-and-llm-feedback-evaluation-framework/83465/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What are the main problems addressed by SEFORA and UNIMATCH?","Question",{"text":75,"@type":76},"The work targets two gaps: few public datasets capture how instructors provide span-anchored inline feedback in real classroom settings, and there is no reliable reference-based method to judge whether LLM-generated feedback matches what an instructor would write.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the SEFORA corpus include?",{"text":80,"@type":76},"SEFORA pairs instructor inline feedback with assignment prompts, rubrics, scores, and multi-draft revisions across multiple college writing genres, containing 564 drafts and 8,240 instructor annotations.",{"name":82,"@type":73,"acceptedAnswer":83},"How does UNIMATCH evaluate LLM feedback?",{"text":84,"@type":76},"UNIMATCH segments feedback into feedback units, scores semantic correspondence under instructor-derived criteria, and aligns instructor and model feedback units via optimal matching to compute interpretable precision, recall, and F1.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]