[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86339-en":3,"doc-seo-86339-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86339,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","A Durability and Cross-Language Transfer Benchmark for a Validated Teaching-Feedback Classification Protocol","The study evaluates whether a previously validated protocol for classifying open-ended teaching-evaluation comments remains effective as representations evolve and whether it transfers across languages. It extends an earlier Spanish setup using documented annotation guidance, intra-annotator reliability, stratified cross-validation, and a held-out evaluation. Re-executions across three representation generations include frozen transformer embeddings and prompted large language models, with sentiment transferred to English using a 45,000-comment balanced corpus.","A Durability and Cross-Language Transfer Benchmark for a Validated Teaching-Feedback Classification Protocol  \nEsteban U. Vega Barajas  \nUniversidad de Guadalajara  \narXiv :2607 . 1 1873v 1 [ cs .CL] 13 Jul 2026  \nAbstract  \nInstitutions collect far more open-ended teaching-evaluation feedback than they read. A prior study introduced a validated protocol for classifying such comments by thematic category and sentiment, built from a documented annotation guide, an intra-annotator reliability measurement, stratified cross-validation, and a held-out evaluation on a Spanish institutional corpus with a frozen-encoder design. Two questions limit its reuse: whether a protocol fixed to 2019-era frozen embeddings stays competitive as representation methods advance, and whether it transfers to a second language. Were-run it on the original Spanish data across three representation generations, sparse lexical features, frozen transformer embeddings, and prompted large language models, and transfer its sentiment task to English with a balanced 45,000-comment corpus checked against an aspect-labeled education dataset. Treating paired comparisons as descriptive, we find the protocol durable: a 2026 frontier model posts the highest thematic F1 on the hardest Spanish task, yet shows no sentiment advantage over a cheap model and no descriptive separation from it on English, so model choice is a deployment decision, not a property of the method.  \nResumen  \nLas instituciones recogen mucho más texto deevaluación docente del que llegan a leer. Unestudio previo presentó un protocolo validado para clasificar esos comentarios por categoríatemática y sentimiento, construido a partir de una guía de anotación documentada, una medición de fiabilidad intra-anotador, validacióncruzada estratificada y una evaluación sobre un conjunto reservado en un corpus institucional en español. Dos preguntas limitan su reutilización: si un protocolo fijado a representaciones congeladas de la generación de 2019 sigue siendo competitivo a medida que avanzan los métodos de representación, y si se transfiere a un segundo  \nidioma. Lo reejecutamos sobre los datos originales en español a través de tres generaciones de representación, rasgos léxicos dispersos, representaciones congeladas de transformadores y modelos de lenguaje grandes, y transferimossu tarea de sentimiento al inglés con un corpus balanceado de 45,000 comentarios contrastado con un conjunto educativo etiquetado por aspectos. Tratando las comparaciones pareadas como descriptivas, encontramos que el protocolo es duradero: un modelo de frontera de 2026 obtiene el F1 temático más alto en la tareamás difícil en español, pero no muestra ventajaen sentimiento sobre un modelo económico ni separación descriptiva frente a él en inglés, demodo que la elección del modelo es una decisión de despliegue y no una propiedad del método.  \nKeywords: teaching-evaluation feedback; text classification; sentiment analysis; cross-lingual transfer; large language models; educational NLP; reproducible benchmark  \nPalabras clave: evaluación docente; clasificaciónde texto; análisis de sentimiento; transferencia translingüe; modelos de lenguaje grandes; PLNeducativo; benchmark reproducible  \n1 Introduction  \nOpen-ended comments on teaching-evaluation surveys are among the most informative and least-used data that educational institutions hold. A single mid-sized faculty generates tens of thousands of free-text comments per academic cycle; the volume routinely exceeds what staff can read, so institutions fall back on the accompanying numeric ratings and the qualitative signal is archived unread. Natural language processing (NLP) can recover that signal, classifying each comment by what it is about and by its affective stance. For an institution to acton the output, the analytical procedure has to be documented, reliability-checked, and validated. An accuracy number on a leaderboard does not by itself  \nmake the procedure something an insti","cbCaii3cxhwxe25c","https://ap.wps.com/l/cbCaii3cxhwxe25c","pdf",232528,3,1,12,"English","en",105,"# Introduction\n## Background and prior protocol\n## Motivation: two objections\n## Paper approach and evaluation setup","[{\"question\":\"What protocol is being tested in this study?\",\"answer\":\"A validated procedure for classifying teaching-feedback comments into thematic categories and sentiment, based on a documented annotation guide, reliability checks, stratified cross-validation, and a held-out evaluation set.\"},{\"question\":\"How is durability evaluated as representation methods advance?\",\"answer\":\"The same protocol is kept fixed while models and representations are varied across three generations, including frozen transformer embeddings and prompted large language models, to assess whether performance remains competitive.\"},{\"question\":\"Does the sentiment task transfer from Spanish to English?\",\"answer\":\"Sentiment is transferred to English using a balanced 45,000-comment corpus validated against an independently annotated aspect-based education dataset, and the results show no sentiment advantage from a frontier model over a cheaper model on descriptive separation in English.\"}]",1784210559,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"a-durability-and-cross-language-transfer-benchmark-for-a-validated-teaching-feedback-classification-protocol","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/a-durability-and-cross-language-transfer-benchmark-for-a-validated-teaching-feedback-classification-protocol/86339/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What protocol is being tested in this study?","Question",{"text":75,"@type":76},"A validated procedure for classifying teaching-feedback comments into thematic categories and sentiment, based on a documented annotation guide, reliability checks, stratified cross-validation, and a held-out evaluation set.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is durability evaluated as representation methods advance?",{"text":80,"@type":76},"The same protocol is kept fixed while models and representations are varied across three generations, including frozen transformer embeddings and prompted large language models, to assess whether performance remains competitive.",{"name":82,"@type":73,"acceptedAnswer":83},"Does the sentiment task transfer from Spanish to English?",{"text":84,"@type":76},"Sentiment is transferred to English using a balanced 45,000-comment corpus validated against an independently annotated aspect-based education dataset, and the results show no sentiment advantage from a frontier model over a cheaper model on descriptive separation in English.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]