[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81938-en":3,"doc-seo-81938-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},81938,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","SCOReD: Student-Aware CoT Optimization for Recommendation Distillation","Chain-of-thought (CoT) distillation is required to train reinforcement learning for recommendation, yet raw teacher traces are misaligned with this task. Large teachers show high reasoning uncertainty and repeatedly verify without changing answers, creating verbose students that rarely revise. SCOReD segments teacher traces, uses the student’s attention to score segment importance, and selects per-segment edits (KEEP/REWRITE/FUSE/PRUNE) via length and log-probability lift. SCOReD pruning preserves dense information, reduces reasoning length by 27.3%, and improves NDCG by 1.56% and Recall@5 by 1.9%.","arXiv :2607 .05734v2 [ cs .IR] 9 Jul 2026  \nSCOReD: Student-Aware CoT Optimization for Recommendation Distillation  \nHaz Sameen Shahgir 1 , Yufei Li2 ,∗ , Xiaohan Wei2 , Yunchen Pu2 , Fei Tian2 , Chonglin Sun2 , Frank Shyu2 , Sandeep Pandey2 , Luke Simon2 , Yue Dong 1 ,†, Xi Liu2 ,†  \n1 University of California Riverside, 2 Meta AI  \n∗ Execution lead, †Joint correspon author  \nChain-of-thought (CoT) distillation in the recommendation domain is a necessary precursor to RL training, but raw teacher traces are ill-suited to this task. Large teachers approach the recommendation task with unusually high reasoning uncertainty, repeatedly rechecking their answers without revising them; supervised fine-tuning on such traces produces verbose students that never revise their initial guess. Furthermore, due to the novelty of the recommendation domain, the teacher’s reasoning traces are highly out-of-distribution for the small student LLM.  \nWe propose Student-Aware CoT Optimization for Recommendation Distillation (SCOReD), a CoT optimization framework tailored to recommendation that first parses each teacher trace into typed segments and uses the student LLM’s attention to score the importance of each segment. Then SCOReD dynamically selects a per-segment edit (KEEP / REWRITE / FUSE / PRUNE) based on the output length and comparative log probability lift of the answer given the edit as per the student. Therefore, SCOReD prunes redundant sections of the reasoning trace while preserving informationdense sections and adapts raw teacher traces to the student’s output distribution. Training on SCOReD-optimized CoTs provides a cleaner learning signal to the student model and improves over baseline SFT by 1.56% NDCG and 1.9% Recall@5, while reducing reasoning length by 27.3% .  \nDate: July 13, 2026  \nCorrespondence: Xi Liu ([xliu1@meta.com](xliu1@meta.com)) and Yue Dong ([yue.dong@ucr.edu](yue.dong@ucr.edu))  \n1 Introduction  \nRecent work has begun to recast recommendation as a generative reasoning problem rather than a purely discriminative ranking task. Deng et al. (2025) propose OneRec, an end-to-end generative recommender that unifies retrieval and ranking by directly generating recommendation outputs. Building on this direction, Liu et al. (2025) argue that generative recommenders should not operate only as implicit predictors, but should also expose explicit in-text reasoning over user intent and item semantics. Liang et al. (2026) similarly frame reranking as a generative reasoning task, where a model reasons over the user history and candidate items before producing the final ranked list. Together, these systems point toward a broader shift: recommendation models are increasingly expected not only to score items, but also to compare, justify, and revise candidate choices through language.  \nThis shift creates a practical bottleneck. The strongest reasoning traces typically come from large LLMs, but real recommender systems often require small models because of latency, memory, and computational cost constraints. Distillation is therefore a necessary precursor for practical generative recommendation: before a small model can be improved with reinforcement learning or deployed as a reranker, it must first learn the basic format of recommendation reasoning from a stronger teacher. Prior work on chain-of-thought distillation shows that teacher rationales can transfer reasoning behavior to smaller models (Hsieh et al. , 2023 ; Li et al. , 2023), while broader LLM distillation work shows that the distillation objective itself must be adapted for generative models (Gu et al. , 2024) .  \nFigure 1 SCOReD recommendation trace compression pipeline. Left) Raw CoT from the teacher LLM is segmented into discrete spans and categorized into 6 stages. Middle) We compute the attention scores over the CoT using the pre-SFT target student LLM with a single forward pass. The average attention score from the \u003C/think> token approximates the student’s perceive","cbCaiccEVJnxWmQj","https://ap.wps.com/l/cbCaiccEVJnxWmQj","pdf",1888505,4,1,31,"English","en",105,"# Introduction\n## Generative recommendation as reasoning\n## Distillation bottleneck for small models\n## Challenges in recommendation CoT distillation\n## Prior work on CoT compression","[{\"question\":\"Why are raw teacher CoT traces ill-suited for recommendation distillation?\",\"answer\":\"Teacher traces contain high reasoning uncertainty and repetitive verification that usually does not change the final answer. This behavior trains students to be verbose and rarely revise their initial guess.\"},{\"question\":\"How does SCOReD decide edits to each segment of the teacher trace?\",\"answer\":\"SCOReD segments the teacher trace into typed spans, scores segment importance using the student LLM’s attention, then selects a per-segment operation (KEEP/REWRITE/FUSE/PRUNE) based on output length and the student’s comparative log-probability lift.\"},{\"question\":\"What improvements does SCOReD bring compared with baseline supervised fine-tuning?\",\"answer\":\"Training on SCOReD-optimized CoTs yields a 1.56% improvement in NDCG and 1.9% improvement in Recall@5, while reducing reasoning length by 27.3%.\"}]","SCOReD: Student-Aware CoT Optimization for Recommendation Distillation | PDF",1784177168,78,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"scored-student-aware-cot-optimization-for-recommendation-distillation","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":20},"https://docshare.wps.com/document/scored-student-aware-cot-optimization-for-recommendation-distillation/81938/",{"url":53,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-29","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why are raw teacher CoT traces ill-suited for recommendation distillation?","Question",{"text":76,"@type":77},"Teacher traces contain high reasoning uncertainty and repetitive verification that usually does not change the final answer. This behavior trains students to be verbose and rarely revise their initial guess.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does SCOReD decide edits to each segment of the teacher trace?",{"text":81,"@type":77},"SCOReD segments the teacher trace into typed spans, scores segment importance using the student LLM’s attention, then selects a per-segment operation (KEEP/REWRITE/FUSE/PRUNE) based on output length and the student’s comparative log-probability lift.",{"name":83,"@type":74,"acceptedAnswer":84},"What improvements does SCOReD bring compared with baseline supervised fine-tuning?",{"text":85,"@type":77},"Training on SCOReD-optimized CoTs yields a 1.56% improvement in NDCG and 1.9% improvement in Recall@5, while reducing reasoning length by 27.3%.","https://schema.org",{"og:url":53,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]