[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85637-en":3,"doc-seo-85637-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85637,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","RWGBench Evaluating Scholarly Positioning in Related Work Generation","Large language models demonstrate strong scientific writing fluency, yet evaluation for related work generation (RWG) remains weak. Current RWG assessments often reuse summarization-oriented metrics that rely on lexical or semantic similarity to reference text, overlooking RWG’s core nature as citation-level scholarly positioning. This can yield fluent outputs that still make critical academic mistakes in citation selection, reference placement, and contextual framing. RWGBench provides a citation-decision-centric benchmark and multidimensional evaluation of selection, context, organization, and framing.","RWGBench: Evaluating Scholarly Positioning in Related Work  \nGeneration  \nAnzhe Xie  \n[xaz25@mails.tsinghua.edu.cn](xaz25@mails.tsinghua.edu.cn)[ ](xaz25@mails.tsinghua.edu.cn)Tsinghua University Beijing, China  \nWeihang Su  \n[swh22@mails.tsinghua.edu.cn](swh22@mails.tsinghua.edu.cn)[ ](swh22@mails.tsinghua.edu.cn)Tsinghua University Beijing, China  \nJiaxin Mao  \n[maojiaxin@gmail.com](maojiaxin@gmail.com)[ ](maojiaxin@gmail.com)Renmin University of China Beijing, China  \nYiqun Liu  \n[yiqunliu@tsinghua.edu.cn](yiqunliu@tsinghua.edu.cn)[ ](yiqunliu@tsinghua.edu.cn)Tsinghua University Beijing, China  \nMin Zhang  \n[z-m@tsinghua.edu.cn](z-m@tsinghua.edu.cn)[ ](z-m@tsinghua.edu.cn)Tsinghua University Beijing, China  \nShaoping Ma  \n[msp@tsinghua.edu.cn](msp@tsinghua.edu.cn)[ ](msp@tsinghua.edu.cn)Tsinghua University Beijing, China  \nQingyao Ai∗ [aiqingyao@gmail.com](aiqingyao@gmail.com)[ ](aiqingyao@gmail.com)Tsinghua University Beijing, China  \narXiv :2606 .24894v 3 [ cs .DL] 11 Jul 2026  \nAbstract  \nLarge language models have shown strong fluency in scientific writing, yet the evaluation of related work generation (RWG) remains limited. Existing RWG evaluations largely inherit summarizationoriented metrics, using lexical or semantic similarity to reference sections as proxies for quality. However, related work writing is fundamentally a citation-level scholarly positioning task: it requires selecting, organizing, and framing prior work to clarify how a target paper relates to, differs from, and contributes beyond existing research. As a result, models may generate coherent and semantically relevant text while exhibiting academically critical failures, such as inappropriate citation selection or misplaced references, that conventional metrics do not capture. To this end, we introduce RWGBench, a benchmark that evaluates RWG from the perspective of citation decision-making rather than text similarity. RWGBench is constructed from a large-scale collection of 40,108 computer science papers and a retrieval corpus of 1.09 million documents, with a peerreviewed test set comprising 100 papers accepted at ICLR, NeurIPS, or ICML and their corresponding author-written related work sections. We propose a multi-dimensional evaluation framework that assesses citation selection, contextual appropriateness, organization, and citation framing. Experiments across representative and frontier generation settings reveal systematic limitations that are obscured by standard evaluations, separating retrieval bottlenecks from generation-level positioning failures. A blinded human study provides supplementary support for the proposed diagnostic metrics. RWGBench offers a citation-centric testbed for developing and evaluating related work generation systems that are better aligned with scholarly writing practices.  \nCCS Concepts  \n• Information systems → Evaluation of retrieval results; Summarization; • Computing methodologies → Natural language generation.  \n∗ Corresponding author.  \nKeywords  \nRelated work generation, citation recommendation, scholarly positioning, benchmark, LLM  \n1 Introduction  \nThe rapid growth of scientific literature has made it increasingly difficult for researchers to situate new work within a rapidly evolving research landscape. In computer science alone, more than 200,000 articles are published annually [4, 21], substantially increasing the burden of identifying relevant prior work, explaining how a study differs from existing research, and articulating its contribution. This challenge has motivated growing interest in automated related work generation (RWG), in which large language models [30, 42, 47] assist authors in drafting related work sections conditioned on a target paper.  \nDespite this growing interest, related work generation is not merely a special case of generic summarization or survey generation. While survey generation typically seeks to synthesize a broad body of topic-level literature with relatively comprehensi","cbCair6Uu9SZSUwq","https://ap.wps.com/l/cbCair6Uu9SZSUwq","pdf",623538,3,1,10,"English","en",105,"# Abstract\n# 1 Introduction\n## Motivation for automated RWG\n## RWG as scholarly positioning, not generic summarization\n## Limitations of existing evaluations","[{\"question\":\"Why is related work generation (RWG) not equivalent to generic summarization?\",\"answer\":\"RWG is anchored to a specific target paper and must rhetorically situate it within prior research. It requires selecting, organizing, and framing prior work to clarify relevance, differentiation, and contribution, which goes beyond topic-level synthesis.\"},{\"question\":\"What shortcomings do existing RWG evaluations have?\",\"answer\":\"Many metrics reuse summarization-style overlap or semantic similarity against reference sections. They can reward outputs that resemble references while missing academically critical citation errors such as poor citation selection, narrow citation coverage, or misrepresented placement contexts.\"},{\"question\":\"How does RWGBench evaluate RWG differently?\",\"answer\":\"RWGBench evaluates RWG from a citation decision-making perspective rather than text similarity. It uses a multidimensional framework covering citation selection, contextual appropriateness, organization, and citation framing, supported by experiments and a blinded human study.\"}]",1784205206,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"rwgbench-evaluating-scholarly-positioning-in-related-work-generation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/rwgbench-evaluating-scholarly-positioning-in-related-work-generation/85637/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is related work generation (RWG) not equivalent to generic summarization?","Question",{"text":75,"@type":76},"RWG is anchored to a specific target paper and must rhetorically situate it within prior research. It requires selecting, organizing, and framing prior work to clarify relevance, differentiation, and contribution, which goes beyond topic-level synthesis.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What shortcomings do existing RWG evaluations have?",{"text":80,"@type":76},"Many metrics reuse summarization-style overlap or semantic similarity against reference sections. They can reward outputs that resemble references while missing academically critical citation errors such as poor citation selection, narrow citation coverage, or misrepresented placement contexts.",{"name":82,"@type":73,"acceptedAnswer":83},"How does RWGBench evaluate RWG differently?",{"text":84,"@type":76},"RWGBench evaluates RWG from a citation decision-making perspective rather than text similarity. It uses a multidimensional framework covering citation selection, contextual appropriateness, organization, and citation framing, supported by experiments and a blinded human study.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":22,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]