[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83713-en":3,"doc-seo-83713-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83713,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","Large-Scale Dataset of Automatically Classified Rhetorical Sections in Scientific Papers","Scientific papers use consistent rhetorical structures that map content into sections such as Introduction, Methods, Results, and Discussion. A large-scale section-level annotation dataset is introduced for millions of papers from the Semantic Scholar Open Research Corpus (S2ORC). A rule-based classifier identifies major sections across 15.6 million papers, then quality filtering yields 13.5 million labeled records. Human and LLM validation shows agreement comparable to human inter-annotator reliability. The dataset supports computational research on scientific discourse and writing patterns.","arXiv :2607 .0338 1v 1 [ cs .DL] 3 Jul 2026  \nLarge-scale dataset of automatically classified rhetorical sections  \nin scientific papers  \nDaniel Verdi 1 ,2 ,3 , Jacob Aarup Dalsgaard2 ,3 , Roberta Sinatra2 ,3 ,4 ,5  \n1 Graduate School of Education, Stanford University, United States  \n2 Center for Social Data Science (SODAS), University of Copenhagen, Denmark  \n3 Networks, Data, and Society (NERDS), IT University of Copenhagen, Denmark  \n4 Pioneer Centre for Artificial Intelligence (P1), Denmark  \n5 Complexity Science Hub, Austria  \n∗ Corresponding authors: [verdi@stanford.edu](verdi@stanford.edu), [jad@sodas.ku.dk](jad@sodas.ku.dk), [robertasinatra@sodas.ku.dk](robertasinatra@sodas.ku.dk)  \nAbstract  \nScientific papers follow rhetorical structures that organize content into sections such as Introduction, Methods, Results, and Discussion. Automatically identifying these sections at scale enables granular analysis of scientific writing patterns. We present a dataset of section-level annotations for millions of scientific papers from the Semantic Scholar Open Research Corpus (S2ORC) . Using a rule-based classification algorithm, we identified and labeled major sections across 15.6 million papers after quality filtering. The dataset covers primarily STEM disciplines, with strong representation in medicine and biology. We provide comprehensive human and LLM-based validation showing that classifier agreement with human annotators is on par with human inter-annotator agreement. This dataset enables large-scale computational studies of scientific discourse and writing patterns.  \n1 Background & Summary  \nUnderstanding how scientists communicate their findings requires analyzing the rhetorical structure of scientific papers. Researchers often organize their work into standardized sections that serve distinct communicative purposes, such as introductions framing the research problem, methods describing experimental procedures, and results presenting findings, many times following the IMRaD format (Introduction, Methods, Results, and Discussion) [1, 2] . Automatically identifying these sections at scale enables new forms of analysis that would be impractical with manual annotation. However, creating such section-level annotations for large corpora presents significant technical challenges due to the variability in section headers across disciplines, publishers, and document structures.  \nSeveral researchers have previously addressed this challenge, though with important limitations. Shahid and Afzal [3] used keyword matching combined with logical section order, such as assuming the introduction appears first. Treeratpituk et al. [4] and Tuarob et al. [5] extended the keyword approach by employing regular expressions to capture broader variations of sectionrelated terms and phrases. Nguyen and Kan [6] developed a Maximum Entropy classifier using four features: section number, relative position within the document, previous section header, and current section header. More recently, Rahman and Finin [7] applied a word-based convolutional neural network model, converting input texts into multi-label one-hot vectors passed through an embedding layer.  \nThese prior approaches share several constraints that limit their applicability to large-scale corpus analysis. Most were evaluated on relatively small datasets [3, 4, 5], often containing only hundreds of papers, making it unclear whether they generalize to the diversity encountered  \nin million-scale corpora. Additionally, none of these research efforts made their code publicly available, preventing replication or extension of their methods.  \nWe address this by developing a classification algorithm adapted to the scale and characteristics of the Semantic Scholar Open Research Corpus (S2ORC) [8], which contains the full text of 15.6 million open-access scientific papers. Although we apply the method to S2ORC due to data availability, the approach is applicable to any scholarly corpus that p","cbCaikTQizJpyzDA","https://ap.wps.com/l/cbCaikTQizJpyzDA","pdf",538950,3,1,15,"English","en",105,"# Background & Summary\n## Rhetorical sections in scientific writing\n## Prior approaches and limitations\n## Proposed scalable classification method\n## Dataset scale, scope, and validation","[{\"question\":\"What problem does the dataset address?\",\"answer\":\"It addresses the difficulty of automatically identifying rhetorical sections in scientific papers at very large scale, despite high variability in section headers across disciplines and document formats.\"},{\"question\":\"How is the dataset created from S2ORC?\",\"answer\":\"A rule-based classification algorithm performs two-pass labeling, followed by label propagation and filtering, to identify and annotate major sections across 15.6 million papers in S2ORC.\"},{\"question\":\"How was the classification quality validated?\",\"answer\":\"Quality was validated using human annotation on 100 papers and LLM-based annotation on 1,000 papers, achieving Krippendorff’s Alpha of 0.61 between human annotators and the algorithm on cleaned labels.\"}]",1784189920,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"large-scale-dataset-of-automatically-classified-rhetorical-sections-in-scientific-papers","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/large-scale-dataset-of-automatically-classified-rhetorical-sections-in-scientific-papers/83713/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the dataset address?","Question",{"text":75,"@type":76},"It addresses the difficulty of automatically identifying rhetorical sections in scientific papers at very large scale, despite high variability in section headers across disciplines and document formats.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the dataset created from S2ORC?",{"text":80,"@type":76},"A rule-based classification algorithm performs two-pass labeling, followed by label propagation and filtering, to identify and annotate major sections across 15.6 million papers in S2ORC.",{"name":82,"@type":73,"acceptedAnswer":83},"How was the classification quality validated?",{"text":84,"@type":76},"Quality was validated using human annotation on 100 papers and LLM-based annotation on 1,000 papers, achieving Krippendorff’s Alpha of 0.61 between human annotators and the algorithm on cleaned labels.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]