[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83049-en":3,"doc-seo-83049-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83049,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","LongCrafter Towards Diverse Long Context Understanding via Evidence Graph Guided Instruction Synthesis","LongCrafter presents a scalable framework for synthesizing long-context supervised fine-tuning (SFT) data to improve large language models’ understanding over extremely long token ranges. The method resolves three recurring issues: narrow task coverage, overly easy instructions, and missing faithfulness supervision. LongCrafter couples a hierarchical task taxonomy with an evidence-grounded pipeline to generate difficulty-controlled, traceable instruction–response pairs strictly grounded in located evidence spans. Models trained on LongCrafter data outperform SFT baselines and official post-trained models on LongBench, LongBench v2, and LooGLE, with the largest gains on difficult tasks and robust evidence locating that mitigates “lost in the middle.”","LONGCRAFTER: Towards Diverse Long-Context Understanding via  \nEvidence-Graph-Guided Instruction Synthesis  \nChenhao Yuan 1 * , Yinhao Xu 1 * , Shuwen Xu 1 , Xizhi Yang 1 , Jiaxiang Liu2 , Chenxi Zhou2 , Shaoping Huang2 , Haolin Ren 1 , Pengfei Cao†2, Jun Zhao2 , Kang Liu†2  \n1University of Chinese Academy of Sciences, Beijing, China  \n2The Key Laboratory of Cognition and Decision Intelligence for Complex Systems,  \nInstitute of Automation, Chinese Academy of Sciences, Beijing, China  \narXiv :2607 .06 160v 1 [ cs .CL] 7 Jul 2026  \nAbstract  \nSynthesizing long-context supervised fine-tuning (SFT) data is a scalable way to enhance the long-context understanding of large language models (LLMs), yet existing approaches share three limitations: narrow task coverage, insufficient instruction difficulty, and a lack of faithfulness supervision. We propose LongCrafter, a structured synthesis framework that couples a hierarchical task taxonomy with an evidencegrounded pipeline. The taxonomy organizes long-context understanding into local/shallow and global/deep levels and yields 32 fine-grained task types that serve as a global generative prior. Guided by this taxonomy, LongCrafter constructs task-aligned long contexts, decomposes them into explicit evidence graphs that model cross-paragraph dependencies, and generates instruction–response pairs strictly grounded in the located evidence spans, ensuring both controllable difficulty and faithful, traceable reasoning. Models fine-tuned on LongCrafter data outperform all SFT baselines and even the official post-trained models on LongBench, LongBench v2, and LooGLE across both Qwen2.5-7B and LLaMA-3.1-8B, with the largest gains on high-difficulty tasks. Further analysis shows that LongCrafter data is more diverse and better spread across difficulty levels, and that the trained models locate evidence robustly regardless of position, effectively mitigating the “lost in the middle” problem.  \n1 Introduction  \nLong-context understanding has emerged as a critical capability for large language models (LLMs), as real-world applications such as question answering, summarization, and complex reasoning require models to recognize and utilize relevant evidence across tens or hundreds of thousands of tokens (Bai et al. 2024b, 2025; Li et al. 2024; Liu et al. 2024; Peng et al. 2026b) . Supervised fine-tuning (SFT) on synthesized long-context instruction data has emerged as a promising and scalable solution (Bai et al. 2024a; Chen et al. 2025; Yang et al. 2025b; Chen et al. 2024; Zhang et al. 2025b; Liet al. 2025; Gao et al. 2025). However, existing long-context SFT datasets suffer from three compounding limitations. 1) Limited task coverage. Without a systematic task taxonomy to guide synthesis, prior work concentrates on a narrow  \nset of task types (e.g., multi-hop QA (Chen et al. 2024; Bai *These authors contributed equally.  \n†Corresponding author.  \net al. 2024a; Yang et al. 2025b; Chen et al. 2025)), leaving diverse real-world capabilities such as temporal reasoning, aggregation, and state tracking insufficiently supervised. 2) Insufficient instruction difficulty. Prior work often generates questions directly from raw documents without modeling evidence structures, where evidence spans may depend on each other in the form of chains, trees, or graphs (Bai et al. 2024a; Gao et al. 2025; Li et al. 2025) . This lack of structural modeling and difficulty stratification naturally biases the generated data toward easy, locally answerable questions, allowing models to exploit shortcuts rather than learn genuine cross-paragraph reasoning. 3) Lack of faithfulness supervision. Without supervision that anchors each reasoning step to source evidence (Xu et al. 2024), models may rely on parametric knowledge rather than the source context, potentially yielding unfaithful reasoning inconsistent with the document.  \nTo this end, we propose LongCrafter, a data synthesis framework that addresses these limitations by c","cbCaiqT3kCiyMTQc","https://ap.wps.com/l/cbCaiqT3kCiyMTQc","pdf",2388135,7,1,17,"English","en",105,"# Abstract\n# Introduction\n## Limitations of Existing Long-Context SFT\n## Proposed Framework: LongCrafter\n### Hierarchical Task Taxonomy\n### Evidence-Constraint Graph Construction\n### Instruction–Response Pair Synthesis","[{\"question\":\"What limitations of existing long-context SFT data does LongCrafter target?\",\"answer\":\"It addresses limited task coverage, insufficient instruction difficulty, and the lack of faithfulness supervision that anchors reasoning to source evidence.\"},{\"question\":\"How does LongCrafter create task diversity for long-context understanding?\",\"answer\":\"It uses a hierarchical task taxonomy that separates local/shallow and global/deep capabilities, producing 32 fine-grained task types as a global generative prior.\"},{\"question\":\"How does LongCrafter ensure instruction difficulty and faithfulness in generated training data?\",\"answer\":\"It constructs an evidence-constraint graph from extracted evidence spans and generates instruction–response pairs conditioned on the located nodes, requiring joint use of key evidence and step-by-step, traceable reasoning.\"}]",1784184870,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"longcrafter-towards-diverse-long-context-understanding-via-evidence-graph-guided-instruction-synthesis","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/longcrafter-towards-diverse-long-context-understanding-via-evidence-graph-guided-instruction-synthesis/83049/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What limitations of existing long-context SFT data does LongCrafter target?","Question",{"text":76,"@type":77},"It addresses limited task coverage, insufficient instruction difficulty, and the lack of faithfulness supervision that anchors reasoning to source evidence.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does LongCrafter create task diversity for long-context understanding?",{"text":81,"@type":77},"It uses a hierarchical task taxonomy that separates local/shallow and global/deep capabilities, producing 32 fine-grained task types as a global generative prior.",{"name":83,"@type":74,"acceptedAnswer":84},"How does LongCrafter ensure instruction difficulty and faithfulness in generated training data?",{"text":85,"@type":77},"It constructs an evidence-constraint graph from extracted evidence spans and generates instruction–response pairs conditioned on the located nodes, requiring joint use of key evidence and step-by-step, traceable reasoning.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]