[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82308-en":3,"doc-seo-82308-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82308,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","WILDTRACE Benchmarking Natural Evidence Trails in Long-Context Reasoning","Answering complex questions over long documents depends on integrating evidence dispersed across distant passages within the same source. Existing benchmarks often rely on needle probes, planted facts, or reverse-engineered chains that may differ in distribution or placement, obscuring whether success reflects real source reasoning or artifacts. WILDTRACE introduces 481 tasks over 214 naturally occurring long-form sources, defining seven source-internal evidence geometries and validating trail necessity, groundedness, rubric fidelity, contamination resistance, and answerability. Evaluations show best models reach 75.3% yet struggle on counterfactual and causal attribution settings.","arXiv :2607 .09328v 1 [ cs .CL] 10 Jul 2026  \nWILDTRACE: Benchmarking Natural Evidence Trails  \nin Long-Context Reasoning  \nZixin Chen 1 ,2 ∗Peng Liu2 Haobo Li 1 Rui Sheng 1 Jianhong Tu2 Xiaodong Deng2 Fei Huang2 Kashun Shum 1 ,2 Dayiheng Liu2 Huamin Qu 1  \n1 Hong Kong University of Science and Technology  \n2 Qwen Team, Alibaba Group  \n[zchendf@connect.ust.hk](zchendf@connect.ust.hk)  \nAbstract  \nAnswering complex questions over long documents frequently requires integrating evidence that the source itself disperses naturally across distant passages. In an incident report, the operating condition, design flaw, and missed safety check that jointly explain a disaster may appear dozens of sections apart; in a novel, a character’s true motive may surface only through scenes far removed from the moment it becomes relevant. This source-internal evidence integration is central to real-world long-document analysis, yet existing benchmarks largely sidestep it. Needle probes, planted facts, and reverse-engineered multi-hop chains embed evidence that may differ from the host text in distribution, placement, or register, making it unclear whether strong performance reflects genuine source reasoning or distributional artifacts. We introduce WILDTRACE, a benchmark of 481 tasks over  \n214 naturally occurring long-form sources such as technical incident reports and lesser-known literary narratives, where all evidence trails arise from the document’sown causal, temporal, and narrative logic. Drawing on Pearl’s causal hierarchy and prior multi-hop reasoning typologies, we define seven source-internal evidence geometries that characterize the distinct relational demands of analytical reading in long documents. A source-first construction pipeline mines candidate trails from document structure before writing questions; each item then undergoes multistage validation covering clue necessity, answer groundedness, rubric fidelity, contamination resistance and answerability. Across 18 frontier systems evaluated under evidence-withheld conditions, the strongest reaches 75.3%, with pronounced weaknesses on reasoning-intensive geometries such as counterfactual branching and causal attribution. These failures persist despite sufficient context capacity. As models are increasingly entrusted with real-world high-stakes analytical tasks, this gap between accessing information and reasoning over naturally dispersed evidence emerges as a defining challenge for the next stage of long-context research.  \n1 Introduction  \nLong-document reasoning is not merely retrieval over more tokens. For a given analytic question, a model must decide which ordinary source fragments are load-bearing evidence, which sameregister details are distractors, and which causal, temporal, comparative, abductive, or counterfactual dependency makes the selected fragments jointly warrant an answer. In literary analysis, a character’s motive may not reside in the scene where the decisive action occurs; it may be warranted only by linking an early promise, a later contradiction, and a delayed consequence that changes how earlier scenes should be read. Structured technical documents pose the same burden in a more procedural register. Even when a report contains findings, headings, tables, and cross-references, a follow-up  \n∗Work done during internship at Qwen Team  \nPreprint.  \nquestion may require linking an operating condition, a design constraint, and a review lapse whose dependency is not packaged as the answer to that query. The relevant passages may all be locatable, but the task is to recover why they jointly support the conclusion.  \nMany existing evaluations measure important adjacent capabilities, but often control the evidence environment in ways that reduce the burden of evidence discovery and relation recovery. Needle-style and controlled long-context probes measure access, position sensitivity, aggregation, and robustness, yet the relevant facts are inserted or otherwise benchm","cbCaivN33WysbJrn","https://ap.wps.com/l/cbCaivN33WysbJrn","pdf",4474980,1,27,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What challenge does WILDTRACE target in long-context reasoning?\",\"answer\":\"WILDTRACE targets the difficulty of reasoning with source-internal evidence that is naturally dispersed across distant passages and must be integrated to jointly warrant an answer.\"},{\"question\":\"How does WILDTRACE design its evaluation tasks and evidence trails?\",\"answer\":\"It defines seven source-internal evidence geometries, constructs tasks by mining candidate trails from document structure before question writing, and applies multistage validation for necessity, answer groundedness, rubric fidelity, contamination resistance, and answerability.\"},{\"question\":\"What performance and failure patterns are reported for frontier models?\",\"answer\":\"Under evidence-withheld conditions, the strongest system reaches 75.3%, with notable weaknesses on reasoning-intensive geometries such as counterfactual branching and causal attribution.\"}]",1784179519,68,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"wildtrace-benchmarking-natural-evidence-trails-in-long-context-reasoning","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/wildtrace-benchmarking-natural-evidence-trails-in-long-context-reasoning/82308/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What challenge does WILDTRACE target in long-context reasoning?","Question",{"text":75,"@type":76},"WILDTRACE targets the difficulty of reasoning with source-internal evidence that is naturally dispersed across distant passages and must be integrated to jointly warrant an answer.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does WILDTRACE design its evaluation tasks and evidence trails?",{"text":80,"@type":76},"It defines seven source-internal evidence geometries, constructs tasks by mining candidate trails from document structure before question writing, and applies multistage validation for necessity, answer groundedness, rubric fidelity, contamination resistance, and answerability.",{"name":82,"@type":73,"acceptedAnswer":83},"What performance and failure patterns are reported for frontier models?",{"text":84,"@type":76},"Under evidence-withheld conditions, the strongest system reaches 75.3%, with notable weaknesses on reasoning-intensive geometries such as counterfactual branching and causal attribution.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]