[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83853-en":3,"doc-seo-83853-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83853,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","A Temporal Reasoning Benchmarking Framework for LRMs via Difficulty-controlled and Dynamic Test Generation","Defining the reasoning boundaries and ensuring the reliability of Large Reasoning Models (LRMs) remains a critical challenge because existing temporal reasoning benchmarks often rely on static data vulnerable to contamination, and evaluate mainly by final outcomes that hide reasoning defects. TRACE introduces temporal reasoning as constraint satisfaction using Allen’s Interval Algebra, enabling fine-grained difficulty control and trace-faithfulness checking through a Trace-Based Verification Oracle. TRACEBench provides 1,200 graded synthesized tests, revealing performance–difficulty negative correlation (Pearson’s ≈ −0.96) and substantial spurious guessing (~28%), plus scale-dependent failure modes.","arXiv :2607 .04784v 1 [ cs . SE] 6 Jul 2026  \nA Temporal Reasoning Benchmarking Framework for LRMs via Difficulty-controlled and Dynamic Test Generation  \nSHIDE ZHOU, Huazhong University of Science and Technology, China  \nKAILONG WANG∗ , Huazhong University of Science and Technology, China and National University of Singapore, Singapore  \nLING SHI, Nanyang Technological University, Singapore  \nHAOYU WANG, Huazhong University of Science and Technology, China  \nDefining the reasoning boundaries and ensuring the reliability of Large Reasoning Models (LRMs) remains a critical challenge. Current benchmarks primarily rely on static datasets susceptible to data contamination or synthetic tasks lacking fine-grained difficulty control. Furthermore, standard outcome-based evaluations often conceal reasoning flaws by neglecting the reasoning process.  \nTo address these limitations, we introduce TRACE, a testing framework that models temporal reasoning as constraint satisfaction problems via Allen’s Interval Algebra. This approach enables precise regulation of logical complexity and incorporates a Trace-Based Verification Oracle to validate reasoning faithfulness. Using this framework, we construct TRACEBench, an extensive benchmark comprising 1,200 synthesized test instances across graded difficulty levels. We employ TRACE to evaluate eight widely used LRMs on TRACEBench. The results confirm a strong negative correlation between model performance and our difficulty metric (Pearson’s 􀁁 ≈ −0 . 96), validating the effectiveness of our difficulty control mechanism. Moreover, our trace-based analysis exposes significant discrepancies between reasoning validity and final answers, revealing a high spurious guessing rate of approximately 28% in mid-sized models. In addition, we diagnose scale-dependent failure modes, ranging from Degenerative Loops in small models to Reasoning Explosion in advanced architectures. TRACE thus provides a robust, automated platform for benchmarking the true temporal reasoning capabilities of LRMs.  \nCCS Concepts: • Computing methodologies → Temporal reasoning; Natural language generation; • Software and its engineering → Software verification and validation.  \nAdditional Key Words and Phrases: Large Reasoning Models, Temporal Reasoning, Automated Testing, Test Generation  \n1 Introduction  \nThe evolution of Large Language Models (LLMs) has culminated in the emergence of Large Reasoning Models (LRMs), such as the DeepSeek-R1 family [7, 8] and OpenAI’s o-series [19], designed specifically for complex problem-solving. In this domain, temporal reasoning is a fundamental capability, demanding strict logical consistency rather than approximate retrieval. While Chain-ofThought(CoT) strategies [15, 28, 30] have yielded significant performance gains, a fundamental question persists: Do these improvements reflect genuine deduction or merely sophisticated pattern matching? This uncertainty complicates reliability assessments [9], underscoring the urgent need for a specialized benchmarking framework.  \n∗ Corresponding author.  \nAuthors’ Contact Information: Shide Zhou, Huazhong University of Science and Technology, Wuhan, China, shidez@ [hust.edu.cn](hust.edu.cn); Kailong Wang, Huazhong University of Science and Technology, Wuhan, China and National University of Singapore, Singapore, Singapore, [wangkl@hust.edu.cn](wangkl@hust.edu.cn); Ling Shi, Nanyang Technological University, Singapore, Singapore, [ling.shi@ntu.edu.sg](ling.shi@ntu.edu.sg); Haoyu Wang, Huazhong University of Science and Technology, Wuhan, China, [haoyuwang@hust.edu](haoyuwang@hust.edu). cn.  \n2026. ACM XXXX-XXXX/2026/7-ART  \n[https://doi.org/10.1145/nnnnnnn.nnnnnnn](https://doi.org/10.1145/nnnnnnn.nnnnnnn)  \n, Vol. 1, No. 1, Article . Publication date: July 2026 .  \n2 Shide Zhou, Kailong Wang, Ling Shi, and Haoyu Wang  \nTable 1 . Comparison of Our Work with Existing Benchmarks.  \n\n|  | TRAM | TimeBench | Test of Time | t-BEN | Our Work |\n| --- | --- | --- | ","cbCaiiqG5MuIPoNj","https://ap.wps.com/l/cbCaiiqG5MuIPoNj","pdf",1074829,4,1,23,"English","en",105,"# Introduction\n# Related Work and Benchmark Limitations\n## Comparison with Existing Benchmarks","[{\"question\":\"What limitations do current temporal reasoning benchmarks have?\",\"answer\":\"They often depend on static datasets that risk data contamination and use synthetic tasks with coarse difficulty control. Outcome-based evaluation can also conceal flaws in the reasoning process by focusing on final answers only.\"},{\"question\":\"How does TRACE model temporal reasoning and control difficulty?\",\"answer\":\"TRACE formulates temporal reasoning as constraint satisfaction problems using Allen’s Interval Algebra. A Difficulty-Aware Constraint Generator builds constraint graphs to regulate logical complexity with fine-grained control.\"},{\"question\":\"What does TRACEBench reveal about LRM performance and reasoning faithfulness?\",\"answer\":\"TRACEBench evaluates eight LRMs and finds a strong negative correlation between performance and the proposed difficulty metric (Pearson’s ≈ −0.96). Trace-based analysis also shows a high spurious guessing rate (~28%) and discrepancies between reasoning validity and final answers, with failure modes varying by model scale.\"}]",1784190996,58,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"a-temporal-reasoning-benchmarking-framework-for-lrms-via-difficulty-controlled-and-dynamic-test-generation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/a-temporal-reasoning-benchmarking-framework-for-lrms-via-difficulty-controlled-and-dynamic-test-generation/83853/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What limitations do current temporal reasoning benchmarks have?","Question",{"text":75,"@type":76},"They often depend on static datasets that risk data contamination and use synthetic tasks with coarse difficulty control. Outcome-based evaluation can also conceal flaws in the reasoning process by focusing on final answers only.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does TRACE model temporal reasoning and control difficulty?",{"text":80,"@type":76},"TRACE formulates temporal reasoning as constraint satisfaction problems using Allen’s Interval Algebra. A Difficulty-Aware Constraint Generator builds constraint graphs to regulate logical complexity with fine-grained control.",{"name":82,"@type":73,"acceptedAnswer":83},"What does TRACEBench reveal about LRM performance and reasoning faithfulness?",{"text":84,"@type":76},"TRACEBench evaluates eight LRMs and finds a strong negative correlation between performance and the proposed difficulty metric (Pearson’s ≈ −0.96). Trace-based analysis also shows a high spurious guessing rate (~28%) and discrepancies between reasoning validity and final answers, with failure modes varying by model scale.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]