[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84076-en":3,"doc-seo-84076-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84076,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","What Resolve Rate Hides: Trajectory Structure Diagnostics for Coding Agents","Coding agents are commonly ranked by resolve rate, yet identical pass/fail outcomes can arise from markedly different processes, leaving developers without explanations for failures or inefficiencies. Trajectory evidence records search, reads, edits, tool calls, validation, and reversions, but raw traces are heterogeneous and hard to compare. TRACEPROBE transforms each run into a canonical nine-type action taxonomy with deterministic effect labels, then applies anti-pattern identification and run alignment to pinpoint divergence under controlled references. On 2,500 trajectories across five production settings on SWE-Bench Verified, it finds that file choice is too coarse, function selection and completion behavior localize effects, and resolved runs still vary in speed and wasted effort.","What Resolve Rate Hides: Trajectory Structure Diagnostics for Coding Agents  \nRui Shu, Chun Yong Chong, Xin Zhou, Yun Peng, Zihan Wu  \nXu Han, Zeyang Zhuang, Guowen Yuan, Yuan Wang  \narXiv :2607 .06 184v 1 [ cs . SE] 7 Jul 2026  \nAbstract—Coding agents are ranked almost entirely by resolve rate: whether their final patch passes the target tests. Yet two agents can reach the same outcome through very different processes, and a single pass/fail label says nothing about why a run failed or why an accepted run spent extra steps, time, or tokens. This process evidence lives in the trajectory, which records a run’s searches, reads, edits, tool calls, validation, and reversions. However, raw traces are heterogeneous and hard to compare across runs. We present TRACEPROBE, a trajectory-diagnostic framework that recovers what resolve rate hides. TRACEPROBE normalizes each raw run into a canonical nine-type action taxonomy with deterministic effect labels, then applies two rule-based modules: INSIGHT names single-trajectory anti-patterns adapted from established debugging practice (e.g., search loops, verification skips), while CONVERGE aligns pairs of runs and classifies where their behavior diverges under controlled references. Applying TRACEPROBE to 2,500 trajectories from five production settings on SWE-Bench Verified, we find that (i) file choice is too coarse to separate success from failure, whereas function selection and completion behavior localize it; (ii) INSIGHT anti-patterns act mainly as corpus-level difficulty clues, with search loops the most stable; and (iii) even resolved runs differ in how quickly they reach relevant code and how much failed work they incur. Trajectory structure thus adds auditable diagnostic context to outcomes by localizing inspection targets, suggesting failure hypotheses, and prioritizing runs for review.  \nIndex Terms—code agents, trajectory analysis, anti-patterns, software engineering, LLM agents  \nI. INTRODUCTION  \nLLM-powered coding agents are now widely used in developer environments [1]–[3] . They inspect repositories, edit files, invoke tools, run tests, and iterate on feedback. As these systems move into production workflows, practitioners need to compare not only whether an agent solved a task, but also how it behaved while attempting the task. As the key metric, resolve rate records whether the final patch passes target tests. This outcome is useful for ranking systems and guiding release decisions, but it hides the process differences that matter to agent developers.  \nThis limitation appears even when two runs reach the same outcome. Figure 1 shows an example where Claude Code (Opus 4.6) and OpenCode (GLM-5) both resolve the same SWE-Bench task. The pass/fail label treats both runs as successes, but their processes differ. Claude Code reaches a targeted fix with no failed actions, while OpenCode takes a longer path with repeated failures and recovery work. This example illustrates why process evidence is required in addition to  \nPreprint. Under review.  \nClaude Code OpenCode  \n10 steps  \nPhases:  understand  implement  validate  \n49 steps  \nfailed recovery  \nfailed recovery  \nordering inefficiency  \noff-anchor read failed recovery  \n debug  report  \nFig. 1: In this example, Claude Code (Opus 4.6) and OpenCode (GLM-5) both RESOLVE the task (SWE-Bench pytest-7982), but their trajectories differ substantially. Claude Code reaches a targeted fix in 10 steps with no failed actions, while OpenCode takes 49 steps and includes repeated failed/recovery spans. Colored dots show workflow phases, and dashed gray lines show actions aligned by our framework. Red brackets provide informal visual annotations for this example, and our analysis later defines the patterns and divergence signals used for measurement.  \npass/fail outcomes. Prior work has approached this need from two directions. One direction questions whether pass/fail outcomes are reliable, showing that generated patches can exploit wea","cbCaiggbmdnM7wgO","https://ap.wps.com/l/cbCaiggbmdnM7wgO","pdf",791048,4,1,12,"English","en",105,"# Introduction\n# Problem Motivation\n# TRACEPROBE Framework\n## Action Taxonomy and Deterministic Labels\n## Anti-Pattern Identification (INSIGHT)\n## Controlled Run Alignment (CONVERGE)\n# Experimental Findings","[{\"question\":\"Why does resolve rate fail to fully explain coding agent performance?\",\"answer\":\"Resolve rate only records whether the final patch passes target tests. It hides differences in how the agent searched, edited, validated, and recovered when failures occurred.\"},{\"question\":\"What is TRACEPROBE and how does it make trajectories comparable?\",\"answer\":\"TRACEPROBE normalizes raw trajectories into a canonical nine-type action taxonomy with deterministic effect labels, preserving local evidence so that runs can be compared across settings.\"},{\"question\":\"What kinds of trajectory signals does TRACEPROBE use to diagnose failures and divergences?\",\"answer\":\"It uses rule-based modules: INSIGHT to name single-trajectory anti-patterns (such as search loops or skipped verification) and CONVERGE to align pairs of runs and classify where their behavior diverges under controlled references.\"}]",1784192542,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"what-resolve-rate-hides-trajectory-structure-diagnostics-for-coding-agents","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/what-resolve-rate-hides-trajectory-structure-diagnostics-for-coding-agents/84076/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does resolve rate fail to fully explain coding agent performance?","Question",{"text":75,"@type":76},"Resolve rate only records whether the final patch passes target tests. It hides differences in how the agent searched, edited, validated, and recovered when failures occurred.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is TRACEPROBE and how does it make trajectories comparable?",{"text":80,"@type":76},"TRACEPROBE normalizes raw trajectories into a canonical nine-type action taxonomy with deterministic effect labels, preserving local evidence so that runs can be compared across settings.",{"name":82,"@type":73,"acceptedAnswer":83},"What kinds of trajectory signals does TRACEPROBE use to diagnose failures and divergences?",{"text":84,"@type":76},"It uses rule-based modules: INSIGHT to name single-trajectory anti-patterns (such as search loops or skipped verification) and CONVERGE to align pairs of runs and classify where their behavior diverges under controlled references.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]