[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81742-en":3,"doc-seo-81742-105":30,"detail-sidebar-cat-0-en-105":84},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81742,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks","Large language models embedded in multi-turn agentic software engineering (SWE) harnesses improve repository-scale coding performance, yet routing every task to a frontier model is inefficient because many issues can be solved with cheaper approaches. Existing LLM routers rely only on task descriptions, imposing an information-theoretic Bayes-error floor in agentic settings. SWE-Router introduces value-based temporal routing that uses a cheap model’s exploratory partial trajectory to decide whether to continue cheaply or escalate to an expensive model. The work proves Bayes-optimality for partial-trajectory conditioning, reports substantial cost reductions, and releases a multi-LLM trajectory dataset for reproducibility.","SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks  \nSeongho Son 1 Sangwoong Yoon 2 Jiahua Tang 3 Shuhan Wang 1 Lorenz Wolf 1 Ilija Bogunovic 1 4  \narXiv :2607 .00053v1 [ cs . SE] 30 Jun 2026  \nAbstract  \nLarge language models (LLMs) embedded in multi-turn agentic harnesses are reshaping software engineering (SWE), but routing every task to a frontier model is wasteful when many issues admit cheap fixes. Existing LLM routers operate on the task description alone, which inherits an information-theoretic Bayes-error floor in agentic settings: a similar issue can hide either a localized typo or a multi-module refactor, and the prompt does not separate the two. We introduce SWERouter, a value-based temporal approach that lets a cheap model run for a few exploratory turns and reads the resulting partial trajectory before deciding whether to continue cheaply or to escalate to an expensive model. We provide a Bayesoptimality theorem showing that conditioning on the partial trajectory never harms routing and is strictly better whenever exploration is informative.  \nAcross the LLM pairs of weak and strong models spanning the contemporary cost–capability frontier, we show that SWE-Router greatly improves the cost efficiency of SWE tasks, while maintaining the majority of the performances of the stronger model. We additionally release a multi-LLM trajectory dataset which allows reproduction of our trajectory-level routing.  \n1. Introduction  \nLarge language models (LLMs) embedded in multi-turn agentic harnesses (Yang et al., 2024a ; Wang et al., 2024b ; Xia et al., 2024 ; Zhang et al., 2024b) are achieving state-of-the-art resolution rates on repository-scale coding benchmarks (Liu et al., 2023 ; Jimenez et al., 2024 ; Zhuo et al., 2024) . However, their per-task inference cost exceeds that of competitive open-weight alternatives by an order of magnitude or more, while only a minority of  \n1University College London, United Kingdom 2Ulsan National Institute of Science and Technology, South Korea 3PSL Research University, France 4University of Basel, Switzerland. Correspondence to: Seongho Son \u003C[seong.son.22@ucl.ac.uk](seong.son.22@ucl.ac.uk)>.  \nThe 5th Deep Learning for Code Workshop, ICML 2026 .  \nthe existing tasks truly require frontier capability. Routing every instance to a frontier model is therefore wasteful in expectation, which motivates LLM routing (Ong et al., 2025 ; Chen et al., 2024a ; Jitkrittum et al., 2025 ; Hu et al., 2024): selecting a model per instance so that expensive inference is reserved for cases in which it is necessary.  \nExisting routers operate as classifiers or calibrated probability estimators over the task description q (Ong et al., 2025 ; Jitkrittum et al., 2025 ; Ding et al., 2024 ; Aggarwal et al., 2023 ; Hu et al., 2024), picking the cheapest model whose cost-adjusted success probability is highest. This is poorly matched to multi-turn agentic SWE: a similar issue description can specify a localized typo or a multi-module refactor, and the distinction is often not identifiable from q alone. The information that resolves this ambiguity is generated by the agent itself : SWE agents follow ReAct-style loops (Yao et al., 2023) of thoughts, actions, and observations, anda substantial fraction of their turns are spent on bug localization rather than patch synthesis (Xia et al., 2024 ; Zhang et al., 2024b) . The intermediate observations —failed tests, retrieved file contents, stack traces—supply structural signal that no prompt-only router can access.  \nSWE-Router, our value-based temporal routing framework, operationalises this insight (Figure 1): a cheap weak model m 1 runs for a few exploratory turns, then a learned value head reads the resulting partial trajectory and predicts whether m 1 will eventually solve the task; if this prediction exceeds a cost-adjusted threshold, m 1 continues, otherwise we escalate to a stronger m2 . The value head is supervised by binary trajectory rewards ","cbCaiaPhF03PHmHX","https://ap.wps.com/l/cbCaiaPhF03PHmHX","pdf",670437,6,1,10,"English","en",105,"# Introduction\n# Preliminary\n## Multi-turn Trajectories in Agentic Tasks","[{\"question\":\"What theoretical guarantee does the document provide for conditioning on the partial trajectory?\",\"answer\":\"It provides a Bayes-optimality theorem stating that conditioning on the partial trajectory never harms routing performance and is strictly better when the exploration is informative.\"}]",1784175768,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":79,"head_meta":81,"extra_data":83,"updated_unix":28},"swe-router-routing-in-multi-turn-agentic-software-engineering-tasks","",{"@graph":36,"@context":78},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/swe-router-routing-in-multi-turn-agentic-software-engineering-tasks/81742/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72],{"name":73,"@type":74,"acceptedAnswer":75},"What theoretical guarantee does the document provide for conditioning on the partial trajectory?","Question",{"text":76,"@type":77},"It provides a Bayes-optimality theorem stating that conditioning on the partial trajectory never harms routing performance and is strictly better when the exploration is informative.","Answer","https://schema.org",{"og:url":52,"og:type":80,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":82,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":85},[86,90,94,98,103,107,112,115,120,123,126],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":87,"show_sort_weight":88,"slug":89},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":91,"show_sort_weight":92,"slug":93},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Exam",70,"exam",{"id":99,"doc_module":4,"doc_module_name":46,"category_name":100,"show_sort_weight":101,"slug":102},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":99,"slug":129},19,"General","general"]