[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86426-en":3,"doc-seo-86426-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86426,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent","Agents-A1 is a 35B Mixture-of-Experts agentic model designed to reach trillion-parameter-level performance by scaling the agent horizon. The work studies horizon scaling through long-horizon trajectory growth and heterogeneous agent ability scaling. It introduces a long-horizon knowledge-action infrastructure that links external knowledge, actions, observations, and verifier outcomes to generate agentic trajectories averaging 45K tokens. Agents-A1 is trained via a three-stage recipe: full-domain SFT, domain-level teacher training, and multi-teacher domain-routed on-policy distillation with vocabulary alignment.","arXiv :2606 .306 16v2 [ cs .CL] 13 Jul 2026  \nScaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent  \nAgents-A1 Team, Shanghai Artificial Intelligence Laboratory  \nWe introduce Agents-A1, a 35B Mixture-of-Experts Agentic Model that reaches trillion-parameterlevel performance by scaling the agent horizon. We investigate agent-horizon scaling from two perspectives: scaling long-horizon trajectories and scaling heterogeneous agent abilities. To support this goal, we build a long-horizon knowledge-action infrastructure that connects external knowledge, actions, observations, and verifier outcomes, producing agentic trajectories with an average length of 45K tokens. Based on this, we train Agents-A1 with a three-stage recipe. First, we perform full-domain supervised fine-tuning to align the base model with broad agentic behaviors. Second, we train domain-level teacher models to capture specialized expertise in each domain. Third, we propose a multi-teacher domain-routed on-policy distillation with salient vocabulary alignment to improve knowledge transfer efficiency across different domains, unifying six heterogeneous domains into one deployable student model.  \nAgents-A1 achieves strong and broad performance for long-horizon agent benchmarks. Compared with 1T-parameter model such as Kimi-K2.6 and DeepSeek-V4-pro, Agents-A1 achieves leading results on SEAL-0 (56.4), IFBench (80.6), HiPhO (46.4), FrontierScience-Olympiad (79.0), and MolBench-Bind (56.8), and remains highly competitive on SciCode (44.3), HLE (47.6) and BrowseComp (75.5). We hope this work provides the community with a practical path for scaling the horizon using a 35B agent that can reach or match the performance of 1T models on long-horizon tasks.  \n Code  Model  \n Agents-A1  Qwen3.6-35B-A3B  Step-3 .5-Flash  Kimi-K2.6  DeepSeek-V4-pro(Max)  gpt-5 .5(xhigh)  \n100  \n80  \n60  \n40  \n20  \n0  \n100  \n80  \n60  \n40  \n20  \n0  \n100  \n80  \n60  \n40  \n20  \n0  \n100  \n80  \n60  \n40  \n20  \n0  \n100  \n80  \n60  \n40  \n20  \n0  \n100  \n80  \n60  \n40  \n20  \n0  \n100  \n80  \n60  \n40  \n20  \n0  \n100  \n80  \n60  \n40  \n20  \n0  \n100  \n80  \n60  \n40  \n20  \n0  \nFrontierScience-Olympiad  \nSEAL-0  \nSciCode  \n100  \n80  \n60  \n40  \n20  \n0  \n100  \n80  \n60  \n40  \n20  \n0  \n100  \n80  \n60  \n40  \n20  \n0  \nFrontierScience-Research  \nGAIA  \nMolBench-Bind  \nFigure 1 | Benchmark performance of Agents-A1  \nContents  \n1 Introduction 3  \n2 Knowledge-Guided General Agent Training with Specialized Teachers 4  \n2.1 Overview 4  \n2.2 Long-Horizon Knowledge-Action Infrastructure 5  \n2.2.1 Knowledge-Action Graph Construction with Atomic Abilities 5  \n2.2.2 Self-play Graph Search and Expansion 6  \n2.3 Domain-Routed On-Policy Distillation with Salient Vocabulary Alignment 6  \n2.3.1 Salient Vocabulary Alignment 7  \n2.3.2 Domain-routed Normalized Objective 7  \n3 Multi-domain Data Pipeline 8  \n3.1 Long-horizon Search 8  \n3.2 Machine Learning Engineering 9  \n3.3 Scientific Reasoning and Research 9  \n3.4 Instruction Following 11  \n3.5 Tool Calling 12  \n4 Three-stage Training Recipe 12  \n4.1 Full-domain Supervised Fine-Tuning 12  \n4.2 Domain-level Teacher Training 14  \n4.2.1 Reinforcement Learning on Search Tasks 14  \n4.2.2 Science-enhanced Supervised Fine-Tuning 15  \n4.2.3 Reinforcement Learning on Instruction Following 16  \n4.2.4 Reinforcement Learning on Tool-calling 16  \n4.3 Multi-teacher On-Policy Distillation 17  \n5 Experimental Results 18  \n5.1 Evaluation Setting 18  \n5.2 Results and Observations 20  \n5.2.1 Full-domain SFT Results 20  \n5.2.2 Results of Domain Teacher Training 20  \n5.2.3 Results of On-policy Distillation 21  \n5.3 Long-Horizon Task Applications 23  \n5.3.1 A 12-Hour Long-Horizon Optimization Run 23  \n5.3.2 Closing the Loop in Earth Science Analysis with Agents-A1 24  \n6 Limitation and Future Work 25  \nReferences 26  \nA Appendix 29  \nA.1 Contributions and Acknowledgments 29  \n1. Introduction  \nRecent progress in LLMs [1, 2, 3, 4, 5, 6] is rapidly pushing AI from passive la","cbCaib5gGDfbjmcY","https://ap.wps.com/l/cbCaib5gGDfbjmcY","pdf",9021449,4,1,29,"English","en",105,"# Introduction\n# Knowledge-Guided General Agent Training with Specialized Teachers\n## Overview\n## Long-Horizon Knowledge-Action Infrastructure\n## Domain-Routed On-Policy Distillation with Salient Vocabulary Alignment\n# Multi-domain Data Pipeline\n# Three-stage Training Recipe\n# Experimental Results\n# Limitation and Future Work\n# References","[{\"question\":\"What is the main idea behind Agents-A1’s performance gains?\",\"answer\":\"Agents-A1 targets trillion-parameter-level capability by scaling the agent horizon rather than solely increasing model size. It focuses on longer decision trajectories and improved transfer across heterogeneous abilities.\"},{\"question\":\"How does the paper enable long-horizon training in practice?\",\"answer\":\"It builds a long-horizon knowledge-action infrastructure that connects external knowledge, agent actions, observations, and verifier outcomes. This produces grounded agentic trajectories with an average length of about 45K tokens.\"},{\"question\":\"What are the three stages of the Agents-A1 training recipe?\",\"answer\":\"The training includes: full-domain supervised fine-tuning for broad agent behaviors, domain-level teacher model training to capture specialized expertise, and multi-teacher domain-routed on-policy distillation with salient vocabulary alignment to improve knowledge transfer.\"}]",1784211674,73,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"scaling-the-horizon-not-the-parameters-reaching-trillion-parameter-performance-with-a-35b-agent","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/scaling-the-horizon-not-the-parameters-reaching-trillion-parameter-performance-with-a-35b-agent/86426/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main idea behind Agents-A1’s performance gains?","Question",{"text":75,"@type":76},"Agents-A1 targets trillion-parameter-level capability by scaling the agent horizon rather than solely increasing model size. It focuses on longer decision trajectories and improved transfer across heterogeneous abilities.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper enable long-horizon training in practice?",{"text":80,"@type":76},"It builds a long-horizon knowledge-action infrastructure that connects external knowledge, agent actions, observations, and verifier outcomes. This produces grounded agentic trajectories with an average length of about 45K tokens.",{"name":82,"@type":73,"acceptedAnswer":83},"What are the three stages of the Agents-A1 training recipe?",{"text":84,"@type":76},"The training includes: full-domain supervised fine-tuning for broad agent behaviors, domain-level teacher model training to capture specialized expertise, and multi-teacher domain-routed on-policy distillation with salient vocabulary alignment to improve knowledge transfer.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]