[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83155-en":3,"doc-seo-83155-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83155,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI","Agentic AI systems often follow a pattern of token maxing: spending increasing tokens on longer reasoning traces, more turns, wider tool payloads, and larger replayed contexts so tokens grow faster than task value, while falling per-token prices hide the underlying inefficiency. A controlled evaluation across six foundation models and 22 enterprise tasks isolates the orchestration layer (harness) and shows 41% lower cost per task, 44% faster runtime, and 38% fewer tokens at parity task quality.","arXiv :2607 .06906v 1 [ cs .AI] 8 Jul 2026  \nThe Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI  \nMuayad Sayed Ali, Aliaksandra Novik, Anji Boddupally, Artem Yavorskyi, Chris Nickerson, Daniel Rica, Emily DuGranrut, Felix Leung, Garrett Prince, Grace Barnett, Heath Robinson, Hosain Al Ahmad, Jesse Resnick, Juan Carlos Farah, Jyothi Swaroop Meruga, Leonid Kuznetsov, Brock Perry, Luke Gorham, Marie Schmoll, Michael Paciullo, Saumya Das, Sharath Sheripally, Tommy Griscom, Mykyta Osadchyi, Neha Mantri, Nick Westrum, Olivia Benowitz, Parikshith Kulkarni, Radik Chernyshov, Rakshith Vasudev, Rohith Nadimpally, Vikas Gangadevi, and Waseem AlShikh  \nWriter, Inc.  \n{muayad, aliaksandra, anji, artem, chris, daniel, emily, felix, garrett, grace, heath, hosain, jesse, juan, jyothi, leonid, brock, luke, marie, michael, saumya, sharath, tommy, mykyta, neha, nick, olivia, parikshith, radik, rakshith, rohith,  \nvikas, [waseem](waseem}@writer.com)[}](waseem}@writer.com)[@writer.com](waseem}@writer.com)  \nJuly 2026  \nAbstract  \nThe dominant pattern in agentic AI development is what we call token maxing: buying capability with tokens—longer reasoning traces, more agent turns, wider tool payloads, larger replayed contexts—so that tokens per task grow faster than task value. Falling per-token prices mask the pattern without fixing it; total spend rises anyway. We argue that the decisive lever against token maxing is the harness: the orchestration layer that assembles context, exposes tools, sequences turns, delegates work, and carries the observability and governance surface an enterprise deployment runs on. To isolate this layer we run a controlled swap: the same 22 locked evaluation tasks on the same six foundation models (Claude Sonnet 4.6, Gemini 3.1, Gemini Flash 3.5, Qwen 3.6, GLM 5.1, and Palmyra X6), changing only the orchestration layer: a conventional production agent loop (the frozen baseline) versus the Writer Agent Harness. Holding models constant, placing the harness at the core of execution cuts blended cost per task by 41%($0.21 → $0.12), median wall-clock by 44%(48 s → 27 s), and tokens per task by 38%(14.2k → 8.8k), while headline task-completion quality holds at parity (0.78 → 0.81, directional at this sample size) . The efficiency gains are model-invariant—every model gets cheaper, by 33% to 61%—while quality gains are capability-dependent: the improvement a model extracts from the harness correlates almost perfectly with its baseline strength (r = 0.99, n = 6), a phenomenon we term harness leverage. Quality per dollar rises 82% and task-completions per million tokens rise from 54.9 to 92.0 . On this workload, the orchestration layer moved cost per task more than switching between the cheapest and most expensive model did. We formalize token economics at the orchestration layer, including an effective-input-price model under prompt caching; define token maxing; detail the six mechanism families behind the effect, from cache-shape discipline to failure-spend governance; compare six widely used agent systems on the same axes; and argue that the harness is the one component whose efficiency multiplies across every model an organization runs—present and future.  \n1 Introduction  \nAn agentic task is not one model call. A single request—“reconcile these two contracts and draft the redline memo”—unfolds into a dozen or more turns: system prompt, tool schemas, retrieval payloads, intermediate reasoning, tool outputs, and, in naive implementations, the full replay of everything above on every subsequent turn. The token bill for the task is the sum over that loop, and the loop is governed not by the model but by the software around it. We call that software the harness: the orchestration layer that decides what enters the context window, which tools are visible, when to retrieve, when to retry, when to delegate, and when to stop.  \nThe industry’s default response to rising agent capability requi","cbCail5G8gXFY7LB","https://ap.wps.com/l/cbCail5G8gXFY7LB","pdf",320793,3,1,21,"English","en",105,"# Abstract\n# 1 Introduction\n## Token maxing as a Jevons dynamic\n## Harness as the orchestration layer\n## Natural experiment design and evaluation approach","[{\"question\":\"What is “token maxing” in agentic AI development?\",\"answer\":\"Token maxing is the practice of buying more capability by spending more tokens—longer reasoning, more turns, larger contexts and tool payloads—so token use grows faster than task value. It is often masked by falling per-token prices while total spend rises anyway.\"},{\"question\":\"How does the paper test whether the orchestration “harness” drives token efficiency?\",\"answer\":\"The study uses a controlled swap: it keeps 22 locked evaluation tasks and six foundation models constant, and changes only the orchestration layer between a conventional agent loop and the Writer Agent Harness. Performance is measured on cost per task, wall-clock time, token usage, and task completion quality.\"},{\"question\":\"What results show that the harness improves token economics without harming quality?\",\"answer\":\"With models held constant, the harness reduces blended cost per task by 41%, median wall-clock time by 44%, and tokens per task by 38%, while task completion quality remains at parity on the reported sample. The paper also finds efficiency gains are model-invariant, while quality gains correlate with baseline model strength (harness leverage).\"}]",1784185656,53,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"the-harness-effect-how-orchestration-design-sets-the-token-economics-of-enterprise-agentic-ai","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/the-harness-effect-how-orchestration-design-sets-the-token-economics-of-enterprise-agentic-ai/83155/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is “token maxing” in agentic AI development?","Question",{"text":75,"@type":76},"Token maxing is the practice of buying more capability by spending more tokens—longer reasoning, more turns, larger contexts and tool payloads—so token use grows faster than task value. It is often masked by falling per-token prices while total spend rises anyway.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper test whether the orchestration “harness” drives token efficiency?",{"text":80,"@type":76},"The study uses a controlled swap: it keeps 22 locked evaluation tasks and six foundation models constant, and changes only the orchestration layer between a conventional agent loop and the Writer Agent Harness. Performance is measured on cost per task, wall-clock time, token usage, and task completion quality.",{"name":82,"@type":73,"acceptedAnswer":83},"What results show that the harness improves token economics without harming quality?",{"text":84,"@type":76},"With models held constant, the harness reduces blended cost per task by 41%, median wall-clock time by 44%, and tokens per task by 38%, while task completion quality remains at parity on the reported sample. The paper also finds efficiency gains are model-invariant, while quality gains correlate with baseline model strength (harness leverage).","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]