[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86274-en":3,"doc-seo-86274-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86274,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","Interaction Scaling Grounding the Third Axis of Test Time Compute","Interaction scaling proposes a third inference-time compute axis beyond reasoning-longer and best-of-N sampling. A frozen model iteratively proposes an artifact, then an external instrument observes its real behavior, and the model revises based on grounded feedback. The gain depends on grounding on both sides: the feedback signal must come from true observation and the evaluation metric must score that same observable outcome. Reasoning and sampling plateau under fixed budgets, while grounded interaction improves without variance and fixes layout defects that screenshot-based judges miss.","arXiv :2607 . 1 1598v 1 [ cs .AI] 13 Jul 2026  \nInteraction Scaling:  \nGrounding the Third Axis of Test-Time Compute  \nBojie Li Noah Shi  \nPine AI University of Washington  \nAbstract  \nThere are two standard ways to spend more compute at test time: let a model reason longer, or sample more attempts and keep one. Both share a hidden limit: they are internal. Every extra token comes from the same frozen weights and the same prompt, so neither can tell the model anything it does not already know. We study a third way, interaction: the model proposes an artifact, an external instrument observes how it actually behaves, and the model revises. Each cycle imports a real observation, so interaction breaks through the ceiling the other two hit.  \nWe argue that a single variable governs this third axis, grounding, and that it must hold on both sides of the loop. The feedback that drives revision must come from an instrument that actually observes the flaw, and so must the metric that scores the result. On hard coding tasks at a fixed token budget, reasoning-only and best-of-N sampling both plateau (the latter even when an oracle picks the best sample), while every interaction strategy keeps improving; our proposer–reviewer harness reaches a perfect 100% pass rate with no run-to-run variance, and the gain holds across three model families. On rendered visual artifacts, the usual judge (a vision–language model, or VLM, reading a screenshot) rates 14 of 15 visibly broken figures “perfect,” because the screenshot hides the flaws before the judge can see them. A tool that measures the real layout instead shows the loop removing 40–74% of defects across four modalities; and that same VLM, used as the reviewer, makes slide layouts worse where the measuring tool repairs them. Interaction scaling is real and distinct from reasoning and sampling, but only visible when both the feedback and the metric are grounded.  \nCode: [https://github.com/19PINE-AI/interaction-scaling](https://github.com/19PINE-AI/interaction-scaling)[ ](https://github.com/19PINE-AI/interaction-scaling)Website: [https://01.me/research/interaction-scaling](https://01.me/research/interaction-scaling)  \nGROUNDED: measure the rendered DOM / run pytest  \ntargeted revision  \nVLM reads a screenshot  \nblind feedback & blind score  \nUNGROUNDED: overflow cropped off-frame, rates broken figures “perfect”  \nFigure 1. Grounding must hold on both sides of the interaction loop. One instrument observation is both (1) grounded feedback (the defect list that drives revision and escapes the internal ceiling) and (2) grounded evaluation, the score that makes the gain measurable. The default VLM-on-a-screenshot judge (orange lane) breaks both: the screenshot drops the defects before the model sees them.  \n1 Introduction  \nOnce a model is trained, the way we make it better on a hard task is no longer to add parameters or data (training-time scaling) but to spend more compute per query on the frozen model: inference-time scaling (OpenAI, 2024; Snell et al., 2025) . On the artifact-producing tasks we study (code, web pages, slides, figures, animations, video edits, research reports), this is the lever we can still pull. Prior work formalizes two ways to pull it. Reasoning scaling (Wei et al., 2022; Yao et al., 2023a; DeepSeek-AI, 2025; Snell et al., 2025) spends more tokens thinking before committing; sampling scaling (Wang et al., 2023b; Brown et al., 2024) draws more attempts and selects one. They look different, but share a property this paper treats as fundamental: both are internal. Every extra token, whether in a longer chain of thought or another sample, comes from the same frozen weights and the same fixed prompt. More internal compute reshuffles what the model already has; it imports nothing new.  \nWe study a third form of inference-time scaling that breaks out of this closed loop: interaction scaling, in which the model queries an external instrument that observes the artifact itself. The m","cbCaipV3Pr032PaY","https://ap.wps.com/l/cbCaipV3Pr032PaY","pdf",1780026,7,1,25,"English","en",105,"# Introduction\n## Inference-time scaling and internal limits\n## Interaction scaling and grounded feedback\n## Grounded evaluation metric","[{\"question\":\"Why do reasoning scaling and best-of-N sampling hit an internal ceiling at test time?\",\"answer\":\"They spend more compute but remain closed to the same prompt and frozen weights, so extra tokens only reshuffle internal beliefs without importing new information about how the artifact actually behaves.\"},{\"question\":\"What is interaction scaling in this work?\",\"answer\":\"The model proposes an artifact and then queries an external instrument that observes the artifact’s real behavior, using those observations to revise in a proposer–reviewer loop.\"},{\"question\":\"What does “grounding” mean and why must it hold on both sides of the loop?\",\"answer\":\"Grounded feedback must come from an instrument that observes the artifact’s actual form or behavior, and grounded evaluation must use the same observable signal for scoring; otherwise the gain becomes invisible or even counterproductive.\"}]",1784209973,63,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"interaction-scaling-grounding-the-third-axis-of-test-time-compute","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/interaction-scaling-grounding-the-third-axis-of-test-time-compute/86274/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why do reasoning scaling and best-of-N sampling hit an internal ceiling at test time?","Question",{"text":76,"@type":77},"They spend more compute but remain closed to the same prompt and frozen weights, so extra tokens only reshuffle internal beliefs without importing new information about how the artifact actually behaves.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is interaction scaling in this work?",{"text":81,"@type":77},"The model proposes an artifact and then queries an external instrument that observes the artifact’s real behavior, using those observations to revise in a proposer–reviewer loop.",{"name":83,"@type":74,"acceptedAnswer":84},"What does “grounding” mean and why must it hold on both sides of the loop?",{"text":85,"@type":77},"Grounded feedback must come from an instrument that observes the artifact’s actual form or behavior, and grounded evaluation must use the same observable signal for scoring; otherwise the gain becomes invisible or even counterproductive.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]