[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86167-en":3,"doc-seo-86167-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86167,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method","Amplitude-Gating (AG) is proposed to improve tool-structured large language model outputs during inference without retraining weights. Motivated by limitations of ORP, which can harm samples by overwriting pretrained FFN directions, AG preserves FFN weight directions and modulates only activation amplitudes. A fine-grained SwiGLU intervention system with deployable sample-level gates is defined and evaluated via an oracle-headroom–separating protocol using task-aware binary/partial-credit metrics. Experiments across Qwen-family models show weak aggregate gains but strongest improvements on tool/structured/agentic tasks, including category-level learned gates and category-specific routing. ","Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method: Gated Evaluation Protocol, and Cross-Model Empirical Results  \nSheng Xu  \n[xusheng.xs@alibaba-inc.com](xusheng.xs@alibaba-inc.com)  \nBoyuan Huang  \n[boyuan.hby@alibaba-inc.com](boyuan.hby@alibaba-inc.com)  \nKe Jia  \n[jiake.jk@alibaba-inc.com](jiake.jk@alibaba-inc.com)[ ](jiake.jk@alibaba-inc.com)Alibaba Cloud  \nJiadun Zhu  \n[jifeng.zjd@taobao.com](jifeng.zjd@taobao.com)  \nZhen Chen  \n[moriarty.cz@alibaba-inc.com](moriarty.cz@alibaba-inc.com)  \nAbstract  \nLarge language models increasingly operate as tool-using agents, where small format, argument, or function-call errors can invalidate an otherwise plausible response. We study inference-time feed-forward network (FFN) intervention as away to improve structured outputs without retraining model weights. The project began with Orthogonal Residual Projection (ORP), an early project-specific inference-time repair attempt that used SVD-derived orthogonal projection to modify residual-stream FFN representations. ORP was scientifically useful but not deployable. It exposed useful structure inside SwiGLU FFNs, including sensitive intervention positions P1, P2, and P3 and non-monotonic dependence on an energy parameter e. However, diagnostic runs showed that harm could exceed recovered fixes, suggesting that pretrained FFN directions encode semantic priors that should not be overwritten at inference time. We therefore propose Amplitude Gating (AG), a non-destructive alternative that preserves pretrained FFN weight directions and modulates only activation magnitudes during generation. We define a fine-grained intervention system for SwiGLU FFNs, covering inherited points P1/P2/P3 and structurally refined points P1s/P2a/P2b, and evaluate entropy-based AG candidates with deployable sample-level gates. A central methodological contribution is an evaluation protocol that separates combination-oracle headroom from fixed configurations and learned gates constrained to online-observable features, avoids pooled multi-configuration accounting, and uses task-aware metrics for strict binary versus partial-credit datasets. Across Qwen-family experiments spanning a hybrid linear/full-attention model (Qwen3 .5-9B) and dense full-attention models (Qwen3-8B and Qwen2 .5-7B), AG is weakly positive in aggregate but strongest on tool-structured tasks. On Qwen3 .5-9B, a category-level learned gate improves tool/structured/agentic tasks from 38.66% to 42.92%(+4 .27 percentage points), with Hermes function-call tasks reaching about +7.6 pp learned-gate lift. On Qwen3-8B, the strongest signal shifts to JSON-structured output, where Hermes JSON mode improves from 41.83% to 53. 19%(+11 .36 pp) . AQwen2 .5-7B follow-up shows that oracle headroom can exist even when current learned gates fail, indicating that AG should be deployed with model-and category-specific routing rather than a universal gate. Fixed-pool comparisons between non-Newton-Schulz entropy AG and Newton-Schulz-windowed AG further show that neither family is uniformly dominant: non-NS variants are sharper current-token amplitude gates, whereas NS variants add local-trajectory diversity whose value changes with model architecture and task category. These results identify tool-structured inference as the most credible first deployment target for safe FFN-level inference optimization. New feature-family ablations with bootstrap confidence intervals support the Qwen3/Qwen3 .5 tool-route gains while confirming Qwen2 .5-7B as a boundary case rather than a deployable positive route; prospective online validation, broader cross-model generalization, and independent overfitting checks remain required follow-up work.  \n1. Introduction  \nTool use changes the error surface of large language models (LLMs) [7] . In ordinary open-ended generation, an answer may remain useful even when phrasing or some minor detail changes. In tool-calling and agentic settings, however, the output is often c","cbCaioemsuiCnSHn","https://ap.wps.com/l/cbCaioemsuiCnSHn","pdf",2094163,3,1,28,"English","en",105,"# Abstract\n# 1. Introduction","[{\"question\":\"Why is tool-structured inference an important target for inference-time optimization?\",\"answer\":\"Tool-structured outputs are consumed by parsers, validators, APIs, or downstream programs. Small formatting or argument errors (e.g., invalid JSON or wrong function endpoints) can turn plausible text into failed actions, so minor inference-time improvements can directly reduce tool-call failures.\"},{\"question\":\"What problem does the paper address with ORP, and how does AG differ?\",\"answer\":\"ORP uses SVD-derived orthogonal projections to modify FFN residual-stream representations, which can sometimes repair outputs but may exceed recovered fixes and cause harm. AG addresses this by preserving pretrained FFN weight directions and only modulating activation magnitudes during generation, making the intervention non-destructive.\"},{\"question\":\"What is the paper’s evaluation protocol designed to achieve?\",\"answer\":\"The protocol separates combination-oracle headroom from fixed configurations and learned gates, avoiding pooled multi-configuration accounting. It constrains learned gates to online-observable features and uses task-aware metrics for strict binary versus partial-credit datasets to assess interventions more rigorously.\"}]",1784209047,71,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"amplitude-only-ffn-intervention-for-tool-structured-llm-inference-method","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/amplitude-only-ffn-intervention-for-tool-structured-llm-inference-method/86167/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is tool-structured inference an important target for inference-time optimization?","Question",{"text":75,"@type":76},"Tool-structured outputs are consumed by parsers, validators, APIs, or downstream programs. Small formatting or argument errors (e.g., invalid JSON or wrong function endpoints) can turn plausible text into failed actions, so minor inference-time improvements can directly reduce tool-call failures.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What problem does the paper address with ORP, and how does AG differ?",{"text":80,"@type":76},"ORP uses SVD-derived orthogonal projections to modify FFN residual-stream representations, which can sometimes repair outputs but may exceed recovered fixes and cause harm. AG addresses this by preserving pretrained FFN weight directions and only modulating activation magnitudes during generation, making the intervention non-destructive.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the paper’s evaluation protocol designed to achieve?",{"text":84,"@type":76},"The protocol separates combination-oracle headroom from fixed configurations and learned gates, avoiding pooled multi-configuration accounting. It constrains learned gates to online-observable features and uses task-aware metrics for strict binary versus partial-credit datasets to assess interventions more rigorously.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]