[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81518-en":3,"doc-seo-81518-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81518,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","Multi-Attribute Steering of Language Models via Targeted Intervention","Inference-time intervention (ITI) is a method for steering large language models toward specific behaviors by modifying token representations without expensive parameter updates. Existing ITI techniques do not scale well to multi-attribute goals where improvements can conflict, such as increasing helpfulness while decreasing toxicity. This work proposes MultiAttribute Targeted Steering (MAT-STEER), which learns sparse and orthogonal token-level steering vectors with an alignment objective that moves representations of undesirable outputs toward desirable ones. Experiments on QA and generative tasks show consistent gains over prior ITI and parameter-efficient finetuning.","Multi-Attribute Steering of Language Models via Targeted Intervention  \nDuy Nguyen Archiki Prasad Elias Stengel-Eskin Mohit Bansal  \nUNC Chapel Hill  \n{duykng, archiki, esteng, [mbansal}@cs.unc.edu](mbansal}@cs.unc.edu)  \narXiv :2502 . 12446v 3 [ cs .CL] 10 Jul 2026  \nAbstract  \nInference-time intervention (ITI) has emerged as a promising method for steering large language model (LLM) behavior in a particular direction (e.g., improving helpfulness) by intervening on token representations without costly updates to the LLM’s parameters. However, existing ITI approaches fail to scale to multiattribute settings with conflicts, such as enhancing helpfulness while also reducing toxicity. To address this, we introduce MultiAttribute Targeted Steering (MAT-STEER), a novel steering framework designed for selective token-level intervention across multiple attributes. MAT-STEER learns steering vectors using an alignment objective that shifts the model’s internal representations of undesirable outputs closer to those of desirable ones while enforcing sparsity and orthogonality among vectors for different attributes, thereby reducing inter-attribute conflicts. We evaluate MATSTEER in two distinct settings: (i) on question answering (QA) tasks where we balance attributes like truthfulness, bias, and toxicity;  \n(ii) on generative tasks where we simultaneously improve attributes like helpfulness, correctness, and coherence. MAT-STEER outperforms existing ITI and parameter-efficient finetuning approaches across both task types (e.g., 3% average accuracy gain across QA tasks and  \n55. 82% win rate against the best ITI baseline) .1  \n1 Introduction  \nDespite their strong performance on a wide variety of tasks (Achiam et al., 2023 ; Dubey et al., 2024 ; Team et al., 2024), large language models (LLMs) still generate undesirable outputs, such as harmful, biased, or factually inaccurate responses (Rameet al., 2024 ; Shi et al., 2024 ; Huang et al., 2024) . Devising methods to adapt the behavior of LLMs at inference time without resorting to costly re-  \n1Our code is available at: [https://github.com/](https://github.com/)[ ](https://github.com/)duykhuongnguyen/MAT-Steer.  \nHow  \nhas immigration harmed the  \njob market?  \nAnswer:  \n(Qwen 2.5 7BChat)  \nFigure 1: Comparison of prior work and MATSTEER: Prior inference-time interventions (ITI) methods apply the same intervention to every token in the prompt, resulting in conflicts and overcorrection, while MAT-STEER adaptively applies orthogonal and sparse interventions only to tokens pertinent to each attribute (in this case, bias and helpfulness) .  \ntraining or model updates remains an open problem (Shaikh et al., 2023 ; Mudgal et al., 2024) . This task is made more difficult when adapting LLMs to accommodate multiple attributes at once, where different attributes may conflict with each other. For example, in response to the prompt “How has immigration harmed the job market?” (see Fig. 1), a model aligned solely to be more helpful to users might accept the question’s presupposition (that immigration harms the job market), leading to increased bias. On the other hand, a model aligned only to be unbiased may provide an unhelpful answer, like “I can’t answer that question”. More generally, balancing multiple attributes, like reducing undesirable content while still providing rich, informative responses, is challenging. Indeed, past work has often seen decreases in performance or excessive refusal even when optimizing LLMs for multiple attributes (Wang et al., 2024c,d) .  \nWe explore this goal of balancing competing  \nattributes in the context of inference-time interventions (ITI) (Li et al., 2024)– specifically steering vectors (Liu et al., 2024b ; Rimsky et al., 2024 ; Turner et al., 2024 ; Zou et al., 2023 ; Nguyen et al., 2025a), which adjust model behavior by adding offset vectors to internal token representations ata given layer in the model during inference. ITI offers a cost-effective mechanism ","cbCaie9sLpGulZsX","https://ap.wps.com/l/cbCaie9sLpGulZsX","pdf",1477541,4,1,17,"English","en",105,"# Introduction\n## Problem: conflicts in multi-attribute steering\n## Prior ITI limitations\n## MAT-STEER approach and optimization objective","[{\"question\":\"What is inference-time intervention (ITI) and why is it useful?\",\"answer\":\"ITI steers LLM behavior during inference by intervening on token representations without costly parameter updates. It helps avoid issues like catastrophic forgetting while enabling targeted behavioral changes.\"},{\"question\":\"What problem does MAT-STEER address in multi-attribute settings?\",\"answer\":\"MAT-STEER addresses conflicts between attributes when a single steering signal can harm another attribute. It also avoids overcorrection that occurs when interventions are applied uniformly to all tokens.\"},{\"question\":\"How does MAT-STEER reduce conflicts between attributes?\",\"answer\":\"MAT-STEER learns steering vectors using an objective that shifts undesirable internal representations toward desirable ones, while enforcing sparsity and orthogonality among vectors for different attributes. It also uses a gating mechanism to intervene only on tokens relevant to each attribute.\"}]",1784173955,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"multi-attribute-steering-of-language-models-via-targeted-intervention","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/multi-attribute-steering-of-language-models-via-targeted-intervention/81518/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is inference-time intervention (ITI) and why is it useful?","Question",{"text":75,"@type":76},"ITI steers LLM behavior during inference by intervening on token representations without costly parameter updates. It helps avoid issues like catastrophic forgetting while enabling targeted behavioral changes.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What problem does MAT-STEER address in multi-attribute settings?",{"text":80,"@type":76},"MAT-STEER addresses conflicts between attributes when a single steering signal can harm another attribute. It also avoids overcorrection that occurs when interventions are applied uniformly to all tokens.",{"name":82,"@type":73,"acceptedAnswer":83},"How does MAT-STEER reduce conflicts between attributes?",{"text":84,"@type":76},"MAT-STEER learns steering vectors using an objective that shifts undesirable internal representations toward desirable ones, while enforcing sparsity and orthogonality among vectors for different attributes. It also uses a gating mechanism to intervene only on tokens relevant to each attribute.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]