[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86404-en":3,"doc-seo-86404-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86404,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","Selective Safety Steering via Value-Filtered Decoding","Large language models can violate safety constraints even after alignment training, so decoding-time steering methods have been developed to improve safety by altering token sampling using a safety reward. Existing approaches may intervene unnecessarily, changing outputs that would have been safe under the base model and harming helpfulness, fluency, style, and coherence. This work proposes value-filtered decoding, which filters tokens with a value-based safety criterion and provides an explicit bound on false interventions. A single threshold controls the trade-off between unnecessary intervention rates and improved safety.","arXiv :2605 . 14746v2 [ cs .LG] 12 Jul 2026  \nSelective Safety Steering via Value-Filtered Decoding  \nBat-Sheva Einbinder∗1, Hen Davidov∗2,3, Yee Whye Teh2 , Yarin Gal3 , and Yaniv Romano 1,4  \n1Department of Electrical and Computer Engineering, Technion IIT, Israel  \n2Department of Statistics, University of Oxford, Oxford, UK 3 OATML, Department of Computer Science, University of Oxford, Oxford, UK 4Department of Computer Science, Technion IIT, Israel  \nAbstract  \nWhile large language models (LLMs) are trained to align with human values, their generations may still violate safety constraints. A growing line of work addresses this problem by modifying the model’s sampling policy at decoding time using a safety reward. However, existing decoding-time steering methods often intervene unnecessarily, modifying generations that would have been safe under the base model. Such unnecessary interventions are undesirable, as they can distort key properties of the base model such as helpfulness, fluency, style, and coherence. We propose a new test-time steering method designed to reduce such unnecessary interventions while improving the safety of unsafe responses. Our approach filters tokens using a value-based safety criterion and provides an explicit bound on the probability of false interventions. A single threshold hyperparameter controls this bound, allowing practitioners to trade off higher rates of unnecessary intervention for better output safety. Across multiple datasets and experiments, we show that our value-filtered decoding method outperforms existing baselines, achieving better trade-offs between safety, helpfulness, and similarity to the base model.  \n1 Introduction  \nLLMs have demonstrated remarkable capabilities across a wide range of tasks, but their deployment in real-world applications requires reliable alignment with human values, intended user goals and safety constraints. The leading approach to LLM alignment is preference-based post-training, where supervised fine-tuning is followed by reinforcement learning from human feedback (RLHF) or closely related methods (1–10) . While effective, fine-tuning is costly, can reduce general capabilities (11, 12), may damage alignment through catastrophic forgetting (13), and must be repeated whenever the objective or base model changes (14, 15) .  \nThis paper focuses on improving the safety of LLM generations via inference time intervention. Inference-time alignment methods steer the decoding policy to balance between safety and similarity to the base model policy (14– 24) . The steering is often done by reweighting the decoding distribution using a reward or safety signal. In this paper, we propose an additional test-time steering method that, as we explain next, offers several advantages over existing alternatives.  \nThe main limitation of most existing decoding-time steering methods is that they provide no guarantee on the rate of unnecessary interventions, that is, interventions on generations that would have been safe under the base model. Any intervention necessarily changes the output relative to the base model’s generation, thus may harm desirable properties not captured by the safety objective, such as helpfulness, coherence, and style. Our goal, therefore, is topreserve the base model’s outputs as much as  possible while ensuring that every generated sequence is safe.  \n*Equal contribution. Order decided by coin flip.  \nIn this paper, we propose a new decoding-time steering method, value-filtered decoding, which samples from a policy that requires each generated token to satisfy a safety constraint. We do this by using a classifier that predicts the value function, defined as the expected safety of the full generation given the prompt and the current partial response. Then, we restrict the policy to sample only tokens whose value exceeds some threshold. We show that this value-filtered policy achieves higher expected safety than the base model. Moreover, we obtain a ","cbCaivqJN5JcL6k9","https://ap.wps.com/l/cbCaivqJN5JcL6k9","pdf",3081555,6,1,40,"English","en",105,"# Abstract\n# Introduction\n## Decoding-time alignment and limitations\n## Proposed value-filtered decoding\n# Notations","[{\"question\":\"Why do decoding-time steering methods sometimes hurt model quality?\",\"answer\":\"They can make interventions even when the base model would have produced a safe output. Any intervention changes the generation and may degrade properties not directly targeted by the safety objective, such as helpfulness, coherence, or style.\"},{\"question\":\"How does value-filtered decoding reduce unnecessary interventions?\",\"answer\":\"It samples using a policy that only allows tokens whose predicted value, defined as expected safety of the full generation given the prompt and current partial response, exceeds a threshold. This prevents disrupting generations that would otherwise remain safe.\"},{\"question\":\"What does the threshold control in the proposed method?\",\"answer\":\"The threshold directly controls the statistical rate of unnecessary interventions (false interventions). By tuning it, users trade off output safety against fidelity to the base model.\"}]",1784211532,101,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"selective-safety-steering-via-value-filtered-decoding","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/selective-safety-steering-via-value-filtered-decoding/86404/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why do decoding-time steering methods sometimes hurt model quality?","Question",{"text":76,"@type":77},"They can make interventions even when the base model would have produced a safe output. Any intervention changes the generation and may degrade properties not directly targeted by the safety objective, such as helpfulness, coherence, or style.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does value-filtered decoding reduce unnecessary interventions?",{"text":81,"@type":77},"It samples using a policy that only allows tokens whose predicted value, defined as expected safety of the full generation given the prompt and current partial response, exceeds a threshold. This prevents disrupting generations that would otherwise remain safe.",{"name":83,"@type":74,"acceptedAnswer":84},"What does the threshold control in the proposed method?",{"text":85,"@type":77},"The threshold directly controls the statistical rate of unnecessary interventions (false interventions). By tuning it, users trade off output safety against fidelity to the base model.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":22,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":107,"slug":137},19,"General","general"]