[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83082-en":3,"doc-seo-83082-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83082,1374391975076,"Riley","https://ap-avatar.wpscdn.com/avatar/14000253ca4ec9f6853?x-image-process=image/resize,m_fixed,w_180,h_180&k=1783305029341752051",8,"Research & Report","DT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety Guardrail","Large language models in open-world deployments need safety guardrails that handle complex risks reliably while meeting low-latency runtime requirements. Existing approaches trade off between efficient classification-based models that struggle with concealed intent and ambiguous judgments, and reasoning-based guards that improve quality but add inference overhead. DT-Guard uses Reasoning-Active Training with Reasoning-Free Inference: reasoning supervision during training, structured safety labels only at inference. It models safety as Intent → Category → Safety and includes Rollout-Guided Progressive Hard-Case Optimization (RG-PHO) for robustness, achieving strong F1 across prompt- and response-side benchmarks.","arXiv :2607 .06326v 1 [ cs .AI ] 7 Jul 2026  \nDIGITAL  \nTECHNOLOGIES  \nDT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety Guardrail  \nHe Liu∗, Changtao Miao∗, Xinjie Yang∗, Tianle Song, Yin Wu, Junchi Chen, Bintao He, Xinyuan Zhang, Bo Zhang†, Shi Yan, Wei Lu, Wei Wang, Danyang Xu, Jiansheng Cai, Zhe Li  \nAnt Digital Technologies, Ant Group  \nAbstract  \nLarge language models deployed in open-world applications require safety guardrails that are both robust to complex risks and efficient enough for low-latency runtime moderation. Existing guardrails face a practical trade-off between lightweight classificationbased models, which are efficient but often struggle with concealed intent, ambiguous semantics, and borderline safety decisions, and reasoning-based guards, which improve judgment quality but introduce additional token generation and inference latency. We present DT-Guard, a content safety guardrail model based on a ReasoningActive Training, Reasoning-Free Inference paradigm. The key idea is to use reasoning supervision during training while emitting only structured safety labels at inference time. DT-Guard formulates safety judgment as a progressive decision process, Intent → Category → Safety, and constructs an intent-driven dataset with intent labels, risk categories, safety labels, and structured reasoning trajectories. To further improve hard-case robustness, we propose Rollout-Guided Progressive Hard-Case Optimization (RG-PHO), which uses multi-rollout consistency to identify stably mastered, persistently failed, and preference-unstable samples, and applies targeted supervised and preference optimization accordingly. At inference time, DT-Guard directly generates structured labels without explicit reasoning traces, preserving deployment efficiency. Experiments on prompt-side and response-side safety benchmarks show that DT-Guard achieves average F1 scores of 0.886 and 0.870, respectively. With only a 4B backbone, it reaches a dual-side average F1 of 0.878, outperforming strong 8B guardrail baselines. These results demonstrate that reasoning supervision can be effectively internalized into low-latency safety discrimination.  \n1. Introduction  \nLarge language models (LLMs) have achieved substantial progress in instruction following, knowledge-intensive question answering, complex reasoning, and multi-turn interaction, and are increasingly deployed in open-ended real-world applications [1, 2, 3] . As their deployment scope expands, LLMs must handle diverse user inputs and generate responses across safetysensitive scenarios. User inputs may contain concealed harmful intent, adversarial jailbreak  \n* Equal contribution.  \n†[Corresponding to boyuan.zb@antgroup.com](Corresponding to boyuan.zb@antgroup.com)  \nOur model Qwen3Guard-8B-Gen  \nQwen3Guard-4B-Gen Qwen3Guard-0.8B-Gen YuFeng-XGuard-Reason-0.6B  \nYuFeng-XGuard-Reason-8B Our model YuFeng-XGuard-Reason-8B  \nYuFeng-XGuard-Reason-0.6BQwen3Guard-8B-Gen  \nQwen3Guard-4B-Gen  \nQwen3Guard-0.8B-Gen Our model  \nYuFeng-XGuard-Reason-8B  \nYuFeng-XGuard-Reason-0.6BQwen3Guard-4B-Gen  \nQwen3Guard-8B-Gen Qwen3Guard-0.8B-Gen  \nSorry-Bench  \nOur model Qwen3Guard-8B-Gen YuFeng-XGuard-Reason-8B  \nQwen3Guard-4B-Gen Qwen3Guard-0.8B-Gen YuFeng-XGuard-Reason-0.6B  \nFigure 1: DT-Guard achieves the top F1 on representative safety benchmarks under reasoningfree inference, outperforming strong guardrail baselines.  \nattempts, or ambiguous borderline requests, while model outputs may include unsafe, harmful, biased, or policy-violating content under specific contexts [4, 5] . Safety guardrail models [6, 7, 8, 9] have therefore become an important runtime safety layer for detecting and intercepting risks before user inputs are passed to the model or before model responses are returned to users.  \nExisting guardrail models typically follow one of two inference paradigms. Classificationbased guardrails [6, 7, 10, 11, 12] directly predict safety labels according to predef","cbCaiof2G3idCBE3","https://ap.wps.com/l/cbCaiof2G3idCBE3","pdf",1888354,2,1,15,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What problem does DT-Guard target in LLM safety guardrails?\",\"answer\":\"DT-Guard addresses the gap between lightweight classification-based guardrails that often fail on concealed intent and borderline cases, and reasoning-based guards that improve decisions but increase token generation and latency.\"},{\"question\":\"How does DT-Guard achieve reasoning-free inference?\",\"answer\":\"DT-Guard applies reasoning supervision during training using reasoning trajectories, but during inference it directly generates structured safety labels without explicit reasoning traces.\"},{\"question\":\"What is RG-PHO and why is it used?\",\"answer\":\"RG-PHO (Rollout-Guided Progressive Hard-Case Optimization) uses multi-rollout consistency to identify samples that are stably mastered, persistently failed, or preference-unstable, then applies targeted supervised and preference optimization to improve hard-case robustness.\"}]",1784185076,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"dt-guard-intent-driven-reasoning-active-training-for-reasoning-free-llm-safety-guardrail","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/dt-guard-intent-driven-reasoning-active-training-for-reasoning-free-llm-safety-guardrail/83082/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does DT-Guard target in LLM safety guardrails?","Question",{"text":75,"@type":76},"DT-Guard addresses the gap between lightweight classification-based guardrails that often fail on concealed intent and borderline cases, and reasoning-based guards that improve decisions but increase token generation and latency.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does DT-Guard achieve reasoning-free inference?",{"text":80,"@type":76},"DT-Guard applies reasoning supervision during training using reasoning trajectories, but during inference it directly generates structured safety labels without explicit reasoning traces.",{"name":82,"@type":73,"acceptedAnswer":83},"What is RG-PHO and why is it used?",{"text":84,"@type":76},"RG-PHO (Rollout-Guided Progressive Hard-Case Optimization) uses multi-rollout consistency to identify samples that are stably mastered, persistently failed, or preference-unstable, then applies targeted supervised and preference optimization to improve hard-case robustness.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]