[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84836-en":3,"doc-seo-84836-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84836,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","BioSecBench-Refusal: Calibrated Refusal and Safe Usefulness in Dual-Use Biology Settings","BioSecBench-Refusal introduces a benchmark for measuring risk identification and refusal behavior in agentic biosecurity evaluation. The benchmark pairs 61 Routine tasks, derived from published life-science research, with 46 Red-Team tasks that use realistic fictional scenarios to conceal a biosecurity hazard. Across 16 model-harness configurations, refusal rates span 7%–74% for Routine tasks and 1%–62% for Red-Team tasks. Refusals frequently stem from provider API filters, while reasoning-enabled models can better surface real threats, supporting safer calibration for agentic biotech R&D.","BioSecBench-Refusal  \nA paired metric for performance and alignment in agentic biosecurity risk assessment  \nEdwin H. Wintermute1, Harmon Bhasin2, Christina M. Agapakis1, Dianzhuo Wang2, Evan Seeyave2, Arjun Banerjee2, Daniel Fulop, Matthew C. Watson, Adam J. Meyer, Sandrine Boissel, Jens H. Kuhn, Rishi Jain, Noah D. Taylor, Helena Shomar, Patrick M. Boyle1,*, Kenny Workman2,*  \n1American Wetware, Boston, MA 2LatchBio, San Francisco, CA  \n[Correspondence: patrick@americanwetware.com](Correspondence: patrick@americanwetware.com), [kenny@latch.bio](kenny@latch.bio)  \n05462v1 [ cs .CR] 6 Jul 2026  \nABSTRACT  \nAs AI agents are incorporated into life science workflows, the capabilities that speed discovery might also enable misuse. We present BioSecBench-Refusal, a benchmark for risk identification and refusal behavior for biological research tasks. The benchmark pairs 61 Routine tasks, legitimate analyses adapted from the published literature, with 46 Red-Team tasks, fictional scenarios that resemble real research but conceal a biosecurity hazard. Across 16 model-harness configurations, refusal rates ranged from 7% to 74% on Routine tasks and 1% to 62% on Red-Team tasks, with many configurations refusing legitimate Routine work at comparable or higher rates than concealed hazards. Refusals were most often triggered by provider API filters applied prior to agentic reasoning. However, models given room to reason showed the potential to identify more real threats. We release BioSecBench-Refusal as a tool for model developers to calibrate capability and caution for agentic biotech R&D.  \narXiv :2607 .  \nIntroduction  \nAI agents are increasingly able to design proteins, plan experiments, and interpret complex data. These capabilities present an acute dual-use problem: they can accelerate beneficial discoveries or, if misused, enable harm [1, 2, 3, 4] . Assessing biosecurity risks requires both technical and normative analysis. To be deployed safely, agentic tools must identify complex biological risksand act upon them in alignment with the values of human communities and institutions[5, 6] .  \nRefusal training is the primary mechanism by which model developers mitigate biosecurity risks in large language models (LLMs) . Refusal operates at two layers: the model itself, trained to decline harmful requests [7, 8], and a provider-side safeguard that screens prompts and outputs before or after the model runs [9, 10] . Recent work has sought to characterize refusal behavior against a range of toxic signals [11] and specifically against biologically relevant signals [12] . This work is motivated, in part, by concerns that certain refusal behaviors (\"over-refusals\") maybe too stringent, disrupting legitimate work and reducing trust in AI tools for productive biological research.  \nAgentic workflows deepen the challenge of calibrating refusal behavior: a refusal at any step of a multi-step reasoning chain can cause the entire task to fail. A model well-calibrated on simple tasks may over-refuse on complex ones. Refusal therefore needs to be measured over the extended reasoning chains where agents actually operate.  \nBenchmarks provide a standardized way to measure LLM performance. Recent benchmarks have been developed to evaluate biosecurity capabilities including in virology [13], synthesis screening [14], or biological design tools [15], among other areas of concern[16, 17] . Unlike strictly technical tasks, a biosecurity refusal evaluation has no single correct answer. Preferred model behavior might range from open access, with minimal safeguards, to complete refusal of biology-related tasks. The choice will depend on a cost-benefit analysis taken in the context of prevailing laws, institutional values and other security measures in place to control access to a particular LLM.  \nBioSecBench-Refusal is a biosecurity-specific benchmark in-  \ntended to help model builders and policymakers navigate this trade-off in refusal behavior. It is comp","cbCaivhphxVA9l2s","https://ap.wps.com/l/cbCaivhphxVA9l2s","pdf",697584,2,1,9,"English","en",105,"# Introduction\n## Refusal in LLM and provider safeguards\n## Need for calibrated refusal in agentic workflows\n## Biosecurity evaluation benchmarks and trade-offs\n# BioSecBench-Refusal benchmark design\n## Red-Team tasks with concealed hazards\n## Routine tasks adapted from published research\n# Results summary\n## Refusal rates across configurations\n## Provider filters vs task-level risk analysis","[{\"question\":\"What does BioSecBench-Refusal benchmark measure?\",\"answer\":\"It measures risk identification ability and refusal behavior for biological research tasks, with emphasis on how models decline requests in dual-use settings.\"},{\"question\":\"How are Routine and Red-Team tasks defined in the benchmark?\",\"answer\":\"Routine tasks are adapted from published life-science analyses that a biologist might perform, while Red-Team tasks are fictional scenarios that resemble real research but conceal a biosecurity hazard.\"},{\"question\":\"What were the observed refusal-rate ranges across model configurations?\",\"answer\":\"Across 16 model-harness configurations, refusal rates ranged from 7% to 74% on Routine tasks and 1% to 62% on Red-Team tasks.\"}]",1784198621,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"biosecbench-refusal-calibrated-refusal-and-safe-usefulness-in-dual-use-biology-settings","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/biosecbench-refusal-calibrated-refusal-and-safe-usefulness-in-dual-use-biology-settings/84836/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does BioSecBench-Refusal benchmark measure?","Question",{"text":75,"@type":76},"It measures risk identification ability and refusal behavior for biological research tasks, with emphasis on how models decline requests in dual-use settings.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are Routine and Red-Team tasks defined in the benchmark?",{"text":80,"@type":76},"Routine tasks are adapted from published life-science analyses that a biologist might perform, while Red-Team tasks are fictional scenarios that resemble real research but conceal a biosecurity hazard.",{"name":82,"@type":73,"acceptedAnswer":83},"What were the observed refusal-rate ranges across model configurations?",{"text":84,"@type":76},"Across 16 model-harness configurations, refusal rates ranged from 7% to 74% on Routine tasks and 1% to 62% on Red-Team tasks.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]