[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83490-en":3,"doc-seo-83490-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83490,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","MolSafeEval: A Benchmark for Uncovering Safety Risks in AI-Generated Molecules","Current molecular generation benchmarks focus on task complexity, molecule novelty, and property alignment, while leaving a critical issue under-addressed: safety risks posed by AI-generated molecules. Many generative models can output compounds that are toxic, reactive, or otherwise hazardous, creating hidden threats. MolSafeEval introduces a dedicated safety benchmark that integrates heterogeneous safety knowledge into a structured molecular safety knowledge graph, enabling LLM-based reasoning to detect and explain unsafe features. It evaluates four representative generation task types with standardized datasets and protocols, guiding safer molecular design.","MolSafeEval: A Benchmark for Uncovering Safety Risks in  \nAI-Generated Molecules  \nTong Xu1,2 , Xinzhe Cao3 , Zhihui Zhu2 , Keyan Ding1,2 , Huajun Chen1,2*  \n1Zhejiang University  \n2ZJU-Hangzhou Global Scientific and Technological Innovation Center  \n3University of Oxford  \n{xtong, dingkeyan, [huajunsir}@zju.edu.cn](huajunsir}@zju.edu.cn)  \narXiv :2607 .00464v 1 [ cs .LG] 1 Jul 2026  \nAbstract  \nCurrent molecular generation benchmarks emphasize task complexity, molecule novelty, and property alignment; they largely overlook a critical concern: the potential safety risks of AI-generated molecules. In practice, many generative models may produce molecules with toxic, reactive, or otherwise hazardous characteristics—posing hidden dangers that remain insufficiently addressed. To address this gap, we introduce MolSafeEval, a benchmark dedicated to evaluating and analyzing the safety risks of molecular generation. Unlike prior approaches that rely on narrow toxicity predictors, MolSafeEval integrates heterogeneous safety knowledge—ranging from toxicological databases to hazard rules—into a structured molecular safety knowledge graph. This graph serves as a foundation for large language model–based reasoning, enabling systematic detection and explanation of unsafe features in generated compounds. We further categorize molecular generative models into four representative task types—unconditional generation, property optimization, target protein–based design, and text-based generation—and provide standardized datasets and safety evaluation protocols for each.By systematically revealing the safety vulnerabilities of current generative approaches, MolSafeEval offers a new lens for benchmarking molecular models and provides essential guidance toward safer, more trustworthy molecular design.  \n1 Introduction  \nDesigning new molecules with desired properties is a central challenge in drug discovery and materials science (Wang et al., 2022) . The chemical compound space, estimated to contain between 1023 and 1080 possible molecules, is so vast that manual exploration is infeasible and resourceintensive (Reymond, 2015) . To address this chal-  \n*Corresponding author.  \nlenge, deep generative models have been increasingly adopted for molecular design, accelerating exploration and yielding promising results (Pei et al., 2023 ; Xu et al., 2023 ; Fang et al., 2024) .  \nThe diversity of molecular generative models, driven by varying tasks and datasets, presents challenges for fair performance comparison. In response, several benchmarks have been developed. For example, Mol-OPT (Gao et al., 2022) standardizes evaluation for molecular optimization, while TARTARUS (Nigam et al., 2023) addresses complex real-world design problems. Despite such progress, a critical gap remains: the safety of generated molecules is rarely assessed. In practice, models may propose compounds that are toxic, reactive, or otherwise hazardous. While prior studies have raised concerns about dual-use risks in AI-powered drug discovery (Urbina et al., 2022), there is still no systematic benchmark for evaluating molecular safety. This gap hampers the transition of generative models from proof-of-concept research to safe and trustworthy deployment.  \nTo address this challenge, we propose MolSafeEval, a benchmark for systematically uncovering the safety risks of AI-generated molecules. A key difficulty in molecular safety evaluation is that risks are heterogeneous—ranging from drug toxicities to chemical hazards—and are scattered across toxicological datasets, regulatory standards, and chemical hazard reports. Simple predictorbased filters cannot capture this diversity. To unify these fragmented sources, MolSafeEval builds MolSafeKG, a structured molecular safety knowledge graph (KG) that integrates over 80,000 hazardous compounds with associated toxicological and hazard annotations. By coupling this resource with large language model (LLM)-based reasoning, we enable both detectio","cbCaiiu6rOyD2Zzy","https://ap.wps.com/l/cbCaiiu6rOyD2Zzy","pdf",2020728,3,1,28,"English","en",105,"# Introduction\n## Motivation and Safety Gap\n## MolSafeEval Overview and MolSafeKG\n## Task Coverage and Evaluation Pipeline\n## Contributions","[{\"question\":\"What safety problem does MolSafeEval target in AI-generated molecules?\",\"answer\":\"MolSafeEval targets the safety risks that are often overlooked by existing molecular generation benchmarks, including toxicity, reactivity, and other hazardous characteristics in generated compounds.\"},{\"question\":\"How does MolSafeEval unify safety knowledge for evaluation?\",\"answer\":\"It builds MolSafeKG, a structured molecular safety knowledge graph that integrates diverse safety sources, including toxicological databases and hazard rules, covering over 80,000 hazardous compounds with annotations.\"},{\"question\":\"Which generation task types are covered by MolSafeEval?\",\"answer\":\"MolSafeEval evaluates four representative task types: unconditional generation, property optimization-based generation, target protein-based design, and text-based generation, each using standardized datasets and a multi-step analysis pipeline.\"}]",1784188387,71,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"molsafeeval-a-benchmark-for-uncovering-safety-risks-in-ai-generated-molecules","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/molsafeeval-a-benchmark-for-uncovering-safety-risks-in-ai-generated-molecules/83490/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What safety problem does MolSafeEval target in AI-generated molecules?","Question",{"text":75,"@type":76},"MolSafeEval targets the safety risks that are often overlooked by existing molecular generation benchmarks, including toxicity, reactivity, and other hazardous characteristics in generated compounds.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does MolSafeEval unify safety knowledge for evaluation?",{"text":80,"@type":76},"It builds MolSafeKG, a structured molecular safety knowledge graph that integrates diverse safety sources, including toxicological databases and hazard rules, covering over 80,000 hazardous compounds with annotations.",{"name":82,"@type":73,"acceptedAnswer":83},"Which generation task types are covered by MolSafeEval?",{"text":84,"@type":76},"MolSafeEval evaluates four representative task types: unconditional generation, property optimization-based generation, target protein-based design, and text-based generation, each using standardized datasets and a multi-step analysis pipeline.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]