[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126718-en":3,"doc-seo-126718-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},126718,962084925636,"Olivia Brown","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Vulnerability Clustering and other Machine Learning Applications of Semantic Vulnerability Embeddings - SRI International report","Cyber-security vulnerabilities are commonly described as short natural-language texts and then enriched over time with labels such as CWE, CPE, and CVSS. This report from the Vulnerability AI project evaluates semantic vulnerability embeddings generated with NLP techniques to build compact representations of the vulnerability space. The embeddings are assessed as a foundation for machine learning applications supporting risk assessment, including clustering, classification, and visualization, as well as a logic-based approach for analyzing theories about vulnerability space structure and evolution.","arXiv :2310 .05935v1 [ cs .CR] 23 Aug 2023  \nVulnerability Clustering and other Machine Learning Applications of Semantic Vulnerability Embeddings  \nMark-Oliver Stehr, Minyoung Kim  \nSRI International, Menlo Park, CA 94025  \nAbstract. Cyber-security vulnerabilities are usually published in form of short natural language descriptions (e.g., in form of MITRE’s CVE list) that over time are further manually enriched with labels such as those defined by the Common Vulnerability Scoring System (CVSS) . In the Vulnerability AI (Analytics and Intelligence) project, we investigated different types of semantic vulnerability embeddings based on natural language processing (NLP) techniques to obtain a concise representation of the vulnerability space. We also evaluated their use as a foundation for machine learning applications that can support cyber-security researchers and analysts in risk assessment and other related activities.  \nThe particular applications we explored and briefly summarize in this report are clustering, classification, and visualization, as well as a new logic-based approach to evaluate theories about the vulnerability space.  \n1 Introduction  \nFor the Vulnerabilities Equities Process (VEP) [36] and more generally for Cyber-Security Risk Assessment, it is important to keep an eye on the big picture of how the cyber-security vulnerability space is structured and how it is evolving over time. Furthermore, it should be possible to quickly place newly discovered vulnerabilities in the space and support subject matter experts with tools to visualize and navigate the space. In this report on the Vulnerability AI (Analytics and Intelligence) project, we discuss a fully data-driven approach that leverages machine learning (especially natural language processing, classification, and clustering) algorithms to serve as a foundation for such tools. Our approach is centered around the generation and evaluation of semantic vulnerability embeddings on which algorithms for visualization, clustering and classification can be built. It also serves as a basis for a novel capability for vulnerability space analysis using logical theories that we briefly explore.  \n2 Vulnerability Datasets and Labels  \nTo employ and validate machine learning algorithms in the proper application context we have identified the National Vulnerability Database (NIST NVD)  \n[27] as the most suitable dataset. It is fed from MITRE’s CVE list [25], and with about 150K entries covering the years 1999 − 2021 the most comprehensive list of publicly disclosed cyber-security vulnerabilities.  \nFrom the NVD dataset, we extract the CVE-identifier (which identifies the vulnerability), the publication date, the natural language description of the vulnerability. We furthermore extract three types of labels: The Common Weakness Enumeration (CWE) [26] label that identifies the underlying software weaknesses (typically only one but sometimes multiple weaknesses are associated with a vulnerability) that the vulnerability exploits, and a reduced version of the Common Platform Enumeration (CPE) [24] label that identifies the vendor and product in which the vulnerability was observed. Furthermore, we extract the Common Vulnerability Scoring System (CVSS) [8] labels related to the most commonly used base metrics, which have components qualitatively summarizing the impact (e.g. the confidentiality, integrity, availability impact, and the potential change of scope) and the exploitability (e.g., attack complexity, attack vector, privilege/authentication requirement, and need for user interaction) . Quantitative assessments on an ad-hoc scale 0 − 10 are also available and sometimes used for our visualizations, but as these are derived scores (and as we will discuss not reflected in empirical data), our machine learning algorithms operate directly on the qualitative features for better precision. We also distinguish between CVSSV2 and V3 labels, 1 but we will gloss over these details ","cbCaibgy7kSJvtrn","https://ap.wps.com/l/cbCaibgy7kSJvtrn","pdf",28072796,1,27,"English","en",105,"# Abstract\n# Introduction\n# Vulnerability Datasets and Labels\n## National Vulnerability Database (NVD)\n## Extracted labels (CWE, CPE, CVSS)\n# Methods and Algorithms\n## NLP stage and ML pipeline\n## Dimensionality reduction, clustering, classification, visualization\n## Logic-based capability for vulnerability space analysis","[{\"question\":\"What is the main goal of the Vulnerability AI project described in this report?\",\"answer\":\"To create semantic vulnerability embeddings from natural-language descriptions and evaluate them as a foundation for machine learning tasks that support cyber-security risk assessment.\"},{\"question\":\"Which dataset is used to build and validate the machine learning models?\",\"answer\":\"The report uses the National Vulnerability Database (NIST NVD), which is fed from MITRE’s CVE list and covers vulnerabilities from 1999 to 2021.\"},{\"question\":\"What machine learning applications does the report explore using semantic vulnerability embeddings?\",\"answer\":\"It explores clustering, classification, and visualization, and also a new logic-based approach to evaluate theories about the vulnerability space.\"}]","Vulnerability Clustering and other Machine Learning Applications of Semantic Vulnerability Embeddings - SRI International report | PDF",1785934388,68,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"vulnerability-clustering-and-other-machine-learning-applications-of-semantic-vulnerability-embeddings-sri-international-report","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/vulnerability-clustering-and-other-machine-learning-applications-of-semantic-vulnerability-embeddings-sri-international-report/126718/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main goal of the Vulnerability AI project described in this report?","Question",{"text":75,"@type":76},"To create semantic vulnerability embeddings from natural-language descriptions and evaluate them as a foundation for machine learning tasks that support cyber-security risk assessment.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which dataset is used to build and validate the machine learning models?",{"text":80,"@type":76},"The report uses the National Vulnerability Database (NIST NVD), which is fed from MITRE’s CVE list and covers vulnerabilities from 1999 to 2021.",{"name":82,"@type":73,"acceptedAnswer":83},"What machine learning applications does the report explore using semantic vulnerability embeddings?",{"text":84,"@type":76},"It explores clustering, classification, and visualization, and also a new logic-based approach to evaluate theories about the vulnerability space.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]