[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126834-en":3,"doc-seo-126834-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},126834,1099523885074,"Ivy","https://ap-avatar.wpscdn.com/davatar_9964176cb1d06d4a9deccf72a44ae3dc",8,"Research & Report","Code Execution Capability as a Metric for Machine Learning - Assisted Software Vulnerability Detection Models - Abstract","This paper studies how the ability to learn code execution tasks influences accuracy on software vulnerability detection (SVD) benchmark datasets. Initial results show models can reach near state-of-the-art SVD accuracy without learning code execution tasks, yet they generalize poorly across SVD benchmarks. The findings suggest dataset bias enabling prediction of non-SVD signals. By combining SVD datasets, SVD accuracy drops while correlation with code execution task accuracy improves, supporting the need for reduced bias collection methods.","Computer Science and Engineering Faculty Publications  \nComputer Science & Engineering  \n2023  \nCode Execution Capability as a Metric for Machine  \nLearning–Assisted Software Vulnerability Detection Models Daniel Grahn  \nLingwei Chen Junjie Zhang  \nFollow this and additional works at: [https://corescholar.libraries.wright.edu/cse](https://corescholar.libraries.wright.edu/cse)  \n Part of the Computer Sciences Commons  \nThis Article is brought to you for free and open access by Wright State University’s CORE Scholar. It has been accepted for inclusion in Computer Science and Engineering Faculty Publications by an authorized administrator of CORE Scholar. For more information, please contact [library-corescholar@wright.edu](library-corescholar@wright.edu).  \nCode Execution Capability as a Metric for Machine Learning–Assisted Software Vulnerability Detection Models  \nDaniel Grahn, Lingwei Chen, Junjie Zhang  \nDepartment of Computer Science and Engineering  \nWright State University  \nDayton, USA  \n{dan.grahn,lingwei.chen,[junjie.zhang}@wright.edu](junjie.zhang}@wright.edu)  \nAbstract—In this paper, we consider how the ability to learn Code Execution Tasks affects a model’s accuracy on software vulnerability detection (SVD) benchmark datasets. We initially find that models can achieve near state-of-the-art accuracy on SVD benchmarks regardless of their ability to learn Code Execution Tasks. However, these models fail to generalize well across SVD benchmarks. The results indicate a bias in the datasets that allows models to predict nonSVD signals. Under the theory that different collection methods will reduce biases, we investigate combining the SVD datasets. When trained on combined datasets, SVD accuracy is reduced but correlation with Code Execution Task accuracy improves. Our contributions are (1) using a reversed curriculum learning to evaluate model capabilities, (2) demonstrating the criticality of code execution understanding to machine learning– assisted software vulnerability detection, (3) evidence that improved diversity of SVD datasets will lead to improved accuracy and generalizability,(4) and benchmarks of recent models across multiple SVD datasets.  \nI. Introduction  \nDetecting software vulnerabilities during development reduces risk and increases the stability of the software. If delayed, undetected vulnerabilities may be exploited by malicious actors. While most vulnerabilities are introduced and removed without incident, others have led to the loss of property, resources, and life. To prevent such damages, significant efforts are organized around the detection and mitigation of vulnerabilities. These endeavors have historically been a resource-intensive process. Pair programming, code reviews, and other best practices take precious human hours. While Source Code Analysis (SCA) may be automated as part of continuous integration, the prevalence of false positives requires programmers to review many reports.  \nAs Machine Learning (ML) becomes capable of performing evermore complex tasks, there is a push to use these capabilities to improve software vulnerability detection (SVD) . Over the last few years, much time and effort have been spent pursuing machine learning–assisted software vulnerability detection (MLAVD) techniques. Previous research has shown significant deficiencies in the datasets used for MLAVD but did not consider the role of the models themselves [10] .  \nUnlike the long-standing application of ML to Computer Vision (CV) and Natural Language Processing (NLP), MLAVD is a relatively new field [11] . As CV and NLP techniques developed, inspiration was drawn from how humans sense and understand the world [9] . But programming languages are different from natural languages [14, 21] . Natural languages evolve across time and space, while programming languages have strict syntax and unambiguous semantics. Likewise, SVD is highly dissimilar to object detection and other CV tasks. Vulnerabilities are not defin","cbCaiixqex3L1kLd","https://ap.wps.com/l/cbCaiixqex3L1kLd","pdf",257708,1,9,"English","en",105,"# Abstract\n# Introduction\n## Motivation for ML-based vulnerability detection\n## Relation to code execution understanding\n# Research Questions and Experimental Setup\n## RQ1: Learning code execution tasks\n## RQ2: Correlation between SVD accuracy and code execution ability\n## RQ3: Effect of combining datasets on SVD accuracy\n## RQ4: Effect of combining datasets on correlation with code execution tasks\n# Main Contributions","[{\"question\":\"How does learning code execution tasks affect SVD accuracy?\",\"answer\":\"Models may achieve near state-of-the-art SVD accuracy even when they cannot learn code execution tasks, but they generalize poorly across benchmarks. Correlation improves when datasets are combined.\"},{\"question\":\"What evidence is there for bias in SVD benchmark datasets?\",\"answer\":\"The paper reports that models lacking code execution learning still perform well on SVD benchmarks but fail to generalize well. This pattern indicates bias that allows predictions based on non-SVD signals.\"},{\"question\":\"What changes when SVD datasets are combined?\",\"answer\":\"Training on merged datasets reduces mean SVD accuracy, yet increases the correlation between SVD performance and code execution task accuracy, suggesting better alignment with code execution understanding.\"}]","Code Execution Capability as a Metric for Machine Learning - Assisted Software Vulnerability Detection Models - Abstract | PDF",1785935113,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"code-execution-capability-as-a-metric-for-machine-learning-assisted-software-vulnerability-detection-models-abstract","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/code-execution-capability-as-a-metric-for-machine-learning-assisted-software-vulnerability-detection-models-abstract/126834/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does learning code execution tasks affect SVD accuracy?","Question",{"text":75,"@type":76},"Models may achieve near state-of-the-art SVD accuracy even when they cannot learn code execution tasks, but they generalize poorly across benchmarks. Correlation improves when datasets are combined.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What evidence is there for bias in SVD benchmark datasets?",{"text":80,"@type":76},"The paper reports that models lacking code execution learning still perform well on SVD benchmarks but fail to generalize well. This pattern indicates bias that allows predictions based on non-SVD signals.",{"name":82,"@type":73,"acceptedAnswer":83},"What changes when SVD datasets are combined?",{"text":84,"@type":76},"Training on merged datasets reduces mean SVD accuracy, yet increases the correlation between SVD performance and code execution task accuracy, suggesting better alignment with code execution understanding.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]