[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121086-en":3,"doc-seo-121086-105":30,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121086,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",6,"Technology","Machine Learning Techniques for Python Source Code - Vulnerability Detection","Software vulnerabilities drive the scale and frequency of cyber attacks, making automated detection a persistent security challenge. This paper applies and compares multiple machine learning algorithms for identifying vulnerabilities in Python source code. An experimental evaluation shows that a Bidirectional Long Short-Term Memory (BiLSTM) model delivers strong results (average Accuracy 98.6%, average F-Score 94.7%, average Precision 96.2%, average Recall 93.3%, average ROC 99.3%), establishing a new performance benchmark for Python vulnerability detection.","Machine Learning Techniques for Python Source Code  \nVulnerability Detection  \nTalaya Farasat  \nUniversity of Passau Passau, Germany  \nJoachim Posegga  \nUniversity of Passau Passau, Germany  \narXiv :2404 .09537v1 [ cs . SE] 15 Apr 2024  \nABSTRACT  \nSoftware vulnerabilities are a fundamental reason for the prevalence of cyber attacks and their identification is a crucial yet challenging problem in cyber security. In this paper, we apply and compare different machine learning algorithms for source code vulnerability detection specifically for Python programming language. Our experimental evaluation demonstrates that our Bidirectional Long Short-Term Memory (BiLSTM) model achieves a remarkable performance (average Accuracy = 98.6%, average F-Score = 94.7%, average Precision = 96.2%, average Recall = 93.3%, average ROC = 99.3%), thereby, establishing a new benchmark for vulnerability detection in Python source code.  \nCCS CONCEPTS  \n• Security and privacy → Software and application security.  \n1 INTRODUCTION  \nCode flaws or vulnerabilities are prevalent in software systems and can potentially lead to system compromise, information leaks, or denial of service. Recognizing the constraints of traditional methods (static & dynamic code analyses), and with the growing accessibility of open-source software repositories, it has been recommended to adopt a data-driven approach for software vulnerability detection. Therefore, various machine learning techniques have been applied to learn vulnerable features of source code, and to automate the process of software vulnerability identification[1–4, 7–9, 11]. Many researchers focus on source code vulnerability detection across various programming languages such as Java, C, and C++ . Some notable studies include [1, 3, 5, 6, 10, 11] . In 2024, Python continues to maintain its prominent position as one of the top programming languages [15], and also a majorly used language on GitHub [14] . Despite its popularity, Python has been relatively overlooked by researchers. Only a few studies [4, 7, 9, 13] focus on vulnerability detection specifically in Python programming language.  \nGiven the abundance of machine learning algorithms and their corresponding hyper-parameters available, there is potential for improved results in this domain. To bridge this gap, we apply and compare five different machine learning models. Notably, our BiLSTM model demonstrates superior performance as compared to all other applied models and also with the approaches presented in [4, 7, 9] . We also open-source all our code and models used in this study for broader dissemination. Our code and models can be accessed at [https://github.com/Tf-arch/Python-Source-Code-Vulnerability](https://github.com/Tf-arch/Python-Source-Code-Vulnerability)Detection/tree/main  \n2 EXPERIMENTAL DESIGN  \nWe’re examining the same software vulnerabilities highlighted in [4], [7], and [9], i.e., SQL injection, cross-site scripting (XSS), command injection, cross-site request forgery (XSRF), path disclosure, remote code execution, and open redirect.  \nDataset: We use the dataset prepared by Wartschinski et al. [4], available at [16], which is compiled by targeting publicly accessible GitHub repositories. GitHub stands out as the largest repository hosting platform for source code globally, making it an ideal resource for this work. Wartschinski et al. [4] gather a distinct dataset for each vulnerability type. The data is collected in the form of commits that contain security-related fixes. Sections of code that are updated or removed in these commits are categorized as vulnerable, along with the surrounding code to provide context. Conversely, the remaining code and the post-fix version are labeled (probably) as not vulnerable. We use 70% data in training, 15% in testing, and 15% in the validation set.  \nWord2vec Embeddings: For the training of machine learning algorithms, it is necessary to represent code tokens as vectors that retain the semantic ","cbCaivtRlYP8XqXw","https://ap.wps.com/l/cbCaivtRlYP8XqXw","pdf",587468,1,3,"English","en",105,"# Introduction\n# Experimental Design\n## Dataset\n## Word2vec Embeddings\n## Machine Learning Algorithms\n## BiLSTM Model","[{\"question\":\"Why is source code vulnerability detection important for cybersecurity?\",\"answer\":\"Vulnerabilities are a core reason for many cyber attacks. Detecting them is crucial yet challenging because vulnerabilities can enable system compromise, information leaks, or denial of service.\"},{\"question\":\"Which machine learning model performs best in the study and how is it evaluated?\",\"answer\":\"The Bidirectional Long Short-Term Memory (BiLSTM) model achieves the highest performance, with average Accuracy 98.6%, F-Score 94.7%, Precision 96.2%, Recall 93.3%, and ROC 99.3%.\"},{\"question\":\"What vulnerabilities are considered in the experimental evaluation?\",\"answer\":\"The experiments focus on SQL injection, cross-site scripting (XSS), command injection, cross-site request forgery (XSRF), path disclosure, remote code execution, and open redirect.\"}]","Machine Learning Techniques for Python Source Code - Vulnerability Detection | PDF",1785733648,8,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":28},"machine-learning-techniques-for-python-source-code-vulnerability-detection","",{"@graph":36,"@context":84},[37,53,67],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":21},"https://docshare.wps.com/document/technology/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/machine-learning-techniques-for-python-source-code-vulnerability-detection/121086/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why is source code vulnerability detection important for cybersecurity?","Question",{"text":74,"@type":75},"Vulnerabilities are a core reason for many cyber attacks. Detecting them is crucial yet challenging because vulnerabilities can enable system compromise, information leaks, or denial of service.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"Which machine learning model performs best in the study and how is it evaluated?",{"text":79,"@type":75},"The Bidirectional Long Short-Term Memory (BiLSTM) model achieves the highest performance, with average Accuracy 98.6%, F-Score 94.7%, Precision 96.2%, Recall 93.3%, and ROC 99.3%.",{"name":81,"@type":72,"acceptedAnswer":82},"What vulnerabilities are considered in the experimental evaluation?",{"text":83,"@type":75},"The experiments focus on SQL injection, cross-site scripting (XSS), command injection, cross-site request forgery (XSRF), path disclosure, remote code execution, and open redirect.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,112,117,121,126,129,133],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":110,"slug":111},50,"technology",{"id":113,"doc_module":4,"doc_module_name":46,"category_name":114,"show_sort_weight":115,"slug":116},7,"Healthcare",40,"healthcare",{"id":29,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},"Research & Report",30,"research-report",{"id":122,"doc_module":4,"doc_module_name":46,"category_name":123,"show_sort_weight":124,"slug":125},9,"Religion & Spirituality",20,"religion-spirituality",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":124,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]