[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118778-en":3,"doc-seo-118778-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118778,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1785132997149421697",8,"Research & Report","Automatic Static Bug Detection for Machine Learning Libraries - Are We There Yet","Automatic detection of software bugs is essential for software security, yet static bug detectors are often evaluated on general programs, leaving uncertainty about their real effectiveness for machine learning libraries. This study analyzes five widely used static bug detectors—Flawfinder, RATS, Cppcheck, Facebook Infer, and Clang static analyzer—on a curated dataset of 410 known bugs collected from four popular ML libraries: mlpack, MXNet, PyTorch, and TensorFlow. Results show only a negligible 0.01% of bugs are detected overall, with Flawfinder and RATS performing best. The work categorizes tool capabilities, highlights strengths and weaknesses, and proposes opportunities to improve ML-focused practical detection.","Automatic Static Bug Detection for Machine Learning Libraries: Are We There Yet?  \nNima Shiri harzevili􀀃 , Jiho Shin􀀃 , Junjie Wang†, Song Wang􀀃 , Nachiappan Nagappan‡  \n􀀃 York University; †Institute of Software, Chinese Academy of Sciences; ‡META  \n{nshiri,jihoshin,[wangsong](wangsong}@yorku.ca)[}](wangsong}@yorku.ca)[@yorku.ca](wangsong}@yorku.ca); [junjie@iscas.ac.cn](junjie@iscas.ac.cn) ; [nachiappan.nagappan@gmail.com](nachiappan.nagappan@gmail.com)  \narXiv :2307 .04080v 1 [ cs . SE] 9 Jul 2023  \nAbstract—Automatic detection of software bugs is a critical task in software security. Many static tools that can help detect bugs have been proposed. While these static bug detectors are mainly evaluated on general software projects call into question their practical effectiveness and usefulness for machine learning libraries. In this paper, we address this question by analyzing 􀀂ve popular and widely used static bug detectors, i.e., Flaw􀀂nder, RATS, Cppcheck, Facebook Infer, and Clang static analyzer on a curated dataset of software bugs gathered from four popular machine learning libraries including Mlpack, MXNet, PyTorch, and TensorFlow with a total of 410 known bugs. Our research provides a categorization of these tools’ capabilities to better understand the strengths and weaknesses of the tools for detecting software bugs in machine learning libraries. Overall, our study shows that static bug detectors 􀀂nd a negligible amount of all bugs accounting for 6/410 bugs (0.01%), Flaw􀀂nder and RATS are the most effective static checker for 􀀂nding software bugs in machine learning libraries. Based on our observations, we further identify and discuss opportunities to make the tools more effective and practical.  \nIndex Terms—Software bugs, static detection, machine learning libraries  \nI. INTRODUCTION  \nProgramming inevitably involves dealing with bugs in software, which is an aspect that can be frustrating for many developers since detecting and 􀀂xing bugs is time-consuming [1],[2] . To help developers 􀀂nd software bugs, many static bug detectors have been developed and are now frequently employed by many industries and open-source projects [3]–[6] . Error Prone from Google [7], Infer from Facebook [8], and SpotBugs [9], the successor to the widely used FindBugs tool [10], are examples of popular static bug detection applications. These tools are usually developed as an analytical framework based on static analysis and are capable of scaling to large applications.  \nPrevious research has examined static bug detectors on traditional projects from different aspects [11]–[16] . There are two major limitations in existing studies. First, the datasets used for the empirical evaluation of static bug detectors are not real-world examples, they are not able to replicate new and sophisticated bug patterns. Second, they do not categorize the type of bugs based on Common Weakness Enumeration(CWE) information. Recently, Lipp et al. [16] addressed the limitations by proposing an empirical evaluation of static bug detectors on real-world datasets collected from CVE records gathered from 27 projects containing 1.15 million lines of code. Their results showed that state-of-the-art tools can detect  \nin-between 20% and 53% of the bugs in a benchmark set of real-world programs. However, it has not been determined if the 􀀂ndings of these studies on conventional projects are applicable to machine learning (ML) libraries. Finding real-world security bugs in ML libraries is critical for a couple of reasons. First, ML libraries have been widely used in many 􀀂elds in the past decades, such as image classi􀀂cation [17], [18], big data analysis [19], pattern recognition [20], autonomous driving [21]–[23], and natural language processing [24]–[26] . Failure to detect bugs in these ML libraries might have devastating implications, such as traf􀀂c accidents [27] . Second, gaining an understanding of the bene􀀂ts and drawbacks of the static bug detectors that are c","cbCair0chur1WH6M","https://ap.wps.com/l/cbCair0chur1WH6M","pdf",246234,1,12,"English","en",105,"# Introduction\n## Research motivation and limitations of prior work\n## Study objective and research methodology\n# (Further sections not fully provided)","[{\"question\":\"Which static bug detectors are evaluated in the study, and on what dataset?\",\"answer\":\"The study evaluates Flawfinder, RATS, Cppcheck, Facebook Infer, and Clang static analyzer on a curated dataset of 410 known bugs collected from mlpack, MXNet, PyTorch, and TensorFlow.\"},{\"question\":\"What is the main finding about how many bugs static tools detect in ML libraries?\",\"answer\":\"Across all tools, static bug detectors find a negligible fraction of bugs, totaling 6 out of 410 bugs (0.01%).\"},{\"question\":\"Why are existing evaluations on general software projects not sufficient for ML libraries?\",\"answer\":\"Existing datasets often fail to replicate new and sophisticated bug patterns, and they do not categorize bug types using CWE information, making it unclear whether results transfer to ML library vulnerabilities.\"}]","Automatic Static Bug Detection for Machine Learning Libraries - Are We There Yet | PDF",1785720200,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"automatic-static-bug-detection-for-machine-learning-libraries-are-we-there-yet","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/automatic-static-bug-detection-for-machine-learning-libraries-are-we-there-yet/118778/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04","2026-08-03",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Which static bug detectors are evaluated in the study, and on what dataset?","Question",{"text":76,"@type":77},"The study evaluates Flawfinder, RATS, Cppcheck, Facebook Infer, and Clang static analyzer on a curated dataset of 410 known bugs collected from mlpack, MXNet, PyTorch, and TensorFlow.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is the main finding about how many bugs static tools detect in ML libraries?",{"text":81,"@type":77},"Across all tools, static bug detectors find a negligible fraction of bugs, totaling 6 out of 410 bugs (0.01%).",{"name":83,"@type":74,"acceptedAnswer":84},"Why are existing evaluations on general software projects not sufficient for ML libraries?",{"text":85,"@type":77},"Existing datasets often fail to replicate new and sophisticated bug patterns, and they do not categorize bug types using CWE information, making it unclear whether results transfer to ML library vulnerabilities.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":122},"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]