[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128233-en":3,"doc-seo-128233-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},128233,2336475104736,"Quinn","https://ap-avatar.wpscdn.com/avatar/22000c4c5e0e5b17e70?x-image-process=image/resize,m_fixed,w_180,h_180&k=1786591360781797222",8,"Research & Report","Identifying Malicious Python Packages - Using Static Indicators and Machine Learning","The sharing culture, especially in open-source, has enabled rapid reuse of Python libraries via the Python Package Index (PyPI). The same openness also allows malicious actors to publish packages with harmful payloads that can infiltrate popular projects and cause significant damage. This master’s thesis increases knowledge of common traits in malicious packages and evaluates how machine learning models perform using static indicators and metadata from a dataset containing 382,712 benign and 7,639 malicious Python packages.","Master’s thesis  \nSimon Rødsbakken Røe  \nIdentifying Malicious Python Packages  \nUsing Static Indicators and Machine Learning  \nMaster’s thesis in Information Security Supervisor: Geir Olav Dyrkolbotn  \nCo-supervisor: Felix Leder June 2023  \nNT NU  \nNorwegian Un iversity of Science and Technology  \nFaculty of Information Techno logy and Electrical Engineering Dept . of Informat ion Security and Communication Technology  \nSimon Rødsbakken Røe  \nIdentifying Malicious Python Packages  \nUsing Static Indicators and Machine Learning  \nMaster’s thesis in Information Security Supervisor: Geir Olav Dyrkolbotn  \nCo-supervisor: Felix Leder June 2023  \nNorwegian University of Science and Technology  \nFaculty of Information Technology and Electrical Engineering Dept. of Information Security and Communication Technology  \nAbstract  \nThe sharing culture, especially Open-Source, has become important with the increased use of technology and the internet. In programming, languages like Python rely on the open-source community to create libraries for people to use. This allows anyone to implement complex application functionality simply by adding a package.  \nHowever, the danger of letting anyone publish such packages is the possibility of evil actors trying to exploit the users by uploading malicious packages, which can cause significant damage if they can sneak into popular projects.  \nThe thesis aims to increase the domain knowledge of the characteristics commonly found in malicious packages in the Python Package Index (PyPi) and to see how common machine learning models perform on this data. We have used static indicators and metadata of a dataset consisting of 382,712 benign and 7,639 malicious Python packages. We have used the feature ranking method InformationGain and common machine learning models from the Python library Scikitlearn to find the most valuable features and identify the expected performance from the models.  \nFrom the experiments conducted, did we find the Neural Network Perceptron model to perform the best with the default options among those we tested with an F1-score of 92% against the verification dataset. We also found the most common features among the malicious packages to be two commonly used libraries,\"requests\" and \"setuptools\". From the results of testing the models, we found that it can identify most of the small malicious samples while the average size of those misclassified was much higher. It needs to be looked closer at improving the detection of larger and potentially more sophisticated packages.  \nWe conclude that static indicators and machine learning could be good for detecting malicious packages. However, more research into optimization is needed, as well as more profound knowledge of what combination of indicators is typically malicious. We have contributed with extended knowledge on the topic and on how some models perform on static indicators on a larger dataset of packages. We have also provided recommendations for what could be focused on further in research on the topic.  \nSammendrag  \nDelingskulturen som finnes på internett og spesielt knyttet til åpen kildekode har vært viktig ved den økende interessen for teknologi. Innenfor programmering og språk slik som Python lener seg på gruppen som produserer og deler sine prosjekter med andre i form av programmer og funksjonalitet.  \nMen, faren ved å stole på andre gjør det mulig for ondsinnede aktører å utnytte vanligebrukere ved å lokke med god funksjonalitet. Ved å legge med skadelig kodei disse pakkene som deles kan føre til store konsekvenser hvis de klarer å snike seg med i større programvareprosjekter.  \nDenne oppgaven har som mål å øke kunnskapsnivået om attributter som er vanlig å finne på slike ondsinnede pakker i Python sitt pakkebibliotek PyPi samt undersøke hva slags prestasjon kan forventes av de vanligste maskinlærings modellene. Vi har gjort eksperimenter på metadata og statiske indikatorer hentet fra ett datasett som inneholder 382,712 l","cbCaia2U9Pfqh4sV","https://ap.wps.com/l/cbCaia2U9Pfqh4sV","pdf",10552107,3,1,107,"English","en",105,"# Abstract\n## Dataset and Features\n## Methods and Models\n## Experimental Results\n## Conclusions and Recommendations","[{\"question\":\"What problem does the thesis address?\",\"answer\":\"It addresses the risk that malicious actors can upload harmful Python packages to PyPI, potentially compromising users and popular projects. The work focuses on detecting such packages using static indicators and machine learning.\"},{\"question\":\"What data and features are used for detection?\",\"answer\":\"The thesis uses a large dataset of Python packages with both benign and malicious samples, relying on static indicators and package metadata. It also applies feature ranking to identify the most valuable features.\"},{\"question\":\"Which machine learning model performed best in the experiments?\",\"answer\":\"The experiments found the Neural Network Perceptron (Multilayer Perceptron) model to perform best by default among the tested options, reaching an F1-score of 92% on the verification dataset.\"}]","Identifying Malicious Python Packages - Using Static Indicators and Machine Learning | PDF",1785945960,270,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"identifying-malicious-python-packages-using-static-indicators-and-machine-learning","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":20},"https://docshare.wps.com/document/research-report/",{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/identifying-malicious-python-packages-using-static-indicators-and-machine-learning/128233/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-27","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the thesis address?","Question",{"text":76,"@type":77},"It addresses the risk that malicious actors can upload harmful Python packages to PyPI, potentially compromising users and popular projects. The work focuses on detecting such packages using static indicators and machine learning.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What data and features are used for detection?",{"text":81,"@type":77},"The thesis uses a large dataset of Python packages with both benign and malicious samples, relying on static indicators and package metadata. It also applies feature ranking to identify the most valuable features.",{"name":83,"@type":74,"acceptedAnswer":84},"Which machine learning model performed best in the experiments?",{"text":85,"@type":77},"The experiments found the Neural Network Perceptron (Multilayer Perceptron) model to perform best by default among the tested options, reaching an F1-score of 92% on the verification dataset.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]