[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82736-en":3,"doc-seo-82736-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82736,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","LLM-Enhanced Hierarchical Heterogeneous Graph Representation Learning for Malicious Python Package Detection","Malicious Python packages pose a major risk to software supply chain ecosystems, especially with the widespread adoption of PyPI and open-source repositories. Existing learning-based detectors often miss hierarchical structure and heterogeneous relationships between program entities. This work proposes an LLM-enhanced hierarchical heterogeneous graph representation learning framework that models heterogeneous entities and structural dependencies, adds LLM-inferred function-level semantic roles, and trains a type-aware graph neural network for package classification. An attribution mechanism localizes suspicious functions and fine-grained malicious behaviors without expert intervention, yielding accurate, robust, and interpretable detection across varied package sizes and dependency complexities.","LLM-Enhanced Hierarchical Heterogeneous Graph Representation Learning for Malicious Python  \nPackage Detection  \nHang Gao 1 ,2 ,∗ , Xiaoyu Chen 1 ,2 ,∗ , Baoquan Cui 1 ,2 , Zhen Tang 1 ,2 , Peng Qiao 1 ,2 , Fengge Wu 1 ,2 ,†, Jian Zhang 1 ,2  \n1Institute of Software, Chinese Academy of Sciences  \n2University of Chinese Academy of Sciences  \n{gaohang, chenxiaoyu2025, qiaopeng, [fengge](fengge}@iscas.ac.cn)[}](fengge}@iscas.ac.cn)[@iscas.ac.cn](fengge}@iscas.ac.cn); {cuibq, [zj](zj}@ios.ac.cn)[}](zj}@ios.ac.cn)[@ios.ac.cn](zj}@ios.ac.cn); [tangzhen12@otcaix.iscas.ac.cn](tangzhen12@otcaix.iscas.ac.cn)  \n∗ Equal contribution †Corresponding author: [fengge@iscas.ac.cn](fengge@iscas.ac.cn)  \narXiv :2607 .03350v 1 [ cs .CR] 3 Jul 2026  \nAbstract—Malicious Python packages have become a major threat to modern software supply chain ecosystems due to the widespread adoption of open-source repositories such as PyPI. Existing learning-based detection methods struggle to capture the hierarchical organization and heterogeneous interactions among different program entities. Although Large Language Models (LLMs) have demonstrated remarkable capabilities in code understanding and semantic reasoning, they are rarely integrated with structural program representations for finegrained malicious behavior analysis. In this paper, we propose an LLM-enhanced hierarchical heterogeneous graph representation learning framework for malicious Python package detection. The proposed framework constructs a hierarchical heterogeneous code graph that explicitly models heterogeneous code entities, together with different types of structural dependencies. To further enrich code representations, LLMs are leveraged to infer function-level semantic roles, introducing an additional layer of semantic heterogeneity. Based on this graph, we develop a hierarchical heterogeneous graph neural network that performs type-aware message passing over different node and edge categories, enabling effective modeling of malicious behavior propagation and accurate package-level classification. Furthermore, the proposed framework incorporates a function-level attribution mechanism which, combined with LLM reasoning, automatically identifies suspicious functions and localizes fine-grained malicious behaviors without requiring human expert intervention. Extensive experiments on real-world datasets demonstrate that the proposed framework consistently outperforms traditional machine learning methods, graph-based detectors, and state-ofthe-art LLMs across packages with varying sizes and dependency complexities, while providing accurate, robust, and interpretable malicious behavior localization. The replication package is available at: [https://github.com/xxy33/malware](https://github.com/xxy33/malware)  \nIndex Terms—Malicious Software Detection, Graph Neural Networks, Large Language Model, Python.  \nI. INTRODUCTION  \nWith the continued shift of modern software development toward modularity and code reuse, Open Source Software (OSS) has become a fundamental pillar of digital infrastructure. As one of the most active software ecosystems, the Python Package Index (PyPI) hosts more than 500,000 packages and serves billions of downloads each month [1] . However, the openness and highly interconnected nature of the ecosystem have also made it a prime target for software supply  \nchain attacks. High-profile security incidents, such as the SolarWinds compromise [2] and the Log4j vulnerability [3], have demonstrated that the compromise of a single component can trigger cascading effects throughout the entire software supply chain. Although industry initiatives such as SLSA [4] and intoto [5] have been proposed to improve supply chain security practices, attacks targeting package management ecosystems remain prevalent. Adversaries exploit the trust model of software repositories through techniques such as typosquatting, dependency confusion, and malicious installation script injection [6, 7] to stea","cbCaiji0qivU3p90","https://ap.wps.com/l/cbCaiji0qivU3p90","pdf",3086444,2,1,13,"English","en",105,"# Introduction\n## Threats and attack techniques in Python package ecosystems\n## Limits of rule-based detection\n## Learning-based approaches and motivation for GNN+LLM","[{\"question\":\"Why do malicious Python package threats matter for software supply chains?\",\"answer\":\"Because ecosystems like PyPI enable attackers to compromise widely used packages, and a single malicious component can trigger cascading effects across the supply chain.\"},{\"question\":\"What core limitation affects existing learning-based detection methods?\",\"answer\":\"They struggle to represent hierarchical organization and heterogeneous interactions among program entities, which weakens fine-grained malicious behavior analysis.\"},{\"question\":\"How does the proposed framework identify suspicious functions without expert involvement?\",\"answer\":\"It uses a function-level attribution mechanism combined with LLM reasoning to automatically pinpoint suspicious functions and localize fine-grained malicious behaviors.\"}]",1784182586,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"llm-enhanced-hierarchical-heterogeneous-graph-representation-learning-for-malicious-python-package-detection","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/llm-enhanced-hierarchical-heterogeneous-graph-representation-learning-for-malicious-python-package-detection/82736/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do malicious Python package threats matter for software supply chains?","Question",{"text":75,"@type":76},"Because ecosystems like PyPI enable attackers to compromise widely used packages, and a single malicious component can trigger cascading effects across the supply chain.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What core limitation affects existing learning-based detection methods?",{"text":80,"@type":76},"They struggle to represent hierarchical organization and heterogeneous interactions among program entities, which weakens fine-grained malicious behavior analysis.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the proposed framework identify suspicious functions without expert involvement?",{"text":84,"@type":76},"It uses a function-level attribution mechanism combined with LLM reasoning to automatically pinpoint suspicious functions and localize fine-grained malicious behaviors.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]