[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125778-en":3,"doc-seo-125778-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125778,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Pitfalls of machine learning models for protein-protein interaction networks - Supplementary methods","Supplementary methods describe how human protein-protein interaction (PPI) data and functional genomics features were curated and processed to train and benchmark machine learning models for PPI networks. Gold-standard interactions were sourced from IntAct, restricted to human heterodimers, and filtered to reduce false positives using complex-expansion removal and MIscore-based quality thresholds. Protein sequences, curated Swiss-Prot features, gene ontology annotations, annotated domains and motifs were encoded as bag-of-words. Gene expression data relied on Bgee and GTEx-derived healthy wild-type patterns, with binary expression presence/absence across anatomical entities and developmental stages. Computational workflows were implemented in Python with Jupyter, Pandas, NumPy, and visualization via Matplotlib and Seaborn, and supporting code and final data were released on GitHub under CC BY 4.0.","SUPPLEMENTARY METHODS  \nB4PPI-Human  \nThe data was obtained from large and professionally curated databases. This limits measurement bias, as each interaction is based on several experiments, and leverages experts’ knowledge in the curation process. On average, each PPI is supported by 3.4 publications (median of 3) and only 4.9% of interactions are obtained from only one source. Standard UniProt IDs are used throughout to ensure maximum compatibilities. Supplementary Figure 11 summarises the benchmarking pipeline described below.  \nMost of the manipulations were done in Python (The Python Language Reference—Python 3.10.1 documentation) with Jupyter Notebooks (Jupyter Project Documentation — Jupyter Documentation 4.1.1 alpha documentation) using the Pandas library (McKinney 2010; Reback et al. 2020) and Numpy (Harris et al. 2020) . The plots were drawn using Matplotlib (Hunter 2007), Seaborn (Waskom 2021) and the MetBrewer colour palettes (GitHub-BlakeRMills/MetBrewer: Color palette package in R inspired by works at the Metropolitan Museum of Art in New York) . All the code and final data are available on GitHub ( [https://github.com/Llannelongue/B4PPI](https://github.com/Llannelongue/B4PPI) ); some intermediary pre-processed datasets are not available online due to file size limits but they can be recreated using the code available. Data is available under Creative Commons Attribution (CC BY 4.0) License.  \nProtein-protein interaction data  \nTo train machine learning algorithms, the quality of the gold standard is paramount. Data on PPIs was obtained from IntAct (Orchard et al. 2014) and downloaded from the EMBLE-EBI FTP server (timestamp: 15/10/2021) . We restricted the data to human heterodimeric protein-protein interactions with UniProt IDs. To reduce the risk of false positives, we removed complex expansions (where the pairwise interactions within a complex are unreliable) and interactions based on colocalisation only. This quality control step leaves 128,790 PPIs, covering 15,506 proteins (out of 20,386 in UniProt) . Based on this dataset, we created an index of the number of recorded interactions per protein and made a list of hubs (highly connected proteins) . In line with the literature, hubs are defined as the 20% of proteins with the most interactions (Jin et al. 2007), which here is equivalent to proteins with more than 21 partners. The quality of the interactions is assessed further by looking at the MIscore (IntAct - User Guide), a quality score based on the manual curation of the interactions and annotations of the HUPO PSI-MI consortium that takes into account the detection method, the interaction type and the number of publications reporting it. In case of PPIs with multiple entries, the highest MIscore was used. When looking at the distribution of the MIscores in the dataset (Figure 1), a threshold of 0.47 is visible, which restrict the dataset to 78,229 interactions, covering 12,026 proteins. We also ran the same analyses without filtering on MIscores (i. e. using all 128,790 interactions) and found that all the results presented here held true.  \nFigure 1 : Distribution of the MIscore in IntAct.  \nFunctional genomics annotations and amino acids sequences  \nProtein sequences in humans are well documented and can be obtained from UniProt, but FG features can be more challenging as they should be diverse (i. e. cover a wide range of properties), of high-quality and have high coverage (i. e. few missing proteins) . For the same reasons as described above, aggregated, manually curated, and professionally reviewed databases are preferred. Based on features that have been successfully used for the task before, it is relevant to include information about cellular and tissue localisation, biological functions and gene expression patterns (Jansen 2003; Ben-Hur and Noble 2005; Zhang et al. 2012; Kotlyar et al. 2015) .  \nOne of the main databases on proteins is UniProt (The UniProt Consortium 2021) and in particular it","cbCaiswjuuyh7epu","https://ap.wps.com/l/cbCaiswjuuyh7epu","pdf",3163240,1,20,"English","en",105,"# Supplementary methods\n## B4PPI-Human data sources and reproducibility\n## Protein-protein interaction data quality control\n## Functional genomics annotations and amino acid sequences\n## Gene expression feature preprocessing","[{\"question\":\"How were high-quality human PPI interactions selected for model training?\",\"answer\":\"Interactions were obtained from IntAct and restricted to human heterodimeric PPIs with UniProt IDs. Complex expansions and colocalisation-only interactions were removed, and an MIscore threshold of 0.47 was applied to retain the higher-confidence subset.\"},{\"question\":\"What role does MIscore play in the dataset construction?\",\"answer\":\"MIscore is an IntAct quality score based on manual curation by the HUPO PSI-MI consortium, considering detection method, interaction type, and supporting publications. For proteins with multiple entries, the highest MIscore was used.\"},{\"question\":\"How were functional genomics features and gene expression data represented?\",\"answer\":\"Swiss-Prot features and gene ontology terms were encoded as bag-of-words sparse vectors. Gene expression was taken from Bgee (using GTEx v6 patterns with additional curation) and converted into binary presence/absence calls per anatomical entity and developmental stage.\"}]","Pitfalls of machine learning models for protein-protein interaction networks - Supplementary methods | PDF",1785901156,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"pitfalls-of-machine-learning-models-for-protein-protein-interaction-networks-supplementary-methods","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/pitfalls-of-machine-learning-models-for-protein-protein-interaction-networks-supplementary-methods/125778/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How were high-quality human PPI interactions selected for model training?","Question",{"text":75,"@type":76},"Interactions were obtained from IntAct and restricted to human heterodimeric PPIs with UniProt IDs. Complex expansions and colocalisation-only interactions were removed, and an MIscore threshold of 0.47 was applied to retain the higher-confidence subset.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What role does MIscore play in the dataset construction?",{"text":80,"@type":76},"MIscore is an IntAct quality score based on manual curation by the HUPO PSI-MI consortium, considering detection method, interaction type, and supporting publications. For proteins with multiple entries, the highest MIscore was used.",{"name":82,"@type":73,"acceptedAnswer":83},"How were functional genomics features and gene expression data represented?",{"text":84,"@type":76},"Swiss-Prot features and gene ontology terms were encoded as bag-of-words sparse vectors. Gene expression was taken from Bgee (using GTEx v6 patterns with additional curation) and converted into binary presence/absence calls per anatomical entity and developmental stage.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":21,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":21,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]