[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83058-en":3,"doc-seo-83058-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83058,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Canopy A Heterograph Foundation Model for Metabolic Engineering","Designing microbial strains that produce high-value chemicals at commercially viable titers is a central bottleneck in metabolic engineering. CANOPY introduces a heterogeneous graph foundation model that unifies ten public and proprietary sources into a 6.9M-node knowledge graph spanning genes, proteins, metabolites, reactions, pathways, strains, and fermentation experiments. A Heterogeneous Graph Transformer with SignNet encodings and four self-supervised objectives learns multi-modal representations from protein, molecule, and biomedical text embeddings. Frozen embeddings improve fermentation titer prediction (R2=0.41) over tabular and homogeneous GNN baselines.","Canopy: A Heterograph Foundation Model for Metabolic Engineering  \nJake Bowden 1 Laurence Legon 1 Satnam Surae 1  \narXiv :2607 .06224v 1 [ cs .LG] 7 Jul 2026  \nAbstract  \nDesigning microbial strains that produce highvalue chemicals at commercially viable titers remains a central challenge in metabolic engineering. Existing computational approaches either rely on stoichiometric constraint-based models that cannot learn from experimental data, or apply tabular machine learning to hand-crafted features that discard the relational structure of biological knowledge. We present CANOPY, a heterogeneous graph foundation model that integrates ten public and proprietary data sources into a unified knowledge graph (KG) of 6.9 M nodes across  \n13 types and 34 edge types, covering genes, proteins, metabolites, reactions, pathways, strains, and fermentation experiments. Node features are encoded through domain-specific foundation models (ESM-2 for protein sequences, MoLFormer for chemical SMILES, and PubMedBERT for biomedical text), yielding a multi-modal representation within a single graph. We pretrain a Heterogeneous Graph Transformer (HGT) augmented with SignNet positional encodings, Jumping Knowledge aggregation, and virtual nodes using four self-supervised objectives (link prediction, masked node modelling, distance prediction, and contrastive experiment clustering), balanced via learned homoscedastic uncertainty weighting. On the downstream task of fermentation titer prediction, frozen CANOPY embeddings achieve R2 = 0 .41 with a lightweight probe, outperforming tabular baselines (best R2 = 0 .24) and homogeneous GNN variants.  \n1. Introduction  \nThe bioeconomy depends on engineered microorganisms that convert renewable feedstocks into fuels, pharmaceu-  \n1Twig Bio, London, United Kingdom. Correspondence to: Jake Bowden \u003C[jake.bowden@twig.bio](jake.bowden@twig.bio) >, Satnam Surae \u003Csat[nam.surae@twig.bio](nam.surae@twig.bio) >.  \nProceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026 . Copyright 2026 by the author(s) .  \nticals, and specialty chemicals. The design–build–test– learn (DBTL) cycle that underpins strain engineering is slow: a single round of genetic modification, fermentation, and analytical characterisation can take weeks to months (Opgenorth et al., 2019), and the combinatorial space of candidate modifications grows exponentially in the number of target genes. Computational tools that prioritise genetic interventions before wet-lab experiments would shorten this loop.  \nThe dominant computational paradigm for strain design is constraint-based modelling via genome-scale metabolic models (GEMs) . Flux balance analysis (FBA) and its extensions predict steady-state fluxes under stoichiometric and thermodynamic constraints, enabling in silico knockout analysis through tools such as OptKnock (Burgard et al., 2003) and StrainDesign (Schneider et al., 2022) . While GEMs encode mechanistic knowledge of metabolism, they cannot incorporate experimental measurements of titer, rate, and yield; they ignore regulatory, expression-level, and environmental effects; and they treat each organism in isolation without using cross-organism transfer.  \nMachine learning offers a complementary approach. Previous work has applied random forests, gradient boosting, and neural networks to predict fermentation titer from features extracted from strain descriptions (Oyetunde et al., 2019 ; Czajka et al., 2021) . These tabular approaches discard the relational structure that links genes to proteins, proteins to reactions, reactions to pathways, and pathways to production phenotypes. Graph neural networks have been applied to metabolic networks for gene essentiality (Hasibi et al., 2024) and site-of-metabolism prediction (Porokhin et al., 2023), but these operate on single-organism reaction graphs rather than on cross-organism knowledge graphs.  \nMeanwhile, the broader ML community has advanced hetero","cbCaitStg1Lnsnp7","https://ap.wps.com/l/cbCaitStg1Lnsnp7","pdf",430011,2,1,14,"English","en",105,"# Abstract\n# Introduction\n# Related Work\n## Foundation models across biological scales","[{\"question\":\"What problem does CANOPY address in metabolic engineering?\",\"answer\":\"CANOPY targets the challenge of designing microbial strains that produce high-value chemicals at commercially viable titers, where existing tools struggle to leverage experimental data effectively.\"},{\"question\":\"How is the CANOPY knowledge graph constructed?\",\"answer\":\"CANOPY integrates ten public and proprietary data sources into a unified heterogeneous knowledge graph with 6.9M nodes across 13 node types and 34 edge types, covering genes, proteins, metabolites, reactions, pathways, strains, and fermentation experiments.\"},{\"question\":\"How does CANOPY achieve better fermentation titer prediction performance?\",\"answer\":\"CANOPY uses a heterogeneous graph foundation model with multi-modal node features (protein, chemical SMILES, biomedical text) and an augmented Heterogeneous Graph Transformer trained with four self-supervised objectives. Frozen CANOPY embeddings with a lightweight probe reach R2=0.41, outperforming tabular baselines (best R2=0.24) and homogeneous GNN variants.\"}]",1784184925,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"canopy-a-heterograph-foundation-model-for-metabolic-engineering","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/canopy-a-heterograph-foundation-model-for-metabolic-engineering/83058/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does CANOPY address in metabolic engineering?","Question",{"text":75,"@type":76},"CANOPY targets the challenge of designing microbial strains that produce high-value chemicals at commercially viable titers, where existing tools struggle to leverage experimental data effectively.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the CANOPY knowledge graph constructed?",{"text":80,"@type":76},"CANOPY integrates ten public and proprietary data sources into a unified heterogeneous knowledge graph with 6.9M nodes across 13 node types and 34 edge types, covering genes, proteins, metabolites, reactions, pathways, strains, and fermentation experiments.",{"name":82,"@type":73,"acceptedAnswer":83},"How does CANOPY achieve better fermentation titer prediction performance?",{"text":84,"@type":76},"CANOPY uses a heterogeneous graph foundation model with multi-modal node features (protein, chemical SMILES, biomedical text) and an augmented Heterogeneous Graph Transformer trained with four self-supervised objectives. Frozen CANOPY embeddings with a lightweight probe reach R2=0.41, outperforming tabular baselines (best R2=0.24) and homogeneous GNN variants.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]