[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83996-en":3,"doc-seo-83996-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83996,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Multimodal Molecular Representation Learning with Graph Neural Networks, Deep & Cross Networks, and SMILES Embeddings","Multimodal molecular property prediction often over-relies on isolated data modalities, leaving continuous 3D GNNs insufficient for long-range topological dependencies and precise macroscopic heuristics. A parameter-efficient Tri-Branch Modular Fusion Neural Network is introduced to synthesize 3D spatial geometry (SchNet), discrete topological grammar (SMILES via ChemBERTa), and explicit physicochemical descriptors (Deep & Cross Network). Late-fusion learning forms a mathematically rigorous multimodal latent space that mitigates arithmetic and oversmoothing limits. Evaluated on QM9 for atomization energy at 0 K, it attains MAE 0.0207 eV with under one million parameters, reducing error by 20.6% versus a controlled geometric baseline, enabling an efficient surrogate for high-throughput virtual screening.","arXiv :2607 .05736v 1 [ cs .LG] 7 Jul 2026  \nMultimodal Molecular Representation Learning with Graph Neural Networks, Deep & Cross Networks, and SMILES  \nEmbeddings  \nQiwei Han 1,2,∗, Chi Zhou2,∗, Ruobing Wang 1,2 , Zheng Ma 1,2  \n1 Department of Chemistry, Duke University, Durham, NC 27708, USA  \n2 Department of Computer Science, Georgia Institute of Technology, Atlanta, GA 30332, USA  \n∗ These authors contributed equally to this work.  \nAbstract  \nMolecular property prediction often relies on isolated data modalities, where continuous 3D graph neural networks (GNNs) struggle to efficiently capture long-range topological dependencies and exact macroscopic heuristics. In this work, we introduce a parameter-efficient Tri-Branch Modular Fusion Neural Network that synthesizes three orthogonal modalities: 3D spatial geometry (SchNet), discrete topological grammar (SMILES via ChemBERTa), and explicit macroscopic physicochemical descriptors (Deep & Cross Network) . By bypassing standard scalar readouts and employing a shared late-fusion architecture, the framework establishes a mathematically rigorous multimodal latent space that effectively resolves the arithmetic andoversmoothing limitations of local message passing. We evaluate the proposed architecture on the QM9 benchmark, targeting the extensive thermodynamic property of atomization energy at 0 K (Uatom0) . Through systematic combinatorial ablation and latent bottleneck optimization (de = 64), the tri-modal framework achieves a validation Mean Absolute Error (MAE) of 0.0207 eV. Operating with fewer than one million parameters, this architecture decisively surpasses the sub-chemical accuracy threshold and yields a substantial 20.6% error reduction over a strictly controlled geometric baseline. Ultimately, our findings demonstrate that integrating orthogonal macroscopic and topological data streams provides a synergistic, O(1) physical shortcut. This multimodal alignment offers a highly efficient alternative to brute-force parameter scaling, establishing a robust surrogate model for high-throughput virtual screening (HTVS) pipelines.  \nKeywords: Multimodal Fusion, Graph Neural Networks (GNNs), SchNet, Deep & Cross Networks (DCNs), Semantic Embeddings, Molecular Property Prediction, QM9 .  \n1 Introduction  \nThe accurate prediction of molecular properties is a fundamental task in computational chemistry, drug discovery, and materials science. Properties such as atomization energy, dipole moments, and electronic characteristics are traditionally calculated using quantum mechanical methods like Density Functional Theory (DFT) [1] . While DFT provides highly reliable predictions, its immense computational cost strictly limits its applicability in large-scale workflows. Consequently, machine learning (ML) has emerged as a critical surrogate, approximating DFT-derived properties at a fraction of the computational expense.  \nThe development of molecular ML has been accelerated by standardized benchmarks like MoleculeNet [2] and generative models that map discrete representations to continuous latent spaces [3] . Recently, geometric deep learning has dominated this space by learning directly from 3D spatial structures. Foundational models like SchNet [4] demonstrated the power of continuous-filter convolutions, while subsequent equivariant networks have achieved remarkable error floors on standard datasets like QM9 [5] .  \nHowever, this pursuit of absolute precision has driven the field toward heavily parameterized architectures. Relying on complex high-order tensors or massive attention mechanisms, state-ofthe-art equivariant models [6, 7] often require millions of parameters. While these architectures achieve exceptional absolute accuracy, their computational overhead and high VRAM requirements present a bottleneck for real-world applications like high-throughput virtual screening (HTVS) . In HTVS pipelines, where chemical libraries frequently exceed millions of candidate compounds, ra","cbCaiuEoAuKLm2Nj","https://ap.wps.com/l/cbCaiuEoAuKLm2Nj","pdf",608750,3,1,14,"English","en",105,"# Introduction\n## Molecular property prediction and the role of ML\n## Benchmarks and geometric deep learning\n## Limitations of parameter-heavy equivariant models\n## Motivation for orthogonal multimodal fusion\n# Proposed Tri-Branch Modular Fusion Framework","[{\"question\":\"What problem does the proposed approach target in molecular property prediction?\",\"answer\":\"It targets the shortcomings of single-modality learning, where continuous 3D GNNs struggle with long-range topological dependencies and exact macroscopic heuristics, often requiring overly large parameterized models for high precision.\"},{\"question\":\"Which three modalities does the Tri-Branch Modular Fusion Neural Network combine?\",\"answer\":\"It combines 3D spatial geometry encoded with SchNet, discrete topological grammar from SMILES via ChemBERTa, and explicit macroscopic physicochemical descriptors modeled with a Deep \\u0026 Cross Network (DCN).\"},{\"question\":\"How is performance evaluated and what results are reported on QM9?\",\"answer\":\"The model is evaluated on QM9 for atomization energy at 0 K (Uatom0). It reports a validation MAE of 0.0207 eV with fewer than one million parameters and achieves a 20.6% error reduction over a strictly controlled geometric baseline.\"}]",1784191927,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"multimodal-molecular-representation-learning-with-graph-neural-networks-deep-cross-networks-and-smiles-embeddings","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/multimodal-molecular-representation-learning-with-graph-neural-networks-deep-cross-networks-and-smiles-embeddings/83996/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the proposed approach target in molecular property prediction?","Question",{"text":75,"@type":76},"It targets the shortcomings of single-modality learning, where continuous 3D GNNs struggle with long-range topological dependencies and exact macroscopic heuristics, often requiring overly large parameterized models for high precision.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which three modalities does the Tri-Branch Modular Fusion Neural Network combine?",{"text":80,"@type":76},"It combines 3D spatial geometry encoded with SchNet, discrete topological grammar from SMILES via ChemBERTa, and explicit macroscopic physicochemical descriptors modeled with a Deep & Cross Network (DCN).",{"name":82,"@type":73,"acceptedAnswer":83},"How is performance evaluated and what results are reported on QM9?",{"text":84,"@type":76},"The model is evaluated on QM9 for atomization energy at 0 K (Uatom0). It reports a validation MAE of 0.0207 eV with fewer than one million parameters and achieves a 20.6% error reduction over a strictly controlled geometric baseline.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]