[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128069-en":3,"doc-seo-128069-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},128069,5909887254083,"Miles","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","ACCELERATING SCIENTIFIC RESEARCH - EMPOWERING AUTOMATED KNOWLEDGE DISCOVERY THROUGH MACHINE LEARNING - Thesis","Automatic knowledge discovery is crucial for advancing scientific research, especially when real-world data contains gaps and when domain language requires structured understanding. This thesis develops machine learning frameworks to generate novel knowledge from massive datasets. It presents GATE, a graph-based variational auto-encoder for missing node feature imputation in networks, addressing limitations of traditional methods. It also introduces a multimodal framework for fine-grained chemical entity typing from literature, leveraging chemical structures and cross-modal attention, with the CHEMET benchmark dataset, showing improved performance over state-of-the-art approaches.","© 2023 Chenkai Sun  \nACCELERATING SCIENTIFIC RESEARCH: EMPOWERING AUTOMATED KNOWLEDGE DISCOVERY THROUGH MACHINE LEARNING  \nBY  \nCHENKAI SUN  \nTHESIS  \nSubmitted in partial fulfillment of the requirements  \nfor the degree of Master of Science in Computer Science  \nin the Graduate College of the  \nUniversity of Illinois Urbana-Champaign, 2023  \nUrbana, Illinois  \nAdvisers:  \nProfessor Heng Ji  \nProfessor ChengXiang Zhai  \nABSTRACT  \nAutomatic knowledge discovery is vital for the progress of scientific research. For example, tackling missing data is especially needed in health care and social network domains, and being able to model complex chemical structures at the text level brings enormous benefits to chemistry and biomedical domains. We aim to develop machine learning-based frameworks that are capable of generating novel knowledge by learning from massive real-world data to help accelerate scientific research. The first part of the thesis introduces GATE, a graphbased variational auto-encoder framework, developed in response to the prevalent issue of missing node features in networks and that traditional methods have fallen short inadequately addressing this challenge. The second part of the thesis tackles the intricate task of predicting fine-grained chemical entity types from chemical literature, a critical component in advancing biomedical and chemical research. It presents a novel multi-modal representation learning framework and a newly created benchmark dataset, CHEMET. This framework leveraged external resources with chemical structures and used cross-modal attention to learn an effective representation of text in the chemistry domain. Experiments show that our approaches significantly outperform state-of-the-art methods.  \nACKNOWLEDGMENTS  \nPart of the research is based upon work supported by the Molecule Maker Lab Institute: An AI Research Institutes program supported by NSF under Award No. 2019897 and NSF No. 2034562. The views and conclusions contained herein are those of the authors and should not be interpreted as necessarily representing the official policies, either expressed or implied, of the U.S. Government. The U.S. Government is authorized to reproduce and distribute reprints for governmental purposes notwithstanding any copyright annotation therein  \nTABLE OF CONTENTS  \nCHAPTER 1 INTRODUCTION ............................ 1  \nCHAPTER 2 MISSING DATA IMPUTATION WITH VARIATIONAL GRAPH NEURAL NETWORKS ................................ 2  \n2.1 Introduction .................................... 2  \n2.2 Related Work ................................... 4  \n2.3 Variational Graph Imputation Nets ....................... 5  \n2.4 Experiments .................................... 11  \n2.5 Summary ..................................... 15  \nCHAPTER 3 FINE-GRAINED CHEMICAL ENTITY TYPING WITH MULTIMODAL KNOWLEDGE REPRESENTATION ................... 17  \n3.1 Introduction .................................... 17  \n3.2 Dataset ...................................... 20  \n3.3 Method ...................................... 23  \n3.4 Experiments .................................... 26  \n3.5 Related Work ................................... 29  \n3.6 Summary ..................................... 31  \nCHAPTER 4 CONCLUSIONS ............................. 32  \nREFERENCES ....................................... 33  \nCHAPTER 1: INTRODUCTION  \nThe rapid growth of scientific research has led to an explosion of data, particularly in the form of complex networks and text. This vast amount of information presents both a challenge and an opportunity. The challenge lies in effectively extracting and analyzing this data to uncover hidden patterns and insights. On the other hand, the opportunity emergesin harnessing this newfound knowledge to propel scientific advancements and tackle tangible, real-world issues.  \nIn this thesis, we address two critical challenges in the realm of knowledge discovery: missing data imputation in networks and fine-grained e","cbCaifeDEKTYrYV5","https://ap.wps.com/l/cbCaifeDEKTYrYV5","pdf",1258247,1,44,"English","en",105,"# Chapter 1 Introduction\n## Knowledge discovery challenges\n## Missing data imputation in networks\n## Fine-grained entity typing in chemical literature\n# Chapter 2 Missing Data Imputation with Variational Graph Neural Networks\n## Introduction\n## Related Work\n## Variational Graph Imputation Nets\n## Experiments\n## Summary\n# Chapter 3 Fine-Grained Chemical Entity Typing with Multimodal Knowledge Representation\n## Introduction\n## Dataset\n## Method\n## Experiments\n## Related Work\n## Summary\n# Chapter 4 Conclusions","[{\"question\":\"What problems does the thesis focus on for automated knowledge discovery?\",\"answer\":\"It focuses on missing data imputation in network datasets and fine-grained entity typing in chemical literature, both of which are important for accelerating scientific progress.\"},{\"question\":\"How does GATE address missing node features in networks?\",\"answer\":\"GATE uses a graph-based variational auto-encoder that incorporates relational information to produce more accurate and reliable imputations than traditional approaches.\"},{\"question\":\"What is the core idea behind the chemical entity typing framework and CHEMET?\",\"answer\":\"The framework learns from chemical literature using multimodal representation learning that leverages chemical structures and cross-modal attention, and it is evaluated on CHEMET to measure fine-grained typing performance.\"}]","ACCELERATING SCIENTIFIC RESEARCH - EMPOWERING AUTOMATED KNOWLEDGE DISCOVERY THROUGH MACHINE LEARNING - Thesis | PDF",1785944648,111,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"accelerating-scientific-research-empowering-automated-knowledge-discovery-through-machine-learning-thesis","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/accelerating-scientific-research-empowering-automated-knowledge-discovery-through-machine-learning-thesis/128069/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problems does the thesis focus on for automated knowledge discovery?","Question",{"text":76,"@type":77},"It focuses on missing data imputation in network datasets and fine-grained entity typing in chemical literature, both of which are important for accelerating scientific progress.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does GATE address missing node features in networks?",{"text":81,"@type":77},"GATE uses a graph-based variational auto-encoder that incorporates relational information to produce more accurate and reliable imputations than traditional approaches.",{"name":83,"@type":74,"acceptedAnswer":84},"What is the core idea behind the chemical entity typing framework and CHEMET?",{"text":85,"@type":77},"The framework learns from chemical literature using multimodal representation learning that leverages chemical structures and cross-modal attention, and it is evaluated on CHEMET to measure fine-grained typing performance.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]