[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-127801-en":3,"doc-seo-127801-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},127801,1099523885074,"Ivy","https://ap-avatar.wpscdn.com/davatar_9964176cb1d06d4a9deccf72a44ae3dc",8,"Research & Report","Scalable Distributed Machine Learning for Knowledge Graphs - Dissertation - Knowledge Graphs","Due to accelerating digitization, massive heterogeneous data sets are often summarized as Big Data, creating the need for robust data integration and analytics. Knowledge Graphs represent diverse sources as directed multi-graphs connected by unique resource identifiers, enabling prediction and analytics. In this work, scalable distributed and explainable machine learning methods are developed for KG data too large for single-machine memory. The approach generates fixed-length numeric feature vectors via graph-kernel-like map-reduce feature extraction, including multi-modal literals, with SPARQL-based reusable, reproducible, and transparent pipelines. Semantic native KG outputs preserve hyper-parameters and explainability. The work extends the SANSA technology stack and supports deployment in notebooks and REST API environments.","Scalable Distributed Machine Learning for Knowledge Graphs  \nDissertation  \nzur  \nErlangung des Doktorgrades (Dr. rer. nat.)  \nder  \nMathematisch-Naturwissenschaftlichen Fakultät  \nder  \nRheinischen Friedrich-Wilhelms-Universität Bonn  \nvorgelegt von  \nCarsten Felix Draschner  \naus  \nBergisch Gladbach, Germany  \nBonn 2022  \nAngefertigt mit Genehmigung der Mathematisch-Naturwissenschaftlichen Fakultät der Rheinischen Friedrich-Wilhelms-Universität Bonn  \n1. Gutachter: Prof. Dr. Jens Lehmann  \n2. Gutachter: Prof. Dr. Stefan Wrobel  \nTag der Promotion: 23.06.2023  \nErscheinungsjahr: 2023  \nAbstract  \nDue to the increasing progress of digitization, immense amounts of data are accumulating, which can be summarized under the term Big Data and form an exciting basis for data analyses. Since the data are heterogeneous and come from many different sources, data integration techniques are beneficial to perform analytics. Knowledge Graphs (KG) link the heterogeneous data within a directed multi-graph by unique resource identifiers. These data can be used for data analytics and prediction methods. One subbranch of Artificial Intelligence (AI) is Machine Learning (ML) . ML models are developed and trained, which, based on the available training data, should approximate the target data as closely as possible. The samples in the training data are usually represented by features. For most data analyticsand ML approaches, these features are fixed-length numeric feature vectors. However, in the context of KGs, there is no native representation within fixed-length numeric feature vectors. Depending on the use case, these problems can also require the concrete use and inclusion of individual actual values from the KG. The sheer size of some large-scale KG data does not fit into the memory of today’s computers. One solution is to use cluster computation through distributed execution, which distributes the data and processing tasks across multiple computers. Both the technologies and the algorithms for this distributed computation must be designated. Due to the possible impact of the results from these data analysis pipelines, special technical implementation of accessible, reproducible, reusable, and explainable approaches is beneficial. These ML and AI development meta-dimensions belong to Ethical AI and Sustainable AI concepts. Within this work, we developed novel approaches for ML on KGs while considering ethical and sustainability dimensions. In particular, we developed technologies that create fixed-length numeric feature vectors. These include methods that, like graph kernels, extract features from the graph in the context of the map-reduce operations relevant for distributed computation. The feature extraction also includes the multi-modal data of KG literals. Accordingly, we have developed methods that enable SPARQL-based feature extraction and assist in creating complex feature-extracting queries. Based on these extracted features, we further contributed scalable, distributed, and explainable ML and data analytics methods such as semantic similarity estimation and classification or regression ML pipelines demonstrating noticeable performance. We support the transparency, reusability, and reproducibility of our novel open-source approaches by results and meta-data semantification. This semantification transfers the original graph data with the hyper-parameter setup and explainability information, in addition to the predicted results of the ML pipelines, into a semantic native KG. Due to the technological complexity, we enable the application of our algorithm technologies through complementary work such as the use in coding notebooks and the use in Rest API-based environments. Our work also describes the multidimensional and interwoven optimization dimensions of ethical and sustainable KG-based ML. We extended the existing technology stack SANSA, which is used for distributed processing and native semantic data handling, by several scientif","cbCaikevLbpBNhwb","https://ap.wps.com/l/cbCaikevLbpBNhwb","pdf",23941451,1,143,"English","en",105,"# Abstract\n## Scalable distributed ML for knowledge graphs\n## Fixed-length feature vectors and SPARQL feature extraction\n## Explainability, reproducibility, and semantic semantification\n## Ethical and sustainable optimization dimensions\n# Acknowledgements","[{\"question\":\"What problem does the dissertation address for knowledge graph machine learning?\",\"answer\":\"Large-scale knowledge graph data often cannot fit into the memory of current computers, and knowledge graph structures lack native fixed-length numeric feature representations. The work addresses scalability through distributed execution and representation through feature-vector construction.\"},{\"question\":\"How are fixed-length numeric feature vectors created in the proposed approach?\",\"answer\":\"The dissertation develops methods that extract features from the knowledge graph using graph-kernel-like ideas aligned with map-reduce operations for distributed computation. It also incorporates multi-modal literal data into the feature extraction.\"},{\"question\":\"How does the work support transparency and explainability of the ML pipelines?\",\"answer\":\"The results include semantic-native knowledge graph outputs that preserve the hyper-parameter setup and explainability information alongside predicted outcomes. The effort is designed to improve accessibility, reusability, and reproducibility of the open-source approaches.\"}]","Scalable Distributed Machine Learning for Knowledge Graphs - Dissertation - Knowledge Graphs | PDF",1785941881,360,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"scalable-distributed-machine-learning-for-knowledge-graphs-dissertation-knowledge-graphs","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/scalable-distributed-machine-learning-for-knowledge-graphs-dissertation-knowledge-graphs/127801/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the dissertation address for knowledge graph machine learning?","Question",{"text":76,"@type":77},"Large-scale knowledge graph data often cannot fit into the memory of current computers, and knowledge graph structures lack native fixed-length numeric feature representations. The work addresses scalability through distributed execution and representation through feature-vector construction.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How are fixed-length numeric feature vectors created in the proposed approach?",{"text":81,"@type":77},"The dissertation develops methods that extract features from the knowledge graph using graph-kernel-like ideas aligned with map-reduce operations for distributed computation. It also incorporates multi-modal literal data into the feature extraction.",{"name":83,"@type":74,"acceptedAnswer":84},"How does the work support transparency and explainability of the ML pipelines?",{"text":85,"@type":77},"The results include semantic-native knowledge graph outputs that preserve the hyper-parameter setup and explainability information alongside predicted outcomes. The effort is designed to improve accessibility, reusability, and reproducibility of the open-source approaches.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]