[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124343-en":3,"doc-seo-124343-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124343,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","Utilizing domain knowledge: Robust machine learning for building energy performance prediction with small, inconsistent datasets","Machine learning in building engineering often struggles when data is sparse or inconsistent, which limits robustness and generalization. The study proposes a knowledge-enhanced strategy that reduces reliance on large datasets by integrating semantic domain knowledge into a Component-Based Machine Learning (CBML) structure. A case experiment compares the approach under extremely low sampling rates (1%–0.0125%) against several typical ML methods. Results show improved robustness, efficient use of diverse records, and effective handling of incomplete data while preserving interpretability and lowering training time. Prerequisites for successful prior-knowledge integration are discussed, and code and datasets are released for reproduction.","Knowledge-Based Systems 294 (2024) 111774  \nContents lists available at ScienceDirect  \nKnowledge-Based Systems  \njournal [homepage: www.elsevier.com/locate/knosys](homepage: www.elsevier.com/locate/knosys)  \n| Utilizing domain knowledge: Robust machine learning for building energy   performance prediction with small, inconsistent datasets\u003Cbr>Xia Chena, *, Manav Mahan Singh b, Philipp Geyera\u003Cbr>a Sustainable Building Systems Group, Institute for Design and Construction, Leibniz University Hannover, Germany b Georg Nemetschek Institute Artificial Intelligence for the Built World, Technical University of Munich, Germany |  |  |\n| --- | --- | --- |\n| A R T I C L E I N F O |  | A B S T R A C T |\n| Keywords:\u003Cbr>Component-based machine learning Compositionality\u003Cbr>Model organization Data utilization Building engineering |  | Machine learning (ML) applications often require large datasets, a requirement that can pose a major challenge in fields where data is sparse or inconsistent. To address this issue, we propose a novel approach that combines prior knowledge with data-driven methods to significantly reduce data dependency. This study represents a disentangled system compositionality knowledge by the method of Component-Based Machine Learning (CBML) in the context of energy-efficient building engineering. In this way, CBML incorporates semantic domain knowledge within the structure of a data-driven model. To understand the advantage of CBML, we conducted a case experiment to assess the effectiveness of this knowledge-encoded ML approach in scenarios with sparse data input (1 % -0.0125 % sampling rate) and several typical ML methods. Our findings reveal three key advantages of this approach over traditional ML methods: 1) It significantly improves the robustness of ML models when dealing with extremely small and inconsistent datasets; 2) It allows for efficient utilization of data from diverse record collections; 3) It can handle incomplete data while maintaining high interpretability and reducing training time. These features offer a promising solution to the challenges associated with deploying data-intensive methods and contribute to more efficient real-world data usage. Additionally, we outline four essential prerequisites to ensure the successful integration of prior knowledge and ML generalization in target scenarios and open-sourced the code and dataset for community reproduction. |\n\n1. Introduction  \nIn a review of the historical path of machine learning (ML) and artificial intelligence (AI), symbolism and connectionism have had many encounters [1,2]. In this study, we delve into their unique interplay. The essence of symbolism lies in its knowledge-based, symbolistic methods, while connectionism thrives in the realm of pattern recognition and prediction, backed by large datasets. Besides traditional knowledge-based, symbolistic methods, the rapid development of AI in the past decade encourages many engineering tasks to use connectionist paradigms for prediction and simulation as decision-making support [3–5]. These models own good generalizations provided in those data-rich and pattern implicit domains, such as image classification [6], object detection [7], natural language processing [8], and even their combination: multimodality [9]. Besides their decent representation learning ability, AI-driven models achieved reliable results, even better than human performance, primarily supported by large datasets [10–12].  \nHowever, the prerequisite of large datasets limits the application of current AI or data-driven methods in domains characterized by complex data-acquisition scenarios or empirical contexts. These domains, often dominated by explicit laws and prior knowledge, have traditionally favored symbolic or semantic-based methods, such as first-principles modeling. Prior knowledge, as a highly compressed abstraction of human experience derived from observation and deduction (e.g., physics), holds great potential for extrap","cbCaic0eCszSHAeL","https://ap.wps.com/l/cbCaic0eCszSHAeL","pdf",2945045,1,10,"English","en",105,"# Introduction\n## Background: symbolism vs. connectionism\n## Motivation: integrating prior knowledge into ML\n## Target application: energy-efficient building engineering\n## Problem framing: simulation limits and sparse data challenge","[{\"question\":\"Why is large-scale data a problem for machine learning in building engineering?\",\"answer\":\"Many building-engineering settings involve complex data acquisition and empirical constraints, leading to sparse or inconsistent datasets. This undermines robustness and generalization of conventional data-driven models.\"},{\"question\":\"What is the proposed method and how does it reduce data dependency?\",\"answer\":\"The approach combines prior semantic domain knowledge with data-driven modeling using Component-Based Machine Learning (CBML). Knowledge is encoded into the model structure so it can learn effectively from limited data.\"},{\"question\":\"What advantages does the study report compared with traditional machine learning methods?\",\"answer\":\"The results indicate improved robustness on extremely small and inconsistent datasets, efficient utilization of data from diverse collections, and the ability to manage incomplete data while keeping high interpretability and reducing training time.\"}]","Utilizing domain knowledge: Robust machine learning for building energy performance prediction with small, inconsistent datasets | PDF",1785821721,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"utilizing-domain-knowledge-robust-machine-learning-for-building-energy-performance-prediction-with-small-inconsistent-datasets","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/utilizing-domain-knowledge-robust-machine-learning-for-building-energy-performance-prediction-with-small-inconsistent-datasets/124343/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is large-scale data a problem for machine learning in building engineering?","Question",{"text":75,"@type":76},"Many building-engineering settings involve complex data acquisition and empirical constraints, leading to sparse or inconsistent datasets. This undermines robustness and generalization of conventional data-driven models.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the proposed method and how does it reduce data dependency?",{"text":80,"@type":76},"The approach combines prior semantic domain knowledge with data-driven modeling using Component-Based Machine Learning (CBML). Knowledge is encoded into the model structure so it can learn effectively from limited data.",{"name":82,"@type":73,"acceptedAnswer":83},"What advantages does the study report compared with traditional machine learning methods?",{"text":84,"@type":76},"The results indicate improved robustness on extremely small and inconsistent datasets, efficient utilization of data from diverse collections, and the ability to manage incomplete data while keeping high interpretability and reducing training time.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":21,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]