[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-127331-en":3,"doc-seo-127331-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},127331,962085570644,"Evangeline","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Yield Prediction with Explainable Machine Learning - Dissertation","Starting from a federal project to predict grapevine yields in Germany, the work addresses five practical challenges to enable machine learning for yield prediction. It tackles small-data training via feature-based remote-sensing representations that improve gradient boosting by 25% in soybean experiments. It adds grouped Shapley explanations aligned with expert knowledge, introduces polynomial-time grouped Shapley computation for random forests, and studies Shapley-based feature selection under explicit conditions. It also mitigates remote-sensing gaps with a U-Net plus partial convolutions interpolation, improving RMSE by 44%, and improves cross-domain prediction using regularized transfer learning with a 16% RMSE gain.","Yield Prediction with Explainable Machine Learning  \nDissertation  \nzur  \nErlangung des Doktorgrades (Dr. rer. nat.)  \nder  \nMathematisch-Naturwissenschaftlichen Fakult¨at  \nder  \nRheinischen Friedrich-Wilhelms-Universit¨at Bonn  \nvorgelegt von Florian Philipp Huber aus  \nMechernich  \nBonn 2024  \nAngefertigt mit Genehmigung der Mathematisch-Naturwissenschaftlichen Fakult¨at  \nder Rheinischen Friedrich-Wilhelms-Universit¨at Bonn  \nGutachter / Betreuer: PD Dr. Volker Steinhage Gutachterin: Prof. Dr. Elena Demidova  \nTag der Promotion: 31.10.2024  \nErscheinungsjahr: 2024  \nAbstract  \nStarting from a federal project to predict grapevine yields in Germany, we faced five challenges to enable machine learning for yield prediction. The first challenge is training on small data sets, as capturing data for yield prediction is very time consuming with most plants following an annual cycle. Providing a feature-based representation of remote sensing data by modeling underlying distributions allows gradient boosting methods to outperform deep learning approaches by 25% in our experiments for soybean yield prediction in the US, one of the biggest datasets for yield prediction that allows international comparability. The second challenge is the need for explanations to show that the model’s decision making is in-line with experts knowledge of the field. For this challenge, we extend the idea of Shapley value feature attributions to predefined groups of features. The groupings are naturally given for yield prediction scenarios and allow for an improved representation of the explanations, as individual features are plentiful and often abstract. We give a novel algorithm to solve the problem of calculating the grouped Shapley values in polynomial time for random forests as they result from the gradient boosting pipeline from challenge one. Third, we work towards better featureselection for yield prediction tasks. The introduction of grouped Shapley values sparks the question of whether Shapley values could be used for feature selection. To address this question, we define four necessary conditions for defining a Shapley value suitable for feature selection. Additionally, we analyze the problem of model averaging where unimportant features are allowed to alter the final feature selection by introducing a novel exhaustive feature selection tool that has no problems with model averaging, and use it to further evaluate Shapley values for feature selection. Our experiments indicate that there is a small loss in accuracy due to model averaging, while the runtime of Shapley values as a heuristic measure for feature selection is superior for random forests. The fourth challenge is handling gaps in remote sensing data. As we need to use remote sensing data to provide consistent coverage for a small research area, clouds that occlude the satellite’s view on the Earth can hide a meaningful amount of data. We approach this challenge by introducing a novel deep interpolation pipeline that uses a U-Net structure together with partial convolutions to gradually fill in remote sensing data in our research area, finally improving previously established statistical methods by 44% in terms of RMSE. Lastly, we worked towards a solution to make predictions for shifting domains, where we used regularized transfer learning to improve yield prediction by transferring knowledge between different domains by 16% in terms of RMSE, compared to not using transfer learning techniques.  \nAcknowledgments  \nForemost, I would like to express my deepest gratitude to PD Dr. Volker Steinhage for offering me the opportunity to work in his research group and for providing invaluable guidance and advice throughout the entire process. His support has been instrumental in the completion of this thesis.  \nI am also grateful to Frank and Timm, with whom I shared an office, for their constant support, numerous discussions, and all the proofreading of my papers. Their collaboration and encour","cbCaih1WfAqcDNFJ","https://ap.wps.com/l/cbCaih1WfAqcDNFJ","pdf",13826564,1,132,"English","en",105,"# 1 Introduction\n## 1.1 The Challenges of Explainable Yield Prediction\n## 1.2 Contributions\n## 1.3 Thesis Structure\n# 2 Related Work\n## 2.1 Yield Prediction\n## 2.2 Explainability and Shapley Values\n## 2.3 Feature Selection\n## 2.4 Filling Gaps in Remote Sensing LST Data","[{\"question\":\"What were the main challenges addressed for explainable yield prediction?\",\"answer\":\"The dissertation targets five challenges: learning from small datasets, providing explanations aligned with expert knowledge, improving feature selection, handling missing remote-sensing data due to occlusions, and making predictions under shifting domains.\"},{\"question\":\"How does the work improve model performance with limited training data?\",\"answer\":\"It uses a feature-based representation of remote sensing data by modeling underlying distributions, enabling gradient boosting methods to outperform deep learning by 25% in experiments for soybean yield prediction.\"},{\"question\":\"How are remote-sensing gaps handled and what improvement is reported?\",\"answer\":\"A deep interpolation pipeline based on a U-Net with partial convolutions gradually fills missing remote-sensing values, improving previously established statistical methods by 44% in RMSE.\"}]","Yield Prediction with Explainable Machine Learning - Dissertation | PDF",1785938325,333,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"yield-prediction-with-explainable-machine-learning-dissertation","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/yield-prediction-with-explainable-machine-learning-dissertation/127331/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What were the main challenges addressed for explainable yield prediction?","Question",{"text":76,"@type":77},"The dissertation targets five challenges: learning from small datasets, providing explanations aligned with expert knowledge, improving feature selection, handling missing remote-sensing data due to occlusions, and making predictions under shifting domains.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the work improve model performance with limited training data?",{"text":81,"@type":77},"It uses a feature-based representation of remote sensing data by modeling underlying distributions, enabling gradient boosting methods to outperform deep learning by 25% in experiments for soybean yield prediction.",{"name":83,"@type":74,"acceptedAnswer":84},"How are remote-sensing gaps handled and what improvement is reported?",{"text":85,"@type":77},"A deep interpolation pipeline based on a U-Net with partial convolutions gradually fills missing remote-sensing values, improving previously established statistical methods by 44% in RMSE.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]