[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118206-en":3,"doc-seo-118206-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118206,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Can machine-learning algorithms improve upon classical palaeoenvironmental reconstruction models?","Classical palaeoenvironmental reconstruction models often embed biological assumptions, especially that taxa in fossil assemblages follow unimodal response functions to environmental variables. Machine-learning methods avoid these biological constraints but require large training datasets to learn relationships between biological assemblages and environment. A two-layer ensemble reconstruction model (MEMLM) is developed using three base ensemble learners and a consensus integration layer, evaluated against weighted-averaging approaches on multiple diatom and pollen training sets and tested on fossil data.","Open Research Online  \nCitation  \nSun, Peng; Holden, Philip and Birks, H John B (2024) . Can machine-learning algorithms improve upon classical palaeoenvironmental reconstruction models? Climate of the Past, 20(10) pp. 2373–2398.  \nURL  \n[https://oro.open.ac.uk/100715/](https://oro.open.ac.uk/100715/)  \nLicense  \n(CC-BY 4.0) Creative Commons: Attribution 4.0  \n[https://creativecommons.org/licenses/by/4.0/](https://creativecommons.org/licenses/by/4.0/)  \nPolicy  \nThis document has been downloaded from Open Research Online, The Open University's repository of research publications. This version is being made available in accordance with Open Research Online policies available from Open Research Online (ORO) Policies  \nVersions  \nIf this document is identified as the Author Accepted Manuscript it is the version after peer review but before type setting, copy editing or publisher branding  \nClim. Past, 20, 2373–2398, 2024  \n[https://doi.org/10.5194/cp-20-2373-2024](https://doi.org/10.5194/cp-20-2373-2024)[ ](https://doi.org/10.5194/cp-20-2373-2024)© Author(s) 2024 . This work is distributed under the Creative Commons Attribution 4 .0 License.  \nCan machine-learning algorithms improve upon classical palaeoenvironmental reconstruction models?  \nPeng Sun 1 , Philip B. Holden2 , and H. John B. Birks3,4  \n1Institute of Environmental Sciences (CML), Leiden University, 2333 CC Leiden, the Netherlands  \n2Environment, Earth and Ecosystem Sciences, The Open University, Walton Hall, Milton Keynes, MK7 6AA, UK  \n3Department of Biological Sciences and Bjerknes Centre for Climate Research, University of Bergen, PO Box 7803, 5020 Bergen, Norway  \n4Environmental Change Research Centre, University College London, London, WC1 6BT, UK Correspondence: Philip B. Holden ([philip.holden@open.ac.uk](philip.holden@open.ac.uk))  \nReceived: 30 August 2023 – Discussion started: 27 September 2023  \nRevised: 5 August 2024 – Accepted: 3 September 2024 – Published: 24 October 2024  \nAbstract. Classical palaeoenvironmental reconstruction models often incorporate biological ideas and commonly assume that the taxa comprising a fossil assemblage exhibit unimodal response functions of the environmental variable of interest. In contrast, machine-learning approaches do not rely upon any biological assumptions but instead need training with large data sets to extract some understanding of the relationships between biological assemblages and their environment. To explore the relative merits of these two approaches, we have developed a two-layered machine-learning reconstruction model MEMLM (Multi Ensemble Machine Learning Model) . The ﬁrst layer applies three different ensemble machine-learning models (random forests, extra random trees, and LightGBM), trained on the modern taxon assemblage and associated environmental data to make reconstructions based on the three different models, while the second layer uses multiple linear regression to integrate these three reconstructions into a consensus reconstruction. We considered three versions of the model: (1) a standard version of MEMLM, which uses only taxon abundance data;  \n(2) MEMLMe, which uses only dimensionally reduced assemblage information, using a natural language-processing model (GloVe), to detect associations between taxa across the training data set; and (3) MEMLMc which incorporates both raw taxon abundance and dimensionally reduced summary (GloVe) data. We trained these MEMLM model variants with three high-quality diatom and pollen training sets and compared their reconstruction performance with three weighted-averaging (WA) approaches (WA-Cla for classi-  \ncal deshrinking, WA-Inv for inverse deshrinking, and WAPLS for partial least squares) . In general, the MEMLM approaches, even when trained on only dimensionally reduced assemblage data, performed substantially better than the WA approaches in the larger training sets, as judged by cross-validatory prediction error. When applied to fossil data, MEMLM vari","cbCaivGaSR7Bdh5e","https://ap.wps.com/l/cbCaivGaSR7Bdh5e","pdf",9434700,1,27,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What is the core difference between classical palaeoenvironmental reconstruction models and machine-learning approaches?\",\"answer\":\"Classical models typically assume biological response shapes (e.g., unimodal taxon-environment functions), while machine-learning approaches do not rely on such biological assumptions and instead learn relationships from large training datasets.\"},{\"question\":\"How does the MEMLM reconstruction model work?\",\"answer\":\"MEMLM uses a first layer with three ensemble models—random forests, extra random trees, and LightGBM—trained on modern taxon and environmental data, then a second layer applies multiple linear regression to combine the three reconstructions into a consensus.\"},{\"question\":\"What do the comparisons and tests show about model performance and reliability?\",\"answer\":\"In larger training sets, MEMLM variants generally outperform weighted-averaging approaches by cross-validated prediction error, but fossil reconstructions can differ qualitatively and may fail under extrapolation. Statistical significance testing identifies cases where reconstructions are not robust to model choice, showing that cross-validation alone is insufficient for transfer-function performance.\"}]","Can machine-learning algorithms improve upon classical palaeoenvironmental reconstruction models? | PDF",1785682181,68,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"can-machine-learning-algorithms-improve-upon-classical-palaeoenvironmental-reconstruction-models","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/can-machine-learning-algorithms-improve-upon-classical-palaeoenvironmental-reconstruction-models/118206/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the core difference between classical palaeoenvironmental reconstruction models and machine-learning approaches?","Question",{"text":75,"@type":76},"Classical models typically assume biological response shapes (e.g., unimodal taxon-environment functions), while machine-learning approaches do not rely on such biological assumptions and instead learn relationships from large training datasets.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the MEMLM reconstruction model work?",{"text":80,"@type":76},"MEMLM uses a first layer with three ensemble models—random forests, extra random trees, and LightGBM—trained on modern taxon and environmental data, then a second layer applies multiple linear regression to combine the three reconstructions into a consensus.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the comparisons and tests show about model performance and reliability?",{"text":84,"@type":76},"In larger training sets, MEMLM variants generally outperform weighted-averaging approaches by cross-validated prediction error, but fossil reconstructions can differ qualitatively and may fail under extrapolation. Statistical significance testing identifies cases where reconstructions are not robust to model choice, showing that cross-validation alone is insufficient for transfer-function performance.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]