[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119901-en":3,"doc-seo-119901-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119901,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","The role of model implementation in neuroscientific applications of machine learning - Thesis","Large-scale machine learning models are increasingly used in modern neuroscience, yet scientific adoption faces unresolved problems in how implementation details shape outcomes. This dissertation studies prediction variability caused by seemingly minor differences such as hardware, operating system, software dependencies, and random seeds. It presents two directions: NeuroCAAS, a cloud platform enabling reproducible, scalable data analysis tool deployment, and large-scale investigations of deep ensembles that use implementation variability for prediction quality and uncertainty characterization.","The role of model implementation in neuroscientific applications of machine learning  \nTaiga Abe  \nSubmitted in partial fulfillment of the  \nrequirements for the degree of  \nDoctor of Philosophy  \nunder the Executive Committee  \nof the Graduate School of Arts and Sciences  \nCOLUMBIA UNIVERSITY  \n© 2023 Taiga Abe All Rights Reserved  \nAbstract  \nThe role of model implementation in neuroscientific applications of machine learning  \nTaiga Abe  \nIn modern neuroscience, large scale machine learning models are becoming increasingly critical components of data analysis. Despite the accelerating adoption of these large scale machine learning tools, there are fundamental challenges to their use in scientific applications that remain largely unaddressed. In this thesis, I focus on one such challenge: variability in the predictions of large scale machine learning models relative to seemingly trivial differences in their implementation. Existing research has shown that the performance of large scale machine learning models (more so than traditional model like linear regression) is meaningfully entangled with design choices such as the hardware components, operating system, software dependencies, and random seed that the corresponding model depends upon. Within the bounds of current practice, there are few ways of controlling this kind of implementation variability across the broad community of neuroscience researchers (making data analysis less reproducible), and little understanding of how data analyses might be designed to mitigate these issues (making data analysis unreliable) . This dissertation will present two broad research directions that address these shortcomings. First, I will describe a novel, cloud-based platform for sharing data analysis tools reproducibly and at scale. This platform, called NeuroCAAS, enables developers of novel data analyses to precisely specify an implementation of their entire data analysis, which can then be used automatically by any other user on custom built cloud resources. I show that this approach is able to efficiently support a wide variety of existing data analysis tools, as well as  \nnovel tools which would not be feasible to build and share outside of a platform like NeuroCAAS. Second, I conduct two large-scale studies on the behavior of deep ensembles. Deep ensembles area class of machine learning model which uses implementation variability to improve the quality of model predictions; in particular, by aggregating the predictions of deep networks over stochastic initialization and training. Deep ensembles simultaneously provide a way to control the impact of implementation variability (by aggregating predictions across random seeds) and also to understand what kind of predictive diversity is generated by this particular form of implementation variability. I present a number of surprising results that contradict widely held intuitions about the performance of deep ensembles as well as the mechanisms behind their success, and show that in many aspects, the behavior of deep ensembles is similar to that of an appropriately chosen single neural network. As a whole, this dissertation presents novel methods and insights focused on the role of implementation variability in large scale machine learning models, and more generally upon the challenges of working with such large models in neuroscience data analysis. I conclude by discussing other ongoing efforts to improve the reproducibility and accessibility of large scale machine learning in neuroscience, as well as long term goals to speed the adoption and reliability of such methods in a scientific context.  \nTable of Contents  \nAcknowledgments . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . xxiii  \nDedication . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . xxv  \nChapter 1: Introduction and Background ........................... 1  \n1.1 Controlling implementation variability with","cbCaidEZkegEZUU9","https://ap.wps.com/l/cbCaidEZkegEZUU9","pdf",12688632,1,232,"English","en",105,"# Table of Contents\n## Chapter 1: Introduction and Background\n## Chapter 2: Cloud based approaches to address issues of reproducibility and scale in neuroscientific data analysis\n## Chapter 3: Properties of deep neural network ensembles on out of distribution data","[{\"question\":\"What problem does the thesis address in machine learning for neuroscience?\",\"answer\":\"It addresses how variability in predictions can arise from trivial implementation differences when using large-scale machine learning models in neuroscience.\"},{\"question\":\"How does NeuroCAAS improve reproducibility and scale?\",\"answer\":\"NeuroCAAS lets developers precisely specify an entire data analysis implementation and run it automatically on custom cloud resources, enabling reproducible sharing at scale.\"},{\"question\":\"What are deep ensembles, and how are they used in the dissertation?\",\"answer\":\"Deep ensembles aggregate predictions from deep networks over stochastic initialization and training. The dissertation studies how this approach both controls the impact of implementation variability and reveals the predictive diversity it generates.\"}]","The role of model implementation in neuroscientific applications of machine learning - Thesis | PDF",1785726912,585,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"the-role-of-model-implementation-in-neuroscientific-applications-of-machine-learning-thesis","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/the-role-of-model-implementation-in-neuroscientific-applications-of-machine-learning-thesis/119901/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04","2026-08-03",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the thesis address in machine learning for neuroscience?","Question",{"text":76,"@type":77},"It addresses how variability in predictions can arise from trivial implementation differences when using large-scale machine learning models in neuroscience.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does NeuroCAAS improve reproducibility and scale?",{"text":81,"@type":77},"NeuroCAAS lets developers precisely specify an entire data analysis implementation and run it automatically on custom cloud resources, enabling reproducible sharing at scale.",{"name":83,"@type":74,"acceptedAnswer":84},"What are deep ensembles, and how are they used in the dissertation?",{"text":85,"@type":77},"Deep ensembles aggregate predictions from deep networks over stochastic initialization and training. The dissertation studies how this approach both controls the impact of implementation variability and reveals the predictive diversity it generates.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]