[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118979-en":3,"doc-seo-118979-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118979,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","Comprehensive Evaluation of Machine Learning Experiments - Algorithm Comparison, Algorithm Performance and Inferential Reproducibility","This doctoral thesis examines key methodological issues in machine learning experimentation to improve how algorithm performance is evaluated, analyzed, and interpreted. It critiques the widely used train-dev-test paradigm by highlighting neglected sources of uncertainty, including algorithm variability and dependence on meta-parameters. The work develops a comprehensive statistical framework based on Linear Mixed Effects Models to model expected risk probabilistically, quantify reliability and performance homogeneity, and connect annotation processes to algorithm effects. It further proposes a unified approach to support inferential reproducibility across datasets.","Comprehensive Evaluation of Machine Learning Experiments  \nAlgorithm Comparison, Algorithm Performance and Inferential Reproducibility  \nInauguraldissertation zur Erlangung des akademischen Grades Doktor der Philosophie der Neuphilologischen Fakultät  \nder Ruprecht-Karls-Universität Heidelberg vorgelegt von  \nMichael Hagmann  \nam  \n07. August 2023  \nHierbei handelt es sich um eine Heidelberger Dissertation.  \nErstgutachter: Prof. Dr. Stefan Riezler  \nInstitut für Computerlinguistik, Universität Heidelberg  \nZweitgutachter: Prof. Dr. Katja Markert  \nInstitut für Computerlinguistik, Universität Heidelberg  \nAbstract  \nThis doctoral thesis addresses critical methodological aspects within machine learning experimentation, focusing on enhancing the evaluation and analysis of algorithm performance. The established \"train-dev-test paradigm\" commonly guides machine learning practitioners, involving nested optimization processes to optimize model parameters and meta-parameters and benchmarking against test data. However, this paradigm overlooks crucial aspects, such as algorithm variability and the intricate relationship between algorithm performance and meta-parameters. This work introduces a comprehensive framework that employs statistical techniques to bridge these gaps, advancing the methodological standards in empirical machine learning research. The foundational premise of this thesis lies in differentiating between algorithms and classifiers, recognizing that an algorithm may yield multiple classifiers due to inherent stochasticity or design choices. Consequently, algorithm performance becomes inherently probabilistic and cannot be captured by a single metric. The contributions of this work are structured around three core themes:  \nAlgorithm Comparison: A fundamental aim of empirical machine learning research is algorithm comparison. To this end, the thesis proposes utilizing Linear Mixed Effects Models (LMEMs) for analyzing evaluation data. LMEMs offer distinct advantages by accommodating complex data structures beyond the typical independent and identically distributed (iid) assumption. Thus LMEMs enable a holistic analysis of algorithm instances and facilitate the construction of nuanced conditional models of expected risk, supporting algorithm comparisons based on diverse data properties.  \nAlgorithm Performance Analysis: Contemporary evaluation practices often treat algorithms and classifiers as black boxes, hindering insights into their performance and parameter dependencies. Leveraging LMEMs, specifically implementing Variance Component Analysis, the thesis introduces methods from psychometrics to quantify algorithm performance homogeneity (reliability) and assess the influence of meta-parameters on  \nperformance. The flexibility of LMEMs allows a granular analysis of this relationship and extends these techniques to analyze data annotation processes linked to algorithm performance.  \nInferential Reproducibility: Building upon the preceding chapters, this section showcases a unified approach to analyze machine learning experiments comprehensively. By leveraging the full range of generated model instances, the analysis provides a nuanced understanding of competing algorithms. The outcomes offer implementation guidelines for algorithmic modifications and consolidate incongruent findings across diverse datasets, contributing to a coherent empirical perspective on algorithmic effects.  \nThis work underscores the significance of addressing algorithmic variability, metaparameter impact, and the probabilistic nature of algorithm performance. This thesis aims to enhance machine learning experiments’ transparency, reproducibility, and interpretability by introducing robust statistical methodologies facilitating extensive empirical analysis. It extends beyond conventional guidelines, offering a principled approach to advance the understanding and evaluation of algorithms in the evolving landscape of machine learning and data scien","cbCaiglGvKiyAGGC","https://ap.wps.com/l/cbCaiglGvKiyAGGC","pdf",3988270,1,140,"English","en",105,"# Introduction\n## Research Question and Contribution\n# Algorithm Comparison\n## The Principles of Statistical Hypothesis Testing\n## Model-Based Algorithm Comparison: Toolbox\n## Model Based Algorithm Comparison: Analyzing an Example\n# Algorithm Performance Analysis\n## Untangling Terminology: Reliability, Agreement etc\n## A Unifying View on Algorithm Performance and Annotation\n## State of the Art Methods for Annotation and Algorithm Performance Analysis\n## Bridging the Gap: Model-based Reliability Analysis\n# Inferential Reproducibility\n## A Scheme for Analyzing Inferential Reproducibility\n## Linear Mixed Effects Models\n## Generalized Likelihood Ratio Tests\n## Variance Component Analysis and Reliability Coefficie","[{\"question\":\"How does the thesis address limitations of the train-dev-test paradigm?\",\"answer\":\"It identifies that the paradigm overlooks algorithm variability and the complex relationship between algorithm performance and meta-parameters. The thesis introduces statistical methods to model these effects more comprehensively.\"},{\"question\":\"What statistical approach is proposed for algorithm comparison?\",\"answer\":\"The thesis proposes using Linear Mixed Effects Models (LMEMs) to analyze evaluation data while accommodating complex data structures. This enables nuanced conditional models of expected risk for comparisons under different data properties.\"},{\"question\":\"How is inferential reproducibility supported in the thesis?\",\"answer\":\"It presents a unified analysis strategy that leverages the full range of generated model instances. The approach consolidates findings across datasets and provides guidance for principled algorithm modifications.\"}]","Comprehensive Evaluation of Machine Learning Experiments - Algorithm Comparison, Algorithm Performance and Inferential Reproducibility | PDF",1785721339,353,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"comprehensive-evaluation-of-machine-learning-experiments-algorithm-comparison-algorithm-performance-and-inferential-reproducibility","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/comprehensive-evaluation-of-machine-learning-experiments-algorithm-comparison-algorithm-performance-and-inferential-reproducibility/118979/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does the thesis address limitations of the train-dev-test paradigm?","Question",{"text":75,"@type":76},"It identifies that the paradigm overlooks algorithm variability and the complex relationship between algorithm performance and meta-parameters. The thesis introduces statistical methods to model these effects more comprehensively.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What statistical approach is proposed for algorithm comparison?",{"text":80,"@type":76},"The thesis proposes using Linear Mixed Effects Models (LMEMs) to analyze evaluation data while accommodating complex data structures. This enables nuanced conditional models of expected risk for comparisons under different data properties.",{"name":82,"@type":73,"acceptedAnswer":83},"How is inferential reproducibility supported in the thesis?",{"text":84,"@type":76},"It presents a unified analysis strategy that leverages the full range of generated model instances. The approach consolidates findings across datasets and provides guidance for principled algorithm modifications.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]