[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125037-en":3,"doc-seo-125037-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125037,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Characterizing the Difficulty of Natural Language Datasets for Machine Learning - Dissertation","Natural language classification can be highly accurate on established tasks, yet generalizing to new tasks remains uncertain. This dissertation examines what makes a task difficult for machine learning models by focusing on data across both model training and downstream evaluation. It studies interactions between a task’s dataset and pretrained representations, evaluation protocols, and pretraining data, including alignment under random labeling and performance on small datasets common in in-context learning tests. It further evaluates whether dataset similarity to pretraining data predicts performance, and presents case studies on challenging benchmarks across multiple domains.","CHARACTERIZING THE DIFFICULTY OF NATURAL LANGUAGE DATASETS FOR MACHINE LEARNING  \nA Dissertation  \nPresented to the Faculty of the Graduate School of Cornell University  \nin Partial Fulfillment of the Requirements for the Degree of Doctor of Philosophy  \nby  \nGregory Yauney  \nAugust 2024  \n© 2024 Gregory Yauney  \nALL RIGHTS RESERVED  \nCHARACTERIZING THE DIFFICULTY  \nOF NATURAL LANGUAGE DATASETS FOR MACHINE LEARNING  \nGregory Yauney, Ph.D.  \nCornell University 2024  \nMachine learning models can now achieve high performance on many natural language classification tasks. But we currently don’t know how well a contemporary large language model will perform on a new task without directly trying it out. What makes a task difficult for machine learning models? We focus on the role of data—both a model’s training data and that of downstream tasks—to go beyond evaluation performance in characterizing the difficulty of natural language tasks. We intervene throughout the language modeling pipeline, examining the interaction between a task’s dataset and a) pretrained representations, b) evaluation, and c) pretraining data. We use random labelingsto contextualize the degree of alignment between a task’s data and a task’s labels under different text representations. We use classifiers that guess uniformly at random, independently across examples, to contextualize a language model’s performance on the small datasets typically used to evaluate in-context learning capabilities. We also examine the extent of evidence for the hypothesis that a downstream dataset’s similarity to a model’s pretraining dataset determines the model’s performance. Finally, we turn to case studies across image-text grounding, literary history, and architectural history where we are specifically interested in a model’s performance on a given challenging dataset. Understanding the interaction between data and model will make our models ever more reliable on datasets that we care about, ultimately meeting text datasets where they are.  \nBIOGRAPHICAL SKETCH  \n[Greg came back to upstate New York for graduate school. He earned an Sc.B. in](Greg came back to upstate New York for graduate school. He earned an Sc.B. in)[ ](Greg came back to upstate New York for graduate school. He earned an Sc.B. in)computer science and an A.B. in the history of art and architecture from Brown University.  \nTo my parents.  \niv  \nACKNOWLEDGEMENTS  \nI have been extremely fortunate to have David Mimno as my advisor. Thankyou, David, for being my constant research mentor and collaborator over the past six years. Thank you for supporting me throughout grad school, especially as my interests have wandered from the digital humanities path we originally laid out.  \nI am thankful to Eshan Chattopadhyay and Austin Benson for serving on my committee. I am grateful to the other professors at Cornell who have improved my grad school experience: Robert Kleinberg for a formative first semester, Nika Haghtalab for teaching me that ML theory can be accessible, Matthew Wilkens for insightful research chats from another angle, Jeff Rzeszotarski for improving my visual sensibility, along with Karthik Sridharan, Adrian Sampson, Anil Damle, and Lillian Lee. I am also grateful to the administrative staff in computer science and information science: Becky Stewart, Shannon Adamsen, and Penny Stewart.  \nEmily Reif and Daphne Ippolito hosted me as a Student Researcher at Google Research in Fall 2022 . Thank you, Emily, for being such a willing and generous collaborator in both research and exploring the Cascades. Working together renewed my excitement for research. Anthony Poon helped me out of a tight spot with a place to live in Ballard that fall.  \nThis dissertation contains papers written in collaboration with David Mimno, Emily Reif, Jack Hessel, and Ted Underwood. My work has also greatly benefited from all of my other co-authors throughout grad school: Maria Antoniak, Daphne Ippolito, Katherine Lee, Shayne Longpr","cbCaifd9Ns1oQ6CT","https://ap.wps.com/l/cbCaifd9Ns1oQ6CT","pdf",12321832,1,178,"English","en",105,"# Introduction\n# Data-Label Alignment via Data-Dependent Complexity\n## Introduction\n## D","[{\"question\":\"What is the main goal of this dissertation?\",\"answer\":\"To characterize what makes natural language tasks difficult for machine learning models by analyzing how data shapes performance across training and evaluation.\"},{\"question\":\"How does the dissertation study the relationship between datasets and task labels?\",\"answer\":\"It uses random labeling to contextualize how aligned a task’s data and labels are under different text representations.\"},{\"question\":\"Does similarity between downstream and pretraining datasets determine model performance?\",\"answer\":\"The dissertation examines the hypothesis that downstream dataset similarity to the model’s pretraining dataset influences performance, and reports evidence through multiple analyses and case studies.\"}]","Characterizing the Difficulty of Natural Language Datasets for Machine Learning - Dissertation | PDF",1785896297,449,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"characterizing-the-difficulty-of-natural-language-datasets-for-machine-learning-dissertation","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/characterizing-the-difficulty-of-natural-language-datasets-for-machine-learning-dissertation/125037/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main goal of this dissertation?","Question",{"text":75,"@type":76},"To characterize what makes natural language tasks difficult for machine learning models by analyzing how data shapes performance across training and evaluation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the dissertation study the relationship between datasets and task labels?",{"text":80,"@type":76},"It uses random labeling to contextualize how aligned a task’s data and labels are under different text representations.",{"name":82,"@type":73,"acceptedAnswer":83},"Does similarity between downstream and pretraining datasets determine model performance?",{"text":84,"@type":76},"The dissertation examines the hypothesis that downstream dataset similarity to the model’s pretraining dataset influences performance, and reports evidence through multiple analyses and case studies.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]