[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117354-en":3,"doc-seo-117354-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117354,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","From Public Health to AI Safety - Improving Machine Learning Approaches by Collecting, Selecting, or Reducing the Need for High-Quality Data","This thesis advances multiple application areas spanning public health and AI safety through novel data-centric machine learning methods. It introduces a training-data selection approach, rigorous large-scale manual collection processes, synthetic data generation when data is scarce, and highly diverse test datasets to measure generalisability. For AI misuse and misalignment, it builds a detector for LLM “lies,” using LLM prompting to produce lie-labeled training data and studying generalisation to unseen settings. It also addresses low-quality abundant data via automatic selection during training, improving speed, accuracy, and compute efficiency. Finally, it collects carefully validated high-quality COVID-19 datasets and fits Bayesian models to estimate NPI effects with robust validation that supports policy decisions.","From Public Health to AI Safety: Improving Machine Learning Approaches by Collecting, Selecting, or Reducing the Need for  \nHigh-Quality Data  \nJan Brauner  \nDepartment of Computer Science University of Oxford  \nA thesis submitted for the degree of Doctor of Philosophy  \nHilary Term 2024  \n2  \nAbstract  \nIn this thesis, we make advances in several application areas, from public health to AI safety. In the process, we develop novel machine learning (ML) methods to tackle various challenges. In each application area, we take a data-centric approach: we develop a method to automatically select high-quality training data, devise and execute rigorous processes for large-scale manual data collection, generate synthetic training data when data is otherwise hard to obtain, and collate highly diverse test data to evaluate the generalisability of our method.  \nFirst, we tackle AI misuse and misalignment by building a detector for “lies”—which we define as particular types of falsehoods—output by large language models (LLMs) . This endeavour is complicated by the difficulty in acquiring appropriate training data. In the case of misuse, an attacker might use a wide range of types of LLM or lie-generation methods; obtaining a sufficiently diverse dataset to cover all situations is challenging. In the case of misalignment, the lies can be subtle, such as convincing but false text that plays to human biases; identifying a sufficient number of examples to build a dataset from can be difficult. Our approach offers a potential remedy: We generate training data by simply prompting an LLM with instructions to lie, making it easy to collect a large dataset. We then collate a large, diverse test dataset and extensively study the generalisation of our lie detector to previously unseen settings. We find that it generalises surprisingly well.  \nIn the second part of the thesis, we deal with a situation where we have abundant training data available, but some of it is of low quality. This is common in modern ML, as deep learning models require vast amounts of data, often scraped from the internet. We develop a method to automatically select the most useful data at any given point in training, enabling us to train neural networks in fewer steps, to higher accuracy, and with lower computational costs.  \nIn the third part of the thesis, a necessity for exceptionally high dataquality arises due to the high stakes of the modelling outcomes. During the COVID-19 pandemic, governments worldwide needed information on which nonpharmaceutical interventions (NPIs) could most effectively curb virus transmission. While this question could be tackled with modelling approaches, existing data on the timing and nature of interventions across countries was incomplete and inaccurate. To guide high-stakes policy decisions, we need carefully validated, high-quality data. We collect two datasets with extensive quality-control measures as the basis of our modelling studies. We then fit multiple Bayesian models to infer the effect of various government interventions on viral transmission. Our effectiveness estimates are backed with rigorous model validation experiments, a step often lacking in other work on this topic. The resulting NPI effectiveness estimates informed policy decisions in various countries.  \nOverall, we make progress in different application areas by leveraging different data-centric approaches: generating synthetic training data, collating diverse test data to evaluate our method, selecting the most useful data when data is abundant, and collecting high-quality data when existing datasets are insufficient.  \nDRAFT Printed on August 3, 2024  \nFrom Public Health to AI Safety: Improving Machine Learning Approaches by Collecting, Selecting, or Reducing the Need for High-Quality Data  \nJan Brauner Department of Computer Science  \nUniversity of Oxford  \nA thesis submitted for the degree of Doctor of Philosophy  \nHilary Term 2024  \nAcknowledgements  \nI thank:  \nMy ","cbCaintNrAySFmm7","https://ap.wps.com/l/cbCaintNrAySFmm7","pdf",22337431,1,330,"English","en",105,"# Abstract\n## AI misuse and misalignment: lie detection data\n## Low-quality abundant data: automatic selection\n## High-stakes public health: COVID-19 NPI data collection and Bayesian modelling\n## Overall data-centric contributions","[{\"question\":\"How does the thesis improve machine learning performance across different application areas?\",\"answer\":\"It applies a data-centric strategy: selecting high-quality training data, running large-scale manual collection, generating synthetic data when necessary, and evaluating with diverse test datasets.\"},{\"question\":\"What is the role of the lie detector in the AI safety part of the thesis?\",\"answer\":\"The thesis builds a detector for specific types of falsehoods generated by large language models and uses LLM prompting to create training data for that detector.\"},{\"question\":\"How does the thesis handle public health questions with limited or unreliable datasets?\",\"answer\":\"It creates two COVID-19 datasets using extensive quality-control measures, then fits multiple Bayesian models and validates them rigorously to infer the effects of government interventions.\"}]","From Public Health to AI Safety - Improving Machine Learning Approaches by Collecting, Selecting, or Reducing the Need for High-Quality Data | PDF",1785675333,832,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"from-public-health-to-ai-safety-improving-machine-learning-approaches-by-collecting-selecting-or-reducing-the-need-for-high-quality-data","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/from-public-health-to-ai-safety-improving-machine-learning-approaches-by-collecting-selecting-or-reducing-the-need-for-high-quality-data/117354/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does the thesis improve machine learning performance across different application areas?","Question",{"text":75,"@type":76},"It applies a data-centric strategy: selecting high-quality training data, running large-scale manual collection, generating synthetic data when necessary, and evaluating with diverse test datasets.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the role of the lie detector in the AI safety part of the thesis?",{"text":80,"@type":76},"The thesis builds a detector for specific types of falsehoods generated by large language models and uses LLM prompting to create training data for that detector.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the thesis handle public health questions with limited or unreliable datasets?",{"text":84,"@type":76},"It creates two COVID-19 datasets using extensive quality-control measures, then fits multiple Bayesian models and validates them rigorously to infer the effects of government interventions.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]