[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117568-en":3,"doc-seo-117568-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117568,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Leveraging Machine Learning for Official Statistics - A Statistical Manifesto - Introduction","The manifesto examines how machine learning should be evaluated and integrated into the production of official statistics, framing the ongoing challenge of impact, utilization, and operational feasibility. It explains why datafication increases dataset size and complexity, making classical statistical analysis harder and motivating algorithmic approaches for unstructured data. The text clarifies distinctions between methodology, methods, and techniques, defining “method” as a systematic procedure and “technique” as a way of carrying out tasks, then positions ML models within the broader production cycle of inference and quality-improving processes.","arXiv :2409 .04365v1 [ stat .ML] 6 Sep 2024  \nLeveraging Machine Learning for Official Statistics: A Statistical Manifesto  \nMarco J.H. Puts 1 , David Salgado2 , and Piet J.H. Daas 1  \n1 Department of Methodology, Statistics Netherlands, CBS-weg 11, 6412 EX, Heerlen, The Netherlands  \n2 Department of Methodology, Statistics Spain, Avenida de  \nManoteras 50-52, 28050 Madrid, Spain February 2024  \n1 Introduction  \n1.1 Production of official statistics and machine learning  \nEvaluating the impact and utilization of machine learning (ML) in the production of official statistics presents an ongoing challenge. ML is a subfield of Artificial Intelligence, which aims at “not just to understand but also to build intelligent entities”(Russell and Norvig, 2010) . It is thus similar to assessing the impact of intelligent human activity on a production system, which is limitless. ML itself is composed of different subfields about how the process of learning is carried out: supervised learning, unsupervised learning, reinforcement learning, etc. (Alpaydin, 2020; Murphy, 2013) .  \nDespite the widespread adoption of ML, implementation still has many challenges (O’Neil, 2016, discusses this subject) . Despite the danger and complexity of ML, the compelling ’datafication’ of our society forces us to look at ML asan addition to our (official) statistical toolbox. Datasets get larger, are more detailed, and become more and more complex. Because of this, it becomes increasingly difficult to perform (classical) statistical analysis on these kinds of data. The need for algorithmic-based approaches, which can handle larger, more complex, and unstructured datasets, is necessary to be able to perform successful analysis; see Breiman (2001) . Without it, we would not be able to extract statistical information from many big data sources, like web scraped data (Daas and van der Doef, 2020) and aerial images (De Jong et al., 2020) .  \nML encompasses various techniques, often referred to as ’a methodology’within the realm of data science. However, most statisticians, as well as most scientists, would disagree with the usage of this term. In social sciences and econometrics,”methodology,””methods,” and ”techniques” carry specific and  \ndistinct meanings, which we advocate for adhering to, at least when talking about ML from an (official) statistical point of view.  \nWhen we look at the term ’methods’ this becomes most apparent. We have observed that the term ’methods’ is incorrectly used in many fields that apply ML. It is often used to merely describe the chosen environment in which a study is performed. So, when summarizing the ML algorithms and hyperparameters used, maybe with some kind of rationale, many data scientists assume this describes the ’method’ used. Such a description, however, falls short when viewed from the statistical standpoint of a methodologist.  \nSo let’s start with the basis. Techniques, and how we combine them, are primarily determined by the ’why’ and ’what’ questions of the application, a.o.:  \n• why are we doing it?  \n• why do we choose certain techniques?  \n• what are we going to do?  \n• what is our ground material?  \n• what is the context?  \nIt is the ’how’ question that is at the core of the techniques themselves: how do we go about performing certain steps? In addition to the algorithmic description, it describes the (pre-and post-)conditions for applying the technique. From this, we can define a method as:  \nA method is a systematic procedure of techniques for accomplishinga certain goal. Most of the time these methods are established.  \nand a technique as:  \nA technique is a way of carrying out a task. Most of the time these techniques are described as algorithms.  \nFor example, preparing a meal involves cutting vegetables, boiling eggs, and grilling steaks. Recipes can be [considered methods. it](considered methods. it) is also possible to consider a method that takes into account the context in which one prepares a meal, a","cbCaipMwHmzCdamj","https://ap.wps.com/l/cbCaipMwHmzCdamj","pdf",717712,1,33,"English","en",105,"# Introduction\n## Production of official statistics and machine learning\n## Methodology vs methods vs techniques\n## ML in the official statistics production cycle","[{\"question\":\"Why is evaluating machine learning in official statistics an ongoing challenge?\",\"answer\":\"Assessing ML impact and utilization remains difficult because ML adoption is widespread but implementation involves danger, complexity, and practical constraints within statistical production.\"},{\"question\":\"How does the document relate machine learning to the official statistical production workflow?\",\"answer\":\"It describes ML as building predictive models applied to new data, supporting both the inference task and complementary quality-improvement tasks such as data collection, coding, and dissemination.\"},{\"question\":\"What distinction does the document make between methodology, methods, and techniques?\",\"answer\":\"“Methodology” is the subject area; a “method” is a systematic procedure of techniques for accomplishing a goal; a “technique” is a way of carrying out a task, often described as algorithms.\"}]","Leveraging Machine Learning for Official Statistics - A Statistical Manifesto - Introduction | PDF",1785677045,83,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"leveraging-machine-learning-for-official-statistics-a-statistical-manifesto-introduction","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/leveraging-machine-learning-for-official-statistics-a-statistical-manifesto-introduction/117568/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is evaluating machine learning in official statistics an ongoing challenge?","Question",{"text":75,"@type":76},"Assessing ML impact and utilization remains difficult because ML adoption is widespread but implementation involves danger, complexity, and practical constraints within statistical production.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the document relate machine learning to the official statistical production workflow?",{"text":80,"@type":76},"It describes ML as building predictive models applied to new data, supporting both the inference task and complementary quality-improvement tasks such as data collection, coding, and dissemination.",{"name":82,"@type":73,"acceptedAnswer":83},"What distinction does the document make between methodology, methods, and techniques?",{"text":84,"@type":76},"“Methodology” is the subject area; a “method” is a systematic procedure of techniques for accomplishing a goal; a “technique” is a way of carrying out a task, often described as algorithms.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]