[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-136023-en":3,"doc-seo-136023-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},136023,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","Dataset vs Reality - Understanding Model Performance from the Perspective of Information Need","Deep learning has produced models that surpass humans on select benchmarks, raising a key concern: whether these models solve real-world tasks when the input/output setup mirrors the benchmark datasets. The work argues that training teaches a model to satisfy the same information need used to build the dataset. By comparing benchmark datasets for question answering and image captioning, the analysis highlights differences in dataset creation procedures and morphosyntactic properties, which reflect task-specific information needs. Researchers are urged to consider information need when using datasets for training and when designing datasets to ensure they represent the intended research task accurately.","arXiv :2212 .02726v2 [ cs .IR] 24 Mar 2023  \nDataset vs Reality: Understanding Model Performance from the  \nPerspective of Information Need  \nMengying Yu  \nSchool of Computer Science and Engineering, Nanyang Technological University  \n50 Nanyang Avenue, Singapore 639798  \n[yume0004@e.ntu.edu.sg](yume0004@e.ntu.edu.sg)  \nAixin Sun*  \nSchool of Computer Science and Engineering, Nanyang Technological University  \n50 Nanyang Avenue, Singapore 639798  \n[axsun@ntu.edu.sg](axsun@ntu.edu.sg)  \nAbstract  \nDeep learning technologies have brought us many models that outperform human beings on a few benchmarks. An interesting question is: can these models well solve real-world problems with similar settings (e.g., identical input/output) to the benchmark datasets? We argue that a model is trained to answer the same information need for which the training dataset is created. Although some datasets may share high structural similarities, e.g., question-answer pairs for the question answering (QA) task and image-caption pairs for the image captioning (IC) task, they may represent di􀀋erent research tasks aiming for answering di􀀋erent information needs. To support our argument, we use the QA task and IC task as two case studies and compare their widely used benchmark datasets. From the perspective of information need in the context of information retrieval, we show the di􀀋erences in the dataset creation processes, and the di􀀋erences in morphosyntactic properties between datasets. The di􀀋erences in these datasets can be attributed to the di􀀋erent information needs of the speciﬁc research tasks. We encourage all researchers to consider the information need the perspective of a research task before utilizing a dataset to train a model. Likewise, while creating a dataset, researchers may also incorporate the information need perspective as a factor to determine the degree to which the dataset accurately reﬂects the research task they intend to tackle.  \n1 Introduction  \nIn the very ﬁrst chapter of the Information Retrieval (IR) book, Manning et al. distinguish information need from query: “An information need is the topic about which the user desires to know more, and a query is what the user conveys to the computer in an attempt to communicate the information need”(Manning, Raghavan, & Sch¨utze, 2008) . When a query is entered into a search engine, the latter provides a list of potential documents that may be relevant to the user's search. The user then evaluates each document to  \n* Corresponding author.  \nFigure 1: Overview of dataset vs reality. Model learns from dataset only and is expected to address the practical task.  \ndetermine if it contains the information he/she is looking for.1 Crafting queries that accurately capture the information need is crucial, and this principle also holds true when it comes to creating datasets for training models, as the ultimate goal is to address practical tasks in the real world.  \nAs illustrated in Figure 1, before we can build a model to address a real-world problem, we need to formally formulate the problem by identifying its input/output, as well as any key constraints. To develop a model, the next essential step is to create a dataset. A dataset is used to simulate the practical task by providing inputs that are the same as, or resemble, those encountered in the real world, along with the expected outputs or ground truth labels. It is expected that the model trained on this dataset will be capable of addressing the practical problem at hand.  \nIt is important to note that the model has no direct access to the practical problem and lacks a literal understanding of its problem deﬁnition. The model is trained to solve the problem that is “deﬁned” by the dataset, utilizing the input and corresponding expected output provided in the dataset. In this sense, a dataset is analogous to a “query” in information retrieval, and the trained model plays the role of a search engine. Let us suppose that we have succ","cbCaivEFPLSb7H1y","https://ap.wps.com/l/cbCaivEFPLSb7H1y","pdf",386671,1,19,"English","en",105,"# Introduction\n## Information need vs query\n## Dataset as a formal task definition\n## Dataset vs reality and benchmark comparison\n## Case studies: QA and image captioning benchmarks","[{\"question\":\"Why can benchmark-trained models fail to solve real-world problems?\",\"answer\":\"Because the model only learns what is defined by the training dataset. If the dataset does not reflect the real task’s information need, performance may not transfer effectively.\"},{\"question\":\"What does the paper mean by “information need” in dataset creation?\",\"answer\":\"It treats information need as the underlying topic the user seeks, and argues that datasets should be designed to correspond to that need so training aligns with the intended task.\"},{\"question\":\"Which datasets are compared as case studies?\",\"answer\":\"For question answering: SQuAD2.0 vs Natural Questions (NQ). For image captioning: MS-COCO, Flickr30K, SBU Captions, and Conceptual Captions.\"}]","Dataset vs Reality - Understanding Model Performance from the Perspective of Information Need | PDF",1787338871,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"dataset-vs-reality-understanding-model-performance-from-the-perspective-of-information-need","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/dataset-vs-reality-understanding-model-performance-from-the-perspective-of-information-need/136023/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-21",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why can benchmark-trained models fail to solve real-world problems?","Question",{"text":75,"@type":76},"Because the model only learns what is defined by the training dataset. If the dataset does not reflect the real task’s information need, performance may not transfer effectively.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the paper mean by “information need” in dataset creation?",{"text":80,"@type":76},"It treats information need as the underlying topic the user seeks, and argues that datasets should be designed to correspond to that need so training aligns with the intended task.",{"name":82,"@type":73,"acceptedAnswer":83},"Which datasets are compared as case studies?",{"text":84,"@type":76},"For question answering: SQuAD2.0 vs Natural Questions (NQ). For image captioning: MS-COCO, Flickr30K, SBU Captions, and Conceptual Captions.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]