[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85522-en":3,"doc-seo-85522-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85522,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Eye-Tracking-while-Reading: A Living Survey of Datasets with Open Library Support","Eye-tracking-while-reading datasets enable research across disciplines, from modeling cognitive processes in reading to machine-learning applications using gaze-based reading comprehension assessment. Over decades, dataset variety has increased in size and in stimulus languages, participant backgrounds, and accompanying psychometric or demographic data. However, fragmented publication and missing sharing standards hinder reuse and interoperability. The work provides an extensive overview of existing datasets, publishes a living online catalog with 55+ features per dataset, and integrates public datasets into the pymovements Python library to strengthen FAIR principles and reproducible science.","arXiv :2602 . 19598v2 [ cs .CL] 13 Jul 2026  \nEye-Tracking-while-Reading: A Living Survey of Datasets with Open Library Support  \nDeborah N. Jakobi 1*, David R. Reich 1, 2 , Paul Prasse2 , Jana M. Hofmann 1 , Lena S. Bolliger 1 , Lena A. J¨ager 1, 3  \n1 Department of Computational Linguistics, University of Zurich.  \n2 Department of Computer Science, University of Potsdam.  \n3 Department of Informatics, University of Zurich.  \n*Corresponding author(s). E-mail(s): [deborahnoemie.jakobi@uzh.ch](deborahnoemie.jakobi@uzh.ch) ; Contributing authors: davidrobert.reich@uzh.ch;  \n[paul.prasse@uni-potsdam.de](paul.prasse@uni-potsdam.de) ; [janamara.hofmann@uzh.ch](janamara.hofmann@uzh.ch) ;  \n[lena.bolliger@uzh.ch](lena.bolliger@uzh.ch) ; [lenaann.jaeger@uzh.ch](lenaann.jaeger@uzh.ch) ;  \nAbstract  \nEye-tracking-while-reading corpora are a valuable resource for many different disciplines and use cases. Use cases range from studying the cognitive processes underlying reading to machine-learning-based applications, such as gaze-based assessments of reading comprehension. The past decades have seen an increase in the number and size of eye-tracking-while-reading datasets as well as increasing diversity with regard to the stimulus languages covered, the linguistic background of the participants, or accompanying psychometric or demographic data. The spread of data across different disciplines and the lack of data sharing standards across the communities lead to many existing datasets that cannot be easily reused due to a lack of interoperability. In this work, we aim at creating more transparency and clarity with regards to existing datasets and their features across different disciplines by i) presenting an extensive overview of existing datasets, ii) simplifying the sharing of newly created datasets by publishing a living overview online, [https://t.uzh.ch/1Yh](https://t.uzh.ch/1Yh), presenting over 55 features for each dataset, and iii) integrating all publicly available datasets into the Python package pymovements which offers an eye-tracking datasets library. By doing so, we aim to strengthen the FAIR principles in eye-tracking-while-reading research and promote good scientific practices, such as reproducing and replicating studies.  \n1  \n1 Introduction  \nFor many people, reading is an everyday process and as such a necessary skill to master a variety of tasks in many areas of human life. The interaction of visual perception and language processing during reading opens a wide range of potential research topics. One of these is the investigation of cognitive processes that are involved in reading to understand how humans process language. Collecting and analyzing eye movements during reading has been shown to provide valuable insights into such processes, and therefore eye movements are considered a gold standard method in reading research (Kliegl, Nuthmann, & Engbert, 2006; Rayner, 1998; Rayner & Carroll, 2018) . The primary use case resulting from this tight connection of eye movements to human cognition is to develop psycholinguistic theories of reading (Just & Carpenter, 1980) . Even beyond this use case, eye movements are extremely valuable when it comes to, for example, investigating why certain individuals experience difficulties in reading in general, or why a given text might be hard to understand. The insights into human language processing can then be further used to, for example, investigate the cognitive plausibility of language models (Beinborn & Hollenstein, 2023) or even enhance them using gaze data (e.g., Deng, Prasse, Reich, Scheffer, & J¨ager, 2024) . The list of use cases of eye-tracking-while-reading data is long and keeps growing, as detailed in Section 2.  \nThe number of use cases is expanded by the diversity of eye-tracking-while-reading data, which comes in many different forms. In general, participants read any type of reading material while their eye movements are being tracked. The first very important charac","cbCaimKgQP3IWa1B","https://ap.wps.com/l/cbCaimKgQP3IWa1B","pdf",560581,3,1,81,"English","en",105,"# Introduction\n## Reading as a research target\n## Diversity of eye-tracking-while-reading stimuli\n## Controlled vs. naturalistic corpora\n## Semi-naturalistic approaches","[{\"question\":\"What is the main purpose of the living survey presented in the document?\",\"answer\":\"It aims to increase transparency about existing eye-tracking-while-reading datasets and their features, while addressing reuse barriers caused by inconsistent sharing and interoperability standards.\"},{\"question\":\"How does the document characterize the types of stimuli used in eye-tracking-while-reading datasets?\",\"answer\":\"It describes stimuli ranging from isolated sentences to paragraphs and full texts, typically categorized as controlled or naturalistic, with controlled materials often using minimal pairs and tightly constrained hypotheses.\"},{\"question\":\"How does the document help researchers share and access newly created or existing datasets?\",\"answer\":\"It publishes a living overview online with more than 55 features per dataset and integrates all publicly available datasets into the Python package pymovements to provide a unified library interface.\"}]",1784204158,204,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"eye-tracking-while-reading-a-living-survey-of-datasets-with-open-library-support","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/eye-tracking-while-reading-a-living-survey-of-datasets-with-open-library-support/85522/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main purpose of the living survey presented in the document?","Question",{"text":75,"@type":76},"It aims to increase transparency about existing eye-tracking-while-reading datasets and their features, while addressing reuse barriers caused by inconsistent sharing and interoperability standards.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the document characterize the types of stimuli used in eye-tracking-while-reading datasets?",{"text":80,"@type":76},"It describes stimuli ranging from isolated sentences to paragraphs and full texts, typically categorized as controlled or naturalistic, with controlled materials often using minimal pairs and tightly constrained hypotheses.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the document help researchers share and access newly created or existing datasets?",{"text":84,"@type":76},"It publishes a living overview online with more than 55 features per dataset and integrates all publicly available datasets into the Python package pymovements to provide a unified library interface.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]