Header

Search

Eye-tracking-while-reading Datasets Overview

On this page, we present our living online overview of eye-tracking-while-reading datasets. In order to create more transparency with regards to existing datasets, we created the overview which shows more than 50 features for each dataset (if available). Please find more information at the links below.

In order to keep the table up to date, we encourage all dataset authors to add their table to the overview. Should there be anything wrong, please contact us as well. Forms to do so are provided below.

Thanks to all authors for your contributions!

Access the dataset review

Please find the dataset overview table at the following link:

https://dili-lab.github.io/datasets.html

Add your dataset to the review

To add a new dataset to our overview table, please fill in the form linked below. We ask authors to  carefully check the existing table for examples to make the integration process as smooth and fast as possible. 

Add your dataset

Edit your dataset

In order to edit a dataset entry in the overview table, we ask you to send a request by filling in a short form linked below. 

Edit a dataset

Citation

The corresponding paper is currently under review. Please cite as follows:

@misc{jakobi2026eyetrackingwhilereadinglivingsurveydatasets,
      title={Eye-Tracking-while-Reading: A Living Survey of Datasets with Open Library Support}, 
      author={Deborah N. Jakobi and David R. Reich and Paul Prasse and Jana M. Hofmann and Lena S. Bolliger and Lena A. Jäger},
      year={2026},
      eprint={2602.19598},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2602.19598}, 
}

Eye-tracking-while-reading corpora are a valuable resource for many different disciplines and use cases. Use cases range from studying the cognitive processes underlying reading to machine-learning-based applications, such as gaze-based assessments of reading comprehension. The past decades have seen an increase in the number and size of eye-tracking-while-reading datasets as well as increasing diversity with regard to the stimulus languages covered, the linguistic background of the participants, or accompanying psychometric or demographic data. The spread of data across different disciplines and the lack of data sharing standards across the communities lead to many existing datasets that cannot be easily reused due to a lack of interoperability. In this work, we aim at creating more transparency and clarity with regards to existing datasets and their features across different disciplines by i) presenting an extensive overview of existing datasets, ii) simplifying the sharing of newly created datasets by publishing a living overview online, https://dili-lab.github.io/datasets.html, presenting over 45 features for each dataset, and iii) integrating all publicly available datasets into the Python package pymovements which offers an eye-tracking datasets library. By doing so, we aim to strengthen the FAIR principles in eye-tracking-while-reading research and promote good scientific practices, such as reproducing and replicating studies.