This is a new version of the repository. Do let us know (lindat-help at ufal.mff.cuni.cz) if you encounter any issues.

Self-paced reading and lexical surprisal on Czech legal and press texts

Please use the following text to cite this item or export to a predefined format:
Cinková, Silvie; Kvapilíková, Ivana; Chládková, Kateřina and Černý, Jan, 2026, Self-paced reading and lexical surprisal on Czech legal and press texts, LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL), http://hdl.handle.net/11234/1-6210.
Date issued
2026-06-30
Size
2214 tokens,
59 items,
1396 mb
Language(s)
Description
Two text collections, legal and press texts. The texts are tokenized. Each token contains the following information: Self-paced-reading time by approx. 100 experiment participants, Universal-Dependencies morphosyntactic information, and Lexical surprisal values from five language models. The Press as well as the Legal data are divided into a train and a test section. The train section is divided into ten folds obtained by two repetitions of a five-fold resampling procedure.
Acknowledgement
This item isPublicly Available
and licensed under:

Files in this item

Name
Legal_TRAIN_FOLDSPLITS_ALLDATA.tsv
Size
583.12 MB
Format
application/octet-stream
Description
MD5
c42cf09fa9e5d138ca0f11101a4409d1
Preview
  File Preview
Name
Legal_test.tsv
Size
3.75 MB
Format
application/octet-stream
Description
MD5
0dda96fa5b7cffaf4d14bb6ab08e1876
Preview
  File Preview
Name
PRESS_test.tsv
Size
6.91 MB
Format
application/octet-stream
Description
MD5
90c80555c52d79702ae6577978732eed
Preview
  File Preview
Name
PRESS_TRAIN_FOLDSPLITS_ALLDATA.tsv
Size
802.18 MB
Format
application/octet-stream
Description
MD5
bcea4e544298aa085255ae03194af5bb
Preview
  File Preview