This is a new version of the repository. Do let us know (lindat-help at ufal.mff.cuni.cz) if you encounter any issues.
What's New
lexicalConceptualResourceLINDAT / CLARIAH-CZ
Author(s):
Description:
The database lists translations of Irish literature into Czech and Slovak from the times of the Czech and Slovak national revivals to the present day, as well as Czech and Slovak reflections on Ireland and Irish culture from the earliest records up until the present (Slovak reflections being covered up to the 1990s). A detailed outline of the coverage of the database and sources of the data is available at https://irbiblio.ff.cuni.cz/about. Originally developed in the early 2000s by Daniel Samek, the database has been maintained and continuously updated by the Centre for Irish Studies of the Faculty of Arts, Charles University. The irbiblio_2026-06-09.csv contains all the data in a flat (one table) CSV format, while irbiblio_2026-06-09.sql is the dump of the same data from a relational database.
This item contains 2 files (1.52 MB).
Publicly Available
Author(s):
Description:
AI Brown is a corpus of English texts generated with large language models (LLMs). Its main purpose is to create a resource for comparing human-written texts with LLM-generated text linguistically. The corpus is multi-genre and rich in terms of topics, authors, and text types, and comparable with existing human-created corpora. The corpus replicates the reference human BE21 corpus by Paul Baker, a modern version of the original Brown Corpus. The new corpus was generated using 25 models from OpenAI, Anthropic, Google, Meta, and DeepSeek, ranging from GPT-3 (davinci-002) to recent frontier models such as GPT-5.4, Claude 4.7, and Gemini 3.1, and is tagged according to the Universal Dependencies standard (i.e., the texts are tokenized, lemmatized, and morphologically and syntactically annotated). The subcorpus size varies according to the model used (on average 959k tokens per subcorpus, 47.0M tokens altogether across 49 subcorpora). The raw data and plain texts are freely available for download under the CC BY 4.0 license, the UD annotated data are under CC BY-NC-SA 4.0 licence. The corpus is also accessible through the KonText search interface of the Czech National Corpus (https://www.korpus.cz/kontext/query?corpname=ai_brown_v2).
This item contains 3 files (760.51 MB).
Publicly Available
Author(s):
Description:
AI Koditex is a corpus of Czech texts generated with large language models (LLMs). Its main purpose is to create a resource for comparing human-written texts with LLM-generated text linguistically. The corpus is multi-genre and rich in terms of topics, authors, and text types, and comparable with existing human-created corpora. The corpus replicates the reference human Koditex corpus that follows the Brown Corpus tradition. The new corpus was generated using 24 models from OpenAI, Anthropic, Google, Meta, and DeepSeek, ranging from GPT-3 (davinci-002) to recent frontier models such as GPT-5.4, Claude 4.7, and Gemini 3.1, and is tagged according to the Universal Dependencies standard (i.e., the texts are tokenized, lemmatized, and morphologically and syntactically annotated). The subcorpus size varies according to the model used (on average 813k tokens per subcorpus, 36.6M tokens altogether across 45 subcorpora). The raw data and plain texts are freely available for download under the CC BY 4.0 license, the UD annotated data are under CC BY-NC-SA 4.0 licence. The corpus is also accessible through the KonText search interface of the Czech National Corpus (https://www.korpus.cz/kontext/query?corpname=ai_koditex_v2).
This item contains 3 files (796.99 MB).
Publicly Available
Most Viewed Items - Last Month
Author(s):
Description:
Tokenizer, POS Tagger, Lemmatizer and Parser models for 94 treebanks of 61 languages of Universal Depenencies 2.5 Treebanks, created solely using UD 2.5 data (http://hdl.handle.net/11234/1-3105). The model documentation including performance can be found at http://ufal.mff.cuni.cz/udpipe/models#universal_dependencies_25_models . To use these models, you need UDPipe binary version at least 1.2, which you can download from http://ufal.mff.cuni.cz/udpipe . In addition to models itself, all additional data and value of hyperparameters used for training are available in the second archive, allowing reproducible training.
This item contains 96 files (2.61 GB).
Publicly Available
Author(s):
show everyone
Description:
Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).
This item contains 3 files (740.61 MB).
Publicly Available
Author(s):
show everyone
Description:
Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).
This item contains 3 files (765.09 MB).
Publicly Available