This is a new version of the repository. Do let us know (lindat-help at ufal.mff.cuni.cz) if you encounter any issues.
What's New
NER-annotated interview transcriptsLINDAT / CLARIAH-CZ
Author(s):
Description:
MalachNER is a collection of transcripts of interviews with 13 Holocaust survivors in 10 languages, covering approximately 37 hours of speech in total, sourced from different archives (Visual History Archive, United States Holocaust Memorial Museum, Milan Šimečka Foundation). Domain experts tagged named entities in the denoised transcripts based on domain-specific annotation guidelines with entity types Person, Organization, Location, Camp, Ghetto, Date, and Miscellaneous. The dataset includes per-language training and test splits and the annotation guidelines.
This item contains 1 file (1.73 MB).
Publicly Available
languageDescriptionLINDAT / CLARIAH-CZ
Author(s):
Description:
This trained model is a commercially licensable variant of the NameTag 3 Multilingual Model 260521. It was trained on a subset of the original training data that excludes datasets restricted to non-commercial use, allowing us to offer the model under a commercial license. NameTag 3 (https://ufal.mff.cuni.cz/nametag/3) is an open-source tool for both flat and nested named entity recognition (NER). It identifies named entities in text and classifies them into predefined categories, such as persons, locations, and organizations. The model was trained jointly on 17 flat named-entity corpora covering 15 languages: Chinese, Croatian, Czech, Danish, English, Greek, Hebrew, Maghrebi Arabic French, Norwegian Bokmål, Norwegian Nynorsk, Portuguese, Serbian, Slovak, Slovenian, and Swedish. For model details and licensing information, see the NameTag 3 Multilingual CL Model documentation (https://ufal.mff.cuni.cz/nametag/3/models#multilingual-CL). If you use NameTag 3 in scientific work, please cite Straková & Straka (2025) (https://aclanthology.org/2025.acl-demo.4.pdf).
This item contains 1 file (1.72 GB).
Publicly Available
Author(s):
Description:
This dataset contains 150 instances of the Spanish cantara paradigm functioning as the simple past tense canté. These instances are drawn from the CORPES XXI corpus (RAE, 2025). First, 1,001 instances of the cantara form were extracted from each year covered by CORPES XXI, ver. 1.4; i.e. 2001-2025 (616 instances for the year 2025). These results were then searched using three AI tools—Gemini Pro, Claude, and ChatGPT—to identify only those instances where cantara functioned as the simple past tense. The results were then manually reviewed. All examples include a bibliographic reference. REAL ACADEMIA ESPAÑOLA, 2025: Banco de datos (CORPES XXI) [online]. Corpus del Español del Siglo XXI (CORPES), ver. 1.4. http://www.rae.es [April 2026]
This item contains 1 file (57.31 KB).
Publicly Available
Most Viewed Items - Last Month
Author(s):
Description:
Tokenizer, POS Tagger, Lemmatizer and Parser models for 94 treebanks of 61 languages of Universal Depenencies 2.5 Treebanks, created solely using UD 2.5 data (http://hdl.handle.net/11234/1-3105). The model documentation including performance can be found at http://ufal.mff.cuni.cz/udpipe/models#universal_dependencies_25_models . To use these models, you need UDPipe binary version at least 1.2, which you can download from http://ufal.mff.cuni.cz/udpipe . In addition to models itself, all additional data and value of hyperparameters used for training are available in the second archive, allowing reproducible training.
This item contains 96 files (2.61 GB).
Publicly Available
Author(s):
show everyone
Description:
Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).
This item contains 3 files (740.61 MB).
Publicly Available
Author(s):
show everyone
Description:
Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).
This item contains 3 files (765.09 MB).
Publicly Available