LINDAT / CLARIAH-CZ Data & Tools
Permanent URI for this collectionhttp://hdl.handle.net/11858/00-097C-0000-0001-4877-A
Data and Tools from partner institutions of LINDAT/CLARIAH-CZ project, formerly LINDAT/CLARIN.
If you are participating in LINDAT/CLARIAH-CZ project, please submit your work to this repository.
Browse
Recent Submissions
Item type: Item , Corpus occurrences of Czech, Slovak, Russian and Belarusian locative adverbs and their combinations with prepositions from a semantic and syntactic point of view (based on comparable web corpora from the Aranea project http://unesco.uniba.sk/)(Aksana Schillová, 2026) Aksana SchillováDataset obsahuje excelové soubory, ve kterých jsou analyticky zpracovány korpusové doklady vybraných a vzájemně korelujících českých, slovenských, ruských a běloruských adverbií a adverbiálně-prepozicionálních spojení s místním významem. Jedná se o následující výrazy: čes. blízko, blízko k, blízko od, blíž(e), blíž(e) k, daleko, daleko od, daleko za, nedaleko, nedaleko od, hluboko, hluboko v, hluboko do, hluboko pod, vpravo, vpravo od, vlevo, vlevo od, vysoko, vysoko v, vysoko do, vysoko nad; slov. blízko, blízko k, blízko od, bližšie, bližšie k, ďaleko, ďaleko od, ďaleko za, neďaleko, neďaleko od, hlboko, hlboko v, hlboko do, hlboko pod, napravo, napravo od, naľavo, naľavo od, vysoko, vysoko v, vysoko do, vysoko nad; rus. близко, близко к, близко от, ближе, ближе к, далеко, далеко от, далеко за, недалеко, недалеко от, глубоко, глубоко в, глубоко под, справа, справа от, слева, слева от, высоко, высоко в, высоко над; běl. блізка, блізка да, блізка ад, бліжэй, бліжэй да, далёка, далёка ад, далёка за, недалёка, недалёка ад, справа, справа ад, злева, злева ад, глыбока, глубока ў, глыбока пад, высока, высока ў, высока над. Doklady užití těchto výrazů strukturované v excelových souborech pocházejí z náhodných vzorků z jejich korpusových konkordancí. Velikost vzorků zpravidla činí 100 řádků. K extrakci vzorků jsou použity srovnatelné webové korpusy Aranea, a to následující: Araneum Bohemicum IV Minus (Czech, 20.03) 125, Araneum Slovacum VII Minus Beta (Slovak, 24.02) 125 M, Araneum Russicum Russicum Minus (Russia-only Russian, 15.03) 120 M, Araneum Albaruthenicum Novum MMXXI (Belarusian, 21.02) 155 M. Výše uvedené korpusy jsou volně přístupné na internetovém portálu http://unesco.uniba.sk/. Po aktualizacích jsou starší verze korpusů Aranea přístupné na vyžádání u tvůrců tohoto portálu. Každý excelový soubor má stejnou strukturu: Je rozvržen do 8 listů seskupujících korpusové doklady podle určitých rysů, a to především lexikálně-sémantického a syntaktického rázu. Tyto excelové listy mají následující pracovní názvy: místní významy_statické (zkoumané výrazy vyjadřují místo na otázku „kde“), místní významy_dynamické (zkoumané výrazy vyjadřují místo na otázku „kam“), predikativní funkce_věty bez subjektu (zkoumané výrazy se používají v predikativní funkci ve větách bez subjektu), nemístní významy (zkoumané výrazy vyjadřují nemístní významy), součást frazémů (zkoumané výrazy vystupují jako obligatorní nebo fakultativní součást frazémů), specifické použití (zkoumané výrazy se používají v rámci pojmenovacích spojení, v rámci titulků, inzerátů, popisků k fotografiím a obrázkům apod.), zmnožené syntaktické pozice_příklady (list dodatečně shrnuje doklady, ve kterých se zkoumané výrazy používají v rámci adordinačních a koordinačních skupin), šum (list registruje množství šumu, který se vyskytl ve vzorku; šum je vždy nahrazen příslušným množstvím použitelných dokladů z původní konkordance). Doklady s místním významem jsou navíc analyzovány podle následujících syntakticko-sémantických charakteristik: syntaktická funkce dominujícího členu a jeho lexikální obsazení; syntaktická funkce adverbia / adverbiálně-prepozicionálního spojení a lexikální obsazení usouvztažněného substantiva; objekt lokalizace a prostorový orientátor a jejich lexikální obsazení; lexikální obsazení členů rozvíjejících adverbiální tvary (jak v rámci adverbiálně-prepozicionálních spojení, tak mimo ně). Výše uvedený rozbor využívá teoreticko-terminologický rámec Mluvnice češtiny 3 (1986). Vícejazyčný korpusový materiál je zpracován plošně na základě jednotného analytického postupu. Dataset má charakter pracovních materiálů a představuje data, na kterých je založen výzkum v rámci disertační práce autorky. Literatura Benko, V. (2014). Aranea: Yet Another Family of (Comparable) Web Corpora. In Petr Sojka, Aleš Horák, Ivan Kopeček and Karel Pala (Eds.): Text, Speech and Dialogue. 17th International Conference, TSD 2014, Brno, Czech Republic, September 8-12, 2014. Proceedings. LNCS 8655. Springer International Publishing Switzerland. Pp. 257–264. Daneš, F., Grepl, M., & Hlavsa, Z. (1987). Mluvnice češtiny. 3, Skladba. Academia.Item type: Item , Possessive Pronoun Preference (2026-05-19) (2026-07-21)(Charles University, Faculty of Arts, 2024-09) Chromá, Anna; Perevozchikova, TatianaThe contribution includes the data frames and the R script (Markdown file) belonging to the paper "Morphological and Pragmatic Conditioning of Reflexivity in Possessive Pronouns: Effects of Number and Form of Address in Czech" submitted to the journal Linguistics: An Interdisciplinary Journal of the Language Sciences in May 2025. During the review process, title of the paper has been changed to "Reflexive vs. non-reflexive possessives in Czech imperative utterances: effects of number and politeness". The 26-05-19 version has major updates to the analytic script and minor updates to the data sets (no changes to the data, just the to the names of conditions).Item type: Item , NameTag 3 Multilingual CL Model 260717(Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL), 2026-07-17) Straková, JanaThis trained model is a commercially licensable variant of the NameTag 3 Multilingual Model 260521. It was trained on a subset of the original training data that excludes datasets restricted to non-commercial use, allowing us to offer the model under a commercial license. NameTag 3 (https://ufal.mff.cuni.cz/nametag/3) is an open-source tool for both flat and nested named entity recognition (NER). It identifies named entities in text and classifies them into predefined categories, such as persons, locations, and organizations. The model was trained jointly on 17 flat named-entity corpora covering 15 languages: Chinese, Croatian, Czech, Danish, English, Greek, Hebrew, Maghrebi Arabic French, Norwegian Bokmål, Norwegian Nynorsk, Portuguese, Serbian, Slovak, Slovenian, and Swedish. For model details and licensing information, see the NameTag 3 Multilingual CL Model documentation (https://ufal.mff.cuni.cz/nametag/3/models#multilingual-CL). If you use NameTag 3 in scientific work, please cite Straková & Straka (2025) (https://aclanthology.org/2025.acl-demo.4.pdf).Item type: Item , Corpus of Journalistic -ra Forms with Indicative Value in Modern Spanish(Filozofická fakulta, Univerzita Karlova, 2026-04) Dana KratochvílováThis dataset contains 150 instances of the Spanish cantara paradigm functioning as the simple past tense canté. These instances are drawn from the CORPES XXI corpus (RAE, 2025). First, 1,001 instances of the cantara form were extracted from each year covered by CORPES XXI, ver. 1.4; i.e. 2001-2025 (616 instances for the year 2025). These results were then searched using three AI tools—Gemini Pro, Claude, and ChatGPT—to identify only those instances where cantara functioned as the simple past tense. The results were then manually reviewed. All examples include a bibliographic reference. REAL ACADEMIA ESPAÑOLA, 2025: Banco de datos (CORPES XXI) [online]. Corpus del Español del Siglo XXI (CORPES), ver. 1.4. http://www.rae.es [April 2026]Item type: Item , Phrasal verbs in the evolution of Italian(Filozofická fakulta, Univerzita Karlova, 2026) Simona PálováThe dataset contains quantitative data on the frequency of occurrence of Italian syntagmatic (phrasal) verbs over the historical development of the Italian language. The data cover the period from the 13th century to the mid-20th century and are divided into five key historical stages. The file serves as a basis for semantic, diamesis and quantitative analysis of the development of this lexical phenomenon. --- Dataset obsahuje kvantitatívne dáta o frekvencii výskytu talianskych syntagmatických (frázových) slovies v priebehu historického vývoja talianskeho jazyka. Dáta pokrývajú obdobie od 13. storočia až po polovicu 20. storočia a sú rozdelené do piatich kľúčových historických etáp. Súbor slúži ako podklad pre sémantickú, diamezickú a kvantitatívnu analýzu vývoja tohto lexikálneho fenoménu.Item type: Item , Self-paced reading and lexical surprisal on Czech legal and press texts(Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL), 2026-06-30) Cinková, Silvie; Kvapilíková, Ivana; Chládková, Kateřina; Černý, JanTwo text collections, legal and press texts. The texts are tokenized. Each token contains the following information: Self-paced-reading time by approx. 100 experiment participants, Universal-Dependencies morphosyntactic information, and Lexical surprisal values from five language models. The Press as well as the Legal data are divided into a train and a test section. The train section is divided into ten folds obtained by two repetitions of a five-fold resampling procedure.Item type: Item , PONK Linguistic Rules: Linguistic Rules and Metrics for PONK(Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL), 2026-07-01) Kraus, Ivan; Stanovský, Arnold; Motalík Hodková, Kateřina; Vidová Hladká, BarboraTool for linguistic analysis of Czech legal text readability, comprising of a set of linguistic rules and stylometric and readability metrics. It is a module of PONK, an advisory tool that contributes to simplification and higher clarity of Czech legal texts, especially those destined to non-expert audience (https://ufal.mff.cuni.cz/ponk).Item type: Item , Old Czech models for UDPipe 2(Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL), 2026-07-15) Krippnerová, LenkaModels and tokenizers for UDPipe 2, which are used for the automated morphological analysis (UPOS, XPOS, lemmas, morphological features) and dependency parsing of Old Czech texts. It includes models for specific historical periods (so-called "etalon" models), a generalized Old Czech model that offers better performance in syntactic analysis at the cost of morphological accuracy, and a contextual language model (fine-tuned RobeCzech).Item type: Item , Czech-Irish Relations: Bibliography of Books and Articles(Center for Irish Studies, Faculty of Arts, Charles University, 2026) Samek, Daniel; Světlík, Martin; Pilný, Ondřej; Tichý, OndřejThe database lists translations of Irish literature into Czech and Slovak from the times of the Czech and Slovak national revivals to the present day, as well as Czech and Slovak reflections on Ireland and Irish culture from the earliest records up until the present (Slovak reflections being covered up to the 1990s). A detailed outline of the coverage of the database and sources of the data is available at https://irbiblio.ff.cuni.cz/about. Originally developed in the early 2000s by Daniel Samek, the database has been maintained and continuously updated by the Centre for Irish Studies of the Faculty of Arts, Charles University. The irbiblio_2026-06-09.csv contains all the data in a flat (one table) CSV format, while irbiblio_2026-06-09.sql is the dump of the same data from a relational database.Item type: Item , AI Brown v2(Charles University, Faculty of Arts, Department of Linguistics, 2026-06-06) Milička, Jiří; Marklová, Anna; Cvrček, VáclavAI Brown is a corpus of English texts generated with large language models (LLMs). Its main purpose is to create a resource for comparing human-written texts with LLM-generated text linguistically. The corpus is multi-genre and rich in terms of topics, authors, and text types, and comparable with existing human-created corpora. The corpus replicates the reference human BE21 corpus by Paul Baker, a modern version of the original Brown Corpus. The new corpus was generated using 25 models from OpenAI, Anthropic, Google, Meta, and DeepSeek, ranging from GPT-3 (davinci-002) to recent frontier models such as GPT-5.4, Claude 4.7, and Gemini 3.1, and is tagged according to the Universal Dependencies standard (i.e., the texts are tokenized, lemmatized, and morphologically and syntactically annotated). The subcorpus size varies according to the model used (on average 959k tokens per subcorpus, 47.0M tokens altogether across 49 subcorpora). The raw data and plain texts are freely available for download under the CC BY 4.0 license, the UD annotated data are under CC BY-NC-SA 4.0 licence. The corpus is also accessible through the KonText search interface of the Czech National Corpus (https://www.korpus.cz/kontext/query?corpname=ai_brown_v2).Item type: Item , AI Koditex v2(Charles University, Faculty of Arts, Department of Linguistics, 2026-06-06) Milička, Jiří; Marklová, Anna; Cvrček, VáclavAI Koditex is a corpus of Czech texts generated with large language models (LLMs). Its main purpose is to create a resource for comparing human-written texts with LLM-generated text linguistically. The corpus is multi-genre and rich in terms of topics, authors, and text types, and comparable with existing human-created corpora. The corpus replicates the reference human Koditex corpus that follows the Brown Corpus tradition. The new corpus was generated using 24 models from OpenAI, Anthropic, Google, Meta, and DeepSeek, ranging from GPT-3 (davinci-002) to recent frontier models such as GPT-5.4, Claude 4.7, and Gemini 3.1, and is tagged according to the Universal Dependencies standard (i.e., the texts are tokenized, lemmatized, and morphologically and syntactically annotated). The subcorpus size varies according to the model used (on average 813k tokens per subcorpus, 36.6M tokens altogether across 45 subcorpora). The raw data and plain texts are freely available for download under the CC BY 4.0 license, the UD annotated data are under CC BY-NC-SA 4.0 licence. The corpus is also accessible through the KonText search interface of the Czech National Corpus (https://www.korpus.cz/kontext/query?corpname=ai_koditex_v2).Item type: Item , TWIN4DEM Corpus of multilingual textual data on legislative processes(Charles University, 2026-06-30) Kocián, Jiří; Kosová, Klára; Buda, Jakab Máté; Casula, Camilla; Gerő, Márton; Körtvélyesi, Zsolt; Meher, Shreyas; Plíštilová, Tereza; Rovera, Marco; Szalai, Alida; Tonelli, Sara; Turula, Ármin; Vedlichová, Klára; Zsigmond, CsillaThese data sets have been created within the TWIN4DEM (Strengthening Democratic Resilience through Digital Twins) project, funded by the European Union under Horizon Europe (Grant Agreement No. 101178061). They contain the multilingual textual corpora and accompanying metadata developed in Work Package 3 to support research on executive aggrandisement and democratic resilience across four national case studies: Czechia, France, Hungary, and the Netherlands. The datasets have been curated according to FAIR and Open Science principles to facilitate reuse, computational analysis, and long-term preservation. The repository contains all datasets that can be openly released under the CC BY 4.0 licence. These include the textual corpora and accompanying metadata collected for the French, Dutch, Hungarian, and Czech case studies, covering legislative texts, parliamentary debates, judicial decisions, advisory opinions, legislative documents, and public profiles of political actors. All datasets have been prepared in machine-readable formats and include metadata enabling further linking between documents, actors, institutions, and legislative procedures. For the French case study, the repository includes Constitutional Court rulings, decisions of the Conseil d'État and other administrative courts, primary and delegated legislation (organic laws, ordinary laws, ordonnances, décrets, and arrêtés), parliamentary debates from both chambers, and legislative files. Data were collected through: French government data portal https://www.data.gouv.fr/ For the Dutch case study, the repository includes parliamentary speeches, bills, amendments, motions, parliamentary questions, Council of State advisory opinions on both bills and draft decrees, Orders in Council (AMvBs), ministerial regulations, and public profiles and metadata of political actors. Data were collected through: Tweede Kamer der Staten-Generaal, https://www.tweedekamer.nl Eerste Kamer der Staten-Generaal, https://www.eerstekamer.nl Raad van State, https://www.raadvanstate.nl Officiele Bekendmakingen (Kennis- en Exploitatiecentrum Officiele Overheidspublicaties, KOOP), https://zoek.officielebekendmakingen.nl OpenKamer, https://openkamer.org ParlaMint (CLARIN), https://www.clarin.eu/parlamint For the Hungarian case study, the repository includes parliamentary speeches, proposed bills, adopted laws, Constitutional Court decisions, and public profiles of Members of Parliament. Government and ministerial decrees (Executive Orders) are not included in the open repository because they are distributed under restricted access in accordance with the licensing agreement with the data provider. Data were collected through: Comparative Agendas Project (CAP), https://www.comparativeagendas.net/datasets_codebooks Website of the Hungarian National Assembly, https://www.parlament.hu/web/guest/iromanyok-m ParlText dataset, https://dataverse.harvard.edu/dataverse/parltext Wolters Kluwer Hungary, https://net.jogtar.hu For the Czech case study, the repository includes parliamentary speeches, parliamentary prints and legislative documents, governmental documents from the ODok system, Constitutional Court decisions from the NALUS database, and their associated metadata. Data were collected through: Poslanecká Sněmovna Parlamentu České republiky, https://www.psp.cz/ NALUS: Vyhlédávání rozhodnutí Ústavního soudu České republiky, https://nalus.usoud.cz/ ODOK Portál, https://odok.gov.cz/portal/veklep/materialyItem type: Item , NameTag 3 Multilingual Model 260521(Charles University in Prague, ÚFAL, 2026-05-21) Straková, JanaThis is a trained model for the supervised machine learning tool NameTag 3 (https://ufal.mff.cuni.cz/nametag/3/). NameTag 3 is an open-source tool for both flat and nested named entity recognition (NER). NameTag 3 identifies proper names in text and classifies them into a set of predefined categories, such as names of persons, locations, organizations, etc. The model was trained jointly on 25 flat NE corpora of 20 languages: Arabic, Chinese, Croatian, Czech, Danish, Dutch, English, German, Greek, Hebrew, Maghrebi Arabic French, Norwegian Bokmål, Norwegian Nynorsk, Portuguese, Serbian, Slovak, Slovenian, Spanish, Swedish, and Ukrainian. The model documentation can be found at https://ufal.mff.cuni.cz/nametag/3/models#multilingual . If you use this tool for scientific work, please give us credit by referencing https://aclanthology.org/2025.acl-demo.4/ .Item type: Item , Possessive Pronoun Preference (2026-05-19)(Charles University, Faculty of Arts, 2024-09) Chromá, Anna; Perevozchikova, TatianaThe contribution includes the data frames and the R script (Markdown file) belonging to the paper "Morphological and Pragmatic Conditioning of Reflexivity in Possessive Pronouns: Effects of Number and Form of Address in Czech" submitted to the journal Linguistics: An Interdisciplinary Journal of the Language Sciences in May 2025. During the review process, title of the paper has been changed to "Reflexive vs. non-reflexive possessives in Czech imperative utterances: effects of number and politeness". The 26-05-19 version has major updates to the analytic script and minor updates to the data sets (no changes to the data, just the to the names of conditions).Item type: Item , Universal Dependencies 2.18(Universal Dependencies Consortium, 2026-05-15) Zeman, Daniel; Nivre, Joakim; Abid, Rimsha; Abrams, Mitchell; Ackermann, Elia; Adolphe, Jephtey; Aepli, Noëmi; Afzal, Muhammad; Agić, Željko; Ahrenberg, Lars; Ajede, Chika Kennedy; Akhundjanova, Arofat; Akkurt, Furkan; Albek, Orly; Aleksandravičiūtė, Gabrielė; Alexandre, Dominick Maia; Alfina, Ika; Algom, Avner; Alnajjar, Khalid; Alzetta, Chiara; Anastasopoulos, Antonios; Andersen, Erik; Andrews, Kirk; Andrews, Matthew; Antonsen, Lene; Aoyama, Tatsuya; Aplonova, Katya; Aquino, Angelina; Aragon, Carolina; Aranes, Glyd; Aranzabe, Maria Jesus; Arıcan, Bilge Nas; Arnardóttir, Þórunn; Aroonmanakun, Wirote; Arutie, Gashaw; Arwidarasti, Jessica Naraiswari; Asadpour, Hiwa; Asahara, Masayuki; Ásgeirsdóttir, Katla; Aslan, Deniz Baran; Asmazoğlu, Cengiz; Ateyah, Luma; Atmaca, Furkan; Attia, Mohammed; Atutxa, Aitziber; Augustinus, Liesbeth; Avelãs, Mariana; Aziz, Salwan; Badmaeva, Elena; Bajorat, Jana; Balasubramani, Keerthana; Ballesteros, Miguel; Banerjee, Esha; Bank, Sebastian; Barbosa, Bryan Khelven da Silva; Barbu Mititelu, Verginica; Barkarson, Starkaður; Basile, Rodolfo; Basmov, Victoria; Batchelor, Colin; Bauer, John; Bedir, Seyyit Talha; Begum, Zarina; Behzad, Shabnam; Beiner, Nathanaël; Belieni, Juan; Bémová, Alevtina; Bengoetxea, Kepa; Benli, İbrahim; Ben Moshe, Yifat; Benzerrak, Marie; Berdicevskis, Aleksandrs; Berg, Ansu; Berk, Gözde; Bernhard, Delphine; Berntsson Ingelstam, Astrid; Bhat, Riyaz Ahmad; Biagetti, Erica; Bick, Eckhard; Bielinskienė, Agnė; Bilgin Taşdemir, Esma Fatıma; Binici, Helin; Bjarnadóttir, Kristín; BK, Samuel; Blaschke, Verena; Blokland, Rogier; Böbel, Nina; Bobicev, Victoria; Boizou, Loïc; Bompolas, Stavros; Bona, Morgane; Bonilla, Johnatan; Borges Völker, Emanuel; Borin, Lars; Börstell, Carl; Bosco, Cristina; Bouma, Gosse; Bowman, Sam; Boyd, Adriane; Braggaar, Anouck; Branco, António; Bras, Myriam; Brauner, Elisheva; Brillet, Théo; Brokaitė, Kristina; Bu, Lanni; Buráňová, Eva; Burchardt, Aljoscha; Cabeza, Carmen; Cáceres Arandia, Natalia; Caftanatov, Olesea; Campos, Marisa; Candito, Marie; Cappello, Caterina Maria; Caron, Bernard; Caron, Gauthier; Carvalheiro, Catarina; Carvalho, Rita; Cassidy, Lauren; Castro, Maria Clara; Castro, Sérgio; Cavalcanti, Tatiana; Cebiroğlu Eryiğit, Gülşen; Cecchini, Flavio Massimiliano; Celano, Giuseppe G. A.; Çepani, Anila; Čéplö, Slavomír; Cesur, Neslihan; Cetin, Savas; Çetinoğlu, Özlem; Chalub, Fabricio; Chamila, Liyanage; Chamoreau, Claudine; Chaudhary, Aditi; Chauhan, Shweta; Chen, Yifei; Chi, Ethan; Chika, Taishi; Cho, Yongseok; Choi, Jinho; Chontaeva, Bermet; Chun, Jayeol; Chung, Juyeon; Cignarella, Alessandra T.; Cinková, Silvie; Cocco, Esther; Collomb, Aurélie; Çöltekin, Çağrı; Connor, Miriam; Corbetta, Claudia; Corbetta, Daniela; Costa, Francisco; Courtin, Marine; Crabbé, Benoît; Cristescu, Mihaela; Cvetkoski, Vladimir; Dahan, Netanel; Dale, Ingerid Løyning; D'Alì, Sabrina; Daniel, Philemon; Danielyan, Anna S.; Daoudi, Khensa; Dash, Bijayalaxmi; Dash, Satya Ranjan; Davidson, Elizabeth; de Alencar, Leonel Figueiredo; Dehouck, Mathieu; de Laurentiis, Martina; de Marneffe, Marie-Catherine; Demir, Ahmet; de Paiva, Valeria; Derin, Mehmet Oguz; de Souza, Elvis; Diaz de Ilarraza, Arantza; Díaz Hernández, Roberto Antonio; Dickerson, Carly; Di Felippo, Ariani; Dinakaramani, Arawinda; Di Nuovo, Elisa; Dione, Bamba; Dirix, Peter; Do, Hoa; Dobrovoljc, Kaja; Dogan, Mahîr; Döhmer, Caroline; Doyle, Adrian; Dozat, Timothy; Droganova, Kira; Duran, Magali Sanches; Dwivedi, Puneet; Dyer, Andrew Thomas; Ebert, Christian; Eckhoff, Hanne; Eguchi, Masaki; Eiche, Sandra; Eiselen, Roald; Eli, Marhaba; Elkahky, Ali; Ephrem, Binyam; Erina, Olga; Erjavec, Tomaž; Esher, Louise; Eslami, Soudabeh; Essaidi, Farah; Etienne, Aline; Evelyn, Wograine; Facundes, Sidney; Farkas, Richárd; Faryad, Ján; Favero, Federica; Ferdaousi, Jannatul; Fernanda, Marília; Fernandez Alcalde, Hector; Fethi, Amal; Foster, Jennifer; Francioni, Barbara; Fransen, Theodorus; Freitas, Cláudia; Fuchs, Shlomit; Fujita, Kazunori; Gajdošová, Katarína; Galbraith, Daniel; Galves, Charlotte Chambelland; Galy, Edith; Gamba, Federica; Garcia, Marcos; García-Miguel, José María; Gärdenfors, Moa; Gaustad, Tanja; Genç, Efe Eren; Gerardi, Fabrício Ferraz; Gerdes, Kim; Gessler, Luke; Ginter, Filip; Godoy, Gustavo; Goenaga, Iakes; Gojenola, Koldo; Gökırmak, Memduh; Goldberg, Yoav; Goldin, Gili; Gómez Guinovart, Xavier; González Saavedra, Berta; Goux, Mathieu; Grand-Clement, Caroline; Griciūtė, Bernadeta; Grioni, Matias; Grobol, Loïc; Grūzītis, Normunds; Guglielmetti, Mario; Guillaume, Bruno; Guiller, Kirian; Guillot-Barbance, Céline; Güngör, Tunga; Gurevich, Vladimir; Habash, Nizar; Hafsteinsson, Hinrik; Hahn, Michael; Hajič, Jan; Hajič jr., Jan; Hajičová, Eva; Hämäläinen, Mika; Hà Mỹ, Linh; Han, Na-Rae; Hanifmuti, Muhammad Yudistira; Harada, Takahiro; Hardwick, Sam; Harris, Kim; Hassert, Naïma; Haug, Dag; Havelka, Jiří; Heinecke, Johannes; Hellwig, Oliver; Hennig, Felix; Hladká, Barbora; Hlaváčová, Jaroslava; Hociung, Florinel; Hoefels, Diana; Hoff, Barbara; Hohle, Petter; Howell, Nick; Huang, Yidi; Huerta Mendez, Marivel; Hwang, Jena; Ikeda, Takumi; Iliadou, Inessa; Ingason, Anton Karl; Ion, Radu; Irimia, Elena; Ishola, Ọlájídé; Islamaj, Artan; Ito, Kaoru; Iurescia, Federica; Ivani, Jessica K.; Jagodzińska, Sandra; Jannat, Siratun; Jelínek, Tomáš; Jha, Apoorva; Jiamsundutsadee, Ratanon; Jiang, Katharine; Job, Sylvanus; Jobanputra, Mayank; Johannsen, Anders; Jónsdóttir, Hildur; Jørgensen, Fredrik; Ju, Zhuoxuan; Juarez Huerta, María Ximena; Juutinen, Markus; Kaşıkara, Hüner; Kabaeva, Nadezhda; Kahane, Sylvain; Kanayama, Hiroshi; Kanerva, Jenna; Kara, Neslihan; Karahóǧa, Ritván; Kárník, Jiří; Kåsen, Andre; Kayadelen, Tolga; Kengatharaiyer, Sarveswaran; Kettnerová, Václava; Khan, Ali Haider; Kharatyan, Lilit; Kirchner, Jesse; Klementieva, Elena; Klironomou, Christina; Klyachko, Elena; Kocharov, Petr; Köhn, Arne; Köksal, Abdullatif; Kolářová, Veronika; Kopacewicz, Kamil; Korkiakangas, Timo; Köse, Mehmet; Koshevoy, Alexey; Kote, Nelda; Kotsyba, Natalia; Kovačić, Barbara; Kovalevskaitė, Jolanta; Kowner, Emmanuelle; Krek, Simon; Krishnamurthy, Parameswari; Kübler, Sandra; Kučová, Lucie; Kuqi, Adrian; Kuriyozov, Elmurod; Kushare, Pranav; Kuyrukçu, Oğuzhan; Kuzgun, Aslı; Kwak, Sookyoung; Kyle, Kris; Laan, Käbi; Laippala, Veronika; Lambertino, Lorenzo; Landau, Israel; Lando, Tatiana; Larasati, Septina Dian; Larrivée, Pierre; Lata, Kusum; Lavrentiev, Alexei; Lee, John; Lê Hồng, Phương; Lenci, Alessandro; Leong, Wei Qi; Lertpradit, Saran; Leung, Herman; Levin, Lori; Levina, Maria; Levine, Lauren; Li, Cheuk Ying; Li, Josie; Li, Keying; Li, Yixuan; Li, Yuan; Lim, KyungTae; Lima Padovani, Bruna; Lin, Yi-Ju Jessica; Lindén, Krister; Lindenbaum, Yitzchak; Liu, Yang Janet; Liu, Zoey; Ljubešić, Nikola; Lobzhanidze, Irina; Loginova, Olga; Lopatková, Markéta; Lopes, Lucelene; Luftiu, Edita; Lukashevskyi, Arsenii; Lusito, Stefano; Lutgen, Anne-Marie; Luthfi, Andry; Luukko, Mikko; Lyashevskaya, Olga; Lynn, Teresa; Macketanz, Vivien; Mahamdi, Menel; Maillard, Jean; Maitreenukul, Punyanuch; Makarchuk, Ilya; Makazhanov, Aibek; Mambrini, Francesco; Mandl, Michael; Manning, Christopher; Manurung, Ruli; Marşan, Büşra; Mărănduc, Cătălina; Mareček, David; Marheinecke, Katrin; Markantonatou, Stella; Márquez Hernández, Ángeles; Martínez Alonso, Héctor; Martín Rodríguez, Lorena; Martins, André; Martins, Cláudia; Masciolini, Arianna; Mašek, Jan; Matlatipov, Sanatbek; Matsuda, Hiroshi; Matsumoto, Yuji; Mauri, Caterina; Mazzei, Alessandro; McDonald, Ryan; McGuinness, Sarah; Mehta, Maitrey; Meiri, Ephraim; Ménard, Pierre André; Mendonça, Gustavo; Merhav, Hilla; Merzhevich, Tatiana; Meurer, Paul; Miekka, Niko; Mikulová, Marie; Milano, Emilia; Miletić, Aleksandra; Miller, Aaron; Min, Junghyun; Minerbi, Yael; Mírovský, Jiří; Mischenkova, Karina; Missilä, Anna; Mititelu, Cătălin; Mitrofan, Maria; Miyao, Yusuke; Mohapatra, Biswakalpita; Molnár, Judit; Moloodi, Amirsaeid; Montemagni, Simonetta; More, Amir; Moreno Romero, Laura; Moretti, Giovanni; Mori, Shinsuke; Morioka, Tomohiko; Moro, Shigeki; Mortensen, Bjartur; Moskalevskyi, Bohdan; Mouzou, Katerina; Muischnek, Kadri; Munro, Robert; Murawaki, Yugo; Mus, Nikolett; Müürisep, Kaili; Nainwani, Pinkey; Nakhlé, Mariam; Nassajian, Minoo; Navarro Horñiacek, Juan Ignacio; Nedoluzhko, Anna; Nešpore-Bērzkalne, Gunta; Nevaci, Manuela; Nguyễn Thị, Lương; Nguyễn Thị Minh, Huyền; Nikaido, Yoshihiro; Nikolaev, Vitaly; Nitisaroj, Rattima; Norrman, Victor; Nourian, Alireza; Novák, Michal; Nunes, Maria das Graças Volpe; Nurmi, Hanna; O'Brien, Colleen Alena; Ojala, Stina; Ojha, Atul Kr.; Óladóttir, Hulda; Olúòkun, Adédayọ̀; Omura, Mai; Onwuegbuzia, Emeka; Ordan, Noam; Osenova, Petya; Östling, Robert; Ott, Annika; Øvrelid, Lilja; Oya, Masanori; Özateş, Şaziye Betül; Özçelik, Merve; Özgür, Arzucan; Öztürk Başaran, Balkız; Paccosi, Teresa; Pajas, Petr; Palakapilly, Thomas; Palmero Aprosio, Alessio; Panevová, Jarmila; Pannitto, Ludovica; Panova, Anastasia; Pardo, Thiago Alexandre Salgueiro; Parida, Shantipriya; Park, Hyunji Hayley; Partanen, Niko; Pascual, Elena; Pasparaki, Thelka; Passarotti, Marco; Patejuk, Agnieszka; Paulino-Passos, Guilherme; Pedonese, Giulia; Peeters, Oggi; Peljak-Łapińska, Angelika; Peng, Siyao; Peng, Siyao Logan; Pereira, Rita; Pereira, Sílvia; Perez, Cenel-Augusto; Perkova, Natalia; Perrier, Guy; Petrov, Slav; Petrova, Daria; Pettersson, Eva; Peverelli, Andrea; Phelan, Jason; Pierre-Louis, Claudel; Piitulainen, Jussi; Pinter, Yuval; Pinto, Clara; Pintucci, Rodrigo; Pirinen, Tommi A; Pitler, Emily; Plamada, Magdalena; Plank, Barbara; Plum, Alistair; Poibeau, Thierry; Polpanumas, Charin; Ponomareva, Larisa; Popel, Martin; Poujade, Clamença; Prangwedana, Rangga Prangwedana; Pretkalniņa, Lauma; Pretorius, Rigardt; Prévost, Sophie; Prokopidis, Prokopis; Przepiórkowski, Adam; Pugh, Robert; Puolakainen, Tiina; Purschke, Christoph; Pyysalo, Sampo; Qi, Peng; Querido, Andreia; Rääbis, Andriela; Rabinovich, Ella; Rademaker, Alexandre; Rahman, Mutee-u; Rahoman, Mizanur; Rama, Taraka; Ramasamy, Loganathan; Ramisch, Carlos; Ramos, Joana; Rashel, Fam; Rasooli, Mohammad Sadegh; Ravishankar, Vinit; Real, Livy; Rebeja, Petru; Reddy, Siva; Regnault, Mathilde; Rehm, Georg; Riabi, Arij; Riabov, Ivan; Rießler, Michael; Rimkutė, Erika; Rinaldi, Larissa; Rituma, Laura; Rizqiyah, Putri; Rocha, Luisa; Rögnvaldsson, Eiríkur; Roksandic, Ivan; Roman, Norton Trevisan; Romanenko, Mykhailo; Romanova, Natalia; Rosa, Rudolf; Roșca, Valentin; Roulon, Paulette; Rovati, Davide; Rozonoyer, Ben; Rudina, Olga; Rueter, Jack; Ruffolo, Paolo; Rúnarsson, Kristján; Rushiti, Rozana; Rutherford, Attapol T.; Sadde, Shoval; Safari, Pegah; Sahala, Aleksi; Sahoo, Kalyanamalini; Sahoo, Saraswati; Saleh, Shadi; Salomoni, Alessio; Samardžić, Tanja; Sammani, Warangana; Sampanis, Konstantinos; Samson, Stephanie; Sánchez-Rodríguez, Xulia; Sandalo, Filomena Spatti; Sanguinetti, Manuela; Sanıyar, Ezgi; Särg, Dage; Sartor, Marta; Sarymsakova, Albina; Sasaki, Mitsuya; Saulīte, Baiba; Savary, Agata; Sawanakunanon, Yanin; Saxena, Shefali; Scannell, Kevin; Scarlata, Salvatore; Schang, Emmanuel; Schikowski, Robert; Schneider, Nathan; Schuster, Sebastian; Schwartz, Lane; Seddah, Djamé; Seeker, Wolfgang; Sellmer, Sven; Sengupta, Kaushik; Seraji, Mojgan; Ševčíková, Magda; Sgall, Petr; Shahzadi, Syeda; Shen, Mo; Shimada, Atsuko; Shin, Gyu-Ho; Shirasu, Hiroyuki; Shishkina, Yana; Shmidman, Avi; Shohibussirri, Muh; Shvedova, Maria; Sibille, Jean; Siewert, Janine; Sigurðsson, Einar Freyr; Silva, João; Silveira, Aline; Silveira, Natalia; Silveira, Sara; Simi, Maria; Simionescu, Radu; Simkó, Katalin; Šimková, Mária; Símonarson, Haukur Barri; Simov, Kiril; Sitchinava, Dmitri; Sither, Ted; Smith, Aaron; Soares-Bastos, Isabela; Solberg, Per Erik; Sollberger, Dolores; Sonnenhauser, Barbara; Sourov, Shafi; Speransky, Nina; Sprugnoli, Rachele; Sriwirote, Panyut; Stamou, Vivian; Steingrímsson, Steinþór; Stella, Antonio; Štěpánek, Jan; Štěpánková, Barbora; Stephen, Abishek; Straka, Milan; Strass, Omer; Strickland, Emmett; Strnadová, Jana; Stymne, Sara; Suhr, Alane; Sulestio, Yogi Lesmana; Sulubacak, Umut; Sun, Jack; Sung, Hakyung; Suzuki, Shingo; Swanson, Daniel; Szántó, Zsolt; Szawerna, Maria Irena; Taguchi, Chihiro; Taji, Dima; Tal, Rachel; Talamo, Luigi; Tamburini, Fabio; Tan, Mary Ann C.; Tanaka, Takaaki; Tanaya, Dipta; Tavoni, Mirko; Teker, Nursena; Tella, Samson; Tellier, Isabelle; Testori, Marinella; Thanyawong, Santhawat; Thomas, Guillaume; Tıraş, Tarık Emre; Tjhi, William Chandra; Tollersrud, Thea; Tomaszek, Kamil; Tonelli, Sara; Torga, Liisi; Toribio, Lucas; Toska, Marsida; Trosterud, Trond; Trukhina, Anna; Tsarfaty, Reut; Tulchynska, Kira; Türk, Utku; Tyers, Francis; Þórðarson, Sveinbjörn; Þorsteinsson, Vilhjálmur; Uematsu, Sumire; Untilov, Roman; Urešová, Zdeňka; Uria, Larraitz; Uszkoreit, Hans; Utka, Andrius; Vagnoni, Elena; Vajjala, Sowmya; Vak, Socrates; Vakirtzian, Socrates; van der Goot, Rob; Vanhove, Martine; van Niekerk, Daniel; van Noord, Gertjan; Varga, Viktor; Vaz, Helena; Vedenina, Uliana; Venturi, Giulia; Vergez-Couret, Marianne; Verkerk, Annemarie; Veronesi, Luiz; Veselovsky, Anna; Vidová Hladká, Barbora; Villemonte de la Clergerie, Eric; Vincze, Veronika; Vissamsetty, Anishka; Vlasova, Natalia; Vligouridou, Eleni; Vogel, Alan; Wakasa, Aya; Wallenberg, Joel C.; Wallin, Lars; Walsh, Abigail; Wang, John; Washington, Jonathan North; Weissweiler, Leonie; Wendt, Maximilan; Widmer, Paul; Wieczorek, Aleksandra; Wigderson, Shira; Wijono, Sri Hartati; Wille, Vanessa Berwanger; Williams, Seyi; Winkler, Miriam; Wintner, Shuly; Wirén, Mats; Wittern, Christian; Witzlack-Makarevich, Alena; Woldemariam, Tsegay; Wong, Tak-sum; Wróblewska, Alina; Wu, Qishen; Yako, Mary; Yamashita, Kayo; Yamazaki, Naoki; Yan, Chunxiao; Yang, Qizhen; Yang, Xiulin; Yasuoka, Koichi; Yavrumyan, Marat M.; Yenice, Arife Betül; Yılandiloğlu, Enes; Yıldız, Olcay Taner; Yu, Zhuoran; Yuliawati, Arlisa; Žabokrtský, Zdeněk; Zahra, Shorouq; Zeldes, Amir; Zhang, Annie; Zhang, Larry; Zhou, He; Zhu, Hanzhi; Zhu, Yilun; Zhuravleva, Anna; Ziane, Rayan; Znotiņš, Artūrs; Zucchini, EleonoraUniversal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).Item type: Item , DeriVallex 1.0(Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL), 2026-05-08) Kettnerová, Václava; Mírovský, Jiří; Kolářová, Veronika; Olbrich, MichalDeriVallex 1.0 is a valency lexicon of automatically generated valency frames of Czech noun and adjectival derivatives the valency of which exhibits systemic correspondences with the valency of their base words. It contains 10,220 derivatives corresponding to 17,288 lexical units (i.e., individual senses). In particular, DeriVallex describes 3,134 nouns corresponding to 5,089 lexical units and 7,086 adjectives corresponding to 12,199 lexical units. DeriVallex was created with the aim of providing information on the valency of nouns and adjectives, which is not sufficiently covered in existing lexical resources. Focusing on nominal and adjectival derivatives that exhibit systematic valency behavior in comparison with their base words, it captures the productive and systemic core of the Czech lexicon, thus laying the foundation for the further extension of current lexical resources. The following word-formation categories are covered: action nouns (e.g., dobytí města nepřáteli ‘conquering the city by enemies’), quality nouns (e.g., učitelova laskavost k dětem ‘the teacher’s kindness to children’), simultaneous action adjectives (e.g., lidé bojující proti bezpráví ‘people fighting against injustice’), anterior action adjectives (e.g., dluh narostlý na 400 milionů ‘a debt that has risen to 400 million’ and muži navrátivší se z války ‘men who have returned from the war’), passive action adjectives (e.g., úspory diktované Evropě konzervativní vládou ‘austerity measures dictated to Europe by a conservative government’), and potentiality adjectives (e.g., dužina oddělitelná od pecky ‘flesh separable from the pit’). In compiling the lexicon, data from the following lexical resources were used: NomVallex 2.6, VALLEX 4.5, and DeriNet 2.3. To satisfy different needs of potential users, the lexicon is distributed (i) online in an HTML version (providing a user-friendly interface allowing human users to search and filter the data) and (ii) in this distribution in a machine-readable form, so that the data can be used in NLP applications. Authors: Václava Kettnerová, Jiří Mírovský, Veronika Kolářová and Michal Olbrich Acknowledgement: The creation of the DeriVallex lexicon has been supported by the LINDAT/CLARIAH-CZ Research Infrastructure (https://lindat.cz), supported by the Ministry of Education, Youth and Sports of the Czech Republic (Project No. LM2023062), and it has been using data and tools provided by this project too. License: DeriVallex is publicly available under the Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International license (CC BY-NC-SA). Its non-commercial use is conditioned by appropriate citation: Kettnerová, Václava and Mírovský, Jiří and Kolářová, Veronika and Olbrich, Michal. 2026. DeriVallex 1.0. LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL). http://hdl.handle.net/11234/1-6109.Item type: Item , SiR 2.0(Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL), 2026-04-30) Mírovský, Jiří; Hladká, Barbora; Kopp, Matyáš; Moravec, VáclavSiR 2.0 is an update of an annotated corpus of Czech articles published on iRozhlas, a news server of a Czech public radio (https://www.irozhlas.cz/). SiR 2.0 is a collection of 1 718 articles (42 890 sentences, 614 995 words) with manually annotated attribution of citation phrases and sources. The sources are classified into several classes of named and unnamed sources. The corpus consists of two parts, depending on the origin of the annotations: (i) expert-annotated articles: 589 articles (13 280 sentences) annotated originally by three or two student annotators and later curated or re-annotated by an expert, (ii) student-annotated articles:: 1 129 articles (29 610 sentences) annotated each by a single student annotator. The data were annotated in the Brat tool (https://brat.nlplab.org/) and are distributed in the Brat native format, i.e. each article is represented by the original plain text and a stand-off annotation file. In total, there are annotated 10 033 citation phrases, 8 960 citation sources and 9 317 links between sources and phrases.Item type: Item , Persian Little Prince annotated with named entities(Charles University in Prague, ÚFAL, 2026-04-30) Nassajian, Minoo; Zeman, DanielThis Corpus is the Persian translation of 'The Little Prince' story, which is annotated by Named Entities Labels. Existing NE recognizers for Persian display significant performance degradation on literary text, identifying systematic errors related to narrative-specific entities, metaphorical language, and discourse structures that challenge conventional NER approaches. This dataset will enable training and/or evaluating NE recognizers for literary text as well.Item type: Item , Uniform Meaning Representation 2.2(UMR Consortium, 2026-04-01) Bonn, Julia; Bonial, Claire; Buchholz, Matt; Cheng, Hsiao-Jung; Chen, Alvin; Chen, Ching-wen; Cowell, Andrew; Croft, William; Denk, Lukas; Elsayed, Ahmed; Fučíková, Eva; Gamba, Federica; Gomez, Carlos; Hajič, Jan; Hajičová, Eva; Havelka, Jiří; Havenmeier, Loden; Hledíková, Hana; Kilgore, Ath; Kolářová, Veronika; Kučová, Lucie; Lai, Kenneth; Li, Bin; Li, Jingyi; Lopatková, Markéta; MacGregor, Marie; Mikulová, Marie; Mírovský, Jiří; Nedoluzhko, Anna; Myers, Skatje; Novák, Michal; O’Gorman, Tim; Pajas, Petr; Palmer, Alexis; Palmer, Martha; Panevová, Jarmila; Post, Benét; Pustejovsky, James; Sgall, Petr; Song, Jialin; Song, Li; Ševčíková, Magda; Štěpánek, Jan; Urešová, Zdeňka; Sun, Haibo; Sun, Yao; Vallejos Yopán, Rosa; Van Gysel, Jens; Vigus, Meagan; Wright‑Bettner, Kristin; Wu, Jiawei; Xue, Nianwen; Xing, Dan; Xu, Keer; Xu, Zhixing; Yue, Liulu; Zeman, Daniel; Zhao, Jin; Zikánová, Šárka; Žabokrtský, ZdeněkNew version of UMR data used in the First Shared Task on UMR Parsing (including submitted systems' outputs).Item type: Item , Human Label Variation in Coreference (Hlava Cor)(Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL), 2026-03-31) Nedoluzhko, Anna; Mírovský, Jiří; Zikánová, Šárka; Hajičová, Eva; Chuffartová Bianca; Dohnalová Šárka; Hartmanová Lucie; Nodlová Eliška; Teska Dominik; Zikánová FrantiškaHuman Label Variation in Coreference (Hlava COR) is a collection of commented multiple annotations (three annotators) of coreferential relations in Czech, i.e. the annotation of expressions that refer to the same extra-linguistic entity, concept, or situation. Given an anaphoric expression, annotators were instructed to identify a coreferential expression in the preceding context (if one exists) and to comment on their decision. The main aim of the annotation is to capture variation in the interpretation of coreference among readers. The dataset includes both written and spoken contexts. For detailed and up-to-date information about the corpus, please visit: https://ufal.mff.cuni.cz/hvar/hlava-cor

