Number of results to display per page
Search Results
112. PARSEME corpora annotated for verbal multiword expressions (version 1.3)
- Creator:
- Savary, Agata, Ramisch, Carlos, Guillaume, Bruno, Hawwari, Abdelati, Walsh, Abigail, Fotopoulou, Aggeliki, Bielinskienė, Agnė, Estarrona, Ainara, Gatt, Albert, Butler, Alexandra, Rademaker, Alexandre, Maldonado, Alfredo, Villavicencio, Aline, Farrugia, Alison, Muscat, Amanda, Gatt, Anabelle, Antić, Anđela, De Santis, Anna, Raffone, Annalisa, Riccio, Anna, Pascucci, Antonio, Gurrutxaga, Antton, Bhatia, Archna, Vaidya, Ashwini, Miral, Ayşenur, QasemiZadeh, Behrang, Priego Sanchez, Belem, Griciūtė, Bernadeta, Erden, Berna, Parra Escartín, Carla, Herrero, Carlos, Carlino, Carola, Pasquer, Caroline, Liebeskind, Chaya, Wang, Chenweng, Ben Khelil, Chérifa, Bonial, Claire, Somers, Clarissa, Aceta, Cristina, Krstev, Cvetana, Bejček, Eduard, Lindqvist, Ellinor, Erenmalm, Elsa, Palka-Binkiewicz, Emilia, Rimkute, Erika, Petterson, Eva, Cap, Fabienne, Hu, Fangyuan, Sangati, Federico, Wick Pedro, Gabriela, Speranza, Giulia, Jagfeld, Glorianna, Blagus, Goranka, Berk, Gözde, Attard, Greta, Eryiğit, Gülşen, Finnveden, Gustav, Martínez Alonso, Héctor, de Medeiros Caseli, Helena, Elyovich, Hevi, Xu, Hongzhi, Xiao, Huangyang, Miranda, Isaac, Jaknić, Isidora, El Maarouf, Ismail, Aduriz, Itziar, Gonzalez, Itziar, Matas, Ivana, Stoyanova, Ivelina, Jazbec, Ivo-Pavao, Busuttil, Jael, Waszczuk, Jakub, Findlay, Jamie, Bonnici, Janice, Šnajder, Jan, Antoine, Jean-Yves, Foster, Jennifer, Chen, Jia, Nivre, Joakim, Monti, Johanna, McCrae, John, Kovalevskaitė, Jolanta, Jain, Kanishka, Simkó, Katalin, Yu, Ke, Azzopardi, Kirsty, Adalı, Kübra, Uria, Larraitz, Zilio, Leonardo, Boizou, Loïc, van der Plas, Lonneke, Galea, Luke, Sarlak, Mahtab, Buljan, Maja, Cherchi, Manuela, Tanti, Marc, Di Buono, Maria Pia, Todorova, Maria, Candito, Marie, Constant, Matthieu, Shamsfard, Mehrnoush, Jiang, Menghan, Boz, Mert, Spagnol, Michael, Onofrei, Mihaela, Li, Minli, Elbadrashiny, Mohamed, Diab, Mona, Rizea, Monica-Mihaela, Hadj Mohamed, Najet, Theoxari, Natasa, Schneider, Nathan, Tabone, Nicole, Ljubešić, Nikola, Vale, Oto, Cook, Paul, Yan, Peiyi, Gantar, Polona, Ehren, Rafael, Fabri, Ray, Ibrahim, Rehab, Ramisch, Renata, Walles, Rinat, Wilkens, Rodrigo, Urizar, Ruben, Sun, Ruilong, Malka, Ruth, Galea, Sara Anne, Stymne, Sara, Louizou, Sevasti, Hu, Sha, Taslimipoor, Shiva, Ratori, Shraddha, Srivastava, Shubham, Cordeiro, Silvio Ricardo, Krek, Simon, Liu, Siyuan, Zeng, Si, Yu, Songping, Arhar Holdt, Špela, Markantonatou, Stella, Papadelli, Stella, Leseva, Svetlozara, Kuzman, Taja, Kavčič, Teja, Lynn, Teresa, Lichte, Timm, Pickard, Thomas, Dimitrova, Tsvetana, Yih, Tsy, Güngör, Tunga, Dinç, Tutkum, Iñurrieta, Uxoa, Tajalli, Vahide, Stefanova, Valentina, Caruso, Valeria, Puri, Vandana, Foufi, Vassiliki, Barbu Mititelu, Verginica, Vincze, Veronika, Kovács, Viktória, Shukla, Vishakha, Giouli, Voula, Ge, Xiaomin, Ha-Cohen Kerner, Yaakov, Öztürk, Yağmur, Yarandi, Yalda, Parmentier, Yannick, Zhang, Yongchen, Zhao, Yun, Urešová, Zdeňka, Yirmibeşoğlu, Zeynep, Qin, Zhenzhen, Stank, Cristescu, Mihaela, Zgreabăn, Bianca-Mădălina, Bărbulescu, Elena-Andreea, and Stanković, Ranka
- Publisher:
- PARSEME
- Type:
- text and corpus
- Subject:
- multiword expressions, verbal multiword expressions, light verb construction, verb-particle constructions, inherently reflexive verbs, verbal idioms, and multi-verb constructions
- Language:
- Arabic, Bulgarian, Czech, German, Modern Greek (1453-), English, Spanish, Basque, Persian, French, Irish, Hebrew, Hindi, Croatian, Hungarian, Lithuanian, Italian, Maltese, Polish, Portuguese, Romanian, Slovenian, Serbian, Swedish, Turkish, and Chinese
- Description:
- This multilingual resource contains corpora in which verbal MWEs have been manually annotated. VMWEs include idioms (let the cat out of the bag), light-verb constructions (make a decision), verb-particle constructions (give up), inherently reflexive verbs (help oneself), and multi-verb constructions (make do). This is the first release of the corpora without an associated shared task. Previous version (1.2) was associated with the PARSEME Shared Task on semi-supervised Identification of Verbal MWEs (2020). The data covers 26 languages corresponding to the combination of the corpora for all previous three editions (1.0, 1.1 and 1.2) of the corpora. VMWEs were annotated according to the universal guidelines. The corpora are provided in the cupt format, inspired by the CONLL-U format. Morphological and syntactic information, including parts of speech, lemmas, morphological features and/or syntactic dependencies, are also provided. Depending on the language, the information comes from treebanks (e.g., Universal Dependencies) or from automatic parsers trained on treebanks (e.g., UDPipe). All corpora are split into training, development and test data, following the splitting strategy adopted for the PARSEME Shared Task 1.2. The annotation guidelines are available online: https://parsemefr.lis-lab.fr/parseme-st-guidelines/1.3 The .cupt format is detailed here: https://multiword.sourceforge.net/cupt-format/
- Rights:
- PARSEME Corpora v. 1.3 - Licence Agreement, https://lindat.mff.cuni.cz/repository/xmlui/page/licence-mwe-1.3, and PUB
113. Plaintext Wikipedia dump 2018
- Creator:
- Rosa, Rudolf
- Publisher:
- Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
- Type:
- text and corpus
- Subject:
- Wikipedia, text corpora, and monolingual corpus
- Language:
- Abkhazian, Achinese, Adyghe, Afrikaans, Akan, Tosk Albanian, Amharic, Old English (ca. 450-1100), Arabic, Official Aramaic (700-300 BCE), Aragonese, Egyptian Arabic, Assamese, Asturian, Atikamekw, Avaric, Aymara, South Azerbaijani, Azerbaijani, Bashkir, Bambara, Bavarian, Central Bikol, Belarusian, Bengali, Bislama, Banjar, Tibetan, Bosnian, Bishnupriya, Breton, Buginese, Bulgarian, Russia Buriat, Catalan, Min Dong Chinese, Cebuano, Czech, Chamorro, Chechen, Cherokee, Church Slavic, Chuvash, Cheyenne, Central Kurdish, Cornish, Corsican, Cree, Crimean Tatar, Kashubian, Welsh, Danish, German, Dinka, Dimli (individual language), Dhivehi, Lower Sorbian, Dzongkha, Modern Greek (1453-), English, Esperanto, Estonian, Basque, Ewe, Extremaduran, Faroese, Persian, Fijian, Finnish, French, Arpitan, Northern Frisian, Western Frisian, Fulah, Friulian, Gagauz, Gan Chinese, Scottish Gaelic, Irish, Galician, Gilaki, Manx, Goan Konkani, Gothic, Guarani, Gujarati, Hakka Chinese, Haitian, Hausa, Hawaiian, Serbo-Croatian, Hebrew, Herero, Fiji Hindi, Hindi, Hiri Motu, Croatian, Upper Sorbian, Hungarian, Armenian, Igbo, Ido, Inuktitut, Interlingue, Iloko, Interlingua (International Auxiliary Language Association), Indonesian, Inupiaq, Icelandic, Italian, Jamaican Creole English, Javanese, Lojban, Japanese, Kara-Kalpak, Kabyle, Kalaallisut, Kannada, Kashmiri, Georgian, Kanuri, Kazakh, Kabardian, Kabiyè, Khmer, Kikuyu, Kinyarwanda, Kirghiz, Komi-Permyak, Komi, Kongo, Korean, Karachay-Balkar, Kölsch, Kurdish, Ladino, Lao, Latin, Latvian, Lak, Lezghian, Ligurian, Limburgan, Lingala, Lithuanian, Lombard, Northern Luri, Latgalian, Luxembourgish, Ganda, Literary Chinese, Marshallese, Maithili, Malayalam, Marathi, Moksha, Eastern Mari, Minangkabau, Macedonian, Malagasy, Maltese, Mongolian, Maori, Western Mari, Malay (macrolanguage), Creek, Mirandese, Burmese, Erzya, Mazanderani, Min Nan Chinese, Neapolitan, Nauru, Navajo, Ndonga, Low German, Nepali (macrolanguage), Newari, Dutch, Norwegian Nynorsk, Norwegian, Novial, Pedi, Nyanja, Occitan (post 1500), Livvi, Oriya (macrolanguage), Oromo, Ossetian, Pangasinan, Pampanga, Panjabi, Papiamento, Picard, Pennsylvania German, Pfaelzisch, Pitcairn-Norfolk, Pali, Piemontese, Western Panjabi, Pontic, Polish, Portuguese, Pushto, Quechua, Vlax Romani, Romansh, Romanian, Rusyn, Rundi, Macedo-Romanian, Russian, Sango, Yakut, Sanskrit, Sicilian, Scots, Samogitian, Sinhala, Slovak, Slovenian, Northern Sami, Samoan, Shona, Sindhi, Somali, Southern Sotho, Spanish, Albanian, Sardinian, Sranan Tongo, Serbian, Swati, Saterfriesisch, Sundanese, Swahili (macrolanguage), Swedish, Silesian, Tahitian, Tamil, Tatar, Tulu, Telugu, Tama (Colombia), Tetum, Tajik, Tagalog, Thai, Tigrinya, Tonga (Tonga Islands), Tok Pisin, Tswana, Tsonga, Turkmen, Tumbuka, Turkish, Twi, Tuvinian, Udmurt, Uighur, Ukrainian, Urdu, Uzbek, Venetian, Venda, Veps, Vietnamese, Vlaams, Volapük, Võro, Waray (Philippines), Walloon, Wolof, Wu Chinese, Kalmyk, Xhosa, Mingrelian, Yiddish, Yoruba, Yue Chinese, Zeeuws, Zhuang, Chinese, Zulu, and Dotyali
- Description:
- Wikipedia plain text data obtained from Wikipedia dumps with WikiExtractor in February 2018. The data come from all Wikipedias for which dumps could be downloaded at [https://dumps.wikimedia.org/]. This amounts to 297 Wikipedias, usually corresponding to individual languages and identified by their ISO codes. Several special Wikipedias are included, most notably "simple" (Simple English Wikipedia) and "incubator" (tiny hatching Wikipedias in various languages). For a list of all the Wikipedias, see [https://meta.wikimedia.org/wiki/List_of_Wikipedias]. The script which can be used to get new version of the data is included, but note that Wikipedia limits the download speed for downloading a lot of the dumps, so it takes a few days to download all of them (but one or a few can be downloaded fast). Also, the format of the dumps changes time to time, so the script will probably eventually stop working one day. The WikiExtractor tool [http://medialab.di.unipi.it/wiki/Wikipedia_Extractor] used to extract text from the Wikipedia dumps is not mine, I only modified it slightly to produce plaintext outputs [https://github.com/ptakopysk/wikiextractor].
- Rights:
- Attribution-ShareAlike 3.0 Unported (CC BY-SA 3.0), http://creativecommons.org/licenses/by-sa/3.0/, and PUB
114. Radvaň n. Dunajom
- Publisher:
- Vojenský zeměpisný ústav
- Format:
- map and 1 mapa : barevná ; 39 x 52 cm na listu 48 x 62 cm
- Type:
- model:map, cartographic, and IMAGE
- Subject:
- udc:913(4), Konspekt:7, udc:912, udc:913(437.6), udc:912.43, udc:(084.3), Konspekt:Geografie Evropy, reálie, cestování, Konspekt:Mapy. Atlasy. Glóby, and czenas:Radvaň nad Dunajom (Slovensko : oblast)
- Language:
- Czech and Hungarian
- Description:
- 4961, Edice dle kladu listů, and (Language) Místní názvy maďarsky
- Rights:
- http://creativecommons.org/publicdomain/mark/1.0/ and policy:public
115. Rožňava
- Publisher:
- Vojenský zeměpisný ústav
- Format:
- map and 1 mapa : barevná ; 39 x 51 cm na listu 47 x 63 cm
- Type:
- model:map, cartographic, and IMAGE
- Subject:
- udc:913(4), Konspekt:7, udc:912, udc:913(437.6), udc:912.43, udc:(084.3), Konspekt:Geografie Evropy, reálie, cestování, Konspekt:Mapy. Atlasy. Glóby, and czenas:Rožňava (Slovensko : oblast)
- Language:
- Czech, Slovak, and Hungarian
- Description:
- 4565, Legenda, Edice dle kladu listů, and (Language) Místní názvy slovensky, částečně maďarsky
- Rights:
- http://creativecommons.org/publicdomain/mark/1.0/ and policy:public
116. Ružový dvor
- Publisher:
- Vojenský zeměpisný ústav
- Format:
- map and 1 mapa : barevná ; 39 x 51 cm na listu 48 x 62 cm
- Type:
- model:map, cartographic, and IMAGE
- Subject:
- udc:913(4), Konspekt:7, udc:912, udc:913(439), udc:912.43, udc:(084.3), Konspekt:Geografie Evropy, reálie, cestování, Konspekt:Mapy. Atlasy. Glóby, and czenas:Gönc (Maďarsko : oblast)
- Language:
- Czech and Hungarian
- Description:
- 4666, Legenda, Edice dle kladu listů, and (Language) Místní názvy maďarsky
- Rights:
- http://creativecommons.org/publicdomain/mark/1.0/ and policy:public
117. Šahy
- Publisher:
- Vojenský zeměpisný ústav
- Format:
- map and 1 mapa : barevná ; 39 x 51 cm na listu 48 x 63 cm
- Type:
- model:map, cartographic, and IMAGE
- Subject:
- udc:913(4), Konspekt:7, udc:912, udc:913(437.6), udc:912.43, udc:(084.3), Konspekt:Geografie Evropy, reálie, cestování, Konspekt:Mapy. Atlasy. Glóby, and czenas:Šahy (Slovensko : oblast)
- Language:
- Czech, Slovak, and Hungarian
- Description:
- 4762, Legenda, Edice dle kladu listů, and (Language) Místní názvy slovensky a maďarsky
- Rights:
- http://creativecommons.org/publicdomain/mark/1.0/ and policy:public
118. Senec
- Publisher:
- Vojenský zeměpisný ústav
- Format:
- map and 1 mapa : barevná ; 39 x 51 cm na listu 47 x 63 cm
- Type:
- model:map, cartographic, and IMAGE
- Subject:
- udc:913(4), Konspekt:7, udc:912, udc:913(437.6), udc:912.43, udc:(084.3), Konspekt:Geografie Evropy, reálie, cestování, Konspekt:Mapy. Atlasy. Glóby, and czenas:Senec (Slovensko : oblast)
- Language:
- Czech, Slovak, and Hungarian
- Description:
- 4759, Legenda, Edice dle kladu listů, and (Language) Místní názvy slovensky a maďarsky
- Rights:
- http://creativecommons.org/publicdomain/mark/1.0/ and policy:public
119. Sklabiná
- Publisher:
- Vojenský zeměpisný ústav
- Format:
- map and 1 mapa : barevná ; 39 x 51 cm na listu 48 x 63 cm
- Type:
- model:map, cartographic, and IMAGE
- Subject:
- udc:913(4), Konspekt:7, udc:912, udc:913(437.6), udc:912.43, udc:(084.3), Konspekt:Geografie Evropy, reálie, cestování, Konspekt:Mapy. Atlasy. Glóby, and czenas:Sklabiná (Slovensko : oblast)
- Language:
- Czech, Slovak, and Hungarian
- Description:
- 4763, Legenda, Edice dle kladu listů, and (Language) Místní názvy slovensky a maďarsky
- Rights:
- http://creativecommons.org/publicdomain/mark/1.0/ and policy:public
120. Šurany
- Publisher:
- Vojenský zeměpisný ústav
- Format:
- map and 1 mapa : barevná ; 39 x 51 cm na listu 48 x 62 cm
- Type:
- model:map, cartographic, and IMAGE
- Subject:
- udc:913(437), Konspekt:7, udc:912, udc:(437.6), udc:912.43, udc:(084.3), Konspekt:Geografie Česka a Slovenska, reálie, cestování, Konspekt:Mapy. Atlasy. Glóby, and czenas:Šurany (Slovensko : oblast)
- Language:
- Czech, Slovak, and Hungarian
- Description:
- 4760, Legenda, Edice dle kladu listů, and (Language) Místní názvy slovensky a maďarsky
- Rights:
- http://creativecommons.org/publicdomain/mark/1.0/ and policy:public