Skip to search
Skip to main content
Skip to first result
Search
Search Results
Publisher:
Universität Bamberg, World Language Documentation Centre
Format:
application/octet-stream
Type:
lexicalConceptualResource
Language:
Afrikaans , Arabic , Basque , Bulgarian , Catalan , Chinese , Czech , Danish , Dutch , English , Esperanto , Estonian , Finnish , French , Galician , Georgian , Modern Greek (1453-) , Hebrew , Hungarian , Icelandic , Indonesian , Interlingua (International Auxiliary Language Association) , Irish , Italian , Japanese , Khmer , Norwegian , Polish , Portuguese , Romanian , Russian , Serbian , Slovak , Spanish , Swedish , Turkish , Ukrainian , and Welsh
Rights:
GFDL or CC and http://www.omegawiki.org/Licensing
Creator:
Jan Patočka
Publisher:
Die Zeit (1977), č. 14 z 25. 3. 1977, str. 8 n. Int. něm. [Písemné odpovědi J. Patočky z 9. března 1977 na otázky Eduarda Neumaiera, dopisovatele týdeníku Die Zeit.]
Type:
Text
Subject:
1977 , 1977/4 , 1979/25 , 1980/27 , 1988/19 , 2006/11 , AS/PD-6 , cs , de , Int. něm. , no , and SS-12/Češi-I
Language:
German , Czech , and Norwegian
Rights:
open access and Rights holder: Archiv Jana Patočky, z.s.
Creator:
Henrik Ibsen and Viktor Šuman
Publisher:
Sfinx
Format:
print , text , regular print , and 121 s. ; 20 cm
Type:
model:monograph and TEXT
Subject:
Germánské literatury , 821.113.5-2 , (0:82-21) , 25 , and 821.11
Language:
Czech and Norwegian
Description:
Henrik Ibsen ; [přeložil Viktor Šuman] and Přeloženo z norštiny
Rights:
http://creativecommons.org/publicdomain/mark/1.0/ and policy:public
Creator:
Rosa, Rudolf
Publisher:
Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:
text and corpus
Subject:
Wikipedia , text corpora , and monolingual corpus
Language:
Abkhazian , Achinese , Adyghe , Afrikaans , Akan , Tosk Albanian , Amharic , Old English (ca. 450-1100) , Arabic , Official Aramaic (700-300 BCE) , Aragonese , Egyptian Arabic , Assamese , Asturian , Atikamekw , Avaric , Aymara , South Azerbaijani , Azerbaijani , Bashkir , Bambara , Bavarian , Central Bikol , Belarusian , Bengali , Bislama , Banjar , Tibetan , Bosnian , Bishnupriya , Breton , Buginese , Bulgarian , Russia Buriat , Catalan , Min Dong Chinese , Cebuano , Czech , Chamorro , Chechen , Cherokee , Church Slavic , Chuvash , Cheyenne , Central Kurdish , Cornish , Corsican , Cree , Crimean Tatar , Kashubian , Welsh , Danish , German , Dinka , Dimli (individual language) , Dhivehi , Lower Sorbian , Dzongkha , Modern Greek (1453-) , English , Esperanto , Estonian , Basque , Ewe , Extremaduran , Faroese , Persian , Fijian , Finnish , French , Arpitan , Northern Frisian , Western Frisian , Fulah , Friulian , Gagauz , Gan Chinese , Scottish Gaelic , Irish , Galician , Gilaki , Manx , Goan Konkani , Gothic , Guarani , Gujarati , Hakka Chinese , Haitian , Hausa , Hawaiian , Serbo-Croatian , Hebrew , Herero , Fiji Hindi , Hindi , Hiri Motu , Croatian , Upper Sorbian , Hungarian , Armenian , Igbo , Ido , Inuktitut , Interlingue , Iloko , Interlingua (International Auxiliary Language Association) , Indonesian , Inupiaq , Icelandic , Italian , Jamaican Creole English , Javanese , Lojban , Japanese , Kara-Kalpak , Kabyle , Kalaallisut , Kannada , Kashmiri , Georgian , Kanuri , Kazakh , Kabardian , Kabiyè , Khmer , Kikuyu , Kinyarwanda , Kirghiz , Komi-Permyak , Komi , Kongo , Korean , Karachay-Balkar , Kölsch , Kurdish , Ladino , Lao , Latin , Latvian , Lak , Lezghian , Ligurian , Limburgan , Lingala , Lithuanian , Lombard , Northern Luri , Latgalian , Luxembourgish , Ganda , Literary Chinese , Marshallese , Maithili , Malayalam , Marathi , Moksha , Eastern Mari , Minangkabau , Macedonian , Malagasy , Maltese , Mongolian , Maori , Western Mari , Malay (macrolanguage) , Creek , Mirandese , Burmese , Erzya , Mazanderani , Min Nan Chinese , Neapolitan , Nauru , Navajo , Ndonga , Low German , Nepali (macrolanguage) , Newari , Dutch , Norwegian Nynorsk , Norwegian , Novial , Pedi , Nyanja , Occitan (post 1500) , Livvi , Oriya (macrolanguage) , Oromo , Ossetian , Pangasinan , Pampanga , Panjabi , Papiamento , Picard , Pennsylvania German , Pfaelzisch , Pitcairn-Norfolk , Pali , Piemontese , Western Panjabi , Pontic , Polish , Portuguese , Pushto , Quechua , Vlax Romani , Romansh , Romanian , Rusyn , Rundi , Macedo-Romanian , Russian , Sango , Yakut , Sanskrit , Sicilian , Scots , Samogitian , Sinhala , Slovak , Slovenian , Northern Sami , Samoan , Shona , Sindhi , Somali , Southern Sotho , Spanish , Albanian , Sardinian , Sranan Tongo , Serbian , Swati , Saterfriesisch , Sundanese , Swahili (macrolanguage) , Swedish , Silesian , Tahitian , Tamil , Tatar , Tulu , Telugu , Tama (Colombia) , Tetum , Tajik , Tagalog , Thai , Tigrinya , Tonga (Tonga Islands) , Tok Pisin , Tswana , Tsonga , Turkmen , Tumbuka , Turkish , Twi , Tuvinian , Udmurt , Uighur , Ukrainian , Urdu , Uzbek , Venetian , Venda , Veps , Vietnamese , Vlaams , Volapük , Võro , Waray (Philippines) , Walloon , Wolof , Wu Chinese , Kalmyk , Xhosa , Mingrelian , Yiddish , Yoruba , Yue Chinese , Zeeuws , Zhuang , Chinese , Zulu , and Dotyali
Description:
Wikipedia plain text data obtained from Wikipedia dumps with WikiExtractor in February 2018.
The data come from all Wikipedias for which dumps could be downloaded at [https://dumps.wikimedia.org/]. This amounts to 297 Wikipedias, usually corresponding to individual languages and identified by their ISO codes. Several special Wikipedias are included, most notably "simple" (Simple English Wikipedia) and "incubator" (tiny hatching Wikipedias in various languages).
For a list of all the Wikipedias, see [https://meta.wikimedia.org/wiki/List_of_Wikipedias].
The script which can be used to get new version of the data is included, but note that Wikipedia limits the download speed for downloading a lot of the dumps, so it takes a few days to download all of them (but one or a few can be downloaded fast).
Also, the format of the dumps changes time to time, so the script will probably eventually stop working one day.
The WikiExtractor tool [http://medialab.di.unipi.it/wiki/Wikipedia_Extractor] used to extract text from the Wikipedia dumps is not mine, I only modified it slightly to produce plaintext outputs [https://github.com/ptakopysk/wikiextractor].
Rights:
Attribution-ShareAlike 3.0 Unported (CC BY-SA 3.0) , http://creativecommons.org/licenses/by-sa/3.0/ , and PUB
Creator:
Jan Patočka
Publisher:
Str. 46–88. Stať.
Type:
Text
Subject:
1975 , 1979/25 , 1981/6 , 1981/7 , 1988/28 , 1988/31 , 1988/32 , 1988/34 , 1994/7 , 1996/4 , 1996/7 , 1998/3 , 1999/8 , 2 , 2001/9 , 2002/21 , 2002/6 , 2006/1 , 2007/1 , 2008/3 , bg , cs , de , en , es , fr , fulltext , hu , it , lt , no , pl , ru , SS-3/PD-III , sv , uk , and v
Language:
Czech , English , Bulgarian , French , Italian , Lithuanian , Hungarian , German , Norwegian , Polish , Russian , Spanish , Swedish , and Ukrainian
Rights:
open access and Rights holder: Archiv Jana Patočky, z.s.
Creator:
Jan Patočka
Publisher:
Str. 1–45. Stať.
Type:
Text
Subject:
1975 , 1979/25 , 1981/6 , 1981/7 , 1988/28 , 1988/31 , 1988/32 , 1988/34 , 1994/7 , 1996/4 , 1996/7 , 1998/3 , 1999/8 , 2001/9 , 2002/1 , 2002/21 , 2002/5 , 2006/1 , 2007/1 , 2008/3 , bg , cs , de , en , es , fr , fulltext , hu , it , lt , no , pl , ru , SS-3/PD-III , sv , and uk
Language:
Czech , English , Bulgarian , French , Italian , Lithuanian , Hungarian , German , Norwegian , Polish , Russian , Spanish , Swedish , and Ukrainian
Rights:
open access and Rights holder: Archiv Jana Patočky, z.s.
Creator:
Jan Patočka
Publisher:
Str. 20 n. Misc.
Type:
Text
Subject:
1977 , 1979/25 , 1988/19 , 2006/11 , AS/PD-6 , cs , no , O , and SS-12/Češi-I
Language:
Czech and Norwegian
Rights:
open access and Rights holder: Archiv Jana Patočky, z.s.
Publisher:
Department of Literature, Area Studies and European Languages, University of Oslo and Department of Linguistics and Nordic Studies, University of Oslo
Type:
corpus
Language:
English , Norwegian , and Russian
Description:
The RuN corpus is a parallel corpus consisting of Norwegian, Russian and English texts. The texts are aligned at the sentence level and have been tagged for grammatical information at the word level.
Rights:
Not specified
Creator:
Bull-Tornöe, J. , Hertzberg, Ebbe , Kraus, Arnošt Vilém , and Ústřední jednota českých hospodářských společenstev v Království českém
Publisher:
Družstvo zemědělců
Format:
print , text , regular print , and 162 s.
Type:
model:monograph and TEXT
Subject:
334.73 , 336.773 , 631.115.8 , 334 , 4 , zemědělství , zemědělská výroba , mezinárodní spolupráce , zemědělské družstevnictví , zemědělská družstva , výrobní družstva , obchodní družstva , družstevní záložny , financování , and Formy organizace a spolupráce v ekonomice
Language:
Czech and Norwegian
Description:
norsky napsali J. Bull-Tornöe a Ebbe Hertzberg ; se svolením spisovatelů přeložil Arnošt Kraus, Vydavatel: Ústřední jednota českých hospodářských společenstev v Království českém, and Converted from MODS 3.5 to DC version 1.8 (EE patch 2015/06/25)
Rights:
http://creativecommons.org/publicdomain/mark/1.0/ and policy:public
Creator:
Lundkvist, Sven,
Type:
text and studie
Subject:
Vojenství. Obrana země. Ozbrojené síly , válka třicetiletá (1618-1648) , bitva u Breitenfeldu (1631) , vojenské operace, války, bitvy , and světové dějiny 1492-1648
Language:
Norwegian
Rights:
unknown