Plaintext Wikipedia dump 2018
Please use the following text to cite this item or export to a predefined format:
Rosa, Rudolf, 2018,
Plaintext Wikipedia dump 2018, LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL),
http://hdl.handle.net/11234/1-2735.
Authors
Item identifier
Date issued
2018-02-25
Size
21 gb,
297 files,
8503209631 words
Language(s)
Achinese ,
Adyghe ,
Akan ,
Amharic ,
Arabic ,
Assamese ,
Asturian ,
Avaric ,
Aymara ,
Bashkir ,
Bambara ,
Bavarian ,
Bengali ,
Bislama ,
Banjar ,
Tibetan ,
Bosnian ,
Breton ,
Buginese ,
Catalan ,
Cebuano ,
Czech ,
Chamorro ,
Chechen ,
Cherokee ,
Chuvash ,
Cheyenne ,
Cornish ,
Corsican ,
Cree ,
Welsh ,
Danish ,
German ,
Dinka ,
Dhivehi ,
Dzongkha ,
English ,
Estonian ,
Basque ,
Ewe ,
Faroese ,
Persian ,
Fijian ,
Finnish ,
French ,
Arpitan ,
Fulah ,
Friulian ,
Gagauz ,
Irish ,
Galician ,
Gilaki ,
Manx ,
Gothic ,
Guarani ,
Gujarati ,
Haitian ,
Hausa ,
Hawaiian ,
Hebrew ,
Herero ,
Hindi ,
HiriMotu ,
Croatian ,
Armenian ,
Igbo ,
Ido ,
Iloko ,
Inupiaq ,
Italian ,
Javanese ,
Lojban ,
Japanese ,
Kabyle ,
Kannada ,
Kashmiri ,
Georgian ,
Kanuri ,
Kazakh ,
Kabiyè ,
Kikuyu ,
Kirghiz ,
Komi ,
Kongo ,
Korean ,
Kölsch ,
Kurdish ,
Ladino ,
Lao ,
Latin ,
Latvian ,
Lak ,
Lezghian ,
Ligurian ,
Lingala ,
Lombard ,
Ganda ,
Maithili ,
Marathi ,
Moksha ,
Malagasy ,
Maltese ,
Maori ,
Creek ,
Burmese ,
Erzya ,
Nauru ,
Navajo ,
Ndonga ,
Nepali ,
Newari ,
Dutch ,
Novial ,
Pedi ,
Nyanja ,
Livvi ,
Oriya ,
Oromo ,
Ossetian ,
Pampanga ,
Panjabi ,
Picard ,
Pali ,
Pontic ,
Polish ,
Pushto ,
Quechua ,
Romansh ,
Romanian ,
Rusyn ,
Rundi ,
Russian ,
Sango ,
Yakut ,
Sanskrit ,
Sicilian ,
Scots ,
Sinhala ,
Slovak ,
Samoan ,
Shona ,
Sindhi ,
Somali ,
Spanish ,
Albanian ,
Serbian ,
Swati ,
Swedish ,
Silesian ,
Tahitian ,
Tamil ,
Tatar ,
Tulu ,
Telugu ,
Tetum ,
Tajik ,
Tagalog ,
Thai ,
Tigrinya ,
TokPisin ,
Tswana ,
Tsonga ,
Turkmen ,
Tumbuka ,
Turkish ,
Twi ,
Tuvinian ,
Udmurt ,
Uighur ,
Urdu ,
Uzbek ,
Venetian ,
Venda ,
Veps ,
Vlaams ,
Volapük ,
Võro ,
Walloon ,
Wolof ,
Kalmyk ,
Xhosa ,
Yiddish ,
Yoruba ,
Zeeuws ,
Zhuang ,
Chinese ,
Zulu ,
Description
Wikipedia plain text data obtained from Wikipedia dumps with WikiExtractor in February 2018.
The data come from all Wikipedias for which dumps could be downloaded at [https://dumps.wikimedia.org/]. This amounts to 297 Wikipedias, usually corresponding to individual languages and identified by their ISO codes. Several special Wikipedias are included, most notably "simple" (Simple English Wikipedia) and "incubator" (tiny hatching Wikipedias in various languages).
For a list of all the Wikipedias, see [https://meta.wikimedia.org/wiki/List_of_Wikipedias].
The script which can be used to get new version of the data is included, but note that Wikipedia limits the download speed for downloading a lot of the dumps, so it takes a few days to download all of them (but one or a few can be downloaded fast).
Also, the format of the dumps changes time to time, so the script will probably eventually stop working one day.
The WikiExtractor tool [http://medialab.di.unipi.it/wiki/Wikipedia_Extractor] used to extract text from the Wikipedia dumps is not mine, I only modified it slightly to produce plaintext outputs [https://github.com/ptakopysk/wikiextractor].
Acknowledgement
Ministerstvo školství, mládeže a tělovýchovy České republiky
Project code:CZ.02.1.01/0.0/0.0/16_013/0001781
Project name:LINDAT/CLARIN - Výzkumná infrastruktura pro jazykové technologie - rozšíření repozitáře a výpočetní kapacity
Subject(s)
Collections
Files in this item
- Name
- get_wikiextractor.sh
- Size
- 67 B
- Format
- application/octet-stream
- Description
- script to download the WikiExtractor tool
- MD5
- e266afce1f727b055d333ec4fb0052c2

The file preview has not been generated yet. Please try again later or contact the system administrator lindat-help@ufal.mff.cuni.cz
- Name
- get_wiki_txt.sh
- Size
- 299 B
- Format
- application/octet-stream
- Description
- script to get text from Wikipedia for the given language code
- MD5
- 3e45b60fe7f36ad78154571a84347e35

The file preview has not been generated yet. Please try again later or contact the system administrator lindat-help@ufal.mff.cuni.cz


