|

Corpora

Below is the list of corpora in the TEITOK/Kontext hybrid set-up, hosted at ÚFAL. To get a larger list of TEITOK projects, see the TEITOK project page. A larger list of Kontext corpora at the UFAL institute can be found in the KonText corpus list, or in the repository. For corpora that have multiple versions in TEITOK, only the most recent version is displayed, but you can click on the version number to see all versions of the corpus. The corpora are listed by corpus type, a description of which can be found here


AcronymLatestToken sizeCorpus TypeCorpus StatusCorpus ContentCorpus Language(s)
infoCzechVerse13MSpecialized CorpuslivePoetryCzech
infoDeltaCorpus1.194MLRL CorpusstableMany
infoEHRI40kSpecialized corpusliveLettersGerman, Czech, English
infoHaCzech18kFacsimile CorpusstableHandwritten textsCzech
infoMaPCorpSpecialized CorpuslivePoetryMacedonian
infoMakoň2020-11-164.2MSpoken CorpusstableTranscribed talksCzech
infoMazon7.9kFacsimile CorpusliveLettersCzech, German, English, French, Russian
infoMigrant Stories400kSpecialized corpusliveMigrant storiesEnglish
infoMuNeCo840MLRL CorpusliveNewspaper articlesMany
infoOCRCZ27MFacsimile CorpusstablePrinted materialCzech
infoPDT-C1.03.9MTreebankstableCzech
infoParCzech4.036MSpoken CorpusstableParliamentary sessionsCzech
infoParlaMint4.11.4GSpecialized corpusstableParliamentary sessionsMany
infoSIR1.0250kSpecialized CorpusstableNewspaper articlesCzech
infoSkript 2015400kLearner CorpusliveCzech
infoUniversal Dependencies2.1432MTreebankstableMany

22 results - showing 1-22 - - click on a value to reduce selection - click on a column to sort - Search