Contributor: Ministerstvo školství, mládeže a tělovýchovy České republiky@@MSM 0021620823@@Český národní korpus a korpusy dalších jazyků@@nationalFunds@@

Start Over Contributor Ministerstvo školství, mládeže a tělovýchovy České republiky@@MSM 0021620823@@Český národní korpus a korpusy dalších jazyků@@nationalFunds@@

1. ORAL2006: Corpus of informal spoken Czech

Creator:: Kopřivová, Marie and Waclawičová, Martina
Publisher:: Faculty of Arts, Institute of the Czech National Corpus, Charles University in Prague
Type:: text and corpus
Subject:: corpus and informal spoken language
Language:: Czech
Description:: Corpus of informal spoken Czech sized 1 MW. It contains transcriptions of 221 recordings made in 2002–2006 in the whole of Bohemia. All the recordings were made in informal situations to ensure prototypically spontaneous spoken language. This means private environment, physical presence of speakers who know each other, unscripted speech and topic not given in advance. The total number of speakers is 754, the metadata include sociolinguistic information about them. The corpus is provided in a (semi-XML) vertical format used as an input to the Manatee query engine. The data thus exactly correspond to the corpus available via query interface to registered users of the CNC. and Výzkumný záměr MSM0021620823 – Český národní korpus a korpusy dalších jazyků
Rights:: Attribution-NonCommercial-ShareAlike 3.0 Unported (CC BY-NC-SA 3.0), http://creativecommons.org/licenses/by-nc-sa/3.0/, and PUB

2. ORAL2008: Balanced corpus of informal spoken Czech

Creator:: Waclawičová, Martina, Kopřivová, Marie, Křen, Michal, and Válková, Lucie
Publisher:: Faculty of Arts, Institute of the Czech National Corpus, Charles University in Prague
Type:: text and corpus
Subject:: informal spoken language and balanced corpus
Language:: Czech
Description:: Balanced corpus of informal spoken Czech sized 1 MW. It contains transcriptions of 297 recordings made in 2002–2007 in the whole of Bohemia. All the recordings were made in informal situations to ensure prototypically spontaneous spoken language. This means private environment, physical presence of speakers who know each other, unscripted speech and topic not given in advance. The total number of speakers is 995, the corpus is balanced in their main sociolinguistic categories (gender, age group, education, region of childhood residence). The corpus is provided in a (semi-XML) vertical format used as an input to the Manatee query engine. The data thus exactly correspond to the corpus available via query interface to registered users of the CNC. and MSM0021620823 – Český národní korpus a korpusy dalších jazyků
Rights:: Attribution-NonCommercial-ShareAlike 3.0 Unported (CC BY-NC-SA 3.0), http://creativecommons.org/licenses/by-nc-sa/3.0/, and PUB

3. SYN2005: balanced corpus of written Czech

Creator:: Čermák, František, Hlaváčová, Jaroslava, Hnátková, Milena, Jelínek, Tomáš, Kocek, Jan, Kopřivová, Marie, Křen, Michal, Novotná, Renata, Petkevič, Vladimír, Schmiedtová, Věra, Skoumalová, Hana, Spoustová, Johanka, Šulc, Michal, and Velíšek, Zdeněk
Publisher:: Faculty of Arts, Institute of the Czech National Corpus, Charles University in Prague
Type:: text and corpus
Subject:: balanced corpus and written language
Language:: Czech
Description:: Balanced corpus of contemporary written Czech sized 100 MW. It was created as a representation of written language from 2000–2004 and thus it contains a wide range of text types and genres (fiction, professional literature, newspapers etc.) in balanced proportions. The corpus is lemmatized and morphologically tagged by a combination of stochastic and rule-based methods. The corpus is provided in a (semi-XML) vertical format used as an input to the Manatee query engine. The data thus correspond to the corpus available via query interface to registered users of the CNC with one important exception: they are shuffled, i.e. divided into blocks sized max. 100 words (respecting the sentence boundaries) whose ordering was randomized within the given document. and MSM0021620823 – Český národní korpus a korpusy dalších jazyků
Rights:: Czech National Corpus (Shuffled Corpus Data), https://lindat.mff.cuni.cz/repository/xmlui/page/license-cnc, and ACA

4. SYN2006PUB: corpus of Czech newspapers

Creator:: Čermák, František, Hlaváčová, Jaroslava, Hnátková, Milena, Jelínek, Tomáš, Kocek, Jan, Kopřivová, Marie, Křen, Michal, Novotná, Renata, Petkevič, Vladimír, Schmiedtová, Věra, Skoumalová, Hana, Spoustová, Johanka, Šulc, Michal, and Velíšek, Zdeněk
Publisher:: Faculty of Arts, Institute of the Czech National Corpus, Charles University in Prague
Type:: text and corpus
Subject:: corpus and written language
Language:: Czech
Description:: Corpus of contemporary Czech newspapers and magazines sized 300 MW. It contains various titles published between the end of 1989 and 2004. The corpus is lemmatized and morphologically tagged by a combination of stochastic and rule-based methods. The corpus is provided in a (semi-XML) vertical format used as an input to the Manatee query engine. The data thus correspond to the corpus available via query interface to registered users of the CNC with one important exception: they are shuffled, i.e. divided into blocks sized max. 100 words (respecting the sentence boundaries) whose ordering was randomized within the given document. and MSM0021620823 – Český národní korpus a korpusy dalších jazyků
Rights:: Czech National Corpus (Shuffled Corpus Data), https://lindat.mff.cuni.cz/repository/xmlui/page/license-cnc, and ACA

5. SYN2009PUB: corpus of Czech newspapers

Creator:: Křen, Michal, Bartoň, Tomáš, Hnátková, Milena, Jelínek, Tomáš, Petkevič, Vladimír, Procházka, Pavel, and Skoumalová, Hana
Publisher:: Faculty of Arts, Institute of the Czech National Corpus, Charles University in Prague
Type:: text and corpus
Subject:: corpus and written language
Language:: Czech
Description:: Corpus of contemporary Czech newspapers and magazines sized 700 MW. It contains various titles published between 1995–2007. The corpus is lemmatized and morphologically tagged by a combination of stochastic and rule-based methods. The corpus is provided in a (semi-XML) vertical format used as an input to the Manatee query engine. The data thus correspond to the corpus available via query interface to registered users of the CNC with one important exception: they are shuffled, i.e. divided into blocks sized max. 100 words (respecting the sentence boundaries) whose ordering was randomized within the given document. and MSM0021620823 – Český národní korpus a korpusy dalších jazyků
Rights:: Czech National Corpus (Shuffled Corpus Data), https://lindat.mff.cuni.cz/repository/xmlui/page/license-cnc, and ACA

6. SYN2010: balanced corpus of written Czech

Creator:: Křen, Michal, Bartoň, Tomáš, Cvrček, Václav, Hnátková, Milena, Jelínek, Tomáš, Kocek, Jan, Novotná, Renata, Petkevič, Vladimír, Procházka, Pavel, Schmiedtová, Věra, and Skoumalová, Hana
Publisher:: Faculty of Arts, Institute of the Czech National Corpus, Charles University in Prague
Type:: text and corpus
Subject:: balanced corpus and written language
Language:: Czech
Description:: Balanced corpus of contemporary written Czech sized 100 MW. It was created as a representation of written language from 2005–2009 and thus it contains a wide range of text types and genres (fiction, professional literature, newspapers etc.) in balanced proportions. The corpus is lemmatized and morphologically tagged by a combination of stochastic and rule-based methods. The corpus is provided in a (semi-XML) vertical format used as an input to the Manatee query engine. The data thus correspond to the corpus available via query interface to registered users of the CNC with one important exception: they are shuffled, i.e. divided into blocks sized max. 100 words (respecting the sentence boundaries) whose ordering was randomized within the given document. and MSM0021620823 – Český národní korpus a korpusy dalších jazyků
Rights:: Czech National Corpus (Shuffled Corpus Data), https://lindat.mff.cuni.cz/repository/xmlui/page/license-cnc, and ACA

1. ORAL2006: Corpus of informal spoken Czech

2. ORAL2008: Balanced corpus of informal spoken Czech

3. SYN2005: balanced corpus of written Czech

4. SYN2006PUB: corpus of Czech newspapers

5. SYN2009PUB: corpus of Czech newspapers

6. SYN2010: balanced corpus of written Czech

Limit your search

Show values starting with

Search

Search Constraints

Search Results

Limit your search

Contributor

Creator

Show values starting with

Language

Publisher

Rights

Subject

Type

Date

Original context has metadata only

Harvested from