1 - 8 of 8
Number of results to display per page
Search Results
2. ORAL2008: Balanced corpus of informal spoken Czech
- Creator:
- Waclawičová, Martina, Kopřivová, Marie, Křen, Michal, and Válková, Lucie
- Publisher:
- Faculty of Arts, Institute of the Czech National Corpus, Charles University in Prague
- Type:
- text and corpus
- Subject:
- informal spoken language and balanced corpus
- Language:
- Czech
- Description:
- Balanced corpus of informal spoken Czech sized 1 MW. It contains transcriptions of 297 recordings made in 2002–2007 in the whole of Bohemia. All the recordings were made in informal situations to ensure prototypically spontaneous spoken language. This means private environment, physical presence of speakers who know each other, unscripted speech and topic not given in advance. The total number of speakers is 995, the corpus is balanced in their main sociolinguistic categories (gender, age group, education, region of childhood residence). The corpus is provided in a (semi-XML) vertical format used as an input to the Manatee query engine. The data thus exactly correspond to the corpus available via query interface to registered users of the CNC. and MSM0021620823 – Český národní korpus a korpusy dalších jazyků
- Rights:
- Attribution-NonCommercial-ShareAlike 3.0 Unported (CC BY-NC-SA 3.0), http://creativecommons.org/licenses/by-nc-sa/3.0/, and PUB
3. ORAL2013: balanced corpus of informal spoken Czech (transcriptions & audio)
- Creator:
- Benešová, Lucie, Křen, Michal, and Waclawičová, Martina
- Publisher:
- Charles University, Faculty of Arts, Institute of the Czech National Corpus
- Type:
- audio and corpus
- Subject:
- balanced corpus, spoken language, and speech corpus
- Language:
- Czech
- Description:
- ORAL2013 is designed as a representation of authentic spoken Czech used in informal situations (private environment, spontaneity, unpreparedness etc.) in the area of the whole Czech Republic. The corpus comprises 835 recordings from 2008–2011 that contain 2 785 189 words (i.e. 3 285 508 tokens including punctuation) uttered by 2 544 speakers, out of which 1 297 speakers are unique. ORAL2013 is balanced in the main sociolinguistic categories of the speakers (gender, age group, education, region of childhood residence). The (anonymized) transcriptions are provided in the Transcriber XML format, audio (with corresponding anonymization beeps) is in uncompressed 16-bit PCM WAV, mono, 16 kHz format. Another format option of the transcriptions is also available under less restrictive CC BY-NC-SA license at http://hdl.handle.net/11234/1-1847
- Rights:
- License Agreement for Czech National Corpus Data, https://lindat.mff.cuni.cz/repository/xmlui/page/license-cnc-data, and ACA
4. ORAL2013: balanced corpus of informal spoken Czech (transcriptions)
- Creator:
- Benešová, Lucie, Křen, Michal, and Waclawičová, Martina
- Publisher:
- Charles University, Faculty of Arts, Institute of the Czech National Corpus
- Type:
- text and corpus
- Subject:
- balanced corpus and spoken language
- Language:
- Czech
- Description:
- ORAL2013 is designed as a representation of authentic spoken Czech used in informal situations (private environment, spontaneity, unpreparedness etc.) in the area of the whole Czech Republic. The corpus comprises 835 recordings from 2008–2011 that contain 2 785 189 words (i.e. 3 285 508 tokens including punctuation) uttered by 2 544 speakers, out of which 1 297 speakers are unique. ORAL2013 is balanced in the main sociolinguistic categories of speakers (gender, age group, education, region of childhood residence). The corpus is provided in a (semi-XML) vertical format used as an input to the Manatee query engine. The data thus correspond to the corpus available via the KonText query engine to registered users of the CNC at http://www.korpus.cz Please note: this item includes only the transcriptions, audio is available under more restrictive non-CC license at http://hdl.handle.net/11234/1-1848
- Rights:
- Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0), http://creativecommons.org/licenses/by-nc-sa/4.0/, and PUB
5. ORTOFON v1: balanced corpus of informal spoken Czech with multi-tier transcription (transcriptions & audio)
- Creator:
- Kopřivová, Marie, Komrsková, Zuzana, Lukeš, David, Poukarová, Petra, and Škarpová, Marie
- Publisher:
- Charles University, Faculty of Arts, Institute of the Czech National Corpus
- Type:
- audio and corpus
- Subject:
- balanced corpus, spoken language, informal language, and Czech
- Language:
- Czech
- Description:
- ORTOFON v1 is designed as a representation of authentic spoken Czech used in informal situations (private environment, spontaneity, unpreparedness etc.) in the area of the whole Czech Republic. The corpus is composed of 332 recordings from 2012–2017 and contains 1 014 786 orthographic words (i.e. a total of 1 236 508 tokens including punctuation); a total of 624 different speakers appear in the probes. ORTOFON v1 is fully balanced regarding the basic sociolinguistic speaker categories (gender, age group, level of education and region of childhood residence). The transcription is linked to the corresponding audio track. Unlike the ORAL-series corpora, the transcription was carried out on two main tiers, orthographic and phonetic, supplemented by an additional metalanguage tier. ORTOFON v1 is lemmatized and morphologically tagged. The (anonymized) transcriptions are provided in the XML Elan Annotation format, audio (with corresponding anonymization beeps) is in uncompressed 16-bit PCM WAV, mono, 16 kHz format. Another format option of the transcriptions is also available under less restrictive CC BY-NC-SA license at http://hdl.handle.net/11234/1-2580
- Rights:
- License Agreement for Czech National Corpus Data, https://lindat.mff.cuni.cz/repository/xmlui/page/license-cnc-data, and ACA
6. ORTOFON v1: balanced corpus of informal spoken Czech with multi-tier transcription (transcriptions)
- Creator:
- Kopřivová, Marie, Komrsková, Zuzana, Lukeš, David, Poukarová, Petra, and Škarpová, Marie
- Publisher:
- Charles University, Faculty of Arts, Institute of the Czech National Corpus
- Type:
- text and corpus
- Subject:
- balanced corpus, spoken language, informal language, and Czech
- Language:
- Czech
- Description:
- ORTOFON v1 is designed as a representation of authentic spoken Czech used in informal situations (private environment, spontaneity, unpreparedness etc.) in the area of the whole Czech Republic. The corpus is composed of 332 recordings from 2012–2017 and contains 1 014 786 orthographic words (i.e. a total of 1 236 508 tokens including punctuation); a total of 624 different speakers appear in the probes. ORTOFON v1 is fully balanced regarding the basic sociolinguistic speaker categories (gender, age group, level of education and region of childhood residence). The transcription is linked to the corresponding audio track. Unlike the ORAL-series corpora, the transcription was carried out on two main tiers, orthographic and phonetic, supplemented by an additional metalanguage tier. ORTOFON v1 is lemmatized and morphologically tagged. The (anonymized) corpus is provided in a (semi-XML) vertical format used as an input to the Manatee query engine. The data thus correspond to the corpus available via the KonText query engine to registered users of the CNC at http://www.korpus.cz Please note: this item includes only the transcriptions, audio (and the transcripts in their original format) is available under more restrictive non-CC license at http://hdl.handle.net/11234/1-2579
- Rights:
- Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0), http://creativecommons.org/licenses/by-nc-sa/4.0/, and PUB
7. SYN2005: balanced corpus of written Czech
- Creator:
- Čermák, František, Hlaváčová, Jaroslava, Hnátková, Milena, Jelínek, Tomáš, Kocek, Jan, Kopřivová, Marie, Křen, Michal, Novotná, Renata, Petkevič, Vladimír, Schmiedtová, Věra, Skoumalová, Hana, Spoustová, Johanka, Šulc, Michal, and Velíšek, Zdeněk
- Publisher:
- Faculty of Arts, Institute of the Czech National Corpus, Charles University in Prague
- Type:
- text and corpus
- Subject:
- balanced corpus and written language
- Language:
- Czech
- Description:
- Balanced corpus of contemporary written Czech sized 100 MW. It was created as a representation of written language from 2000–2004 and thus it contains a wide range of text types and genres (fiction, professional literature, newspapers etc.) in balanced proportions. The corpus is lemmatized and morphologically tagged by a combination of stochastic and rule-based methods. The corpus is provided in a (semi-XML) vertical format used as an input to the Manatee query engine. The data thus correspond to the corpus available via query interface to registered users of the CNC with one important exception: they are shuffled, i.e. divided into blocks sized max. 100 words (respecting the sentence boundaries) whose ordering was randomized within the given document. and MSM0021620823 – Český národní korpus a korpusy dalších jazyků
- Rights:
- Czech National Corpus (Shuffled Corpus Data), https://lindat.mff.cuni.cz/repository/xmlui/page/license-cnc, and ACA
8. SYN2010: balanced corpus of written Czech
- Creator:
- Křen, Michal, Bartoň, Tomáš, Cvrček, Václav, Hnátková, Milena, Jelínek, Tomáš, Kocek, Jan, Novotná, Renata, Petkevič, Vladimír, Procházka, Pavel, Schmiedtová, Věra, and Skoumalová, Hana
- Publisher:
- Faculty of Arts, Institute of the Czech National Corpus, Charles University in Prague
- Type:
- text and corpus
- Subject:
- balanced corpus and written language
- Language:
- Czech
- Description:
- Balanced corpus of contemporary written Czech sized 100 MW. It was created as a representation of written language from 2005–2009 and thus it contains a wide range of text types and genres (fiction, professional literature, newspapers etc.) in balanced proportions. The corpus is lemmatized and morphologically tagged by a combination of stochastic and rule-based methods. The corpus is provided in a (semi-XML) vertical format used as an input to the Manatee query engine. The data thus correspond to the corpus available via query interface to registered users of the CNC with one important exception: they are shuffled, i.e. divided into blocks sized max. 100 words (respecting the sentence boundaries) whose ordering was randomized within the given document. and MSM0021620823 – Český národní korpus a korpusy dalších jazyků
- Rights:
- Czech National Corpus (Shuffled Corpus Data), https://lindat.mff.cuni.cz/repository/xmlui/page/license-cnc, and ACA