dc.contributor.author | Kopřivová, Marie |
dc.contributor.author | Komrsková, Zuzana |
dc.contributor.author | Lukeš, David |
dc.contributor.author | Poukarová, Petra |
dc.contributor.author | Škarpová, Marie |
dc.date.accessioned | 2018-01-02T12:21:53Z |
dc.date.available | 2018-01-02T12:21:53Z |
dc.date.issued | 2017-12-28 |
dc.identifier.uri | http://hdl.handle.net/11234/1-2580 |
dc.description | ORTOFON v1 is designed as a representation of authentic spoken Czech used in informal situations (private environment, spontaneity, unpreparedness etc.) in the area of the whole Czech Republic. The corpus is composed of 332 recordings from 2012–2017 and contains 1 014 786 orthographic words (i.e. a total of 1 236 508 tokens including punctuation); a total of 624 different speakers appear in the probes. ORTOFON v1 is fully balanced regarding the basic sociolinguistic speaker categories (gender, age group, level of education and region of childhood residence). The transcription is linked to the corresponding audio track. Unlike the ORAL-series corpora, the transcription was carried out on two main tiers, orthographic and phonetic, supplemented by an additional metalanguage tier. ORTOFON v1 is lemmatized and morphologically tagged. The (anonymized) corpus is provided in a (semi-XML) vertical format used as an input to the Manatee query engine. The data thus correspond to the corpus available via the KonText query engine to registered users of the CNC at http://www.korpus.cz Please note: this item includes only the transcriptions, audio (and the transcripts in their original format) is available under more restrictive non-CC license at http://hdl.handle.net/11234/1-2579 |
dc.language.iso | ces |
dc.publisher | Charles University, Faculty of Arts, Institute of the Czech National Corpus |
dc.relation.isformatof | http://hdl.handle.net/11234/1-2579 |
dc.relation.isreplacedby | http://hdl.handle.net/11234/1-5687 |
dc.rights | Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) |
dc.rights.uri | http://creativecommons.org/licenses/by-nc-sa/4.0/ |
dc.source.uri | http://wiki.korpus.cz/doku.php/en:cnk:ortofon |
dc.subject | balanced corpus |
dc.subject | spoken language |
dc.subject | informal language |
dc.subject | Czech |
dc.title | ORTOFON v1: balanced corpus of informal spoken Czech with multi-tier transcription (transcriptions) |
dc.type | corpus |
metashare.ResourceInfo#ContentInfo.mediaType | text |
dc.rights.label | PUB |
has.files | yes |
branding | LINDAT / CLARIAH-CZ |
contact.person | David Lukeš david.lukes@ff.cuni.cz Charles University, Faculty of Arts, Institute of the Czech National Corpus |
sponsor | Ministerstvo školství, mládeže a tělovýchovy LM2015044 Český národní korpus nationalFunds |
size.info | 1000000 words |
files.size | 13525766 |
files.count | 1 |
Files in this item
This item is
Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0)
Publicly Available
and licensed under:Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0)
- Name
- ortofon_v1_vert.gz
- Size
- 12.9 MB
- Format
- application/x-gzip
- Description
- the data
- MD5
- 3a363ef604e7ea7fbb8af44d88009bc7