This is a new version of the repository. Do let us know (lindat-help at ufal.mff.cuni.cz) if you encounter any issues.
Please use the following text to cite this item or export to a predefined format:
Kopřivová, Marie; Komrsková, Zuzana; Lukeš, David; Poukarová, Petra and Škarpová, Marie, 2017, ORTOFON v1: balanced corpus of informal spoken Czech with multi-tier transcription (transcriptions & audio), LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL), http://hdl.handle.net/11234/1-2579.
dc.contributor.authorKopřivová, Marie
dc.contributor.authorKomrsková, Zuzana
dc.contributor.authorLukeš, David
dc.contributor.authorPoukarová, Petra
dc.contributor.authorŠkarpová, Marie
dc.date.accessioned2018-01-02T12:24:15Z
dc.date.available2018-01-02T12:24:15Z
dc.date.issued2017-12-28
dc.descriptionORTOFON v1 is designed as a representation of authentic spoken Czech used in informal situations (private environment, spontaneity, unpreparedness etc.) in the area of the whole Czech Republic. The corpus is composed of 332 recordings from 2012–2017 and contains 1 014 786 orthographic words (i.e. a total of 1 236 508 tokens including punctuation) a total of 624 different speakers appear in the probes. ORTOFON v1 is fully balanced regarding the basic sociolinguistic speaker categories (gender, age group, level of education and region of childhood residence). The transcription is linked to the corresponding audio track. Unlike the ORAL-series corpora, the transcription was carried out on two main tiers, orthographic and phonetic, supplemented by an additional metalanguage tier. ORTOFON v1 is lemmatized and morphologically tagged. The (anonymized) transcriptions are provided in the XML Elan Annotation format, audio (with corresponding anonymization beeps) is in uncompressed 16-bit PCM WAV, mono, 16 kHz format. Another format option of the transcriptions is also available under less restrictive CC BY-NC-SA license at http://hdl.handle.net/11234/1-2580
dc.identifier.urihttp://hdl.handle.net/11234/1-2579
dc.language.isoces
dc.publisherCharles University, Faculty of Arts, Institute of the Czech National Corpus
dc.relation.isreplacedbyhttp://hdl.handle.net/11234/1-5686
dc.rightsLicense Agreement for Czech National Corpus Data
dc.rights.labelACA
dc.rights.urihttps://lindat.mff.cuni.cz/repository/static/license-cnc-data.html
dc.source.urihttp://wiki.korpus.cz/doku.php/en:cnk:ortofon
dc.subjectbalanced corpus
dc.subjectspoken language
dc.subjectinformal language
dc.subjectCzech
dc.titleORTOFON v1: balanced corpus of informal spoken Czech with multi-tier transcription (transcriptions & audio)
dc.typecorpus
local.brandingLINDAT / CLARIAH-CZ
local.contact.personDavid Lukeš david.lukes@ff.cuni.cz Charles University, Faculty of Arts, Institute of the Czech National Corpus
local.files.count1
local.files.size10309980762
local.has.filesyes
local.language.nameCzech
local.size.info1000000 words
local.sponsornationalFunds LM2015044 Ministerstvo školství, mládeže a tělovýchovy Český národní korpus
metashare.ResourceInfo#ContentInfo.mediaTypeaudio
This item isAcademic Use
and licensed under:

Files in this item

Name
ortofon_v1.tar.gz
Size
9.6 GB
Format
application/x-gzip
Description
Neznámý
MD5
51e4f525fe2a478e7f5f73ef7af61c42
Preview
  File Preview