This is a new version of the repository. Do let us know (lindat-help at ufal.mff.cuni.cz) if you encounter any issues.
Please use the following text to cite this item or export to a predefined format:
Šebesta, Karel; et al., 2017, CzeSL Grammatical Error Correction Dataset (CzeSL-GEC), LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL), http://hdl.handle.net/11234/1-2143.
dc.contributor.authorŠebesta, Karel
dc.contributor.authorBedřichová, Zuzanna
dc.contributor.authorŠormová, Kateřina
dc.contributor.authorŠtindlová, Barbora
dc.contributor.authorHrdlička, Milan
dc.contributor.authorHrdličková, Tereza
dc.contributor.authorHana, Jiří
dc.contributor.authorPetkevič, Vladimír
dc.contributor.authorJelínek, Tomáš
dc.contributor.authorŠkodová, Svatava
dc.contributor.authorJaneš, Petr
dc.contributor.authorLundáková, Kateřina
dc.contributor.authorSkoumalová, Hana
dc.contributor.authorSládek, Šimon
dc.contributor.authorPierscieniak, Piotr
dc.contributor.authorToufarová, Dagmar
dc.contributor.authorStraka, Milan
dc.contributor.authorRosen, Alexandr
dc.contributor.authorNáplava, Jakub
dc.contributor.authorPoláčková, Marie
dc.date.accessioned2017-05-03T08:08:33Z
dc.date.available2017-05-03T08:08:33Z
dc.date.issued2017-04-30
dc.descriptionCzeSL-GEC is a corpus containing sentence pairs of original and corrected versions of Czech sentences collected from essays written by both non-native learners of Czech and Czech pupils with Romani background. To create this corpus, unreleased CzeSL-man corpus (http://utkl.ff.cuni.cz/learncorp/) was utilized. All sentences in the corpus are word tokenized.
dc.identifier.urihttp://hdl.handle.net/11234/1-2143
dc.language.isoces
dc.publisherCharles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
dc.relation.isreplacedbyhttp://hdl.handle.net/11234/1-3057
dc.rightsAttribution-ShareAlike 3.0 Unported (CC BY-SA 3.0)
dc.rights.labelPUB
dc.rights.urihttp://creativecommons.org/licenses/by-sa/3.0/
dc.subjectnatural language correction
dc.subjectgrammatical error correction
dc.titleCzeSL Grammatical Error Correction Dataset (CzeSL-GEC)
dc.typecorpus
local.brandingLINDAT / CLARIAH-CZ
local.contact.personMilan Straka straka@ufal.mff.cuni.cz Charles University, UFAL
local.contact.personJakub Náplava naplava@ufal.mff.cuni.cz Charles University, UFAL
local.files.count1
local.files.size5326473
local.has.filesyes
local.language.nameCzech
local.size.info108067 sentences
local.size.info48 files
local.sponsornationalFunds LM2015071 Ministerstvo školství, mládeže a tělovýchovy České republiky LINDAT/CLARIN: Institut pro analýzu, zpracování a distribuci lingvistických dat
local.sponsornationalFunds GAČR 16-10185S Grantová agentura České republiky Čeština nerodilých mluvčích z pohledu teoretického a komputačního / Non-native Czech from the Theoretical and Computational Perspective
metashare.ResourceInfo#ContentInfo.mediaTypetext
This item isPublicly Available
and licensed under:

Files in this item

Name
2017-czesl-gec.zip
Size
5.08 MB
Format
application/zip
Description
corpus data and metadata, zipped
MD5
49dba121e7bf8deb180e673693410cc9
Preview
  File Preview
  • word2simword
    • a1_targets_train.txt524 kB
    • a1_targets_test.txt35 kB
    • a2_targets_train.txt326 kB
    • a1_targets_dev.txt33 kB
    • a2_inputs_train.txt324 kB
    • a1_inputs_test.txt34 kB
    • a2_targets_dev.txt33 kB
    • a1_inputs_dev.txt32 kB
    • a2_inputs_test.txt34 kB
    • a2_inputs_dev.txt32 kB
    • a1_inputs_train.txt521 kB
    • a2_targets_test.txt35 kB
  • word2words
    • a1_targets_train.txt1 MB
    • a1_targets_test.txt73 kB
    • a2_targets_train.txt667 kB
    • a1_targets_dev.txt70 kB
    • a2_inputs_train.txt647 kB
    • a1_inputs_test.txt71 kB
    • a2_targets_dev.txt70 kB
    • a1_inputs_dev.txt69 kB
    • a2_inputs_test.txt71 kB
    • a2_inputs_dev.txt69 kB
    • a1_inputs_train.txt1 MB
    • a2_targets_test.txt73 kB
  • word2word
    • a1_targets_train.txt598 kB
    • a1_targets_test.txt37 kB
    • a2_targets_train.txt368 kB
    • a1_targets_dev.txt38 kB
    • a2_inputs_train.txt366 kB
    • a1_inputs_test.txt37 kB
    • a2_targets_dev.txt38 kB
    • a1_inputs_dev.txt38 kB
    • a2_inputs_test.txt37 kB
    • a2_inputs_dev.txt38 kB
    • a1_inputs_train.txt593 kB
    • a2_targets_test.txt37 kB
  • sent2sent
    • a1_targets_train.txt1 MB
    • a1_targets_test.txt79 kB
    • a2_targets_train.txt653 kB
    • a1_targets_dev.txt71 kB
    • a2_inputs_train.txt638 kB
    • a1_inputs_test.txt78 kB
    • a2_targets_dev.txt71 kB
    • a1_inputs_dev.txt70 kB
    • a2_inputs_test.txt78 kB
    • a2_inputs_dev.txt70 kB
    • a1_inputs_train.txt1 MB
    • a2_targets_test.txt79 kB
    • README.md2 kB
    • LICENSE.txt21 kB