This is a new version of the repository. Do let us know (lindat-help at ufal.mff.cuni.cz) if you encounter any issues.

Many Czech References for 50 Sentences Selected from WMT11 Data

Please use the following text to cite this item or export to a predefined format:
Bojar, Ondřej; Macháček, Matouš; Tamchyna, Aleš and Zeman, Daniel, 2013, Many Czech References for 50 Sentences Selected from WMT11 Data, LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL), http://hdl.handle.net/11858/00-097C-0000-0023-10B2-F.
Date issued
2013-09-01
Size
15431447 sentences
Language(s)
Description
This dataset contains the whole set of very many Czech translations for 50 English source sentences coming from WMT11 test set (http://www.statmt.org/wmt11). In total, there are 15431447 Czech sentences, i.e. 300k reference translations per source English sentence on average, but the exact number greatly varies across sentences. You can find more details in included README file. If you use this dataset, please cite the following paper which describes the technique used to construct the Czech translations: Bojar Ondřej, Macháček Matouš, Tamchyna Aleš, Zeman Daniel: Scratching the Surface of Possible Translations. Lecture Notes in Computer Science, Vol. 8082, Text, Speech and Dialogue: 16th International Conference, TSD 2013. Proceedings, Copyright © Springer Verlag, Berlin / Heidelberg, ISBN 978-3-642-40584-6, ISSN 0302-9743, pp. 465-474, 2013, DOI: 10.1007/978-3-642-40585-3_59
Acknowledgement
This item isPublicly Available
and licensed under:

Files in this item

Name
many-czech-references.zip
Size
116.86 MB
Format
application/zip
Description
zip archive containing many references, english source sentences, official wmt11 czech reference translations and README
MD5
f26e17e4d25333948cea1fa0cbb79e96
Preview
  File Preview
Name
README
Size
1.78 KB
Format
application/octet-stream
Description
copy of the README file from the main archive
MD5
c46fd2011b3dac0df3bee31044b02914
Preview
  File Preview