This is a new version of the repository. Do let us know (lindat-help at ufal.mff.cuni.cz) if you encounter any issues.

WMT16 Tuning Shared Task Models (English-to-Czech)

Please use the following text to cite this item or export to a predefined format:
Kamran, Amir; Jawaid, Bushra; Bojar, Ondřej and Stanojevic, Milos, 2016, WMT16 Tuning Shared Task Models (English-to-Czech), LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL), http://hdl.handle.net/11372/LRT-1672.
Date issued
2016-03-21
Language(s)
Description
This item contains models to tune for the WMT16 Tuning shared task for English-to-Czech. CzEng 1.6pre (http://ufal.mff.cuni.cz/czeng/czeng16pre) corpus is used for the training of the translation models. The data is tokenized (using Moses tokenizer), lowercased and sentences longer than 60 words and shorter than 4 words are removed before training. Alignment is done using fast_align (https://github.com/clab/fast_align) and the standard Moses pipeline is used for training. Two 5-gram language models are trained using KenLM: one only using the CzEng Czech data and the other is trained using all available Czech mono data for WMT except Common Crawl. Also included are two lexicalized bidirectional reordering models, word based and hierarchical, with msd conditioned on both source and target of processed CzEng.
Acknowledgement

Files in this item

Name
wmt16.mono.blm.cs.tgz
Size
18.98 GB
Format
application/x-gzip
Description
kenlm 5-gram language model (binarized) trained on all Czech mono data available for WMT except Common Crawl (see the Makefile for the details of mono data used)
MD5
6347aa8e420db60142cb1384ea1cab0d
Preview
  File Preview
    • wmt16.mono.blm.cs32 GB
Name
en2cs_model.tgz
Size
37.56 GB
Format
application/x-gzip
Description
Contains Lexical Models (lex.e2f and lex.f2e), Phrase Table (phrase-table.gz), word based reordering model (reordering-table.wbe-msd-bidirectional-fe.gz), hierarchical reordering model (reordering-table.hier-msd-bidirectional-fe.gz) and the moses.ini file
MD5
bed88bbeef3afc454c3f02845ab72769
Preview
  File Preview
    • reordering-table.wbe-msd-bidirectional-fe.gz7 GB
    • moses.ini1 kB
    • lex.f2e668 MB
    • reordering-table.hier-msd-bidirectional-fe.gz7 GB
    • lex.e2f668 MB
    • phrase-table.gz21 GB
Name
wmt16.czeng.blm.cs.tgz
Size
9.28 GB
Format
application/x-gzip
Description
kenlm 5-gram language model (binarized) trained only on the Czech side of CzEng parallel data used
MD5
de338fd4ba04b82631aab9488c468cd6
Preview
  File Preview
    • wmt16.czeng.blm.cs15 GB
Name
Makefile
Size
16.96 KB
Format
application/octet-stream
Description
You can recreate the models using this Makefile
MD5
5f56434491ccb9591c35d8fe20fb8aa9
Preview
  File Preview
Name
moses.ini
Size
1.29 KB
Format
application/octet-stream
Description
The moses.ini file for tuning
MD5
8c60e67f303419ad03fee1fe00aef8cf
Preview
  File Preview