This is not the latest version of this item. The latest version can be found here.
English-Hindi Parallel Corpus
Please use the following text to cite this item or export to a predefined format:
Bojar, Ondřej; Straňák, Pavel; Zeman, Daniel; Jain, Gaurav and Damani, Om Prakesh, 2010,
English-Hindi Parallel Corpus, LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL),
http://hdl.handle.net/11858/00-097C-0000-0001-BD17-1.
Authors
Item identifier
Date issued
2010-05-11
Description
English-Hindi parallel corpus collected from several sources. Tokenized and sentence-aligned. A part of the data is our patch for the Emille parallel corpus.
Acknowledgement
European Union
Project code:FP7-ICT-2007-3-231720
Project name:EuroMatrix Plus
Ministerstvo školství, mládeže a tělovýchovy České republiky
Project code:7E09003
Project name:EuroMatrixPlus – Bringing Machine Translation for European Languages to the User
Subject(s)
Collections
Version History
This item isPublicly Available
and licensed under:
Files in this item
- Name
- English-Hindi-without-Emille.tgz
- Size
- 12.16 MB
- Format
- application/x-gzip
- Description
- The complete parallel data, including the patch for the Emille corpus
- MD5
- fbe1e19c0e80fd7792e900656ce4c1a9

- UMC002-English-Hindi
- wikipedia-named-entities-2008
- en.tok.gz5 kB
- hi.tok.gz6 kB
- agrocorpus
- en.tok.gz17 kB
- README693 B
- hi.tok.gz15 kB
- shabdanjali-dictionary
- en.tok.gz76 kB
- README1 kB
- hi.tok.gz213 kB
- en.filtered.tok.gz4 kB
- hi.filtered.tok.gz6 kB
- tides-cleaned-by-ufal
- hi.test.tok.gz3 MB
- hi.train.tok.gz3 MB
- en.test.tok.gz2 MB
- hi.dev.tok.gz66 kB
- en.dev.tok.gz47 kB
- en.train.tok.gz2 MB
- README680 B
- danielpipes
- en.tok.gz354 kB
- README51 B
- hi.tok.gz319 kB
- acl-2005-shared-task
- wikipedia-named-entities-2009
- en.tok.gz4 kB
- hi.tok.gz5 kB
- wikipedia-named-entities-2008

