English-Urdu Religious Parallel Corpus
Please use the following text to cite this item or export to a predefined format:
Jawaid, Bushra and Zeman, Daniel, 2010,
English-Urdu Religious Parallel Corpus, LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL),
http://hdl.handle.net/11234/1-2582.
Authors
Item identifier
Project URL
Referenced by
Date issued
2010
Size
14371 sentences
Description
English-Urdu parallel corpus is a collection of religious texts (Quran, Bible) in English and Urdu language with sentence alignments. The corpus can be used for experiments with statistical machine translation. Our modifications of crawled data include but are not limited to the following:
1- Manually corrected sentence alignment of the corpora.
2- Our data split (training-development-test) so that our published experiments can be reproduced.
3- Tokenization (optional, but needed to reproduce our experiments).
4- Normalization (optional) of e.g. European vs. Urdu numerals, European vs. Urdu punctuation, removal of Urdu diacritics.
Acknowledgement
Ministerstvo školství, mládeže a tělovýchovy České republiky
Project code:MSM0021620838
Project name:Modern Methods, Structures and Systems of Computer Science
Grantová agentura České republiky
Project code:GAP406/11/1499
Project name:Čeština ve věku strojového překladu
Charles University in Prague
Project code:SVV 261314/2010
Project name:SVV 261 314
Subject(s)
Collections
This item isPublicly Available
and licensed under:
Files in this item
- Name
- en-ur-parallel-corpus.zip
- Size
- 3.51 MB
- Format
- application/zip
- Description
- Unknown
- MD5
- 8440be07c883b4c0289961ba577a634b


