Prague Czech-English Dependency Treebank 2.0
Please use the following text to cite this item or export to a predefined format:
Hajič, Jan; et al., 2012,
Prague Czech-English Dependency Treebank 2.0, LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL),
http://hdl.handle.net/11858/00-097C-0000-0015-8DAF-4.
Authors
Hajič, Jan ; et al.
Item identifier
Project URL
Date issued
2012
Size
49208 sentences
Description
Texts
The Prague Czech-English Dependency Treebank 2.0 (PCEDT 2.0) is a major update of the Prague Czech-English Dependency Treebank 1.0 (LDC2004T25). It is a manually parsed Czech-English parallel corpus sized over 1.2 million running words in almost 50,000 sentences for each part.
Data
The English part contains the entire Penn Treebank - Wall Street Journal Section (LDC99T42). The Czech part consists of Czech translations of all of the Penn Treebank-WSJ texts. The corpus is 1:1 sentence-aligned. An additional automatic alignment on the node level (different for each annotation layer) is part of this release, too. The original Penn Treebank-like file structure (25 sections, each containing up to one hundred files) has been preserved. Only those PTB documents which have both POS and structural annotation (total of 2312 documents) have been translated to Czech and made part of this release.
Each language part is enhanced with a comprehensive manual linguistic annotation in the PDT 2.0 style (LDC2006T01, Prague Dependency Treebank 2.0). The main features of this annotation style are:
dependency structure of the content words and coordinating and similar structures (function words are attached as their attribute values)
semantic labeling of content words and types of coordinating structures
argument structure, including an argument structure ("valency") lexicon for both languages
ellipsis and anaphora resolution.
This annotation style is called tectogrammatical annotation and it constitutes the tectogrammatical layer in the corpus. For more details see below and documentation.
Annotation of the Czech part
Sentences of the Czech translation were automatically morphologically annotated and parsed into surface-syntax dependency trees in the PDT 2.0 annotation style. This annotation style is sometimes called analytical annotation; it constitutes the analytical layer of the corpus. The manual tectogrammatical (deep-syntax) annotation was built as a separate layer above the automatic analytical (surface-syntax) parse. A sample of 2,000 sentences was manually annotated on the analytical layer.
Annotation of the English part
The resulting manual tectogrammatical annotation was built above an automatic transformation of the original phrase-structure annotation of the Penn Treebank into surface dependency (analytical) representations, using the following additional linguistic information from other sources:
PropBank (LDC2004T14)
VerbNet
NomBank (LDC2008T23)
flat noun phrase structures (by courtesy of D. Vadas and J.R. Curran)
For each sentence, the original Penn Treebank phrase structure trees are preserved in this corpus together with their links to the analytical and tectogrammatical annotation.
Acknowledgement
Ministerstvo školství, mládeže a tělovýchovy České republiky
Project code:MSM 0021620838
Project name:Moderní metody, struktury a systémy informatiky
Ministerstvo školství, mládeže a tělovýchovy České republiky
Project code:LC536
Project name:Centrum komputační lingvistiky
Ministerstvo školství, mládeže a tělovýchovy České republiky
Project code:ME09008
Project name:Mnohojazyčná univerzální anotace lingvistických dat
Ministerstvo školství, mládeže a tělovýchovy České republiky
Project code:LM2010013
Project name:LINDAT/CLARIN: Institut pro analýzu, zpracování a distribuci lingvistických dat
Ministerstvo školství, mládeže a tělovýchovy České republiky
Project code:7E09003
Project name:EuroMatrixPlus – Bringing Machine Translation for European Languages to the User
Ministerstvo školství, mládeže a tělovýchovy České republiky
Project code:7E11051
Project name:EuroMatrixPlus - Enlarged European Union Bringing Machine Translation for European Languages to the User
Ministerstvo školství, mládeže a tělovýchovy České republiky
Project code:7E11041
Project name:Feedback Analysis for User Adaptive Statistical Translation
Grantová agentura České republiky
Project code:GAP406/10/0875
Project name:Komputační lingvistika: Explicitní popis jazyka a anotovaná data se zřetelem na češtinu
Grantová agentura České republiky
Project code:GPP406/10/P193
Project name:Nástroje pro revizi a tektogramatickou anotaci českého závislostního korpusu
Grantová agentura České republiky
Project code:GA405/09/0729
Project name:Od struktury věty k textovým vztahům
Grantová agentura Akademie věd České republiky
Project code:1ET101120503
Project name:Integrace jazykových zdrojů za účelem extrakce informací z přirozených textů
Grantová agentura Univerzity Karlovy v Praze
Project code:GAUK 116310/2010
Project name:Anglicko-český strojový překlad s využitím hloubkové syntaxe
European Union
Project code:FP6-IST-5-034434-IP
Project name:Companions IP
Grantová agentura Univerzity Karlovy v Praze
Project code:GAUK 3537/2011
Project name:Detekce větné polarity v počítačovém korpusu
European Union
Project code:FP6-IST-5-034291-STP
Project name:Euromatrix
European Union
Project code:FP7-ICT-2007-3-231720
Project name:EuroMatrix Plus
European Union
Project code:FP7-ICT-2009-4-247762
Project name:Faust
Grantová agentura Univerzity Karlovy v Praze
Project code:GAUK 1580/2010
Project name:Značkování aktuálního členění věty v paralelním anglicko-českém závislostním korpusu
Collections
Version History
Files in this item
- Name
- PCEDT-2-doc-with-trees.zip
- Size
- 625.61 MB
- Format
- application/zip
- Description
- Documentation, including visualisation of all trees
- MD5
- 489c19ab2ce4cec967919bb0e12e3c58

The file preview has not been generated yet. Please try again later or contact the system administrator lindat-help@ufal.mff.cuni.cz
- Name
- PCEDT-2-full-DVD.zip
- Size
- 1.06 GB
- Format
- application/zip
- Description
- Data + doc + tools
- MD5
- fe874b94a84ffea87a788592169b343f

The file preview has not been generated yet. Please try again later or contact the system administrator lindat-help@ufal.mff.cuni.cz
- Name
- PCEDT-2-data-only.zip
- Size
- 263.41 MB
- Format
- application/zip
- Description
- Only the data files (incl. valency lexicons)
- MD5
- 7a7d3395752cd1074826685c62bc35aa

The file preview has not been generated yet. Please try again later or contact the system administrator lindat-help@ufal.mff.cuni.cz

