Universal Dependencies 2.0 – CoNLL 2017 Shared Task Development and Test Data
Please use the following text to cite this item or export to a predefined format:
Nivre, Joakim; et al., 2017,
Universal Dependencies 2.0 – CoNLL 2017 Shared Task Development and Test Data, LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL),
http://hdl.handle.net/11234/1-2184.
Authors
Nivre, Joakim ; et al.
Item identifier
Project URL
Demo URL
Date issued
2017-05-18
Size
156582 sentences,
2807034 tokens,
2842220 words
Language(s)
Urdu,
Description
Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).
This release contains the test data used in the CoNLL 2017 shared task on parsing Universal Dependencies. Due to the shared task the test data was held hidden and not released together with the training and development data of UD 2.0. Therefore this release complements the UD 2.0 release (http://hdl.handle.net/11234/1-1983) to a full release of UD treebanks. In addition, the present release contains 18 new parallel test sets and 4 test sets in surprise languages. The present release also includes the development data already released with UD 2.0. Unlike regular UD releases, this one uses the folder-file structure that was visible to the systems participating in the shared task.
Publisher
Acknowledgement
Grantová agentura České republiky
Project code:15-10472S
Project name:Morphologically and Syntactically Annotated Corpora of Many Languages
Collections


