This is not the latest version of this item. The latest version can be found here.
Annotated corpora and tools of the PARSEME Shared Task on Semi-Supervised Identification of Verbal Multiword Expressions (edition 1.2)
Please use the following text to cite this item or export to a predefined format:
Ramisch, Carlos; et al., 2020,
Annotated corpora and tools of the PARSEME Shared Task on Semi-Supervised Identification of Verbal Multiword Expressions (edition 1.2), LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL),
http://hdl.handle.net/11234/1-3367.
Authors
Ramisch, Carlos ; et al.
Item identifier
Project URL
Date issued
2020-07-09
Size
279785 sentences,
5517910 tokens,
68503 multiWordUnits
Description
This multilingual resource contains corpora in which verbal MWEs have been manually annotated, gathered at the occasion of the 1.2 edition of the PARSEME Shared Task on semi-supervised Identification of Verbal MWEs (2020).
VMWEs include idioms (let the cat out of the bag), light-verb constructions (make a decision), verb-particle constructions (give up), inherently reflexive verbs (help oneself), and multi-verb constructions (make do).
For the 1.2 shared task edition, the data covers 14 languages, for which VMWEs were annotated according to the universal guidelines. The corpora are provided in the cupt format, inspired by the CONLL-U format.
Morphological and syntactic information – not necessarily using UD tagsets – including parts of speech, lemmas, morphological features and/or syntactic dependencies are also provided. Depending on the language, the information comes from treebanks (e.g., Universal Dependencies) or from automatic parsers trained on treebanks (e.g., UDPipe).
This item contains training, development and test data, as well as the evaluation tools used in the PARSEME Shared Task 1.2 (2020). The annotation guidelines are available online: http://parsemefr.lif.univ-mrs.fr/parseme-st-guidelines/1.2
Publisher
Collections
Version History
Files in this item
- Name
- GA.tgz
- Size
- 802.56 KB
- Format
- application/x-gzip
- Description
- Irish files
- MD5
- b554c8424df0d989c4580ad7ab43ce56

- GA
- test.blind.cupt1 MB
- dev-stats.md195 B
- README.md3 kB
- train-stats.md197 B
- dev.cupt474 kB
- test-stats.md379 B
- train.cupt417 kB
- test.cupt1 MB
- Name
- trial.tgz
- Size
- 113.47 KB
- Format
- application/x-gzip
- Description
- Trial files (English)
- MD5
- 17c8e72d5cd58194868598f0579ab524

- trial
- EN-trial.test.cupt8 kB
- README.md1 kB
- EN-trial.raw.conllu511 kB
- EN-trial.train.cupt7 kB
- EN-trial.test.pred.cupt8 kB
- EN-trial.test.blind.cupt8 kB
- Name
- SV.tgz
- Size
- 1.22 MB
- Format
- application/x-gzip
- Description
- Swedish files
- MD5
- 868c003f369f6af324a5699dd4be0726

- SV
- test.blind.cupt2 MB
- dev-stats.md178 B
- README.md3 kB
- train-stats.md203 B
- dev.cupt755 kB
- test-stats.md368 B
- train.cupt2 MB
- test.cupt2 MB
- Name
- IT.tgz
- Size
- 5.82 MB
- Format
- application/x-gzip
- Description
- Italian files
- MD5
- 4b6dc92ccf29768d80e7d136285efa09

- IT
- test.blind.cupt5 MB
- dev-stats.md224 B
- README.md8 kB
- train-stats.md253 B
- dev.cupt1 MB
- test-stats.md398 B
- train.cupt16 MB
- test.cupt5 MB
- Name
- DE.tgz
- Size
- 2.73 MB
- Format
- application/x-gzip
- Description
- German files
- MD5
- 55d0f986e358739e56573c221c617282

- DE
- test.blind.cupt2 MB
- dev-stats.md197 B
- README.md4 kB
- train-stats.md210 B
- dev.cupt856 kB
- test-stats.md367 B
- train.cupt9 MB
- test.cupt2 MB
- Name
- HE.tgz
- Size
- 6.45 MB
- Format
- application/x-gzip
- Description
- Hebrew files
- MD5
- 384e4e150b958ebc98d2d420feb1f4d8

- HE
- test.blind.cupt5 MB
- dev-stats.md165 B
- README.md3 kB
- train-stats.md175 B
- dev.cupt1 MB
- test-stats.md333 B
- train.cupt22 MB
- test.cupt5 MB
- Name
- EU.tgz
- Size
- 2.95 MB
- Format
- application/x-gzip
- Description
- Basque files
- MD5
- a3cdb2376ab2e000820a98e1cf22ba87

- EU
- test.blind.cupt5 MB
- dev-stats.md149 B
- README.md4 kB
- train-stats.md153 B
- dev.cupt1 MB
- test-stats.md318 B
- train.cupt4 MB
- test.cupt5 MB
- Name
- PT.tgz
- Size
- 10.15 MB
- Format
- application/x-gzip
- Description
- Portuguese files
- MD5
- 4827ddfb24df8f634543be85b5939609

- PT
- test.blind.cupt8 MB
- dev-stats.md174 B
- README.md7 kB
- train-stats.md184 B
- dev.cupt2 MB
- test-stats.md345 B
- train.cupt34 MB
- test.cupt8 MB
- Name
- TR.tgz
- Size
- 5.21 MB
- Format
- application/x-gzip
- Description
- Turkish files
- MD5
- 7ed6b16e8fcc30d646d04791ebb54b4a

- TR
- test.blind.cupt4 MB
- dev-stats.md142 B
- README.md4 kB
- train-stats.md149 B
- dev.cupt1 MB
- test-stats.md310 B
- train.cupt22 MB
- test.cupt4 MB
- Name
- HI.tgz
- Size
- 771.32 KB
- Format
- application/x-gzip
- Description
- Hindi files
- MD5
- 8ea2dd1a1f9082f53234c597e330eaad

- HI
- test.blind.cupt2 MB
- dev-stats.md140 B
- README.md2 kB
- train-stats.md161 B
- dev.cupt598 kB
- test-stats.md327 B
- train.cupt549 kB
- test.cupt2 MB
- Name
- ZH.tgz
- Size
- 8.21 MB
- Format
- application/x-gzip
- Description
- Chinese files
- MD5
- 052a11ca5136136be9d46935765b7a2a

- ZH
- test.blind.cupt2 MB
- dev-stats.md181 B
- README.md3 kB
- train-stats.md192 B
- dev.cupt884 kB
- test-stats.md348 B
- train.cupt27 MB
- test.cupt2 MB
- Name
- EL.tgz
- Size
- 9.41 MB
- Format
- application/x-gzip
- Description
- Greek files
- MD5
- 4dfafa58e1f504f3600d54401beccf24

- EL
- test.blind.cupt6 MB
- dev-stats.md177 B
- README.md3 kB
- train-stats.md190 B
- dev.cupt2 MB
- test-stats.md346 B
- train.cupt42 MB
- test.cupt6 MB
- Name
- FR.tgz
- Size
- 7.43 MB
- Format
- application/x-gzip
- Description
- French files
- MD5
- a3f1331707d34c31b0dc221799a5e8b6

- FR
- test.blind.cupt7 MB
- dev-stats.md176 B
- README.md5 kB
- train-stats.md186 B
- dev.cupt2 MB
- test-stats.md346 B
- train.cupt22 MB
- test.cupt7 MB
- Name
- PL.tgz
- Size
- 8.29 MB
- Format
- application/x-gzip
- Description
- Polish files
- MD5
- f1db1c5bd299d7d4ed1eaa4d76ab91a8

- PL
- test.blind.cupt7 MB
- dev-stats.md163 B
- README.md9 kB
- train-stats.md172 B
- dev.cupt2 MB
- test-stats.md331 B
- train.cupt30 MB
- test.cupt7 MB
- Name
- RO.tgz
- Size
- 20.59 MB
- Format
- application/x-gzip
- Description
- Romanian files
- MD5
- 2972d6276e83a15596ec1d527c8689b0

- RO
- test.blind.cupt50 MB
- dev-stats.md164 B
- README.md2 kB
- train-stats.md169 B
- dev.cupt9 MB
- test-stats.md337 B
- train.cupt14 MB
- test.cupt50 MB
- Name
- bin.tgz
- Size
- 19.7 KB
- Format
- application/x-gzip
- Description
- Evaluation scripts
- MD5
- 456e2a812566cc791a0d6be38f507bdd

- bin
- validate_cupt.py4 kB
- bmc_munkres
- LICENSE561 B
- README.md1 kB
- munkres.py23 kB
- evaluate.py23 kB
- average_of_evaluations.py6 kB
- tsvlib.py12 kB
- tsvlib_usage_example.py1 kB
- Name
- README.md
- Size
- 6.7 KB
- Format
- application/octet-stream
- Description
- General README file
- MD5
- a8b7e1ba4c2b8b09cf76c040fb5d41ab

The file preview has not been generated yet. Please try again later or contact the system administrator lindat-help@ufal.mff.cuni.cz

