This is a new version of the repository. Do let us know (lindat-help at ufal.mff.cuni.cz) if you encounter any issues.

Annotated corpora and tools of the PARSEME Shared Task on Semi-Supervised Identification of Verbal Multiword Expressions (edition 1.2)

Please use the following text to cite this item or export to a predefined format:
Ramisch, Carlos; et al., 2020, Annotated corpora and tools of the PARSEME Shared Task on Semi-Supervised Identification of Verbal Multiword Expressions (edition 1.2), LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL), http://hdl.handle.net/11234/1-3367.
Authors
show everyone
Date issued
2020-07-09
Size
279785 sentences,
5517910 tokens,
68503 multiWordUnits
Description
This multilingual resource contains corpora in which verbal MWEs have been manually annotated, gathered at the occasion of the 1.2 edition of the PARSEME Shared Task on semi-supervised Identification of Verbal MWEs (2020). VMWEs include idioms (let the cat out of the bag), light-verb constructions (make a decision), verb-particle constructions (give up), inherently reflexive verbs (help oneself), and multi-verb constructions (make do). For the 1.2 shared task edition, the data covers 14 languages, for which VMWEs were annotated according to the universal guidelines. The corpora are provided in the cupt format, inspired by the CONLL-U format. Morphological and syntactic information ­­­­– not necessarily using UD tagsets – including parts of speech, lemmas, morphological features and/or syntactic dependencies are also provided. Depending on the language, the information comes from treebanks (e.g., Universal Dependencies) or from automatic parsers trained on treebanks (e.g., UDPipe). This item contains training, development and test data, as well as the evaluation tools used in the PARSEME Shared Task 1.2 (2020). The annotation guidelines are available online: http://parsemefr.lif.univ-mrs.fr/parseme-st-guidelines/1.2
Publisher
This item isPublicly Available
and licensed under:

Files in this item

Name
GA.tgz
Size
802.56 KB
Format
application/x-gzip
Description
Irish files
MD5
b554c8424df0d989c4580ad7ab43ce56
Preview
  File Preview
  • GA
    • test.blind.cupt1 MB
    • dev-stats.md195 B
    • README.md3 kB
    • train-stats.md197 B
    • dev.cupt474 kB
    • test-stats.md379 B
    • train.cupt417 kB
    • test.cupt1 MB
Name
trial.tgz
Size
113.47 KB
Format
application/x-gzip
Description
Trial files (English)
MD5
17c8e72d5cd58194868598f0579ab524
Preview
  File Preview
  • trial
    • EN-trial.test.cupt8 kB
    • README.md1 kB
    • EN-trial.raw.conllu511 kB
    • EN-trial.train.cupt7 kB
    • EN-trial.test.pred.cupt8 kB
    • EN-trial.test.blind.cupt8 kB
Name
SV.tgz
Size
1.22 MB
Format
application/x-gzip
Description
Swedish files
MD5
868c003f369f6af324a5699dd4be0726
Preview
  File Preview
  • SV
    • test.blind.cupt2 MB
    • dev-stats.md178 B
    • README.md3 kB
    • train-stats.md203 B
    • dev.cupt755 kB
    • test-stats.md368 B
    • train.cupt2 MB
    • test.cupt2 MB
Name
IT.tgz
Size
5.82 MB
Format
application/x-gzip
Description
Italian files
MD5
4b6dc92ccf29768d80e7d136285efa09
Preview
  File Preview
  • IT
    • test.blind.cupt5 MB
    • dev-stats.md224 B
    • README.md8 kB
    • train-stats.md253 B
    • dev.cupt1 MB
    • test-stats.md398 B
    • train.cupt16 MB
    • test.cupt5 MB
Name
DE.tgz
Size
2.73 MB
Format
application/x-gzip
Description
German files
MD5
55d0f986e358739e56573c221c617282
Preview
  File Preview
  • DE
    • test.blind.cupt2 MB
    • dev-stats.md197 B
    • README.md4 kB
    • train-stats.md210 B
    • dev.cupt856 kB
    • test-stats.md367 B
    • train.cupt9 MB
    • test.cupt2 MB
Name
HE.tgz
Size
6.45 MB
Format
application/x-gzip
Description
Hebrew files
MD5
384e4e150b958ebc98d2d420feb1f4d8
Preview
  File Preview
  • HE
    • test.blind.cupt5 MB
    • dev-stats.md165 B
    • README.md3 kB
    • train-stats.md175 B
    • dev.cupt1 MB
    • test-stats.md333 B
    • train.cupt22 MB
    • test.cupt5 MB
Name
EU.tgz
Size
2.95 MB
Format
application/x-gzip
Description
Basque files
MD5
a3cdb2376ab2e000820a98e1cf22ba87
Preview
  File Preview
  • EU
    • test.blind.cupt5 MB
    • dev-stats.md149 B
    • README.md4 kB
    • train-stats.md153 B
    • dev.cupt1 MB
    • test-stats.md318 B
    • train.cupt4 MB
    • test.cupt5 MB
Name
PT.tgz
Size
10.15 MB
Format
application/x-gzip
Description
Portuguese files
MD5
4827ddfb24df8f634543be85b5939609
Preview
  File Preview
  • PT
    • test.blind.cupt8 MB
    • dev-stats.md174 B
    • README.md7 kB
    • train-stats.md184 B
    • dev.cupt2 MB
    • test-stats.md345 B
    • train.cupt34 MB
    • test.cupt8 MB
Name
TR.tgz
Size
5.21 MB
Format
application/x-gzip
Description
Turkish files
MD5
7ed6b16e8fcc30d646d04791ebb54b4a
Preview
  File Preview
  • TR
    • test.blind.cupt4 MB
    • dev-stats.md142 B
    • README.md4 kB
    • train-stats.md149 B
    • dev.cupt1 MB
    • test-stats.md310 B
    • train.cupt22 MB
    • test.cupt4 MB
Name
HI.tgz
Size
771.32 KB
Format
application/x-gzip
Description
Hindi files
MD5
8ea2dd1a1f9082f53234c597e330eaad
Preview
  File Preview
  • HI
    • test.blind.cupt2 MB
    • dev-stats.md140 B
    • README.md2 kB
    • train-stats.md161 B
    • dev.cupt598 kB
    • test-stats.md327 B
    • train.cupt549 kB
    • test.cupt2 MB
Name
ZH.tgz
Size
8.21 MB
Format
application/x-gzip
Description
Chinese files
MD5
052a11ca5136136be9d46935765b7a2a
Preview
  File Preview
  • ZH
    • test.blind.cupt2 MB
    • dev-stats.md181 B
    • README.md3 kB
    • train-stats.md192 B
    • dev.cupt884 kB
    • test-stats.md348 B
    • train.cupt27 MB
    • test.cupt2 MB
Name
EL.tgz
Size
9.41 MB
Format
application/x-gzip
Description
Greek files
MD5
4dfafa58e1f504f3600d54401beccf24
Preview
  File Preview
  • EL
    • test.blind.cupt6 MB
    • dev-stats.md177 B
    • README.md3 kB
    • train-stats.md190 B
    • dev.cupt2 MB
    • test-stats.md346 B
    • train.cupt42 MB
    • test.cupt6 MB
Name
FR.tgz
Size
7.43 MB
Format
application/x-gzip
Description
French files
MD5
a3f1331707d34c31b0dc221799a5e8b6
Preview
  File Preview
  • FR
    • test.blind.cupt7 MB
    • dev-stats.md176 B
    • README.md5 kB
    • train-stats.md186 B
    • dev.cupt2 MB
    • test-stats.md346 B
    • train.cupt22 MB
    • test.cupt7 MB
Name
PL.tgz
Size
8.29 MB
Format
application/x-gzip
Description
Polish files
MD5
f1db1c5bd299d7d4ed1eaa4d76ab91a8
Preview
  File Preview
  • PL
    • test.blind.cupt7 MB
    • dev-stats.md163 B
    • README.md9 kB
    • train-stats.md172 B
    • dev.cupt2 MB
    • test-stats.md331 B
    • train.cupt30 MB
    • test.cupt7 MB
Name
RO.tgz
Size
20.59 MB
Format
application/x-gzip
Description
Romanian files
MD5
2972d6276e83a15596ec1d527c8689b0
Preview
  File Preview
  • RO
    • test.blind.cupt50 MB
    • dev-stats.md164 B
    • README.md2 kB
    • train-stats.md169 B
    • dev.cupt9 MB
    • test-stats.md337 B
    • train.cupt14 MB
    • test.cupt50 MB
Name
bin.tgz
Size
19.7 KB
Format
application/x-gzip
Description
Evaluation scripts
MD5
456e2a812566cc791a0d6be38f507bdd
Preview
  File Preview
  • bin
    • validate_cupt.py4 kB
    • bmc_munkres
      • LICENSE561 B
      • README.md1 kB
      • munkres.py23 kB
    • evaluate.py23 kB
    • average_of_evaluations.py6 kB
    • tsvlib.py12 kB
    • tsvlib_usage_example.py1 kB
Name
README.md
Size
6.7 KB
Format
application/octet-stream
Description
General README file
MD5
a8b7e1ba4c2b8b09cf76c040fb5d41ab
Preview
  File Preview