UDify Pretrained Model
Please use the following text to cite this item or export to a predefined format:
Kondratyuk, Dan and Straka, Milan, 2019,
UDify Pretrained Model, LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL),
http://hdl.handle.net/11234/1-3042.
Authors
Item identifier
Project URL
Referenced by
Date issued
2019-08-26
Type
Language(s)
Arabic ,
Basque ,
Croatian ,
Czech ,
Danish ,
Dutch ,
English ,
Estonian ,
Finnish ,
French ,
German ,
Gothic ,
Hebrew ,
Hindi ,
Irish ,
Italian ,
Japanese ,
Latin ,
Persian ,
Polish ,
Romanian ,
Spanish ,
Swedish ,
Tamil ,
Catalan ,
Chinese ,
Galician ,
Kazakh ,
Latvian ,
Russian ,
Turkish ,
Coptic ,
Sanskrit ,
Slovak ,
Uighur ,
Korean ,
Urdu ,
Marathi ,
Serbian ,
Telugu ,
Amharic ,
Armenian ,
Breton ,
Faroese ,
Tagalog ,
Thai ,
Warlpiri ,
Yoruba ,
Akkadian ,
Bambara ,
Erzya ,
Description
Pretrained model weights for the UDify model, and extracted BERT weights in pytorch-transformers format. Note that these weights slightly differ from those used in the paper.
Acknowledgement
Education, Audiovisual and Culture Executive Agency
Project code:EMJMD
Project name:Erasmus Mundus, Language and Communication Technologies
Ministerstvo školství, mládeže a tělovýchovy České republiky
Project code:LM2015071
Project name:LINDAT/CLARIN: Institut pro analýzu, zpracování a distribuci lingvistických dat
Subject(s)
Collections
This item isPublicly Available
and licensed under:
Files in this item
- Name
- udify-model.tar.gz
- Size
- 759.07 MB
- Format
- application/x-gzip
- Description
- The UDify model weights
- MD5
- 42aacc00e0ed6272b31ca7329055c108

- vocabulary
- upos.txt79 B
- xpos.txt245 kB
- non_padded_namespaces.txt37 B
- head_tags.txt2 kB
- tokens.txt18 MB
- lemmas.txt2 MB
- token_characters.txt32 kB
- feats.txt1 MB
-
- config.json5 kB
- weights.th809 MB

