UDPipe is a trainable pipeline for tokenization, tagging, lemmatization and dependency parsing of CoNLL-U files. UDPipe is language-agnostic and can be trained given annotated data in CoNLL-U format. Trained models are provided for nearly all UD treebanks. UDPipe is available as a binary for Linux/Windows/OS X, as a library for C++, Python, Perl, Java, C#, and as a web service. Third-party R CRAN package also exists.

UDPipe is a free software distributed under the Mozilla Public License 2.0 and the linguistic models are free for non-commercial use and distributed under the CC BY-NC-SA license, although for some models the original data used to create the model may impose additional licensing conditions. UDPipe is versioned using Semantic Versioning.

Copyright 2017 by Institute of Formal and Applied Linguistics, Faculty of Mathematics and Physics, Charles University, Czech Republic.

Basic info
Authors Milan Straka
Homepage http://ufal.mff.cuni.cz/udpipe
Development repository http://github.com/ufal/udpipe
Status Stable, version 1.2, 2.0
OS Linux, Windows, OS X
License of the library Mozilla Public License 2.0
License of the models CC BY-NC-SA
Contact straka@ufal.mff.cuni.cz