Subject: morphological annotation - LINDAT/CLARIAH-CZ Catalog Search Results

Start Over Subject morphological annotation

Publisher:: Copenhagen Business School
Format:: application/octet-stream
Type:: corpus
Subject:: parallel treebank, POS annotation, discourse annotation, morphological annotation, syntactic annotation, and semantic annotation
Language:: Danish, English, German, Italian, and Spanish
Description:: Parallel treebanks with annotation of syntax, discourse, coreference, morphology, and semantics. Version 3 also includes the Danish Dependency Treebank (version 1) and the Danish-English Parallel Dependency Treebank (version 2).
Rights:: GNU General Public License

Creator:: Skoumalová, Hana
Publisher:: Charles University, Faculty of Arts, Institute of Theoretical and Computational Linguistics
Type:: text and corpus
Subject:: annotated corpus and morphological annotation
Language:: Czech
Description:: Etalon is a manually annotated corpus of contemporary Czech. The corpus contains 1,885,589 words (2,265,722 tokens) and is annotated in the same way as SYN2020 of the Czech National Corpus. The corpus includes fiction (ca 24%), professional and scientific literature (ca 40%) and newspapers (ca 36%). The corpus is provided in a vertical format, where sentence boundaries are marked with a blank line. Every word form is written on a separate line, followed by five tab-separated attributes: syntactic word, lemma, sublemma, tag and verbtag. The texts are shuffled in random chunks of 100 words at maximum (respecting sentence boundaries).
Rights:: Creative Commons - Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0), http://creativecommons.org/licenses/by-nc-sa/4.0/, and PUB

Creator:: Křivan, Jan and Šindlerová, Jana
Format:: bez média and svazek
Type:: model:article and TEXT
Subject:: lemmatization, tokenization, morphological annotation, verbal morphology, lemma variants, lemmatizace, tokenizace, morfologická anotace, slovesná morfologie, and varianty lemmatu
Language:: Czech
Description:: This paper introduces some major conceptual enhancements to the morphological annotation of the SYN series corpora of the Czech National Corpus. Apart from minor changes in tokenization and in the positional tagset, three major conceptual changes have been applied which affect the representation of various lexical and grammatical patterns. In the paper, we present the actual impact of the changes in linguistic data and search for possibilities in three linguistic areas. First, the treatment of phonic, graphemic, and morphological variants via a two-tier lemma structure is discussed; second, a new approach to periphrastic verb forms, auxiliaries, participles and the interpretation of verbal grammatical categories through a new attribute, called verbtag, is explained; and third, a complex multi-value treatment of multiword tokens is introduced.
Rights:: http://creativecommons.org/licenses/by-nc-sa/4.0/ and policy:public

Limit your search