iula_preprocess
Please use the following text to cite this item or export to a predefined format:
Institut Universitari de Lingüística Aplicada, Universitat Pompeu Fabra, 2014,
iula_preprocess, LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL),
http://hdl.handle.net/11372/LRT-1413.
Item identifier
Date issued
2014-07-30
Type
Description
Text preprocess (this preprocess service requires that the input text be in plain text format (file .txt) and UTF-8).
Basically, it carries out: (i) text segmentation into minor structural units (titles, paragraphs, sentences, etc.); (ii) detection of entities not found in dictionaries (numbers, abbreviations, URLs, emails, proper nouns, etc.); and (iii) the keeping of sequences of two or more words in a single block (dates, phrases, proper nouns, etc.).
Collections

