OpenLegalData (2022 - Corpus)
Please use the following text to cite this item or export to a predefined format:
Rüdiger, Jan Oliver, 2023,
OpenLegalData (2022 - Corpus), LINDAT/CLARIAH-CZ digital library at the Institute of Formal and Applied Linguistics (ÚFAL),
http://hdl.handle.net/11372/LRT-5196.
Authors
Item identifier
Project URL
Date issued
2023-07-31
Size
194 files,
610739824 tokens,
39597426 sentences,
169216 texts
Language(s)
Description
OpenLegalData is a free and open platform that makes legal documents and information available to the public. The aim of this platform is to improve the transparency of jurisprudence with the help of open data and to help people without legal training to understand the justice system. The project is committed to the Open Data principles and the Free Access to Justice Movement.
OpenLegalData's DUMP as of 2022-10-18 was used to create this corpus. The data was cleaned, automatically annotated (TreeTagger: POS & Lemma) and grouped based on the metadata (jurisdiction - BundeslandID - sub-size if applicable - ex: Verwaltungsgerichtsbarkeit_11_05.cec6.gz - jurisdiction: administrative jurisdiction, BundeslandID = 11 - sub-corpus = 05). Sub-corpora are randomly split into 50 MB each.
Corpus data is available in CEC6 format. This can be converted into many different corpus formats - use the software www.CorpusExplorer.de if necessary.
Publisher
Subject(s)
Collections
This item isPublicly Available
and licensed under:


