Subject: natural language generation - LINDAT/CLARIAH-CZ Catalog Search Results

Start Over Subject natural language generation

1. Alex Context NLG Dataset

Creator:: Dušek, Ondřej and Jurčíček, Filip
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: text and corpus
Subject:: dialogue system, natural language generation, dialogue alignment, and entrainment
Language:: English
Description:: A dataset intended for fully trainable natural language generation (NLG) systems in task-oriented spoken dialogue systems (SDS), covering the English public transport information domain. It includes preceding context (user utterance) along with each data instance (pair of source meaning representation and target natural language paraphrase to be generated). Taking the form of the previous user utterance into account for generating the system response allows NLG systems trained on this dataset to entrain (adapt) to the preceding utterance, i.e., reuse wording and syntactic structure. This should presumably improve the perceived naturalness of the output, and may even lead to a higher task success rate. Crowdsourcing has been used to obtain natural context user utterances as well as natural system responses to be generated.
Rights:: Creative Commons - Attribution-ShareAlike 4.0 International (CC BY-SA 4.0), http://creativecommons.org/licenses/by-sa/4.0/, and PUB

2. Czech restaurant information dataset for NLG

Creator:: Dušek, Ondřej, Jurčíček, Filip, Dvořák, Josef, Grycová, Petra, Hejda, Matěj, Olivová, Jana, Starý, Michal, and Štichová, Eva
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL)
Type:: text and corpus
Subject:: natural language generation, dialogue system, and morphological generation
Language:: Czech
Description:: This is a dataset for natural language generation (NLG) in task-oriented spoken dialogue systems with Czech as the target language. It originated as a translation of the English San Francisco Restaurants dataset by Wen et al. (2015). It includes input dialogue acts and the corresponding output natural language paraphrases in Czech. Since the dataset is intended for recurrent neural network based NLG systems using delexicalization, inflection tables for all slot values appearing verbatim in the text are provided.
Rights:: Creative Commons - Attribution-ShareAlike 4.0 International (CC BY-SA 4.0), http://creativecommons.org/licenses/by-sa/4.0/, and PUB

3. THEaiTRobot 1.0

Creator:: Rosa, Rudolf, Dušek, Ondřej, Kocmi, Tom, Mareček, David, Musil, Tomáš, Schmidtová, Patrícia, Jurko, Dominik, Bojar, Ondřej, Hrbek, Daniel, Košťák, David, Kinská, Martina, Nováková, Marie, Doležal, Josef, and Vosecká, Klára
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL), The Švanda Theatre in Smíchov, and The Academy of Performing Arts in Prague, Theatre Faculty (DAMU)
Type:: tool and toolService
Subject:: theatre and natural language generation
Language:: English and Czech
Description:: The THEaiTRobot 1.0 tool allows the user to interactively generate scripts for individual theatre play scenes. The tool is based on GPT-2 XL generative language model, using the model without any fine-tuning, as we found that with a prompt formatted as a part of a theatre play script, the model usually generates continuation that retains the format. We encountered numerous problems when generating the script in this way. We managed to tackle some of the problems with various adjustments, but some of them remain to be solved in a future version. THEaiTRobot 1.0 was used to generate the first THEaiTRE play, "AI: Když robot píše hru" ("AI: When a robot writes a play").
Rights:: The MIT License (MIT), http://opensource.org/licenses/mit-license.php, and PUB

4. THEaiTRobot 2.0

Creator:: Rosa, Rudolf, Dušek, Ondřej, Kocmi, Tom, Mareček, David, Musil, Tomáš, Schmidtová, Patrícia, Jurko, Dominik, Bojar, Ondřej, Hrbek, Daniel, Košťák, David, Kinská, Martina, Nováková, Marie, Doležal, Josef, Vosecká, Klára, Zakhtarenko, Alisa, and Obaid, Saad
Publisher:: Charles University, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics (UFAL), The Švanda Theatre in Smíchov, and The Academy of Performing Arts in Prague, Theatre Faculty (DAMU)
Type:: tool and toolService
Subject:: theatre and natural language generation
Language:: English and Czech
Description:: The THEaiTRobot 2.0 tool allows the user to interactively generate scripts for individual theatre play scenes. The previous version of the tool (http://hdl.handle.net/11234/1-3507) was based on GPT-2 XL generative language model, using the model without any fine-tuning, as we found that with a prompt formatted as a part of a theatre play script, the model usually generates continuation that retains the format. The current version also uses vanilla GPT-2 by default, but can also instead use a GPT-2 medium model fine-tuned on theatre play scripts (as well as film and TV series scripts). Apart from the basic "flat" generation using a theatrical starting prompt and the script model, the tool also features a second, hierarchical variant, where in the first step, a play synopsis is generated from its title using a synopsis model (GPT-2 medium fine-tuned on synopses of theatre plays, as well as film, TV series and book synopses). The synopsis is then used as input for the second stage, which uses the script model. The choice of models to use is done by setting the MODEL variable in start_server.sh and start_syn_server.sh THEaiTRobot 2.0 was used to generate the second THEaiTRE play, "Permeation/Prostoupení".
Rights:: The MIT License (MIT), http://opensource.org/licenses/mit-license.php, and PUB

1. Alex Context NLG Dataset

2. Czech restaurant information dataset for NLG

3. THEaiTRobot 1.0

4. THEaiTRobot 2.0

Limit your search

Show values starting with

Search

Search Constraints

Search Results

Limit your search

Contributor

Creator

Show values starting with

Language

Publisher

Rights

Subject

Type

Original context has metadata only

Harvested from