Automatic induction of language model data for a spoken dialogue system.
Saved in:
| Title: | Automatic induction of language model data for a spoken dialogue system. |
|---|---|
| Authors: | Chao Wang1 wangc@csail.mit.edu, Chung, Grace2 gchung@cnri.reston.va.us, Seneff, Stephanie1 seneff@csail.mit.edu |
| Source: | Language Resources & Evaluation. Feb2006, Vol. 40 Issue 1, p25-46. 22p. |
| Subjects: | Dialogue analysis, Interpersonal communication, Simulation methods & models, Recognition (Psychology), Information resources, Language & languages |
| Abstract: | In this paper, we address the issue of generating in-domain language model training data when little or no real user data are available. The two-stage approach taken begins with a data induction phase whereby linguistic constructs from out-of-domain sentences are harvested and integrated with artificially constructed in-domain phrases. After some syntactic and semantic filtering, a large corpus of synthetically assembled user utterances is induced. In the second stage, two sampling methods are explored to filter the synthetic corpus to achieve a desired probability distribution of the semantic content, both on the sentence level and on the class level. The first method utilizes user simulation technology, which obtains the probability model via an interplay between a probabilistic user model and the dialogue system. The second method synthesizes novel dialogue interactions from the raw data by modelling after a small set of dialogues produced by the developers during the course of system refinement. Evaluation is conducted on recognition performance in a restaurant information domain. We show that a partial match to usage-appropriate semantic content distribution can be achieved via user simulations. Furthermore, word error rate can be reduced when limited amounts of in-domain training data are augmented with synthetic data derived by our methods. [ABSTRACT FROM AUTHOR] |
| Copyright of Language Resources & Evaluation is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Links: – Type: pdflink Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 23218135 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Automatic induction of language model data for a spoken dialogue system. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Chao+Wang%22">Chao Wang</searchLink><relatesTo>1</relatesTo><i> wangc@csail.mit.edu</i><br /><searchLink fieldCode="AR" term="%22Chung%2C+Grace%22">Chung, Grace</searchLink><relatesTo>2</relatesTo><i> gchung@cnri.reston.va.us</i><br /><searchLink fieldCode="AR" term="%22Seneff%2C+Stephanie%22">Seneff, Stephanie</searchLink><relatesTo>1</relatesTo><i> seneff@csail.mit.edu</i> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Language+Resources+%26+Evaluation%22">Language Resources & Evaluation</searchLink>. Feb2006, Vol. 40 Issue 1, p25-46. 22p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Dialogue+analysis%22">Dialogue analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Interpersonal+communication%22">Interpersonal communication</searchLink><br /><searchLink fieldCode="DE" term="%22Simulation+methods+%26+models%22">Simulation methods & models</searchLink><br /><searchLink fieldCode="DE" term="%22Recognition+%28Psychology%29%22">Recognition (Psychology)</searchLink><br /><searchLink fieldCode="DE" term="%22Information+resources%22">Information resources</searchLink><br /><searchLink fieldCode="DE" term="%22Language+%26+languages%22">Language & languages</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: In this paper, we address the issue of generating in-domain language model training data when little or no real user data are available. The two-stage approach taken begins with a data induction phase whereby linguistic constructs from out-of-domain sentences are harvested and integrated with artificially constructed in-domain phrases. After some syntactic and semantic filtering, a large corpus of synthetically assembled user utterances is induced. In the second stage, two sampling methods are explored to filter the synthetic corpus to achieve a desired probability distribution of the semantic content, both on the sentence level and on the class level. The first method utilizes user simulation technology, which obtains the probability model via an interplay between a probabilistic user model and the dialogue system. The second method synthesizes novel dialogue interactions from the raw data by modelling after a small set of dialogues produced by the developers during the course of system refinement. Evaluation is conducted on recognition performance in a restaurant information domain. We show that a partial match to usage-appropriate semantic content distribution can be achieved via user simulations. Furthermore, word error rate can be reduced when limited amounts of in-domain training data are augmented with synthetic data derived by our methods. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Language Resources & Evaluation is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=23218135 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1007/s10579-006-9007-3 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 22 StartPage: 25 Subjects: – SubjectFull: Dialogue analysis Type: general – SubjectFull: Interpersonal communication Type: general – SubjectFull: Simulation methods & models Type: general – SubjectFull: Recognition (Psychology) Type: general – SubjectFull: Information resources Type: general – SubjectFull: Language & languages Type: general Titles: – TitleFull: Automatic induction of language model data for a spoken dialogue system. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Chao Wang – PersonEntity: Name: NameFull: Chung, Grace – PersonEntity: Name: NameFull: Seneff, Stephanie IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 02 Text: Feb2006 Type: published Y: 2006 Identifiers: – Type: issn-print Value: 1574020X Numbering: – Type: volume Value: 40 – Type: issue Value: 1 Titles: – TitleFull: Language Resources & Evaluation Type: main |
| ResultId | 1 |