Automatic induction of language model data for a spoken dialogue system.

Saved in:
Bibliographic Details
Title: Automatic induction of language model data for a spoken dialogue system.
Authors: Chao Wang1 wangc@csail.mit.edu, Chung, Grace2 gchung@cnri.reston.va.us, Seneff, Stephanie1 seneff@csail.mit.edu
Source: Language Resources & Evaluation. Feb2006, Vol. 40 Issue 1, p25-46. 22p.
Subjects: Dialogue analysis, Interpersonal communication, Simulation methods & models, Recognition (Psychology), Information resources, Language & languages
Abstract: In this paper, we address the issue of generating in-domain language model training data when little or no real user data are available. The two-stage approach taken begins with a data induction phase whereby linguistic constructs from out-of-domain sentences are harvested and integrated with artificially constructed in-domain phrases. After some syntactic and semantic filtering, a large corpus of synthetically assembled user utterances is induced. In the second stage, two sampling methods are explored to filter the synthetic corpus to achieve a desired probability distribution of the semantic content, both on the sentence level and on the class level. The first method utilizes user simulation technology, which obtains the probability model via an interplay between a probabilistic user model and the dialogue system. The second method synthesizes novel dialogue interactions from the raw data by modelling after a small set of dialogues produced by the developers during the course of system refinement. Evaluation is conducted on recognition performance in a restaurant information domain. We show that a partial match to usage-appropriate semantic content distribution can be achieved via user simulations. Furthermore, word error rate can be reduced when limited amounts of in-domain training data are augmented with synthetic data derived by our methods. [ABSTRACT FROM AUTHOR]
Copyright of Language Resources & Evaluation is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Links:
  – Type: pdflink
Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 23218135
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Automatic induction of language model data for a spoken dialogue system.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Chao+Wang%22">Chao Wang</searchLink><relatesTo>1</relatesTo><i> wangc@csail.mit.edu</i><br /><searchLink fieldCode="AR" term="%22Chung%2C+Grace%22">Chung, Grace</searchLink><relatesTo>2</relatesTo><i> gchung@cnri.reston.va.us</i><br /><searchLink fieldCode="AR" term="%22Seneff%2C+Stephanie%22">Seneff, Stephanie</searchLink><relatesTo>1</relatesTo><i> seneff@csail.mit.edu</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Language+Resources+%26+Evaluation%22">Language Resources & Evaluation</searchLink>. Feb2006, Vol. 40 Issue 1, p25-46. 22p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Dialogue+analysis%22">Dialogue analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Interpersonal+communication%22">Interpersonal communication</searchLink><br /><searchLink fieldCode="DE" term="%22Simulation+methods+%26+models%22">Simulation methods & models</searchLink><br /><searchLink fieldCode="DE" term="%22Recognition+%28Psychology%29%22">Recognition (Psychology)</searchLink><br /><searchLink fieldCode="DE" term="%22Information+resources%22">Information resources</searchLink><br /><searchLink fieldCode="DE" term="%22Language+%26+languages%22">Language & languages</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: In this paper, we address the issue of generating in-domain language model training data when little or no real user data are available. The two-stage approach taken begins with a data induction phase whereby linguistic constructs from out-of-domain sentences are harvested and integrated with artificially constructed in-domain phrases. After some syntactic and semantic filtering, a large corpus of synthetically assembled user utterances is induced. In the second stage, two sampling methods are explored to filter the synthetic corpus to achieve a desired probability distribution of the semantic content, both on the sentence level and on the class level. The first method utilizes user simulation technology, which obtains the probability model via an interplay between a probabilistic user model and the dialogue system. The second method synthesizes novel dialogue interactions from the raw data by modelling after a small set of dialogues produced by the developers during the course of system refinement. Evaluation is conducted on recognition performance in a restaurant information domain. We show that a partial match to usage-appropriate semantic content distribution can be achieved via user simulations. Furthermore, word error rate can be reduced when limited amounts of in-domain training data are augmented with synthetic data derived by our methods. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Language Resources & Evaluation is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=23218135
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1007/s10579-006-9007-3
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 22
        StartPage: 25
    Subjects:
      – SubjectFull: Dialogue analysis
        Type: general
      – SubjectFull: Interpersonal communication
        Type: general
      – SubjectFull: Simulation methods & models
        Type: general
      – SubjectFull: Recognition (Psychology)
        Type: general
      – SubjectFull: Information resources
        Type: general
      – SubjectFull: Language & languages
        Type: general
    Titles:
      – TitleFull: Automatic induction of language model data for a spoken dialogue system.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Chao Wang
      – PersonEntity:
          Name:
            NameFull: Chung, Grace
      – PersonEntity:
          Name:
            NameFull: Seneff, Stephanie
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 02
              Text: Feb2006
              Type: published
              Y: 2006
          Identifiers:
            – Type: issn-print
              Value: 1574020X
          Numbering:
            – Type: volume
              Value: 40
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Language Resources & Evaluation
              Type: main
ResultId 1