Survey on automatic lip-reading in the era of deep learning.

Saved in:
Bibliographic Details
Title: Survey on automatic lip-reading in the era of deep learning.
Authors: Fernandez-Lopez, Adriana1 adriana.fernandez@upf.edu, Sukno, Federico M.1 federico.sukno@upf.edu
Source: Image & Vision Computing. Oct2018, Vol. 78, p53-72. 20p.
Subjects: Braille education, Reading mobile apps, Reading standards, Reading strategies, Reading motivation
Abstract: Abstract In the last few years, there has been an increasing interest in developing systems for Automatic Lip-Reading (ALR). Similarly to other computer vision applications, methods based on Deep Learning (DL) have become very popular and have permitted to substantially push forward the achievable performance. In this survey, we review ALR research during the last decade, highlighting the progression from approaches previous to DL (which we refer to as traditional) toward end-to-end DL architectures. We provide a comprehensive list of the audio-visual databases available for lip-reading, describing what tasks they can be used for, their popularity and their most important characteristics, such as the number of speakers, vocabulary size, recording settings and total duration. In correspondence with the shift toward DL, we show that there is a clear tendency toward large-scale datasets targeting realistic application settings and large numbers of samples per class. On the other hand, we summarize, discuss and compare the different ALR systems proposed in the last decade, separately considering traditional and DL approaches. We address a quantitative analysis of the different systems by organizing them in terms of the task that they target (e.g. recognition of letters or digits and words or sentences) and comparing their reported performance in the most commonly used datasets. As a result, we find that DL architectures perform similarly to traditional ones for simpler tasks but report significant improvements in more complex tasks, such as word or sentence recognition, with up to 40% improvement in word recognition rates. Hence, we provide a detailed description of the available ALR systems based on end-to-end DL architectures and identify a tendency to focus on the modeling of temporal context as the key to advance the field. Such modeling is dominated by recurrent neural networks due to their ability to retain context at multiple scales (e.g. short- and long-term information). In this sense, current efforts tend toward techniques that allow a more comprehensive modeling and interpretability of the retained context. [ABSTRACT FROM AUTHOR]
Copyright of Image & Vision Computing is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 131729291
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Survey on automatic lip-reading in the era of deep learning.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Fernandez-Lopez%2C+Adriana%22">Fernandez-Lopez, Adriana</searchLink><relatesTo>1</relatesTo><i> adriana.fernandez@upf.edu</i><br /><searchLink fieldCode="AR" term="%22Sukno%2C+Federico+M%2E%22">Sukno, Federico M.</searchLink><relatesTo>1</relatesTo><i> federico.sukno@upf.edu</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Image+%26+Vision+Computing%22">Image & Vision Computing</searchLink>. Oct2018, Vol. 78, p53-72. 20p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Braille+education%22">Braille education</searchLink><br /><searchLink fieldCode="DE" term="%22Reading+mobile+apps%22">Reading mobile apps</searchLink><br /><searchLink fieldCode="DE" term="%22Reading+standards%22">Reading standards</searchLink><br /><searchLink fieldCode="DE" term="%22Reading+strategies%22">Reading strategies</searchLink><br /><searchLink fieldCode="DE" term="%22Reading+motivation%22">Reading motivation</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Abstract In the last few years, there has been an increasing interest in developing systems for Automatic Lip-Reading (ALR). Similarly to other computer vision applications, methods based on Deep Learning (DL) have become very popular and have permitted to substantially push forward the achievable performance. In this survey, we review ALR research during the last decade, highlighting the progression from approaches previous to DL (which we refer to as traditional) toward end-to-end DL architectures. We provide a comprehensive list of the audio-visual databases available for lip-reading, describing what tasks they can be used for, their popularity and their most important characteristics, such as the number of speakers, vocabulary size, recording settings and total duration. In correspondence with the shift toward DL, we show that there is a clear tendency toward large-scale datasets targeting realistic application settings and large numbers of samples per class. On the other hand, we summarize, discuss and compare the different ALR systems proposed in the last decade, separately considering traditional and DL approaches. We address a quantitative analysis of the different systems by organizing them in terms of the task that they target (e.g. recognition of letters or digits and words or sentences) and comparing their reported performance in the most commonly used datasets. As a result, we find that DL architectures perform similarly to traditional ones for simpler tasks but report significant improvements in more complex tasks, such as word or sentence recognition, with up to 40% improvement in word recognition rates. Hence, we provide a detailed description of the available ALR systems based on end-to-end DL architectures and identify a tendency to focus on the modeling of temporal context as the key to advance the field. Such modeling is dominated by recurrent neural networks due to their ability to retain context at multiple scales (e.g. short- and long-term information). In this sense, current efforts tend toward techniques that allow a more comprehensive modeling and interpretability of the retained context. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Image & Vision Computing is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=131729291
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1016/j.imavis.2018.07.002
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 20
        StartPage: 53
    Subjects:
      – SubjectFull: Braille education
        Type: general
      – SubjectFull: Reading mobile apps
        Type: general
      – SubjectFull: Reading standards
        Type: general
      – SubjectFull: Reading strategies
        Type: general
      – SubjectFull: Reading motivation
        Type: general
    Titles:
      – TitleFull: Survey on automatic lip-reading in the era of deep learning.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Fernandez-Lopez, Adriana
      – PersonEntity:
          Name:
            NameFull: Sukno, Federico M.
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 10
              Text: Oct2018
              Type: published
              Y: 2018
          Identifiers:
            – Type: issn-print
              Value: 02628856
          Numbering:
            – Type: volume
              Value: 78
          Titles:
            – TitleFull: Image & Vision Computing
              Type: main
ResultId 1