NResNet: nested residual network based on channel and frequency domain attention mechanism for speaker verification in classroom.

Saved in:
Bibliographic Details
Title: NResNet: nested residual network based on channel and frequency domain attention mechanism for speaker verification in classroom.
Authors: Zheng, Qiuyu1 (AUTHOR) qiuyu@mails.ccnu.edu.cn, Chen, Zengzhao1 (AUTHOR) zzchen@ccnu.edu.cn, Jiang, Xinxing1 (AUTHOR) xinxingj@mails.ccnu.edu.cn, Lin, Mengting1 (AUTHOR) linmengting@mails.ccnu.edu.cn, Wang, Mengke1 (AUTHOR) moco@mails.ccnu.edu.cn, Lu, Yuanyuan1 (AUTHOR) yuanyuan.lu@whxy.edu.cn
Source: Multimedia Tools & Applications. May2025, Vol. 84 Issue 15, p14235-14251. 17p.
Subjects: Artificial intelligence, Smart speakers, Student speech, Image processing, Error rates
Abstract: With the development of deep learning technology, the pattern of artificial intelligence in education has attracted more and more attention. However, most of the existing verbal interaction analysis methods utilized in the classroom are still in the semi-artificial stage, which lacks intelligence and normality. Therefore, we propose a nested residual network with multi-scale aggregation and speaker attention mechanism, which can distinguish the speech of teachers and students by identifying audio clips in the classroom. Thus, the teaching mode can be analyzed by the verbal interaction between teachers and students. However, the existing method of speaker verification cannot be adapted to the classroom scene, one reason is that the language environment is inconsistent, and the other is the difference in speaker distribution. Therefore, a deep multi-scale aggregation residual network model was proposed, which can ensure the validity of voiceprint information to the greatest extent. A speaker attention mechanism that includes channel-domain and frequency-domain information were introduced to obtain the differences in pronunciation habits and voiceprint amplitude of teachers and students. Experimental results demonstrate that the proposed method achieves outstanding performance with significant learning-capacity, outperforming the state-of-the-art methods. The proposed method obtained a 6.20% accuracy improvement over the compared methods with a 4.00% equal error rate improvement on the English public dataset LibriSpeech. In order to adapt to Chinese classroom, we also proved that the proposed method has good cross-language adaptability through training performance on the Chinese dataset AISHELL. The Experimental results in Chinese classroom shown that the proposed method got a highest improvement 22.70% than other. Our project will be publicly available at http://ecourse.nercel.com. [ABSTRACT FROM AUTHOR]
Copyright of Multimedia Tools & Applications is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: egs
DbLabel: Engineering Source
An: 185424321
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: NResNet: nested residual network based on channel and frequency domain attention mechanism for speaker verification in classroom.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Zheng%2C+Qiuyu%22">Zheng, Qiuyu</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> qiuyu@mails.ccnu.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Chen%2C+Zengzhao%22">Chen, Zengzhao</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> zzchen@ccnu.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Jiang%2C+Xinxing%22">Jiang, Xinxing</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> xinxingj@mails.ccnu.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Lin%2C+Mengting%22">Lin, Mengting</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> linmengting@mails.ccnu.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Wang%2C+Mengke%22">Wang, Mengke</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> moco@mails.ccnu.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Lu%2C+Yuanyuan%22">Lu, Yuanyuan</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> yuanyuan.lu@whxy.edu.cn</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Multimedia+Tools+%26+Applications%22">Multimedia Tools & Applications</searchLink>. May2025, Vol. 84 Issue 15, p14235-14251. 17p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Artificial+intelligence%22">Artificial intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Smart+speakers%22">Smart speakers</searchLink><br /><searchLink fieldCode="DE" term="%22Student+speech%22">Student speech</searchLink><br /><searchLink fieldCode="DE" term="%22Image+processing%22">Image processing</searchLink><br /><searchLink fieldCode="DE" term="%22Error+rates%22">Error rates</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: With the development of deep learning technology, the pattern of artificial intelligence in education has attracted more and more attention. However, most of the existing verbal interaction analysis methods utilized in the classroom are still in the semi-artificial stage, which lacks intelligence and normality. Therefore, we propose a nested residual network with multi-scale aggregation and speaker attention mechanism, which can distinguish the speech of teachers and students by identifying audio clips in the classroom. Thus, the teaching mode can be analyzed by the verbal interaction between teachers and students. However, the existing method of speaker verification cannot be adapted to the classroom scene, one reason is that the language environment is inconsistent, and the other is the difference in speaker distribution. Therefore, a deep multi-scale aggregation residual network model was proposed, which can ensure the validity of voiceprint information to the greatest extent. A speaker attention mechanism that includes channel-domain and frequency-domain information were introduced to obtain the differences in pronunciation habits and voiceprint amplitude of teachers and students. Experimental results demonstrate that the proposed method achieves outstanding performance with significant learning-capacity, outperforming the state-of-the-art methods. The proposed method obtained a 6.20% accuracy improvement over the compared methods with a 4.00% equal error rate improvement on the English public dataset LibriSpeech. In order to adapt to Chinese classroom, we also proved that the proposed method has good cross-language adaptability through training performance on the Chinese dataset AISHELL. The Experimental results in Chinese classroom shown that the proposed method got a highest improvement 22.70% than other. Our project will be publicly available at http://ecourse.nercel.com. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Multimedia Tools & Applications is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=185424321
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1007/s11042-024-19588-9
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 17
        StartPage: 14235
    Subjects:
      – SubjectFull: Artificial intelligence
        Type: general
      – SubjectFull: Smart speakers
        Type: general
      – SubjectFull: Student speech
        Type: general
      – SubjectFull: Image processing
        Type: general
      – SubjectFull: Error rates
        Type: general
    Titles:
      – TitleFull: NResNet: nested residual network based on channel and frequency domain attention mechanism for speaker verification in classroom.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Zheng, Qiuyu
      – PersonEntity:
          Name:
            NameFull: Chen, Zengzhao
      – PersonEntity:
          Name:
            NameFull: Jiang, Xinxing
      – PersonEntity:
          Name:
            NameFull: Lin, Mengting
      – PersonEntity:
          Name:
            NameFull: Wang, Mengke
      – PersonEntity:
          Name:
            NameFull: Lu, Yuanyuan
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 21
              M: 05
              Text: May2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 13807501
          Numbering:
            – Type: volume
              Value: 84
            – Type: issue
              Value: 15
          Titles:
            – TitleFull: Multimedia Tools & Applications
              Type: main
ResultId 1