NResNet: nested residual network based on channel and frequency domain attention mechanism for speaker verification in classroom.
Saved in:
| Title: | NResNet: nested residual network based on channel and frequency domain attention mechanism for speaker verification in classroom. |
|---|---|
| Authors: | Zheng, Qiuyu1 (AUTHOR) qiuyu@mails.ccnu.edu.cn, Chen, Zengzhao1 (AUTHOR) zzchen@ccnu.edu.cn, Jiang, Xinxing1 (AUTHOR) xinxingj@mails.ccnu.edu.cn, Lin, Mengting1 (AUTHOR) linmengting@mails.ccnu.edu.cn, Wang, Mengke1 (AUTHOR) moco@mails.ccnu.edu.cn, Lu, Yuanyuan1 (AUTHOR) yuanyuan.lu@whxy.edu.cn |
| Source: | Multimedia Tools & Applications. May2025, Vol. 84 Issue 15, p14235-14251. 17p. |
| Subjects: | Artificial intelligence, Smart speakers, Student speech, Image processing, Error rates |
| Abstract: | With the development of deep learning technology, the pattern of artificial intelligence in education has attracted more and more attention. However, most of the existing verbal interaction analysis methods utilized in the classroom are still in the semi-artificial stage, which lacks intelligence and normality. Therefore, we propose a nested residual network with multi-scale aggregation and speaker attention mechanism, which can distinguish the speech of teachers and students by identifying audio clips in the classroom. Thus, the teaching mode can be analyzed by the verbal interaction between teachers and students. However, the existing method of speaker verification cannot be adapted to the classroom scene, one reason is that the language environment is inconsistent, and the other is the difference in speaker distribution. Therefore, a deep multi-scale aggregation residual network model was proposed, which can ensure the validity of voiceprint information to the greatest extent. A speaker attention mechanism that includes channel-domain and frequency-domain information were introduced to obtain the differences in pronunciation habits and voiceprint amplitude of teachers and students. Experimental results demonstrate that the proposed method achieves outstanding performance with significant learning-capacity, outperforming the state-of-the-art methods. The proposed method obtained a 6.20% accuracy improvement over the compared methods with a 4.00% equal error rate improvement on the English public dataset LibriSpeech. In order to adapt to Chinese classroom, we also proved that the proposed method has good cross-language adaptability through training performance on the Chinese dataset AISHELL. The Experimental results in Chinese classroom shown that the proposed method got a highest improvement 22.70% than other. Our project will be publicly available at http://ecourse.nercel.com. [ABSTRACT FROM AUTHOR] |
| Copyright of Multimedia Tools & Applications is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Text: Availability: 1 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 185424321 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: NResNet: nested residual network based on channel and frequency domain attention mechanism for speaker verification in classroom. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Zheng%2C+Qiuyu%22">Zheng, Qiuyu</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> qiuyu@mails.ccnu.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Chen%2C+Zengzhao%22">Chen, Zengzhao</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> zzchen@ccnu.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Jiang%2C+Xinxing%22">Jiang, Xinxing</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> xinxingj@mails.ccnu.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Lin%2C+Mengting%22">Lin, Mengting</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> linmengting@mails.ccnu.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Wang%2C+Mengke%22">Wang, Mengke</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> moco@mails.ccnu.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Lu%2C+Yuanyuan%22">Lu, Yuanyuan</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> yuanyuan.lu@whxy.edu.cn</i> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Multimedia+Tools+%26+Applications%22">Multimedia Tools & Applications</searchLink>. May2025, Vol. 84 Issue 15, p14235-14251. 17p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Artificial+intelligence%22">Artificial intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Smart+speakers%22">Smart speakers</searchLink><br /><searchLink fieldCode="DE" term="%22Student+speech%22">Student speech</searchLink><br /><searchLink fieldCode="DE" term="%22Image+processing%22">Image processing</searchLink><br /><searchLink fieldCode="DE" term="%22Error+rates%22">Error rates</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: With the development of deep learning technology, the pattern of artificial intelligence in education has attracted more and more attention. However, most of the existing verbal interaction analysis methods utilized in the classroom are still in the semi-artificial stage, which lacks intelligence and normality. Therefore, we propose a nested residual network with multi-scale aggregation and speaker attention mechanism, which can distinguish the speech of teachers and students by identifying audio clips in the classroom. Thus, the teaching mode can be analyzed by the verbal interaction between teachers and students. However, the existing method of speaker verification cannot be adapted to the classroom scene, one reason is that the language environment is inconsistent, and the other is the difference in speaker distribution. Therefore, a deep multi-scale aggregation residual network model was proposed, which can ensure the validity of voiceprint information to the greatest extent. A speaker attention mechanism that includes channel-domain and frequency-domain information were introduced to obtain the differences in pronunciation habits and voiceprint amplitude of teachers and students. Experimental results demonstrate that the proposed method achieves outstanding performance with significant learning-capacity, outperforming the state-of-the-art methods. The proposed method obtained a 6.20% accuracy improvement over the compared methods with a 4.00% equal error rate improvement on the English public dataset LibriSpeech. In order to adapt to Chinese classroom, we also proved that the proposed method has good cross-language adaptability through training performance on the Chinese dataset AISHELL. The Experimental results in Chinese classroom shown that the proposed method got a highest improvement 22.70% than other. Our project will be publicly available at http://ecourse.nercel.com. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Multimedia Tools & Applications is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=185424321 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1007/s11042-024-19588-9 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 17 StartPage: 14235 Subjects: – SubjectFull: Artificial intelligence Type: general – SubjectFull: Smart speakers Type: general – SubjectFull: Student speech Type: general – SubjectFull: Image processing Type: general – SubjectFull: Error rates Type: general Titles: – TitleFull: NResNet: nested residual network based on channel and frequency domain attention mechanism for speaker verification in classroom. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Zheng, Qiuyu – PersonEntity: Name: NameFull: Chen, Zengzhao – PersonEntity: Name: NameFull: Jiang, Xinxing – PersonEntity: Name: NameFull: Lin, Mengting – PersonEntity: Name: NameFull: Wang, Mengke – PersonEntity: Name: NameFull: Lu, Yuanyuan IsPartOfRelationships: – BibEntity: Dates: – D: 21 M: 05 Text: May2025 Type: published Y: 2025 Identifiers: – Type: issn-print Value: 13807501 Numbering: – Type: volume Value: 84 – Type: issue Value: 15 Titles: – TitleFull: Multimedia Tools & Applications Type: main |
| ResultId | 1 |