Speech enhancement using neural free attention with multi-stage squeeze temporal convolutional networks.

Saved in:
Bibliographic Details
Title: Speech enhancement using neural free attention with multi-stage squeeze temporal convolutional networks.
Authors: Jannu, Chaitanya1 (AUTHOR) pvspj3@gmail.com, Burra, Manaswini2 (AUTHOR) manaswini.burra@gmail.com, Vanambathina, Sunny Dayal3 (AUTHOR) sunny.dayal@vitap.ac.in, Parisae, Veeraswamy1 (AUTHOR) veera2u@gmail.com
Source: Multimedia Tools & Applications. Mar2026, Vol. 85 Issue 3, p1-23. 23p.
Abstract: Speech enhancement is a fundamental task in many speech processing systems. In recent years, researchers have increasingly focused on boosting performance by capturing long-range contextual dependencies within speech signals. A widely adopted solution is multi-stage learning, where several deep learning components are arranged in sequence to refine the results step by step. Likewise, attention-based mechanisms have proven highly effective in improving speech quality, especially when combined with convolutional neural networks (CNNs). Nevertheless, most conventional attention designs rely on fully connected and convolutional operations, which significantly increase both the number of parameters and computational demands. To address this, the present study introduces a multi-stage speech enhancement framework that integrates Squeeze Temporal Convolutional Modules (STCM) with exponentially increasing dilation rates and a Neural-Free Attention (NFA) mechanism at each stage. At every phase, an intermediate estimate is generated and further refined in subsequent phases, with a Feature Fusion Module (FFM) reintroducing original information at the start of each stage. This design allows the intermediate outputs to undergo step-by-step improvements through successive STCMs, ultimately enabling precise spectral estimation. The NFA, a lightweight and easily integrable component, enhances the model’s ability to capture fine-grained energy distributions across frequency channels by generating attention weights with a trainable Gaussian function. The proposed system is evaluated on the VCTK and LibriSpeech datasets, showing superior performance compared to state-of-the-art deep learning methods in terms of PESQ, STOI, CSIG, CBAK, and COVL metrics. [ABSTRACT FROM AUTHOR]
Copyright of Multimedia Tools & Applications is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 191919750
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Speech enhancement using neural free attention with multi-stage squeeze temporal convolutional networks.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Jannu%2C+Chaitanya%22">Jannu, Chaitanya</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> pvspj3@gmail.com</i><br /><searchLink fieldCode="AR" term="%22Burra%2C+Manaswini%22">Burra, Manaswini</searchLink><relatesTo>2</relatesTo> (AUTHOR)<i> manaswini.burra@gmail.com</i><br /><searchLink fieldCode="AR" term="%22Vanambathina%2C+Sunny+Dayal%22">Vanambathina, Sunny Dayal</searchLink><relatesTo>3</relatesTo> (AUTHOR)<i> sunny.dayal@vitap.ac.in</i><br /><searchLink fieldCode="AR" term="%22Parisae%2C+Veeraswamy%22">Parisae, Veeraswamy</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> veera2u@gmail.com</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Multimedia+Tools+%26+Applications%22">Multimedia Tools & Applications</searchLink>. Mar2026, Vol. 85 Issue 3, p1-23. 23p.
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Speech enhancement is a fundamental task in many speech processing systems. In recent years, researchers have increasingly focused on boosting performance by capturing long-range contextual dependencies within speech signals. A widely adopted solution is multi-stage learning, where several deep learning components are arranged in sequence to refine the results step by step. Likewise, attention-based mechanisms have proven highly effective in improving speech quality, especially when combined with convolutional neural networks (CNNs). Nevertheless, most conventional attention designs rely on fully connected and convolutional operations, which significantly increase both the number of parameters and computational demands. To address this, the present study introduces a multi-stage speech enhancement framework that integrates Squeeze Temporal Convolutional Modules (STCM) with exponentially increasing dilation rates and a Neural-Free Attention (NFA) mechanism at each stage. At every phase, an intermediate estimate is generated and further refined in subsequent phases, with a Feature Fusion Module (FFM) reintroducing original information at the start of each stage. This design allows the intermediate outputs to undergo step-by-step improvements through successive STCMs, ultimately enabling precise spectral estimation. The NFA, a lightweight and easily integrable component, enhances the model’s ability to capture fine-grained energy distributions across frequency channels by generating attention weights with a trainable Gaussian function. The proposed system is evaluated on the VCTK and LibriSpeech datasets, showing superior performance compared to state-of-the-art deep learning methods in terms of PESQ, STOI, CSIG, CBAK, and COVL metrics. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Multimedia Tools & Applications is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=191919750
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1007/s11042-026-21428-x
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 23
        StartPage: 1
    Titles:
      – TitleFull: Speech enhancement using neural free attention with multi-stage squeeze temporal convolutional networks.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Jannu, Chaitanya
      – PersonEntity:
          Name:
            NameFull: Burra, Manaswini
      – PersonEntity:
          Name:
            NameFull: Vanambathina, Sunny Dayal
      – PersonEntity:
          Name:
            NameFull: Parisae, Veeraswamy
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 03
              Text: Mar2026
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 13807501
          Numbering:
            – Type: volume
              Value: 85
            – Type: issue
              Value: 3
          Titles:
            – TitleFull: Multimedia Tools & Applications
              Type: main
ResultId 1