Object and setting identification in natural auditory scenesa).

Saved in:
Bibliographic Details
Title: Object and setting identification in natural auditory scenesa).
Authors: McMullin, Margaret A.1 (AUTHOR) margaretmcmullin@ufl.edu, Higgins, Nathan C.2 (AUTHOR), Kumar, Rohit3 (AUTHOR), Elhilali, Mounya3 (AUTHOR), Snyder, Joel S.4 (AUTHOR) joel.snyder@unlv.edu
Source: Journal of the Acoustical Society of America. May2026, Vol. 159 Issue 5, p4722-4735. 14p.
Subjects: Auditory perception, Auditory scene analysis, Object recognition algorithms, Cognitive psychology
Abstract: Auditory scene perception allows listeners to identify both the setting and the objects within a scene, supporting decision-making and situational awareness. While visual scene and object recognition are well-studied, less is known about how listeners identify settings (e.g., forest) and objects (e.g., birds) in complex auditory environments. This study examined how scene duration influences listeners' ability to identify settings (e.g., café) and objects (e.g., music, talking, espresso machines) in natural auditory scenes. Participants listened to scenes of varying durations (1, 2, and 4 s) and reported the setting and objects present in each scene. Object identification was more accurate than setting identification across durations, but performance on both tasks benefited from increased durations. Several low- and mid-level acoustic features significantly predicted performance, although these models explained relatively little variance overall. These findings suggest that while acoustic structure contributes to performance, differences between object and setting identification likely reflect interacting perceptual and cognitive processes. Setting identification may depend more on integrating information across multiple sound sources and global scene properties (e.g., openness or naturalness), whereas object identification may rely more on segregation and recognition of individual sound sources and their features. [ABSTRACT FROM AUTHOR]
Copyright of Journal of the Acoustical Society of America is the property of American Institute of Physics and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Description
Abstract:Auditory scene perception allows listeners to identify both the setting and the objects within a scene, supporting decision-making and situational awareness. While visual scene and object recognition are well-studied, less is known about how listeners identify settings (e.g., forest) and objects (e.g., birds) in complex auditory environments. This study examined how scene duration influences listeners' ability to identify settings (e.g., café) and objects (e.g., music, talking, espresso machines) in natural auditory scenes. Participants listened to scenes of varying durations (1, 2, and 4 s) and reported the setting and objects present in each scene. Object identification was more accurate than setting identification across durations, but performance on both tasks benefited from increased durations. Several low- and mid-level acoustic features significantly predicted performance, although these models explained relatively little variance overall. These findings suggest that while acoustic structure contributes to performance, differences between object and setting identification likely reflect interacting perceptual and cognitive processes. Setting identification may depend more on integrating information across multiple sound sources and global scene properties (e.g., openness or naturalness), whereas object identification may rely more on segregation and recognition of individual sound sources and their features. [ABSTRACT FROM AUTHOR]
ISSN:00014966
DOI:10.1121/10.0043922