Local Time-Frequency Feature Fusion Using Cross-Attention for Acoustic Scene Classification
preprint
OA: closed
Abstract
To address the interdependence of local time-frequency information in audio scene recognition, a segment-based time-frequency feature fusion method based on Cross-Attention is proposed. Since audio scene recognition is highly sensitive to individual sound events within a scene, the input features are segmented into multiple segments along the time dimension to obtain local features, allowing the subsequent attention mechanism to focus on the time slices of key sound events. Furthermore, to leverage the advantages of both Convolutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs), which are mainstream structures in audio scene recognition tasks, this paper employs a parallel structure to separately obtain the time-frequency features output by CNNs and RNNs, and then fuses the two sets of features using Cross-Attention. Experiments on the TUT2018, TAU2019, and TAU2020 datasets demonstrate that the performance of this algorithm improves the official baseline results by 17.78%, 15.95%, and 20.13%, respectively.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.
Source provenance
- europepmc
- last seen: 2026-05-20T01:45:00.602351+00:00