Resnet-conformer network with shared weights and attention mechanism for sound event localization, detection, and distance estimation

Quoc Thinh Vo; David Han

doi:10.48550/arxiv.2507.17941

Back

Resnet-conformer network with shared weights and attention mechanism for sound event localization, detection, and distance estimation

Preprint

Open access

Resnet-conformer network with shared weights and attention mechanism for sound event localization, detection, and distance estimation

Quoc Thinh Vo and David Han

23 Jul 2025

DOI: https://doi.org/10.48550/arxiv.2507.17941

Files and links (1)

url

https://arxiv.org/pdf/2507.17941View

Open

Abstract

Computer Science - Sound

This technical report outlines our approach to Task 3A of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2024, focusing on Sound Event Localization and Detection (SELD). SELD provides valuable insights by estimating sound event localization and detection, aiding in various machine cognition tasks such as environmental inference, navigation, and other sound localization-related applications. This year's challenge evaluates models using either audio-only (Track A) or audiovisual (Track B) inputs on annotated recordings of real sound scenes. A notable change this year is the introduction of distance estimation, with evaluation metrics adjusted accordingly for a comprehensive assessment. Our submission is for Task A of the Challenge, which focuses on the audio-only track. Our approach utilizes log-mel spectrograms, intensity vectors, and employs multiple data augmentations. We proposed an EINV2-based [1] network architecture, achieving improved results: an F-score of 40.2%, Angular Error (DOA) of 17.7 degrees, and Relative Distance Error (RDE) of 0.32 on the test set of the Development Dataset [2 ,3].

Metrics

10 Record Views

Details

Title: Resnet-conformer network with shared weights and attention mechanism for sound event localization, detection, and distance estimation
Creators: Quoc Thinh Vo
David Han
Resource Type: Preprint
Language: English
Academic Unit: Electrical and Computer Engineering
Other Identifier: 991022065153404721

Resnet-conformer network with shared weights and attention mechanism for sound event localization, detection, and distance estimation

Files and links (1)

Abstract

Metrics

Details

Drexel University Social media