Strahl S, Müller M (2026)
Publication Language: English
Publication Type: Conference contribution, Conference Contribution
Publication year: 2026
Pages Range: 16087-16091
Conference Proceedings Title: Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP)
Event location: Barcelona, Spain
DOI: 10.1109/ICASSP55912.2026.11463185
Fundamental frequency (F0) estimation is a key task in audio signal processing, traditionally addressed with digital signal processing (DSP) methods based on spectrogram or autocorrelation analysis. More recent deep learning approaches have improved accuracy and robustness, but often at the cost of high complexity and limited interpretability. We propose a hybrid approach that extracts mid-level features from classical DSP-based methods and fuses them using a neural network, thereby leveraging the strengths of model-based F0 estimators without relying on their hard decisions. Specifically, we use soft time–frequency representations derived from YIN, SWIPE, and the cepstrum alongside spectrograms. While spectrograms contain strong components at the F0 and higher harmonics, the other representations emphasize the F0 and subharmonics, thus providing complementary information. These features are then fused by a lightweight convolutional architecture with 6.5k trainable parameters. Cross-dataset experiments demonstrate that our method yields robust and accurate F0 estimates, achieving competitive performance compared to purely data-driven methods while largely preserving the interpretability of classical approaches.
APA:
Strahl, S., & Müller, M. (2026). Robust And Lightweight F0 Estimation Through Mid-Level Fusion of DSP-Informed Features. In Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) (pp. 16087-16091). Barcelona, Spain, ES.
MLA:
Strahl, Sebastian, and Meinard Müller. "Robust And Lightweight F0 Estimation Through Mid-Level Fusion of DSP-Informed Features." Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), Barcelona, Spain 2026. 16087-16091.
BibTeX: Download