Audio-vision Contrastive Learning for Phonological Class Recognition

Liu D, Arias Vergara T, Hutter J, Maier A, Perez Toro PA (2026)


Publication Type: Conference contribution, Abstract of lecture

Publication year: 2026

Journal

Publisher: Springer Science and Business Media Deutschland GmbH

Pages Range: 450-450

Conference Proceedings Title: Informatik aktuell

Event location: Lübeck DE

ISBN: 9783658510992

DOI: 10.1007/978-3-658-51100-5_88

Abstract

Real-time magnetic resonance imaging (rtMRI) enables detailed visualization of articulatory structures during speech production, making it invaluable for analysing articulatory-phonological features and advancing clinical speech technologies. While MRI captures the anatomical dynamics of articulation, concurrent audio signals provide complementary acoustic information that enhances temporal resolution in speech processing. Nevertheless, deriving meaningful phonological representations from rtMRI data remains difficult when audio signals are unavailable – situations that commonly arise during MRI scanning due to acoustic noise interference or in cases involving speech disorders such as those seen in glossectomy patients. To address this limitation, we propose a contrastive learning framework for automatically classifying three fundamental articulatory dimensions from MRI data: manner of articulation, place of articulation, and voicing. During training, paired MRI frames and speech segments are encoded separately using vision transformer (ViT) and Wav2Vec2 architectures, respectively, with contrastive learning employed to maximize cross-modal alignment between visual and acoustic representations. Critically, only MRI data is required during inference, enabling phonological classification without audio input. We evaluated four experimental configurations on the USC-TIMIT dataset: unimodal rtMRI, unimodal audio, multimodal middle fusion, and our contrastive learning-based approach. Results show that contrastive learning achieves state-of-the-art performance with an average F1-score of 0.81 across 15 phonological classes, representing absolute improvements of 0.23 over the unimodal baseline and 0.09 over multimodal fusion, thereby confirming the efficacy of cross-modal contrastive representation learning for MRI-based articulatory analysis when audio signals are unavailable [1].

Authors with CRIS profile

How to cite

APA:

Liu, D., Arias Vergara, T., Hutter, J., Maier, A., & Perez Toro, P.A. (2026). Audio-vision Contrastive Learning for Phonological Class Recognition. Paper presentation at Bildverarbeitung für die Medizin Workshop, BVM 2026, Lübeck, DE.

MLA:

Liu, Daiqi, et al. "Audio-vision Contrastive Learning for Phonological Class Recognition." Presented at Bildverarbeitung für die Medizin Workshop, BVM 2026, Lübeck Ed. Heinz Handels, Katharina Breininger, Thomas Deserno, Andreas Maier, Klaus Maier-Hein, Christoph Palm, Thomas Tolxdorff, Springer Science and Business Media Deutschland GmbH, 2026.

BibTeX: Download