Cross-modal enhancement of speech representations via textual supervision for paralinguistic analysis

Parra-Gallego LF, Magimai.-Doss M, Orozco-Arroyave JR (2026)


Publication Type: Journal article

Publication year: 2026

Journal

Book Volume: 183

Article Number: 103447

DOI: 10.1016/j.specom.2026.103447

Abstract

Speech conveys both what is said (linguistic content) and how it is said (acoustic cues). However, most automatic unimodal systems model these dimensions separately, which is insufficient because each carries fundamentally different information. Multimodal fusion strategies address this limitation by leveraging the complementary strengths of both modalities. However, these systems require automatic transcripts in the language modality, which increases system complexity and raises privacy concerns. To overcome these limitations, this study builds upon Speech-to-BERT (Sunder et al., 2022), a framework that transfers linguistic knowledge from a BERT-based model into speech representations during training. As a result, linguistically informed features can be extracted directly from audio without requiring transcripts at inference time. We extend the original Speech-to-BERT study by systematically evaluating multiple pretrained speech encoders (Wav2Vec2, HuBERT, WavLM, and Whisper) and analyzing the impact of task-specific teacher specialization. We further validate its effectiveness across four paralinguistic tasks: customer satisfaction (CS), dementia assessment (DA), spoken-intent recognition (SIR), and speech emotion recognition (SER). This comprehensive evaluation provides new empirical insights into how linguistic knowledge transfer interacts with modern speech encoders. Results show improvements over strong speech-only baselines, with the magnitude of gains varying across tasks. The approach maintains low computational cost and reduces reliance on external ASR systems. Our analyses revealed that encoders with natural word- and phoneme-level modeling, particularly Whisper, HuBERT, and WavLM, benefited the most from language supervision. These results demonstrate that such linguistic knowledge can be distilled into speech representations without fine-tuning large speech encoders. This provides an efficient approach that reduces reliance on external ASR systems and limits data exposure during inference for paralinguistic analysis across diverse speech-based tasks.

Involved external institutions

How to cite

APA:

Parra-Gallego, L.F., Magimai.-Doss, M., & Orozco-Arroyave, J.R. (2026). Cross-modal enhancement of speech representations via textual supervision for paralinguistic analysis. Speech Communication, 183. https://doi.org/10.1016/j.specom.2026.103447

MLA:

Parra-Gallego, Luis Felipe, Mathew Magimai.-Doss, and Juan Rafael Orozco-Arroyave. "Cross-modal enhancement of speech representations via textual supervision for paralinguistic analysis." Speech Communication 183 (2026).

BibTeX: Download