Abstract
Healthy-to-Pathological Voice Conversion (H2P-VC) is a vital area of Voice Conversion (VC) that serves as a tool for clinical decision support and a data augmentation strategy to improve dysarthric speech recognition systems. However, pathological speech often exhibits high acoustic variability and atypical articulatory patterns, while limited data availability poses challenges for model development. To address these issues, we introduce a phonetics-based feature derived from vowel space compression, an acoustic marker of articulatory degradation, and incorporate it into a source-filter-based synthesis architecture. Specifically, we construct a vowel space representation using corner vowel distributions and integrate it into the model to guide severity-dependent variation. This enables controllable generation of pathological speech while preserving articulatory and phonatory abnormalities across severity levels. Experiments on the UASpeech dataset show that the generated speech closely matches real dysarthric speech in intelligibility, speaker identity, severity progression, and vowel space structure. Ablation studies further confirm the importance of vowel space information in modeling speech degradation. To our knowledge, this is the first work to incorporate vowel space compression into pathological speech modeling, offering a structured and interpretable framework for future research in speech disorder analysis.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Audio, Speech and Language Processing |
| DOIs | |
| Publication status | Accepted/In press - 2026 |
All Science Journal Classification (ASJC) codes
- Acoustics and Ultrasonics
- Electrical and Electronic Engineering
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver