Skip to main navigation Skip to search Skip to main content

Vowel Space-Guided DDSP for Severity-Preserved Healthy-to-Pathological Voice Conversion

Research output: Contribution to journalArticlepeer-review

Abstract

Healthy-to-Pathological Voice Conversion (H2P-VC) is a vital area of Voice Conversion (VC) that serves as a tool for clinical decision support and a data augmentation strategy to improve dysarthric speech recognition systems. However, pathological speech often exhibits high acoustic variability and atypical articulatory patterns, while limited data availability poses challenges for model development. To address these issues, we introduce a phonetics-based feature derived from vowel space compression, an acoustic marker of articulatory degradation, and incorporate it into a source-filter-based synthesis architecture. Specifically, we construct a vowel space representation using corner vowel distributions and integrate it into the model to guide severity-dependent variation. This enables controllable generation of pathological speech while preserving articulatory and phonatory abnormalities across severity levels. Experiments on the UASpeech dataset show that the generated speech closely matches real dysarthric speech in intelligibility, speaker identity, severity progression, and vowel space structure. Ablation studies further confirm the importance of vowel space information in modeling speech degradation. To our knowledge, this is the first work to incorporate vowel space compression into pathological speech modeling, offering a structured and interpretable framework for future research in speech disorder analysis.

Original languageEnglish
JournalIEEE Transactions on Audio, Speech and Language Processing
DOIs
Publication statusAccepted/In press - 2026

All Science Journal Classification (ASJC) codes

  • Acoustics and Ultrasonics
  • Electrical and Electronic Engineering

Cite this