Two-level hierarchical alignment for semi-coupled HMM-based audiovisual emotion recognition with temporal course

Chung Hsien Wu, Jen Chun Lin, Wen Li Wei

Research output: Contribution to journalArticlepeer-review

38 Citations (Scopus)


A complete emotional expression typically contains a complex temporal course in face-to-face natural conversation. To address this problem, a bimodal hidden Markov model (HMM)-based emotion recognition scheme, constructed in terms of sub-emotional states, which are defined to represent temporal phases of onset, apex, and offset, is adopted to model the temporal course of an emotional expression for audio and visual signal streams. A two-level hierarchical alignment mechanism is proposed to align the relationship within and between the temporal phases in the audio and visual HMM sequences at the model and state levels in a proposed semi-coupled hidden Markov model (SC-HMM). Furthermore, by integrating a sub-emotion language model, which considers the temporal transition between sub-emotional states, the proposed two-level hierarchical alignment-based SC-HMM (2H-SC-HMM) can provide a constraint on allowable temporal structures to determine an optimal emotional state. Experimental results show that the proposed approach can yield satisfactory results in both the posed MHMC and the naturalistic SEMAINE databases, and shows that modeling the complex temporal structure is useful to improve the emotion recognition performance, especially for the naturalistic database (i.e., natural conversation). The experimental results also confirm that the proposed 2H-SC-HMM can achieve an acceptable performance for the systems with sparse training data or noisy conditions.

Original languageEnglish
Article number6542683
Pages (from-to)1880-1895
Number of pages16
JournalIEEE Transactions on Multimedia
Issue number8
Publication statusPublished - 2013 Dec 2

All Science Journal Classification (ASJC) codes

  • Signal Processing
  • Media Technology
  • Computer Science Applications
  • Electrical and Electronic Engineering

Fingerprint Dive into the research topics of 'Two-level hierarchical alignment for semi-coupled HMM-based audiovisual emotion recognition with temporal course'. Together they form a unique fingerprint.

Cite this