Abstract
Sarcasm detection is essential for accurately interpreting communication in applications such as dialogue systems. However, most existing approaches treat sarcasm as a single phenomenon and ignore the linguistic distinction between illocutionary, embedded, and propositional sarcasm, each of which relies on different contextual and emotional cues. In this work, we propose a type-aware multimodal framework that explicitly models these three sarcasm types by training dedicated representations for each and combining them within a unified architecture. To encourage the learning of emotionally informative features, we incorporate explicit and implicit emotion prediction as auxiliary regularization tasks in a multitask learning setting. Experiments on the multimodal sarcasm dataset MUStARD++ show that our system achieves an F1 score of 82.65% in sarcasm detection. These results demonstrate that type-specific modeling improves multimodal sarcasm detection.
| Original language | English |
|---|---|
| Pages (from-to) | 3606-3617 |
| Number of pages | 12 |
| Journal | IEEE Access |
| Volume | 14 |
| DOIs | |
| Publication status | Published - 2026 |
All Science Journal Classification (ASJC) codes
- General Computer Science
- General Materials Science
- General Engineering
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver