arXiv · 2510.07096
Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
Abstract
Sarcasm is a pragmatic phenomenon in which speakers convey meanings that diverge from literal content, relying on an interaction between semantics and prosodic expression. However, how these cues jointly contribute to the recognition of sarcasm remains poorly understood. We propose a computational framework that models sarcasm as the integration of semantic interpretation and prosodic realization. Semantic cues are derived from an LLaMA 3 model fine-tuned to capture discourse-level markers of sarcastic intent, while prosodic cues are extracted through semantically aligned utterances drawn from a database of sarcastic speech, providing prosodic exemplars of sarcastic delivery. Using a speech synthesis testbed, perceptual evaluations show that semantic and prosodic cues enhance perceived sarcasm, with the combined system achieving the best downstream F1 while maintaining high subjective sarcasm ratings. These findings highlight the complementary roles of semantics and prosody in pragmatic interpretation and illustrate how modeling can shed light on the mechanisms underlying sarcastic communication.
Explore related subjects
Keep this discovery
Zhu Li, Yuqing Zhang, Xiyuan Gao, Shekhar Nayak, Matt Coler. 2025-10-08. Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework. https://arxiv.org/abs/2510.07096
Cite the original work for its findings. Save a collection to share your selection of sources.