SARCASM IN SYNTHETIC SPEECH: EVALUATING ELEVENLABS’ TTS PROSODY FOR PRAGMATICS INSTRUCTION

Authors

  • M. Faisal Afif Zawawi Universitas Negeri Malang
  • Evynurul Laily Zen Universitas Negeri Malang
  • Nurenzia Yannuar Universitas Negeri Malang

DOI:

https://doi.org/10.22460/eltin.v14i2.p359-372

Keywords:

Sarcasm, Prosody, Text-to-speech, Acoustic features, Pragmatics

Abstract

This study examines how sarcasm is acoustically signaled by text-to-speech (TTS) and human encoders by analyzing key acoustic parameters including mean fundamental frequency (F0), F0 SD, F0 range, mean amplitude, amplitude range, and speech rate. The study compares sarcasm encoding in ElevenLabs, a commercially available advanced TTS model, and native American English speakers. A total of twelve encoders participated in the study, consisting of six TTS voices from ElevenLabs and six native American English speakers recruited through Fiverr. The dataset comprises 240 recordings, with each encoder generating both sincere and sarcastic versions of the same sentences. Acoustic parameters were extracted using Praat and computed using descriptive statistics and Linear Mixed Models (LMMs). The results reveal that sarcasm is consistently associated with reduced pitch, pitch variability, and slower speech rate, while amplitude does not differentiate significantly between attitudes. Subsequently, differences between TTS and human encoders are not uniform, which indicates that the TTS model approximates general prosodic patterns but differs in the magnitude of how these features are signaled. Overall, the findings suggest that TTS presents a partial approximation of human sarcastic prosody, capturing general acoustic patterns while remaining limited in expressive variability in sarcastic contexts.

References

Aini, N., & Basthomi, Y. (2025). Integration of Artificial Intelligence (AI) in learning English writing in higher education. Journal of Learning for Development, 12(2), 364–371. https://doi.org/10.56059/jl4d.v12i2.1596

Al-Jarf, R. (2022). Text-To-Speech software for promoting EFL freshman students’ decoding skills and pronunciation accuracy. Journal of Computer Science and Technology Studies, 4(2), 19–30. https://doi.org/10.32996/jcsts.2022.4.2.4

Amaliah, S., Wahyuni, I. Y., Murdia, M., & Wahid, A. (2025). AI-assisted Podcast Creation in EFL Learning: Enhancing Speaking Fluency and Reducing Foreign Language Anxiety among Gen Z Learners. Klasikal Journal of Education Language Teaching and Science, 7(2), 1174–1189. https://doi.org/10.52208/klasikal.v7i2.1566

Amin, E. A. (2024). EFL Students’ Perception of Using AI Text-to-Speech Apps in Learning Pronunciation. SSRN Electronic Journal. https://doi.org/10.2139/ssrn.4746800

Anolli, L., Ciceri, R., & Infantino, M. G. (2000). Irony as a game of implicitness: acoustic profiles of ironic communication. Journal of Psycholinguistic Research, 29(3), 275–311. https://doi.org/10.1023/a:1005100221723

Bione, T., Grimshaw, J., & Cardoso, W. (2017). An evaluation of TTS as a pedagogical tool for pronunciation instruction: the ‘foreign’ language context. In K. Borthwick, L. Bradley & S. Thouësny (Eds), CALL in a climate of change: adapting to turbulent global conditions – short papers from EUROCALL 2017 (pp. 56-61). Research-publishing.net. https://doi.org/10.14705/rpnet.2017.eurocall2017.689

Blasko, D. G., Kazmerski, V. A., & Dawood, S. S. (2021). Saying what you don’t mean: A cross-cultural study of perceptions of sarcasm. Canadian Journal of Experimental Psychology/Revue Canadienne De Psychologie Expérimentale, 75(2), 114–119. https://doi.org/10.1037/cep0000258

Boersma, P., & Weenink, D. (2025). Praat: doing phonetics by computer [Computer program]. Version 6.4.39, retrieved 13 July 2025 from https://praat.org

Budianto, L., Zulaiha, S., Ar-Rahma, A. P., & Addiyaullami, N. D. (2024). Students' responses on ElevenLabs AI for English-speaking learning. In Proceeding International Seminar on Islamic Education and Peace, 4, 80–93.

Cardoso, W. (2018). Learning L2 pronunciation with a text-to-speech synthesizer. In P. Taalas, J. Jalkanen, L. Bradley & S. Thouësny (Eds), Future-proof CALL: language learning as exploration and encounters – short papers from EUROCALL 2018 (pp. 16-21). Research-publishing.net. https://doi.org/10.14705/rpnet.2018.26.806

Cheang, H. S., & Pell, M. D. (2008). The sound of sarcasm. Speech Communication, 50(5), 366–381. https://doi.org/10.1016/j.specom.2007.11.003

Cheang, H. S., & Pell, M. D. (2009). Acoustic markers of sarcasm in Cantonese and English. The Journal of the Acoustical Society of America, 126(3), 1394–1405. https://doi.org/10.1121/1.3177275

Chiang, H. (2019). A Comparison between Teacher-Led and Online Text-to-Speech Dictation for Students’ Vocabulary Performance. English Language Teaching, 12(3), 77. https://doi.org/10.5539/elt.v12n3p77

Chuan, C. W., Pau, L. C., Hui, O. Y., & Yunus, M. M. (2025). Leveraging AI-Generated Audio in The Metaverse for Enhanced English Listening Skills: An Innovative Approach. Zenodo (CERN European Organization for Nuclear Research). https://doi.org/10.5281/zenodo.14854932

Creswell, J.W. (2014) Research Design: Qualitative, Quantitative and Mixed Methods Approaches. Sage Publications Ltd.

Dabkowski, P., & Staniszewski, M. (2025, June 3). Introducing Eleven v3 (alpha). ElevenLabs. Retrieved April 2, 2026, from https://elevenlabs.io/blog/eleven-v3

Dewatri, R. A. F., Al Aqthar, A. Z., Pradana, H., Anugerah, B., & Nurcahyo, W. H. (2023). Potential tools to support learning: OpenAI and ElevenLabs integration. ODELIA: Southeast Asia Journal on Open and Distance Learning, 1(2), 59–69.

Dress, M. L., Kreuz, R. J., Link, K. E., & Caucci, G. M. (2008). Regional variation in the use of sarcasm. Journal of Language and Social Psychology, 27(1), 71–85. https://doi.org/10.1177/0261927x07309512

Field, A.P. (2018) Discovering Statistics Using IBM SPSS Statistics. 5th Edition, Sage, Newbury Park.

Floris, F. D., Widiati, U., Renandya, W. A., & Basthomi, Y. (2024). Artificial intelligence in English Language Teaching: Fostering joint enterprise in Online communities. JEES (Journal of English Educators Society), 9(1), 12–21. https://doi.org/10.21070/jees.v9i1.1825

Gibbs, R. W., Jr. (2000). Irony in talk among friends. Metaphor and Symbol, 15(1-2), 5–27. https://doi.org/10.1207/S15327868MS151&2_2

Haverkate, H. (1990). A speech act analysis of irony. Journal of Pragmatics, 14(1), 77–109. https://doi.org/10.1016/0378-2166(90)90065-L

Hastomo, T., Sari, A. S., Widiati, U., Ivone, F. M., Zen, E. L., & Andianto, A. (2025a). Exploring EFL teachers’ strategies in employing AI chatbots in writing instruction to enhance student engagement. World Journal of English Language, 15(7), 93. https://doi.org/10.5430/wjel.v15n7p93

Hastomo, T., Sari, A. S., Widiati, U., Ivone, F. M., Zen, E. L., & Kholid, M. F. N. (2025b). Does Student Engagement with Chatbots Enhance English Proficiency? ELOPE English Language Overseas Perspectives and Enquiries, 22(1), 93–109. https://doi.org/10.4312/elope.22.1.93-109

Hastomo, T., Widiati, U., Ivone, F. M., & Zen, E. L. (2026). Integrating Generative AI in Primary English Material Design: Insights from Indonesian Teachers’ Perceptions. Revija Za Elementarno Izobraževanje, 19(1). https://doi.org/10.18690/rei5409

Hastomo, T., Widiati, U., Ivone, F. M., Zen, E. L., Hasbi, M., & Khulel, B. (2025c). AI-powered conversational agents and intercultural learning: Insights from Indonesian EFL students. Intercultural Communication Education, 8(1), 103217. https://doi.org/10.29140/ice.v8n1.103127

Inworld AI. (2026, February 13). Best AI Voice Generators for Realistic, Low-Latency TTS (2026 Comparison + Benchmarks). Retrieved April 9, 2026, from https://inworld.ai/resources/best-ai-voice-generators?utm_source=chatgpt.com

Jiang, W. (2024). Synthesis of sarcastic speech: Research on adjusting pitch and energy at keyword level using FastSpeech 2 (Master’s thesis). University of Groningen – Campus Fryslan

Kamal, A., & Imran, M. C. (2026). Integrating GenAI in creating digital storytelling as STEAM-based English tenses material on social media. IDEAS: Journal of Language Teaching and Learning, Linguistics and Literature, 14(1), 225–241. https://doi.org/10.24256/ideas.v14i1.8481

Koh, J., Lee, S., & Lee, J. M. (2022). L2 pragmatic comprehension of aural sarcasm: Tone, context, and literal meaning. System, 105, 102724. https://doi.org/10.1016/j.system.2022.102724

Kreuz, R. J., & Glucksberg, S. (1989). How to be sarcastic: The echoic reminder theory of verbal irony. Journal of Experimental Psychology: General, 118(4), 374–386. https://doi.org/10.1037/0096-3445.118.4.374

Lameris, H., Szekely, E., & Gustafson, J. (2024). The Role of Creaky Voice in Turn Taking and the Perception of Speaker Stance: Experiments Using Controllable TTS. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 16058–16065, Torino, Italia. ELRA and ICCL.

Le, P. N., Vu, H. M. L., & Tran, M. N. (2022). Improving EFL Students’ Intonation in Text Using Shadowing Technique with The Implementation of Google Text-to-Speech. AsiaCALL Online Journal, 13(1), 93121.EOI:http://eoi.citefactor.org/10.11251/acoj.13.01.006

Li, Z., Gao, X., Nayak, S., & Coler, M. (2023). SarcasticSpeech: Speech Synthesis for Sarcasm in Low-Resource Scenarios. In Proceedings 12th ISCA Speech Synthesis Workshop (SSW2023) (pp. 242-243). ISCA.

Lœvenbruck, H., Jannet, M.A.B., D'Imperio, M., Spini, M., Champagne-Lavau, M. (2013) Prosodic cues of sarcastic speech in French: slower, higher, wider.In Proceedings Interspeech 2013, 3537-3541, doi: 10.21437/Interspeech.2013-761

Lubis, F. K., & Bahri, S. (2023). Sarcasm in Indonesian television show Pesbukers. Randwick International of Social Science Journal, 4(1), 91–99. https://doi.org/10.47175/rissj.v4i1.626

Mubarok, A. F., Salim, M. A., & Deliany, Z. (2025). AI Based Applications to Foster Pronunciation Accuracy. Essence Journal of English Language Teaching Linguistics and Literature, 2(1). https://doi.org/10.33367/essence.v2i1.7479

Niebuhr, O. (2014) “A little more ironic” Voice quality and segmental reduction differences between sarcastic and neutral utterances. in Proceedings Speech Prosody 2014, 608-612, doi: 10.21437/SpeechProsody.2014-110

Nuzhnyy, S. (2026, April 8). Best Text-to-Speech AI 2026: Top picks & In-depth reviews. Retrieved April 9, 2026, from https://aimlapi.com/blog/best-text-to-speech-ai?utm_source=chatgpt.com

Oktalia, D., & Drajatı, N. A. (2018). English teachers’ perceptions of text to speech software and Google site in an EFL Classroom: What English teachers really think and know. International Journal of Education and Development Using Information and Communication Technology, 14(3), 183–192. http://files.eric.ed.gov/fulltext/EJ1201565.pdf

Önder, F. (2025). The effectiveness of online Text-To-Speech tools in improving EFL students’ pronunciation. Journal of Educational Studies and Multidisciplinary Approaches, 5(1). https://doi.org/10.51383/jesma.2025.116

Putri, S., GE, H., Hart, A., Yip, V., & Chen, A. (2019). The effect of explicit training on comprehension of English focus-to-prosody mapping by Indonesian learners of English. In Proceedings of the 19th International Congress of Phonetic Sciences, 1937-1941.

Rahardi, R. K., Handoko, H., Rahmat, W., & Setyaningsih, Y. (2024). Javanese silly gags on daily communication on social media: Pragmatic Meanings and Functions approach. JURNAL ARBITRER, 11(1), 49–59. https://doi.org/10.25077/ar.11.1.49-59.2024

Rao, R. (2013). Prosodic Consequences of Sarcasm Versus Sincerity in Mexican Spanish. Concentric: Studies in Linguistics, 39(2), 33-59. https://doi.org/10.6241/concentric.ling.39.2.02

Rockwell, P. (2000). Lower, Slower, Louder: Vocal Cues of Sarcasm. Journal of Psycholinguistic Research, 29(5), 483–495. https://doi.org/10.1023/a:1005120109296

Rockwell, P., & Theriot, E. M. (2001). Culture, gender, and gender mix in encoders of sarcasm: A self‐assessment analysis. Communication Research Reports, 18(1), 44–52. https://doi.org/10.1080/08824090109384781

Rockwell, P. (2007). Vocal features of Conversational Sarcasm: A comparison of methods. Journal of Psycholinguistic Research, 36(5), 361–369. https://doi.org/10.1007/s10936-006-9049-0

Rotaeta Perez, P. (2025). Can AI speak like us? An analysis of Accent Replication in Speech Synthesis. In Independent Degree Project at Undergraduate Level [Thesis].

Syafruddin, S., Thaba, A., Rahim, A. R., Munirah, M., & Syahruddin, S. (2021). Indonesian people’s sarcasm culture: an ethnolinguistic research. Linguistics and Culture Review, 5(1), 160–179. https://doi.org/10.21744/lingcure.v5n1.1150

Van Duong, T. (2022). The Effects of using Online Text-To Speech Tools on EFL Students’ Perceptions in Learning Pronunciation. International Journal of Science and Management Studies (IJSMS), 270–276. https://doi.org/10.51386/25815946/ijsms-v5i4p129

Xiao, Y. (2025). The impact of AI-driven speech recognition on EFL listening comprehension, flow experience, and anxiety: a randomized controlled trial. Humanities and Social Sciences Communications, 12(1). https://doi.org/10.1057/s41599-025-04672-8

Yang, S. (2021). Listener’s ratings and acoustic analyses of voice qualities associated with English and Korean sarcastic utterances. Speech Communication, 129, 1–6. https://doi.org/10.1016/j.specom.2021.02.002

Downloads

Published

2026-09-02