SARCASM IN SYNTHETIC SPEECH: EVALUATING ELEVENLABS’ TTS PROSODY FOR PRAGMATICS INSTRUCTION
DOI:
https://doi.org/10.22460/eltin.v14i2.p359-372Keywords:
Sarcasm, Prosody, Text-to-speech, Acoustic features, PragmaticsAbstract
This study examines how sarcasm is acoustically signaled by text-to-speech (TTS) and human encoders by analyzing key acoustic parameters including mean fundamental frequency (F0), F0 SD, F0 range, mean amplitude, amplitude range, and speech rate. The study compares sarcasm encoding in ElevenLabs, a commercially available advanced TTS model, and native American English speakers. A total of twelve encoders participated in the study, consisting of six TTS voices from ElevenLabs and six native American English speakers recruited through Fiverr. The dataset comprises 240 recordings, with each encoder generating both sincere and sarcastic versions of the same sentences. Acoustic parameters were extracted using Praat and computed using descriptive statistics and Linear Mixed Models (LMMs). The results reveal that sarcasm is consistently associated with reduced pitch, pitch variability, and slower speech rate, while amplitude does not differentiate significantly between attitudes. Subsequently, differences between TTS and human encoders are not uniform, which indicates that the TTS model approximates general prosodic patterns but differs in the magnitude of how these features are signaled. Overall, the findings suggest that TTS presents a partial approximation of human sarcastic prosody, capturing general acoustic patterns while remaining limited in expressive variability in sarcastic contexts.
References
Aini, N., & Basthomi, Y. (2025). Integration of Artificial Intelligence (AI) in learning English writing in higher education. Journal of Learning for Development, 12(2), 364–371. https://doi.org/10.56059/jl4d.v12i2.1596
Al-Jarf, R. (2022). Text-To-Speech software for promoting EFL freshman students’ decoding skills and pronunciation accuracy. Journal of Computer Science and Technology Studies, 4(2), 19–30. https://doi.org/10.32996/jcsts.2022.4.2.4
Amaliah, S., Wahyuni, I. Y., Murdia, M., & Wahid, A. (2025). AI-assisted Podcast Creation in EFL Learning: Enhancing Speaking Fluency and Reducing Foreign Language Anxiety among Gen Z Learners. Klasikal Journal of Education Language Teaching and Science, 7(2), 1174–1189. https://doi.org/10.52208/klasikal.v7i2.1566
Amin, E. A. (2024). EFL Students’ Perception of Using AI Text-to-Speech Apps in Learning Pronunciation. SSRN Electronic Journal. https://doi.org/10.2139/ssrn.4746800
Anolli, L., Ciceri, R., & Infantino, M. G. (2000). Irony as a game of implicitness: acoustic profiles of ironic communication. Journal of Psycholinguistic Research, 29(3), 275–311. https://doi.org/10.1023/a:1005100221723
Bione, T., Grimshaw, J., & Cardoso, W. (2017). An evaluation of TTS as a pedagogical tool for pronunciation instruction: the ‘foreign’ language context. In K. Borthwick, L. Bradley & S. Thouësny (Eds), CALL in a climate of change: adapting to turbulent global conditions – short papers from EUROCALL 2017 (pp. 56-61). Research-publishing.net. https://doi.org/10.14705/rpnet.2017.eurocall2017.689
Blasko, D. G., Kazmerski, V. A., & Dawood, S. S. (2021). Saying what you don’t mean: A cross-cultural study of perceptions of sarcasm. Canadian Journal of Experimental Psychology/Revue Canadienne De Psychologie Expérimentale, 75(2), 114–119. https://doi.org/10.1037/cep0000258
Boersma, P., & Weenink, D. (2025). Praat: doing phonetics by computer [Computer program]. Version 6.4.39, retrieved 13 July 2025 from https://praat.org
Budianto, L., Zulaiha, S., Ar-Rahma, A. P., & Addiyaullami, N. D. (2024). Students' responses on ElevenLabs AI for English-speaking learning. In Proceeding International Seminar on Islamic Education and Peace, 4, 80–93.
Cardoso, W. (2018). Learning L2 pronunciation with a text-to-speech synthesizer. In P. Taalas, J. Jalkanen, L. Bradley & S. Thouësny (Eds), Future-proof CALL: language learning as exploration and encounters – short papers from EUROCALL 2018 (pp. 16-21). Research-publishing.net. https://doi.org/10.14705/rpnet.2018.26.806
Cheang, H. S., & Pell, M. D. (2008). The sound of sarcasm. Speech Communication, 50(5), 366–381. https://doi.org/10.1016/j.specom.2007.11.003
Cheang, H. S., & Pell, M. D. (2009). Acoustic markers of sarcasm in Cantonese and English. The Journal of the Acoustical Society of America, 126(3), 1394–1405. https://doi.org/10.1121/1.3177275
Chiang, H. (2019). A Comparison between Teacher-Led and Online Text-to-Speech Dictation for Students’ Vocabulary Performance. English Language Teaching, 12(3), 77. https://doi.org/10.5539/elt.v12n3p77
Chuan, C. W., Pau, L. C., Hui, O. Y., & Yunus, M. M. (2025). Leveraging AI-Generated Audio in The Metaverse for Enhanced English Listening Skills: An Innovative Approach. Zenodo (CERN European Organization for Nuclear Research). https://doi.org/10.5281/zenodo.14854932
Creswell, J.W. (2014) Research Design: Qualitative, Quantitative and Mixed Methods Approaches. Sage Publications Ltd.
Dabkowski, P., & Staniszewski, M. (2025, June 3). Introducing Eleven v3 (alpha). ElevenLabs. Retrieved April 2, 2026, from https://elevenlabs.io/blog/eleven-v3
Dewatri, R. A. F., Al Aqthar, A. Z., Pradana, H., Anugerah, B., & Nurcahyo, W. H. (2023). Potential tools to support learning: OpenAI and ElevenLabs integration. ODELIA: Southeast Asia Journal on Open and Distance Learning, 1(2), 59–69.
Dress, M. L., Kreuz, R. J., Link, K. E., & Caucci, G. M. (2008). Regional variation in the use of sarcasm. Journal of Language and Social Psychology, 27(1), 71–85. https://doi.org/10.1177/0261927x07309512
Field, A.P. (2018) Discovering Statistics Using IBM SPSS Statistics. 5th Edition, Sage, Newbury Park.
Floris, F. D., Widiati, U., Renandya, W. A., & Basthomi, Y. (2024). Artificial intelligence in English Language Teaching: Fostering joint enterprise in Online communities. JEES (Journal of English Educators Society), 9(1), 12–21. https://doi.org/10.21070/jees.v9i1.1825
Gibbs, R. W., Jr. (2000). Irony in talk among friends. Metaphor and Symbol, 15(1-2), 5–27. https://doi.org/10.1207/S15327868MS151&2_2
Haverkate, H. (1990). A speech act analysis of irony. Journal of Pragmatics, 14(1), 77–109. https://doi.org/10.1016/0378-2166(90)90065-L
Hastomo, T., Sari, A. S., Widiati, U., Ivone, F. M., Zen, E. L., & Andianto, A. (2025a). Exploring EFL teachers’ strategies in employing AI chatbots in writing instruction to enhance student engagement. World Journal of English Language, 15(7), 93. https://doi.org/10.5430/wjel.v15n7p93
Hastomo, T., Sari, A. S., Widiati, U., Ivone, F. M., Zen, E. L., & Kholid, M. F. N. (2025b). Does Student Engagement with Chatbots Enhance English Proficiency? ELOPE English Language Overseas Perspectives and Enquiries, 22(1), 93–109. https://doi.org/10.4312/elope.22.1.93-109
Hastomo, T., Widiati, U., Ivone, F. M., & Zen, E. L. (2026). Integrating Generative AI in Primary English Material Design: Insights from Indonesian Teachers’ Perceptions. Revija Za Elementarno Izobraževanje, 19(1). https://doi.org/10.18690/rei5409
Hastomo, T., Widiati, U., Ivone, F. M., Zen, E. L., Hasbi, M., & Khulel, B. (2025c). AI-powered conversational agents and intercultural learning: Insights from Indonesian EFL students. Intercultural Communication Education, 8(1), 103217. https://doi.org/10.29140/ice.v8n1.103127
Inworld AI. (2026, February 13). Best AI Voice Generators for Realistic, Low-Latency TTS (2026 Comparison + Benchmarks). Retrieved April 9, 2026, from https://inworld.ai/resources/best-ai-voice-generators?utm_source=chatgpt.com
Jiang, W. (2024). Synthesis of sarcastic speech: Research on adjusting pitch and energy at keyword level using FastSpeech 2 (Master’s thesis). University of Groningen – Campus Fryslan
Kamal, A., & Imran, M. C. (2026). Integrating GenAI in creating digital storytelling as STEAM-based English tenses material on social media. IDEAS: Journal of Language Teaching and Learning, Linguistics and Literature, 14(1), 225–241. https://doi.org/10.24256/ideas.v14i1.8481
Koh, J., Lee, S., & Lee, J. M. (2022). L2 pragmatic comprehension of aural sarcasm: Tone, context, and literal meaning. System, 105, 102724. https://doi.org/10.1016/j.system.2022.102724
Kreuz, R. J., & Glucksberg, S. (1989). How to be sarcastic: The echoic reminder theory of verbal irony. Journal of Experimental Psychology: General, 118(4), 374–386. https://doi.org/10.1037/0096-3445.118.4.374
Lameris, H., Szekely, E., & Gustafson, J. (2024). The Role of Creaky Voice in Turn Taking and the Perception of Speaker Stance: Experiments Using Controllable TTS. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 16058–16065, Torino, Italia. ELRA and ICCL.
Le, P. N., Vu, H. M. L., & Tran, M. N. (2022). Improving EFL Students’ Intonation in Text Using Shadowing Technique with The Implementation of Google Text-to-Speech. AsiaCALL Online Journal, 13(1), 93121.EOI:http://eoi.citefactor.org/10.11251/acoj.13.01.006
Li, Z., Gao, X., Nayak, S., & Coler, M. (2023). SarcasticSpeech: Speech Synthesis for Sarcasm in Low-Resource Scenarios. In Proceedings 12th ISCA Speech Synthesis Workshop (SSW2023) (pp. 242-243). ISCA.
Lœvenbruck, H., Jannet, M.A.B., D'Imperio, M., Spini, M., Champagne-Lavau, M. (2013) Prosodic cues of sarcastic speech in French: slower, higher, wider.In Proceedings Interspeech 2013, 3537-3541, doi: 10.21437/Interspeech.2013-761
Lubis, F. K., & Bahri, S. (2023). Sarcasm in Indonesian television show Pesbukers. Randwick International of Social Science Journal, 4(1), 91–99. https://doi.org/10.47175/rissj.v4i1.626
Mubarok, A. F., Salim, M. A., & Deliany, Z. (2025). AI Based Applications to Foster Pronunciation Accuracy. Essence Journal of English Language Teaching Linguistics and Literature, 2(1). https://doi.org/10.33367/essence.v2i1.7479
Niebuhr, O. (2014) “A little more ironic” Voice quality and segmental reduction differences between sarcastic and neutral utterances. in Proceedings Speech Prosody 2014, 608-612, doi: 10.21437/SpeechProsody.2014-110
Nuzhnyy, S. (2026, April 8). Best Text-to-Speech AI 2026: Top picks & In-depth reviews. Retrieved April 9, 2026, from https://aimlapi.com/blog/best-text-to-speech-ai?utm_source=chatgpt.com
Oktalia, D., & Drajatı, N. A. (2018). English teachers’ perceptions of text to speech software and Google site in an EFL Classroom: What English teachers really think and know. International Journal of Education and Development Using Information and Communication Technology, 14(3), 183–192. http://files.eric.ed.gov/fulltext/EJ1201565.pdf
Önder, F. (2025). The effectiveness of online Text-To-Speech tools in improving EFL students’ pronunciation. Journal of Educational Studies and Multidisciplinary Approaches, 5(1). https://doi.org/10.51383/jesma.2025.116
Putri, S., GE, H., Hart, A., Yip, V., & Chen, A. (2019). The effect of explicit training on comprehension of English focus-to-prosody mapping by Indonesian learners of English. In Proceedings of the 19th International Congress of Phonetic Sciences, 1937-1941.
Rahardi, R. K., Handoko, H., Rahmat, W., & Setyaningsih, Y. (2024). Javanese silly gags on daily communication on social media: Pragmatic Meanings and Functions approach. JURNAL ARBITRER, 11(1), 49–59. https://doi.org/10.25077/ar.11.1.49-59.2024
Rao, R. (2013). Prosodic Consequences of Sarcasm Versus Sincerity in Mexican Spanish. Concentric: Studies in Linguistics, 39(2), 33-59. https://doi.org/10.6241/concentric.ling.39.2.02
Rockwell, P. (2000). Lower, Slower, Louder: Vocal Cues of Sarcasm. Journal of Psycholinguistic Research, 29(5), 483–495. https://doi.org/10.1023/a:1005120109296
Rockwell, P., & Theriot, E. M. (2001). Culture, gender, and gender mix in encoders of sarcasm: A self‐assessment analysis. Communication Research Reports, 18(1), 44–52. https://doi.org/10.1080/08824090109384781
Rockwell, P. (2007). Vocal features of Conversational Sarcasm: A comparison of methods. Journal of Psycholinguistic Research, 36(5), 361–369. https://doi.org/10.1007/s10936-006-9049-0
Rotaeta Perez, P. (2025). Can AI speak like us? An analysis of Accent Replication in Speech Synthesis. In Independent Degree Project at Undergraduate Level [Thesis].
Syafruddin, S., Thaba, A., Rahim, A. R., Munirah, M., & Syahruddin, S. (2021). Indonesian people’s sarcasm culture: an ethnolinguistic research. Linguistics and Culture Review, 5(1), 160–179. https://doi.org/10.21744/lingcure.v5n1.1150
Van Duong, T. (2022). The Effects of using Online Text-To Speech Tools on EFL Students’ Perceptions in Learning Pronunciation. International Journal of Science and Management Studies (IJSMS), 270–276. https://doi.org/10.51386/25815946/ijsms-v5i4p129
Xiao, Y. (2025). The impact of AI-driven speech recognition on EFL listening comprehension, flow experience, and anxiety: a randomized controlled trial. Humanities and Social Sciences Communications, 12(1). https://doi.org/10.1057/s41599-025-04672-8
Yang, S. (2021). Listener’s ratings and acoustic analyses of voice qualities associated with English and Korean sarcastic utterances. Speech Communication, 129, 1–6. https://doi.org/10.1016/j.specom.2021.02.002
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
The author is responsible for acquiring the permission(s) to reproduce any copyrighted figures, tables, data, or text that are being used in the submitted paper. Authors should note that text quotations of more than 250 words from a published or copyrighted work will require grant of permission from the original publisher to reprint. The written permission letter(s) must be submitted together with the manuscript.