Context, Perception, Production: A Model of Vocal Persona
preprint
OA: closed
CC-BY-4.0
Abstract
We present a contextualized production-perception model of vocal persona based on a deductive thematic analysis of interviews with voice and performance experts. The model formalizes how the vocal persona frames and bounds expressive vocal interactions (both biological and synthesized), and centers a person's agency over their communicative role. This article provides insights into opportunities for improvement in Voice User Interfaces (VUI) and Augmentative \& Assistive Communications (AAC) technologies based on the study's results. The proposed contextualized production-perception model fills an important gap in the literature on expressive and interactive speech technologies by identifying a key missing mechanism in contextualized vocal communication models and providing a resource for more nuanced approaches to expressive speech synthesis methods. Incorporating vocal persona into expressive vocal synthesis has the potential to significantly enhance the level of agency and embodiment experienced by VUI and AAC users during communication, resulting in a heightened sense of authenticity and an improved relationship with their environment.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.
Source provenance
- europepmc
- last seen: 2026-05-19T01:45:01.086888+00:00
- unpaywall
- last seen: 2026-06-02T02:00:03.124865+00:00
License: CC-BY-4.0