Reverse-Engineering Speech and Music Categorization from a Single Sound Source

preprint OA: closed
View at publisher

Abstract

Classifying whether an auditory signal is music or speech is important for both humans and computational systems. Although previous literature suggests that music and speech are easily separable categories, common experimental approaches may bias findings toward this distinction by relying on stimuli from different sound sources and predefined response labels. Here, we use stimulus material from the dùndún drum–a speech surrogate that can signal either speech-related or musical content. We first replicate standard speech-music categorization results (N=108). Then, we depart from the typical experimental procedure by asking new participants (N=180) to sort and label the stimulus material, without predefined categories. Hierarchical clustering of participants’ stimulus groupings reveals multiple organizing dimensions, with the speech–music distinction reliably present but secondary under label-free conditions. By reverse-engineering the relationship between sorting behavior, acoustic features, and semantic labels, we characterize how speech–music categorization relates to other salient perceptual dimensions and how its behavioral prominence depends on task constraints.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00