Extracting extended vocal units from two neighborhoods in the embedding plane
preprint
OA: closed
Abstract
Annotating and proofreading data sets of complex natural behaviors are tedious tasks because instances of a given behavior need to be correctly segmented from background noise and must be classified with minimal false positive error rate. Low-dimensional embeddings have proven very useful for this task because they provide a visually appealing overview of a data set in which relevant clusters appear spontaneously. However, low-dimensional embeddings introduce errors because they fail to preserve high dimensional distances; and embeddings represent only objects of fixed dimensionality, which conflicts with natural objects such as vocalizations that have variable dimensions stemming from their variable durations. To mitigate these issues, we introduce a semi-supervised method for simultaneous segmentation and clustering of vocalizations. We define vocal units of a given type in terms of two density-based regions in low-dimensional embedding space, one associated with onsets and the other with offsets. We demonstrate our approach on the task of clustering adult zebra finch vocalizations embedded into the 2d plane with UMAP. We show that two-neighborhood (2N) extraction allows the identification of short and long vocal renditions from continuous data streams without initially committing to a particular segmentation of the data. Also, 2N vocal extraction achieves much lower false positive error rate than approaches based on a single defining region.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.
Source provenance
- europepmc
- last seen: 2026-05-19T01:45:01.086888+00:00