The Alignment Between Language Properties and Computational Algorithms Enhances Statistical Word Segmentation: Evidence from Korean Child-Directed Speech
preprint
OA: closed
Abstract
This study tests the hypothesis that language input directed at young children exhibits enhanced segmentability by examining the interplay between language properties in different speech registers and the performance of segmentation algorithms. Employing a speaker-matched corpus of naturalistic child-directed speech (CDS) and adult-directed speech (ADS) in Korean, we observed that Korean CDS is characterized by shorter utterances and words, lower lexical diversity, fewer hapax legomena and interjections, a more child-like nature, a higher frequency of one-word utterances, and greater lexical ambiguity than ADS. We applied multiple computational algorithms to assess the impact of these linguistic properties on word segmentation performance in both registers. The results show a significantly higher word segmentation F-score for CDS than ADS, suggesting that child-oriented linguistic adaptations in CDS facilitate easier segmentation. This observation is further supported by statistical modeling, which indicates that the enhanced segmentability in CDS is modulated by the linguistic properties of the register. Nuanced implications of the linguistic properties and their implications on the performance of segmentation algorithms are discussed.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.
Source provenance
- europepmc
- last seen: 2026-05-20T01:45:00.602351+00:00