Collecting a High-Quality Dataset for Neural Text-to-Speech in the Kyrgyz Language

preprint OA: closed
View at publisher

Abstract

This work introduces the first thoroughly developed, high-quality Kyrgyz-language speech dataset intended for modern neural text-to-speech (TTS) systems. While significant progress has been made in speech synthesis for widely represented languages, Turkic low-resource languages such as Kyrgyz lack curated corpora suitable for high-fidelity neural models. To address this gap, we designed a full-scale dataset creation pipeline—from linguistic analysis and text preparation to professional recording, post-processing, and metadata structuring. The initial corpus consisted of approximately ten hours of raw studio audio, which was systematically reduced to four hours of polished, noise-free, transcription-aligned material. We present a detailed phonetic analysis of Kyrgyz, highlighting challenging vowels ``ү'', ``ө'', and the velar nasal ``ң''. A multi-stage audio refinement pipeline including silence trimming, segmentation, error removal, normalization, and quality control was applied. Three state-of-the-art TTS architectures (Tacotron~2, FastPitch, and VITS) were trained and evaluated using the dataset. Subjective MOS tests demonstrate that the resulting models achieve high naturalness, confirming the suitability of the dataset for neural synthesis research. This work provides both a methodology and a calibrated dataset framework to advance speech technologies in Kyrgyz.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00