A Flexible Semi-Synthetic Data Generator for Risky Drinking Behavior

preprint OA: closed
View at publisher

Abstract

Machine intelligence has garnered immense attention owing to its ability to discover hidden patterns in abstract and high-dimensional datasets. However, its success is often limited by the fundamental bottleneck of data scarcity. In this work, we offer a universal data augmentation solution to resolve this impasse. We first discovered the hidden knowledge within the existing scarce dataset using the machine learning (ML) technique and then synthetically augmented the dataset according to its feature importance. In principle, scarce and augmented datasets should share a common statistical property. Using this property, we specifically study the scarce dataset representing the binge-drinking behavior of university students and show that our method is effective in augmenting a limited dataset with high fidelity. The current work challenges the status quo in data scarcity with rule-less-based ML, which removes the ostensible barrier that prevents the application of data-driven techniques to the data scarce clinical research.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00