Preserving Data Distribution in Sampling and Instance Selection with Renyi's Divergence
preprint
OA: closed
CC-BY-4.0
Abstract
Abstract This paper presents a novel method for sampling based on a distribution function. By leveraging Renyi's divergence criterion, a recursive formulation is directly derived from the disparity between the original and estimated distributions. The Gradient descent algorithm is employed to achieve this recursive equation, and then new samples are derived. Since this method uses the distribution of the original data, it has the ability to preserve the data distribution in the sampled dataset. In instances where the original distribution is unknown, Kernel Density Estimation is employed for estimation. Experimental results demonstrate that the proposed method effectively preserves the data distribution and prevents any alteration in the data concept while maintaining the predictability of the model over selected instances. More precisely, this method has successfully reduced the dataset size by approximately $70\%$, while the accuracy of the learning algorithm has remained unaffected.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.
Source provenance
- europepmc
- last seen: 2026-05-19T01:45:01.086888+00:00
- unpaywall
- last seen: 2026-05-24T02:00:01.246996+00:00
License: CC-BY-4.0