Reusing ML Models in Dynamic Data Environments: A Data Similarity-Based Approach for Efficient MLOps

preprint OA: closed CC-BY-4.0
🔓 Open OA copy View at publisher

Abstract

The rapid integration of Machine Learning (ML) in organizational practices has driven demand for substantial computational resources, incurring both high economic costs and environmental impact, particularly from energy consumption. This challenge is amplified in dynamic data environments, where ML models must be frequently retrained to adapt to evolving data patterns. To address this, more sustainable MLOps pipelines are needed that reduce environmental impacts while maintaining model accuracy. In this paper, we propose a model reuse approach based on data similarity metrics, allowing organizations to leverage previously trained models where applicable. We introduce a tailored set of meta-features to characterize data windows, enabling efficient similarity assessment between historical and new data. The effectiveness of the proposed method is validated across multiple ML tasks using cosine and Bray-Curtis distance functions, evaluating both model reuse rates and the performance of reused models relative to newly trained alternatives. The results indicate that the proposed approach can significantly reduce the frequency of model retraining while maintaining or even improving predictive performance, contributing to more resource-efficient and sustainable MLOps practices.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-24T02:00:01.246996+00:00
License: CC-BY-4.0