Exploring CrossFit Performance Prediction and Analysis via Extensive Data and Machine Learning

preprint OA: closed
View at publisher

Abstract

(1) Background: The analysis of athletic performance has always aroused great interest from sport scientist. This study utilized machine learning methods to build predictive models using a comprehensive CrossFit (CF) dataset, aiming to reveal valuable insights into the factors influencing performance and emerging trends.; (2) Methods: The study used Random Forest (RF) and Multiple Linear Regression (MLR) models to predict performance in four key weightlifting exercises within CF: clean & jerk, snatch, back squat, and deadlift. Performance was evaluated using R-squared (R2) values and Mean Squared Error (MSE). Feature importance analysis was conducted using RF, XGBoost, and AdaBoost models.; (3) Results: The RF model excelled in deadlift performance prediction (R2 = 0.80), while the MLR model demonstrated remarkable accuracy in clean & jerk (R2 = 0.93). Across exercises, clean & jerk consistently emerged as a crucial predictor. The feature importance analysis revealed intricate relationships among exercises, with gender significantly impacting deadlift performance.; (4) Conclusions: This research advances our understanding of performance prediction in CF through machine learning techniques. It provides actionable insights for practitioners, optimize performance, and demonstrates the potential for future advancements in data-driven sports analytics.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00