A Machine Learning Approach for Estimating Person Counts Using Anonymous WiFi Data in a University Library

preprint OA: closed
View at publisher

Abstract

Accurately estimating indoor occupancy is essential for managing building spaces and infrastructure, with applications ranging from ensuring safe distancing and adequate ventilation during health crises to optimizing energy use and resource allocation. Howev-er, no existing technology simultaneously provides accurate, low-cost, and priva-cy-preserving indoor occupancy measurement. This work explores the use of existing WiFi infrastructure as a non-intrusive sensing system, where access points act as soft sensors by passively collecting anonymized connection metadata as proxies for human presence. We validated the approach in a university library over eight months, training supervised machine learning regression models on WiFi data and comparing predictions against computer-vision ground truth. The best-performing models (SVR, Ridge, and MLP) consistently achieved R² ≈ 0.95, with mean absolute errors of about 8 persons and relative errors (SMAPE) below 10% at medium-to-high occupancies. Tree-based ensem-bles, particularly XGBoost, exhibited weaker generalization at extreme capacity ranges, likely due to data sparsity and sensitivity to hyperparameters. Importantly, no temporal degradation was observed across the 8-month horizon, confirming the long-term stability of the method. Overall, the results demonstrate that WiFi-based occupancy estimation can provide a robust, low-cost, and privacy-preserving solution for real-world deployments.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00