A Lightweight Hybrid Vision Transformer Network for Radar-based Human Activity Recognition
preprint
OA: closed
Abstract
Radar-based human activity recognition (HAR) offers a non-contact technique with privacy protection and lighting robustness for many advanced applications. The introduction of artificial intelligence (AI) technology encouraged the effort to develop deep learning (DL) for radar-based HAR. These methods employ the micro-Doppler signal because of its unique correspondence with human activity characteristics. Traditional methods give insufficient consideration of the HAR network's lightweight while focusing on recognition accuracy. In this paper, an efficient network based on a lightweight hybrid Vision Transformer (LH-ViT) is proposed to address the HAR accuracy and network lightweight simultaneously. Feature Pyramid architecture is applied for the multi-scale feature extraction for the micro-Doppler map. Feature enhancement is executed by the stacked Radar-ViT subsequently, in which the fold and unfold operations are added to lower the computational load of the attention mechanism. The convolution operator in the LH-ViT is replaced by the RES-SE block, an efficient structure that combines the residual learning framework with the Squeeze-and-Excitation network. Experiments based on two human activity datasets indicate our method's advantages in terms of expressiveness and computing efficiency over traditional methods.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.
Source provenance
- europepmc
- last seen: 2026-05-19T01:45:01.086888+00:00