Exploring Temporal Context in Human Pose Estimation: Benefits of 2D vs. 3D Input Data

preprint OA: closed CC-BY-4.0
📄 Open PDF View at publisher

Abstract

The human body pose estimation is a challenging task that has been tackled by many different approaches over the last decades. The importance of solving this problem is rooted in an ability to automatically identify the pose of a subject in visual data. Even though numerous methods have been proposed to address the task, there is still ambiguity over what input data format is more efficient to use. Moreover, studies have shown that incorporating temporal information can be beneficial in one setting while leading to excessive computational complexity in others. In this paper, we investigate the effect of spatio-temporal context on the accuracy of the final estimations in end-to-end pose estimation methods, most importantly, whether it offers any additional value over the single-frame pose data. Furthermore, to address the question of the appropriate input data format, we compare experiments using the most widely used 2D and 3D visual input data structures.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-24T02:00:01.246996+00:00
License: CC-BY-4.0