EgoExoMoCap
Distributed Ego-Exo Human Motion Capture
ECCV 2026 SpotlightAbstract
Human motion capture from head-mounted devices (HMDs) offers a scalable way to acquire real-world human motion and interaction data, which is crucial for applications in embodied AI and VR/AR. Existing approaches focus on either egocentric body tracking, estimating the motion of the subject wearing the device, or exocentric tracking, capturing the movements of people in the wearer’s surroundings. So far, these two paradigms have largely been explored in isolation. In this paper, we propose a novel distributed framework that jointly leverages ego- and exocentric multi-modal signals for human motion estimation from HMDs. Unlike traditional motion capture systems requiring bulky multi-camera setups or obtrusive mocap suits, our approach, EgoExoMoCap, is as simple as two (or more) people, each wearing a pair of smart glasses. The method leverages head (plus potentially wrist) tracking signals for accurate estimation of global motion in the 3D world and combines context-aware image features based on DINOv3 to achieve robustness in the presence of noise and occlusions. Extensive experiments on two in-the-wild datasets show that our approach can robustly reconstruct motion even in challenging scenarios.
Reference
Jiaxi Jiang, Bharat Lal Bhatnagar, Nan Yang, Lingni Ma, Sebastian Starke, Robin Kips, Nadine Bertsch, Christian Holz, and Federica Bogo. EgoExoMoCap: Distributed Ego-Exo Human Motion Capture. In European Conference on Computer Vision 2026 (ECCV).
BibTeX citation
@article{jiang2026egoexomocap, title={EgoExoMoCap: Distributed Ego-Exo Human Motion Capture}, author={Jiaxi Jiang and Bharat Lal Bhatnagar and Nan Yang and Lingni Ma and Sebastian Starke and Robin Kips and Nadine Bertsch and Christian Holz and Federica Bogo}, journal={arXiv preprint arXiv:2607.15868}, year={2026} }
Heading

Figure 2: Overview of our method. Given an egocentric and one or more exocentric streams from HMDs, we first roughly estimate 3D body poses from egocentric streams (EgoNet) to identify regions of interest in exocentric frames. From these, ViTPose-extracted 2D keypoints are unprojected into 3D rays and softly weighted by DINOv3-based confidence scores to form Exo Tokens. A Spatial Transformer fuses Ego and Exo tokens into View-Aggregated (VA) Tokens, followed by a Temporal Transformer for smoothness to output final full-body motions.

Figure 3: Qualitative comparison of EgoExoMoCap versus baselines. The first four rows compare diverse activities from Nymeria, while the last two rows are drawn from the EgoHumans dataset and show two players playing tennis together.






