EgoExoMoCap

Distributed Ego-Exo Human Motion Capture

ECCV 2026 Spotlight
Jiaxi Jiang1,2, Bharat Lal Bhatnagar1, Nan Yang1, Lingni Ma1, Sebastian Starke1, Robin Kips1, Nadine Bertsch1, Christian Holz2, and Federica Bogo1
1Meta Reality Labs2Department of Computer Science, ETH Zürich, Switzerland
EgoExoMoCap

EgoExoMoCap is a lightweight, distributed human motion capture system. Given two or more subjects wearing head-mounted devices, the approach combines continuous egocentric signals with intermittent exocentric camera views to robustly handle challenges such as out-of-view motions (a) and severe body occlusions (b).

Abstract

Human motion capture from head-mounted devices (HMDs) offers a scalable way to acquire real-world human motion and interaction data, which is crucial for applications in embodied AI and VR/AR. Existing approaches focus on either egocentric body tracking, estimating the motion of the subject wearing the device, or exocentric tracking, capturing the movements of people in the wearer’s surroundings. So far, these two paradigms have largely been explored in isolation. In this paper, we propose a novel distributed framework that jointly leverages ego- and exocentric multi-modal signals for human motion estimation from HMDs. Unlike traditional motion capture systems requiring bulky multi-camera setups or obtrusive mocap suits, our approach, EgoExoMoCap, is as simple as two (or more) people, each wearing a pair of smart glasses. The method leverages head (plus potentially wrist) tracking signals for accurate estimation of global motion in the 3D world and combines context-aware image features based on DINOv3 to achieve robustness in the presence of noise and occlusions. Extensive experiments on two in-the-wild datasets show that our approach can robustly reconstruct motion even in challenging scenarios.

Reference

Jiaxi Jiang, Bharat Lal Bhatnagar, Nan Yang, Lingni Ma, Sebastian Starke, Robin Kips, Nadine Bertsch, Christian Holz, and Federica Bogo. EgoExoMoCap: Distributed Ego-Exo Human Motion Capture. In European Conference on Computer Vision 2026 (ECCV).

BibTeX citation

@article{jiang2026egoexomocap, title={EgoExoMoCap: Distributed Ego-Exo Human Motion Capture}, author={Jiaxi Jiang and Bharat Lal Bhatnagar and Nan Yang and Lingni Ma and Sebastian Starke and Robin Kips and Nadine Bertsch and Christian Holz and Federica Bogo}, journal={arXiv preprint arXiv:2607.15868}, year={2026} }

Heading

method overview

Figure 2: Overview of our method. Given an egocentric and one or more exocentric streams from HMDs, we first roughly estimate 3D body poses from egocentric streams (EgoNet) to identify regions of interest in exocentric frames. From these, ViTPose-extracted 2D keypoints are unprojected into 3D rays and softly weighted by DINOv3-based confidence scores to form Exo Tokens. A Spatial Transformer fuses Ego and Exo tokens into View-Aggregated (VA) Tokens, followed by a Temporal Transformer for smoothness to output final full-body motions.

EgoExoMoCap Qualitative results

Figure 3: Qualitative comparison of EgoExoMoCap versus baselines. The first four rows compare diverse activities from Nymeria, while the last two rows are drawn from the EgoHumans dataset and show two players playing tennis together.