PROCEEDINGS OF THE 2026 ACM MULTIMEDIA SYSTEM CONFERENCE, MMSYS 2026
Abstract
Current Virtual Reality (VR) systems rely on inside-out visual-inertial tracking, which enables accurate localization but provides only a partial representation of the user's body. This limitation restricts embodiment and interaction fidelity in interactive VR scenarios requiring full-body awareness and expressive gestures. To capture both global body motion and fine-grained interaction cues within a single sensing framework, we introduce MultiSenseVR, the first open multimodal dataset that jointly captures synchronized millimeter-wave (mmWave) Wi-Fi, Surface Electromyography (sEMG), inertial signals, and high-precision 3D motion capture for ground truth in an immersive VR setting. The dataset includes recordings from 24 participants interacting with a custom fast-food simulation designed to elicit natural full-body movement. In addition to objective sensing data, MultiSenseVR provides subjective measures of presence and cybersickness. Baseline evaluations show that mmWave Wi-Fi sensing supports 3D pose estimation with accuracy comparable to camera-based approaches, while sEMG enables accurate subject-specific grasp classification. The dataset and supporting code are publicly available at https://osf.io/f6r7d.