arXiv preprint · 2607.22973

mmSimPrior: Learning Simulation Priors for Data-Efficient Real-World Generalizable Radar-Based Human Motion Reconstruction

Cheng Guo1, Qiming Cao2, Shengkai Xu1, Haoyu Xie1, Kaixiang Su1, Pu Wang1, Hongfei Xue1

1University of North Carolina at Charlotte 2Purdue University

Learning transferable signal, motion, and mapping priors from large-scale simulation for real-world radar-based human motion reconstruction.

Classification for constrained zero-shot transfer Regression for flexible limited-data adaptation

Jumping Jack, Basketball Shot, Squat, and Throw. Radar inputs, ground-truth motion, and reconstructions from the complementary mmSimPrior-Cls and mmSimPrior-Reg modes.

Abstract

Millimeter-wave (mmWave) radar offers privacy-preserving and lighting-robust sensing for human motion reconstruction, but learning models that generalize across real deployments requires diverse paired radar–motion data that are costly to collect. Simulation provides scalable supervision, yet models trained on clean synthetic signals transfer poorly because of multipath, clutter, response statistics, and resolution degradation.

We present mmSimPrior, a simulation-pretrained framework that factorizes transferable knowledge into signal, motion, and radar-to-motion mapping priors. A multi-modal signal encoder is pretrained with a physics-informed domain-randomization curriculum that emulates propagation- and acquisition-level variations, while a joint-temporal tokenizer learns a discrete prior over plausible human motion. A shared mapping prior supports classification over a learned motion codebook for constrained zero-shot reconstruction and continuous regression for flexible limited-data adaptation.

We further construct a 4.2M-frame, 31K-sequence dataset suite and introduce a No-Overlap Setting that excludes repeated complete subject–environment–location–motion configurations across adaptation and test. Experiments on mmSimPrior-Real and RT-Pose demonstrate consistent gains: with only 24 paired real sequences, mmSimPrior-Reg reduces MPJPE by 24.7–39.0% over the strongest baseline across the three environments, while mmSimPrior-Cls reduces zero-shot MPJPE by 8.5% without finetuning.

Overview

mmSimPrior shifts paired supervision from scarce real recordings to large-scale simulated radar–motion data, then transfers the learned priors to real deployments.

Overview of mmSimPrior simulation priors, transfer modes, dataset scale, and No-Overlap Setting
Overview of mmSimPrior. Physics-informed domain randomization supports transferable signal, motion, and mapping priors for zero-shot reconstruction and motion-prior-guided real-world adaptation.

Minute-level Real-world Adaptation

With only 24 paired real sequences, approximately one minute of supervision, mmSimPrior-Reg reduces MPJPE by 27.0%, 24.7%, and 39.0% in Hall, Lab, and Studio relative to the strongest adapted baseline.

Zero-shot Transfer

Without target-benchmark adaptation, the frozen motion codebook constrains uncertain radar features toward plausible motion. On unseen environments, mmSimPrior-Cls improves MPJPE and PA-MPJPE by 8.5% and 7.5% over the strongest baseline.

Generalization to Held-out Factors

The No-Overlap Setting prevents any adaptation–test pair from sharing the same complete subject, environment, sensing location, and motion configuration. We separately isolate environment, subject-identity, and motion-category shifts below.

Held-out factor 01

Unseen Environment

Hall-to-Lab adaptation evaluates a deployment environment absent from finetuning, while zero-shot evaluation uses no target-benchmark adaptation. The shift introduces unseen clutter, multipath patterns, and reflective structures.

Held-out factor 02

Unseen Subject

Nine-fold leave-one-subject-out adaptation reserves one participant for testing while using 24 sequences from the remaining participants. Over 1,013 held-out-subject recordings, mmSimPrior-Reg reaches 72.6 mm MPJPE and both modes remain robust to unseen body geometry and motion style.

Held-out factor 03

Unseen Motion

Motion categories in this split are absent from the 24-sequence adaptation set. Across 454 recordings from Hall, Lab, and Studio, the simulation priors preserve action-specific articulation and temporal plausibility beyond the observed motion vocabulary.

External Validation on RT-Pose

On the public RT-Pose benchmark, mmSimPrior-Reg improves MPJPE, PA-MPJPE, and MPVPE over the strongest adapted baseline by 9.1 mm, 3.5 mm, and 12.6 mm, demonstrating transfer beyond the collected benchmark.

Comparison with Radar Baselines

Direct side-by-side comparisons include ground truth, mmSimPrior-Cls, mmSimPrior-Reg, mmDiff, mmGPE, and RAPTR. The four 96-frame locomotion clips expose temporal drift, global-orientation errors, and foot sliding.