Fréchet Gesture Distance
Released AESKConv_240_100 features over 32-frame-aligned SMPL-X rotations with the official covariance regularization.
Audio-conditioned co-speech motion generation on the fixed BEAT2 English speaker-2 protocol. Every ranked row is generated and evaluated inside Motius; paper-reported values are shown only as parity targets.
High-resolution, camera-controllable GT, Language of Motion, EMAGE, and DiffuseStyleGesture+ previews across all 15 clips. Full native SMPL-X predictions remain downloadable as NPZ.
Lower is better for FGD and paired errors; higher is better for BC and Diversity.
| Method | Status | Official BEAT2 metrics | Joint-only uTMR | Paired diagnostics | |||||
|---|---|---|---|---|---|---|---|---|---|
| FGD ↓ | BC ↑ | Diversity ↑ | FID ↓ | Paired Dist. ↓ | Rotation ↓ | Expression ↓ | Translation ↓ | ||
One immutable population, official metric assets, and explicit checkpoint provenance.
Released AESKConv_240_100 features over 32-frame-aligned SMPL-X rotations with the official covariance regularization.
Audio onsets against upper-body velocity minima, excluding the official 60-frame margin at each end.
Mean absolute deviation of generated 55-joint positions after the official SMPL-X metric forward kinematics.
Universal SMPL-H joints66 motion embeddings over balanced ≤20-second windows. FID uses per-sample L2 normalization; paired distance and diversity use native embeddings.
Paper-reported values are audit targets and never participate in ranking.
| Row | FGD ↓ | BC ↑ | Diversity ↑ |
|---|