GestureLSM: Latent Shortcut based Co-Speech Gesture Generation with Spatial-Temporal Modeling

  1. Upload your audio.
    2. Wait for the rendering to happen (1-4 minutes).
    3. View the generated gesture video.
    4. The face animation is fixed; only body motion is generated.
Examples

Project: GestureLSM | Paper: arXiv:2501.18898