GestureLSM: Latent Shortcut based Co-Speech Gesture Generation with Spatial-Temporal Modeling
Upload your audio. 2. Wait for the rendering to happen (1-4 minutes). 3. View the generated gesture video. 4. The face animation is fixed; only body motion is generated.