Tag
ReMoMask-2 improves text-to-motion generation by embedding retrieval directly into the generator's latent space, eliminating representation gaps and achieving state-of-the-art results on benchmarks like KIT-ML and SnapMoGen.
NVIDIA's text-to-motion model has been ported to run on CPU without requiring NVIDIA's stack through kimodo.cpp, a GGML/C++ implementation that is open-source and rapidly gaining traction.
PRISM is a 1.4 billion parameter text-to-motion model that can generate precise motion sequences from text commands, returning SMPL-X parameters for use in 3D rigs.
MUGEN introduces a unified motion-language framework that avoids discrete codebooks and iterative decoding, using a single adaptive-length autoencoder with continuous latent slots and one-shot generation to achieve efficient, high-quality text-to-motion and motion-to-text performance across HumanML3D and SnapMoGen benchmarks.