Abstract
Text-conditioned human motion generation enables synthesizing realistic humanoid movements directly from natural language descriptions for applications in animation, virtual agents, and robot motion planning. Although diffusion-based models have recently achieved strong performance in this area, objective comparison across methods is still challenging due to differences in preprocessing, normalization conventions, and evaluation protocols. This paper presents a unified and controlled comparative study for three representative text-to-motion diffusion models: MDM, MotionDiffuse, and ReMoDiffuse, evaluated on the HumanML3D dataset under identical conditions. Alongside standard generative metrics such as FID, Diversity, Multi-Modality, and Multi-Modal Distance, we incorporate two physically motivated indicators, Average Jerk and Foot Sliding, to assess kinematic smoothness and contact stability. In addition to global evaluation, a six-category per-action analysis covering locomotion and basic object interactions is performed, which reveals task-specific behaviors that remain hidden in aggregate scores. The results provide a consistent comparison of current text-to-motion diffusion models and offer insights into their realism, diversity, semantic alignment, and physical plausibility. This benchmarking framework establishes a reliable reference point for future research in text-driven human motion generation and humanoid control.
| Original language | English |
|---|---|
| Title of host publication | 2026 IEEE Conference on Artificial Intelligence, CAI 2026 |
| Publisher | Institute of Electrical and Electronics Engineers Inc. |
| Pages | 1140-1145 |
| Number of pages | 6 |
| ISBN (Electronic) | 9798331560393 |
| DOIs | |
| Publication status | Published - 2026 |
| Event | 4th IEEE Conference on Artificial Intelligence, CAI 2026 - Granada, Spain Duration: 8 May 2026 → 10 May 2026 |
Publication series
| Name | 2026 IEEE Conference on Artificial Intelligence, CAI 2026 |
|---|
Conference
| Conference | 4th IEEE Conference on Artificial Intelligence, CAI 2026 |
|---|---|
| Country/Territory | Spain |
| City | Granada |
| Period | 8/05/26 → 10/05/26 |
Bibliographical note
Publisher Copyright:© 2026 IEEE.
Fingerprint
Dive into the research topics of 'A Comparative Study of Text-Driven Diffusion Models for Generative Human-Like Motion'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver