Skip to main navigation Skip to search Skip to main content

A Comparative Study of Text-Driven Diffusion Models for Generative Human-Like Motion

  • Ali Ihsan Ozcetin*
  • , Burak Tantay
  • , Kadir Yavuz Kurt
  • , Hakan Temeltas
  • *Corresponding author for this work
  • Istanbul Technical University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Text-conditioned human motion generation enables synthesizing realistic humanoid movements directly from natural language descriptions for applications in animation, virtual agents, and robot motion planning. Although diffusion-based models have recently achieved strong performance in this area, objective comparison across methods is still challenging due to differences in preprocessing, normalization conventions, and evaluation protocols. This paper presents a unified and controlled comparative study for three representative text-to-motion diffusion models: MDM, MotionDiffuse, and ReMoDiffuse, evaluated on the HumanML3D dataset under identical conditions. Alongside standard generative metrics such as FID, Diversity, Multi-Modality, and Multi-Modal Distance, we incorporate two physically motivated indicators, Average Jerk and Foot Sliding, to assess kinematic smoothness and contact stability. In addition to global evaluation, a six-category per-action analysis covering locomotion and basic object interactions is performed, which reveals task-specific behaviors that remain hidden in aggregate scores. The results provide a consistent comparison of current text-to-motion diffusion models and offer insights into their realism, diversity, semantic alignment, and physical plausibility. This benchmarking framework establishes a reliable reference point for future research in text-driven human motion generation and humanoid control.

Original languageEnglish
Title of host publication2026 IEEE Conference on Artificial Intelligence, CAI 2026
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages1140-1145
Number of pages6
ISBN (Electronic)9798331560393
DOIs
Publication statusPublished - 2026
Event4th IEEE Conference on Artificial Intelligence, CAI 2026 - Granada, Spain
Duration: 8 May 202610 May 2026

Publication series

Name2026 IEEE Conference on Artificial Intelligence, CAI 2026

Conference

Conference4th IEEE Conference on Artificial Intelligence, CAI 2026
Country/TerritorySpain
CityGranada
Period8/05/2610/05/26

Bibliographical note

Publisher Copyright:
© 2026 IEEE.

Fingerprint

Dive into the research topics of 'A Comparative Study of Text-Driven Diffusion Models for Generative Human-Like Motion'. Together they form a unique fingerprint.

Cite this