Article detail · 2026
A Comparative Study of Text-Driven Diffusion Models for Generative Human-Like Motion
- Year
- 2026
- Type
- conference-paper
Abstract
OpenAlex · English
Text-conditioned human motion generation enables synthesizing realistic humanoid movements directly from natural language descriptions for applications in animation, virtual agents, and robot motion planning. Although diffusion-based models have recently achieved strong performance in this area, objective comparison across methods is still challenging due to differences in preprocessing, normalization conventions, and evaluation protocols. This paper presents a unified and controlled comparative study for three representative text-to-motion diffusion models: MDM, MotionDiffuse, and ReMoDiffuse, evaluated on the HumanML3D dataset under identical conditions. Alongside standard generative metrics such as FID, Diversity, Multi-Modality, and Multi-Modal Distance, we incorporate two physically motivated indicators, Average Jerk and Foot Sliding, to assess kinematic smoothness and contact stability. In addition to global evaluation, a six-category per-action analysis covering locomotion and basic object interactions is performed, which reveals task-specific behaviors that remain hidden in aggregate scores. The results provide a consistent comparison of current text-to-motion diffusion models and offer insights into their realism, diversity, semantic alignment, and physical plausibility. This benchmarking framework establishes a reliable reference point for future research in text-driven human motion generation and humanoid control.
Topics
Citations
OpenAlex cited_by_count. Not a WoS or Scopus citation count; those sources have no separate column here.
0 citations
OpenAlex cited_by_count (cache / database)