<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">80992</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2026.080992</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Global-Local Embedding Gating Network for Part-Wise Text-to-Motion Generation</article-title>
<alt-title alt-title-type="left-running-head">Global-Local Embedding Gating Network for Part-Wise Text-to-Motion Generation</alt-title>
<alt-title alt-title-type="right-running-head">Global-Local Embedding Gating Network for Part-Wise Text-to-Motion Generation</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Kim</surname><given-names>Chanyoung</given-names></name></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Kim</surname><given-names>Jion</given-names></name></contrib>
<contrib id="author-3" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Shin</surname><given-names>Byeong-Seok</given-names></name><email>bsshin@inha.ac.kr</email></contrib>
<aff id="aff-1">
<institution>Department of Electrical and Computer Engineering, Inha University</institution>, <addr-line>Incheon</addr-line>, <country>Republic of Korea</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Byeong-Seok Shin. Email: <email>bsshin@inha.ac.kr</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>15</day><month>06</month><year>2026</year>
</pub-date>
<volume>88</volume>
<issue>2</issue>
<elocation-id>40</elocation-id>
<history>
<date date-type="received">
<day>20</day>
<month>02</month>
<year>2026</year>
</date>
<date date-type="accepted">
<day>20</day>
<month>04</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_80992.pdf"></self-uri>
<abstract>
<p>Diffusion-based methods have substantially improved the performance of full-body Text-to-Motion (T2M) generation from natural language descriptions. Despite this progress, accurately capturing the fine-grained semantics of composite prompts remains challenging. Approaches that rely solely on a single global text condition often fail to retain part-specific semantic cues, leading to deviations in the motions of certain body parts from the intended descriptions. Recent methods have attempted to address this by incorporating both global and local conditions, yet these are typically combined using fixed ratios or applied in separate stages, which restricts their adaptability to evolving semantic requirements during generation. To address these constraints, this work proposes the Embedding Gating Network (EGN), which dynamically modulates the contributions of global and local information according to the current noisy motion state and the diffusion timestep. By conditioning the gating mechanism on the intermediate noisy motion estimate, EGN adjusts the relative importance of global and local information to emphasize semantics that remain underrepresented at each denoising step. The conditioned signals are processed through independent part-wise generation pathways to minimize semantic interference, while a lightweight fusion module enables inter-part information exchange to preserve structural coherence across the full body. Experiments on the HumanML3D benchmark show that the proposed method consistently improves text-motion alignment over existing full-body and part-based baselines, without compromising motion quality or diversity. Analysis of the learned gating coefficients reveals that local conditions primarily contribute to the formation of part-wise structural outlines during early denoising stages, whereas global conditions become increasingly influential, integrating cross-part semantics and refining full-body consistency as denoising advances. These findings indicate that dynamically modulating conditioning signals during generation is an effective alternative to fixed-ratio conditioning.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Motion generation</kwd>
<kwd>diffusion model</kwd>
<kwd>human motion synthesis</kwd>
<kwd>text-to-motion</kwd>
<kwd>condition embedding</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Institute of Information &#x0026; Communications Technology Planning &#x0026; Evaluation (IITP) grant funded by the Korea government (MSIT)</funding-source>
<award-id>RS-2025-02214780</award-id>
</award-group>
<award-group id="awg2">
<funding-source>National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIT)</funding-source>
<award-id>RS-2026-25479030</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Text-to-Motion (T2M) synthesis aims to generate coherent and realistic 3D human motion sequences from natural language descriptions, serving as a foundational technology across virtual reality, animation, robotics, and gaming [<xref ref-type="bibr" rid="ref-1">1</xref>&#x2013;<xref ref-type="bibr" rid="ref-3">3</xref>]. The primary challenge arises from the intricate spatiotemporal structures of human motion and the inherent flexibility of natural language, both of which hinder precise semantic alignment [<xref ref-type="bibr" rid="ref-4">4</xref>&#x2013;<xref ref-type="bibr" rid="ref-6">6</xref>].</p>
<p>Recent advances in T2M research have been driven by diverse generative paradigms, including diffusion-based approaches [<xref ref-type="bibr" rid="ref-7">7</xref>,<xref ref-type="bibr" rid="ref-8">8</xref>], autoregressive token-based approaches [<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-10">10</xref>], and mask-based token modeling approaches [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-12">12</xref>]. These approaches have improved full-body motion quality, temporal coherence, and text-motion alignment, as well as the ability to process complex natural language descriptions. Nevertheless, higher generation quality does not guarantee fine-grained controllability. Most current methods condition on text using a single global embedding [<xref ref-type="bibr" rid="ref-7">7</xref>,<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-13">13</xref>], which often fails to preserve part-level semantics present in composite sentences. Consequently, although the generated motion may appear natural overall, the movements of specific body parts can diverge from the intended text. Recent approaches have sought to address this limitation through part-wise generation [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-14">14</xref>] or by employing finer-grained condition separation strategies [<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-16">16</xref>]. However, these methods typically combine global and part-level conditions with fixed contributions or treat part-wise generation and global integration as separate stages. Such designs implicitly assume that the relative importance of global context and part-level information remains constant throughout the generation process. In practice, when generating composite actions, the balance between maintaining overall motion structure and capturing part-specific semantics varies over time. Therefore, a conditioning mechanism capable of flexibly adjusting the contributions of global context and part-level semantics throughout the generation process is required.</p>
<p>To address this, we propose a diffusion-based T2M framework incorporating an Embedding Gating Network (EGN), which dynamically adjusts the relative contributions of global and part-level embeddings based on the current generation state. In contrast to existing methods that combine global and local conditions with fixed weights or process them in separate stages, EGN ensures that the global context continuously informs part-wise generation throughout the entire process. The gating mechanism is conditioned on both static text embeddings and the current noisy motion state, enabling dynamic adjustment of the global-local balance based on semantics that remain underrepresented in the current motion estimate. We further observe that part-level semantics contribute not only to late-stage refinement but also to early structure formation. Based on this, the framework incorporates local embeddings from the earliest generation stages. The modulated embeddings are routed through dedicated part-wise generation pathways, which limit inter-part interference inherent in shared pathways and allow each generator to specialize in its assigned part-level semantics. Simultaneously, the global context is maintained through EGN, supporting both part-wise specialization and full-body coherence.</p>
<p>Our main contributions are as follows:<list list-type="bullet">
<list-item>
<p>We introduce the EGN, which dynamically adjusts the contributions of global and part-level embeddings to reflect the current generation state, achieving precise part-level semantic alignment while preserving global motion consistency.</p></list-item>
<list-item>
<p>We demonstrate that local embeddings, when explicitly incorporated from the early generation stages, contribute to initial structure formation beyond late-stage refinement.</p></list-item>
<list-item>
<p>We design dedicated part-wise generation pathways with a fusion architecture that enables each generator to focus on part-specific semantics without inter-part interference, while maintaining structurally coherent full-body motion through controlled inter-part information exchange.</p></list-item>
</list></p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<sec id="s2_1">
<label>2.1</label>
<title>Text-to-Motion Generation</title>
<p>T2M generation aims to synthesize 3D human motion sequences from natural language descriptions. Early studies relied on conditioning with simple action classes or text labels [<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-18">18</xref>]. As the field expanded to include free-form natural language instructions and required precise semantic alignment between text and motion, controllability became a key challenge. Initial approaches employed variational autoencoders (VAEs) [<xref ref-type="bibr" rid="ref-19">19</xref>] to align text and motion in a shared embedding space [<xref ref-type="bibr" rid="ref-5">5</xref>,<xref ref-type="bibr" rid="ref-6">6</xref>] or to learn text&#x2013;motion correspondence via conditional decoding [<xref ref-type="bibr" rid="ref-4">4</xref>]. Tevet et al. [<xref ref-type="bibr" rid="ref-7">7</xref>] used visual priors by encoding motion descriptions with a frozen Contrastive Language-Image Pre-training (CLIP) [<xref ref-type="bibr" rid="ref-20">20</xref>] text encoder and aligning the motion latent space accordingly. More recently, diffusion models [<xref ref-type="bibr" rid="ref-21">21</xref>] have been widely adopted in T2M due to their strong generation quality and temporal stability. The Motion Diffusion Model (MDM) [<xref ref-type="bibr" rid="ref-7">7</xref>] introduced a Transformer-based [<xref ref-type="bibr" rid="ref-22">22</xref>] denoising network that directly models temporal dependencies. Chen et al. [<xref ref-type="bibr" rid="ref-13">13</xref>] improved efficiency and quality by leveraging latent-space diffusion with optimized sampling. However, these methods often apply the same text condition across all body parts, which makes it difficult to capture accurate part-level semantic mappings and leads to mismatches in part motion [<xref ref-type="bibr" rid="ref-14">14</xref>,<xref ref-type="bibr" rid="ref-15">15</xref>]. In parallel, token-based approaches tokenize continuous motion into codebook indices using vector quantization (VQ) [<xref ref-type="bibr" rid="ref-23">23</xref>] and treat it as a discrete sequence modeling problem. Guo et al. [<xref ref-type="bibr" rid="ref-9">9</xref>] formulated text-to-motion as a discrete sequence generation problem, and Zhang et al. [<xref ref-type="bibr" rid="ref-10">10</xref>] improved token prediction via autoregressive generation. More recently, Guo et al. [<xref ref-type="bibr" rid="ref-11">11</xref>] combined hierarchical quantization with bidirectional masked prediction, mitigating the error accumulation inherent in sequential autoregressive generation. However, tokenization can result in information loss when compressing continuous motion into a limited codebook, potentially restricting motion diversity and fidelity [<xref ref-type="bibr" rid="ref-24">24</xref>]. Although generation architectures have evolved in diverse directions, from a conditioning perspective, most methods still share the limitation of relying on fixed-length global embeddings, such as the CLIP [CLS] token. This reliance dilutes fine-grained semantics for specific body parts in composite prompts, motivating the need for conditioning mechanisms that faithfully reflect part-level semantics while maintaining expressiveness.</p>
<p>Ghosh et al. [<xref ref-type="bibr" rid="ref-3">3</xref>] began with coarse divisions such as upper and lower body, treating each part through a separate generation pathway. Athanasiou et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] automated spatial composition by using GPT-3 to extract action-body part mappings and combining independently generated part-wise motions post-hoc. Subsequent research has focused on increasing the number of segmented parts to achieve finer motion representation. Part-Coordinating (ParCo) [<xref ref-type="bibr" rid="ref-14">14</xref>] decomposes the full body into multiple semantic parts and performs part-wise generation while considering inter-part coordination. More recent work has combined part-based modeling with text-based motion generation, directly connecting detailed elements to part-level representations. The Local-to-Global pipeline for Text-to-Motion generation (LGTM) [<xref ref-type="bibr" rid="ref-15">15</xref>] decomposes global text into part-specific descriptions to strengthen semantic alignment at the part level. Wang et al. [<xref ref-type="bibr" rid="ref-25">25</xref>] extract body-part-related cues through large language model (LLM)-based semantic parsing and reflect sentence structure to condition detailed semantics more precisely. Fan et al. [<xref ref-type="bibr" rid="ref-26">26</xref>] identify salient body parts and enforce semantic alignment for interaction motion. Improvements have also been attempted from the perspective of the conditioning mechanism itself. Chang et al. [<xref ref-type="bibr" rid="ref-27">27</xref>] introduce a composite-aware text encoder and a text-motion aligner, enabling dynamic word-level correspondence rather than fixed-length global embeddings. Li and Feng [<xref ref-type="bibr" rid="ref-28">28</xref>] improve the generation accuracy of composite actions by simultaneously leveraging coarse- and fine-grained descriptions. These studies supply fine-grained cues that would otherwise be lost in a global summary embedding, enabling body-part actions specified in text to be more faithfully reflected in the generated motion. However, existing methods combine global and local conditions at fixed ratios or perform partial generation and global optimization as separate stages, which can cause one condition to overshadow the other. Additionally, post-hoc composition approaches do not incorporate inter-part coordination into the generation process itself, and learning-based part-wise generation methods that inject part conditions independently through dedicated pathways struggle to capture full-body context. A mechanism that dynamically adjusts the relative contributions of global and local conditions as generation progresses has yet to be explored.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Conditioning Mechanism in Generative Models</title>
<p>The challenge of fine-grained semantics being diluted under global conditioning has been widely recognized in conditional generative modeling. In text-to-image synthesis, efforts to address this issue have included decomposing composite prompts into per-concept diffusion processes [<xref ref-type="bibr" rid="ref-29">29</xref>] and leveraging linguistic structure to restructure cross-attention [<xref ref-type="bibr" rid="ref-30">30</xref>]. Zarei et al. [<xref ref-type="bibr" rid="ref-31">31</xref>] demonstrated that the output space of CLIP is suboptimal for compositional prompts, showing that attention contributions from unrelated tokens are mixed into the final token embeddings, leading to failures in attribute-object binding. In the facial synthesis domain, Song et al. [<xref ref-type="bibr" rid="ref-32">32</xref>] improved attribute-level alignment under multi-attribute conditions by introducing a module that dynamically balances global text features and local attribute features through learnable gating. It has also been observed that the role of conditioning changes qualitatively across stages of the diffusion process. Balaji et al. [<xref ref-type="bibr" rid="ref-33">33</xref>] empirically showed that text-to-image diffusion models rely heavily on text conditioning during early denoising but largely ignore it in later stages. More recently, Cho et al. [<xref ref-type="bibr" rid="ref-34">34</xref>] argued that static conditioning cannot flexibly adapt to the dynamic nature of multi-stage denoising, which evolves from coarse structure to fine detail, and proposed TC-LoRA to dynamically adjust conditioning based on both the denoising timestep and the control signal. These studies consistently demonstrate that hierarchical semantic decomposition and dynamic conditioning are critical to generation quality, yet these insights remain largely unexplored in T2M. In this paper, we combine such adaptive conditioning strategies with part-level semantic decomposition for T2M and show that timestep-conditioned gating improves motion-text alignment.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Methodology</title>
<sec id="s3_1">
<label>3.1</label>
<title>Part-Wise Motion Representation</title>
<p>Recent text-to-motion research, particularly following HumanML3D [<xref ref-type="bibr" rid="ref-4">4</xref>], typically represents full-body motion using a canonical pose that incorporates root information and joint features relative to the root. A motion sequence consists of F frames with J joints, where each frame <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msup><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msup><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>r</mml:mi><mml:mo>&#x02D9;</mml:mo></mml:mover></mml:mrow><mml:mi>a</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mi mathvariant="bold">r</mml:mi></mml:mrow><mml:mo>&#x02D9;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mi>z</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mi>h</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">j</mml:mi></mml:mrow><mml:mi>p</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">j</mml:mi></mml:mrow><mml:mi>r</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">j</mml:mi></mml:mrow><mml:mi>v</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">c</mml:mi></mml:mrow><mml:mi>f</mml:mi></mml:msub><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>. In this formulation, <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mrow><mml:mover><mml:mi>r</mml:mi><mml:mo>&#x02D9;</mml:mo></mml:mover></mml:mrow><mml:mi>a</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mn>1</mml:mn></mml:msup></mml:math></inline-formula> denotes the root yaw angular velocity, <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mi mathvariant="bold">r</mml:mi></mml:mrow><mml:mo>&#x02D9;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mi>z</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:math></inline-formula> denotes the root linear velocity on the XZ plane, and <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mi>r</mml:mi><mml:mi>h</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mn>1</mml:mn></mml:msup></mml:math></inline-formula> denotes the root height. The features <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mrow><mml:mi mathvariant="bold">j</mml:mi></mml:mrow><mml:mi>p</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>J</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mrow><mml:mi mathvariant="bold">j</mml:mi></mml:mrow><mml:mi>r</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mn>6</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>J</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula>, and <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mrow><mml:mi mathvariant="bold">j</mml:mi></mml:mrow><mml:mi>v</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>J</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> represent joint position, rotation, and velocity in the root coordinate system, respectively, while <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mrow><mml:mi mathvariant="bold">c</mml:mi></mml:mrow><mml:mi>f</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mn>4</mml:mn></mml:msup></mml:math></inline-formula> is a binary feature indicating foot-ground contact. This representation encodes both global trajectory and relative joint motion, and is widely adopted for stable modeling of diverse motions. Building on this canonical representation, our approach reorganizes motion into semantic parts, reflecting the fact that text descriptions frequently target specific body parts. We partition joints into six groups (root, backbone, left arm, right arm, left leg, right leg), and construct part-wise motion by selecting only the joint features (position, rotation, velocity, and foot contact) for each group.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Embedding Gating Network</title>
<p>We propose the EGN, which consists of a transformation module and a <italic>GlocalGate</italic>, as illustrated in <xref ref-type="fig" rid="fig-1">Fig. 1a</xref>. In diffusion-based T2M models, text is typically condensed into a single global embedding, which can obscure fine-grained cues in complex sentences. Motivated by this, we separate text conditions into global and local components. The global text <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mrow><mml:mi mathvariant="bold">T</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msub></mml:math></inline-formula> represents the original full-body description, providing overall action identity and coarse context. Part-specific local texts <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">T</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>P</mml:mi></mml:msubsup></mml:math></inline-formula> are derived by decomposing <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mrow><mml:mi mathvariant="bold">T</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msub></mml:math></inline-formula> into short action phrases for each of the six body parts using an LLM. For instance, given the global description <italic>&#x201C;a person bends forward and picks up an object in their right hand,&#x201D;</italic> the LLM assigns <italic>&#x201C;bends forward&#x201D;</italic> to the pelvis and <italic>&#x201C;picks up an object&#x201D;</italic> to the right arm, while the remaining parts receive a null descriptor indicating no part-specific action. Each text input is independently encoded by the pretrained CLIP text encoder, extracting the end-of-sequence token representation as a fixed-dimensional embedding. This process yields a global embedding <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">e</mml:mi></mml:mrow><mml:mi>g</mml:mi><mml:mrow><mml:mtext>CLIP</mml:mtext></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and part-specific local embeddings <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">e</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mtext>CLIP</mml:mtext></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>D</mml:mi></mml:mrow></mml:msup><mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>P</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>. Since the text encoder&#x2019;s embedding space may not align with the conditioning space required for motion generation, a transformation module is introduced to expand the representational capacity of both global and local embeddings. This module is implemented as a residual feed-forward adapter that learns motion-relevant correction terms on top of the original CLIP embeddings. We formulate this transformation as a residual function, allowing the embeddings to be progressively adapted for motion conditioning while retaining the original CLIP embeddings as a semantic anchor. The final conditioning embeddings are obtained as
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mrow><mml:mi mathvariant="bold">e</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">e</mml:mi></mml:mrow><mml:mi>g</mml:mi><mml:mrow><mml:mrow><mml:mtext>CLIP</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>g</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">e</mml:mi></mml:mrow><mml:mi>g</mml:mi><mml:mrow><mml:mrow><mml:mtext>CLIP</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msub><mml:mrow><mml:mi mathvariant="bold">e</mml:mi></mml:mrow><mml:mi>l</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">e</mml:mi></mml:mrow><mml:mi>l</mml:mi><mml:mrow><mml:mrow><mml:mtext>CLIP</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">e</mml:mi></mml:mrow><mml:mi>l</mml:mi><mml:mrow><mml:mrow><mml:mtext>CLIP</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>f</mml:mi><mml:mi>g</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>f</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:math></inline-formula> denote the learnable nonlinear mappings for global and local conditions, respectively.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>The overall architecture of the proposed framework. (<bold>a</bold>) The final embeddings computed by the EGN are applied to part-wise generation modules. (<bold>b</bold>) The generated part-wise motions are fused within PartFuse blocks to maintain overall motion consistency.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_80992-fig-1.tif"/>
</fig>
<p>Diffusion models are known to recover global structures at high noise levels and fine-grained details at low noise levels [<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-36">36</xref>]. We therefore hypothesize that the relative importance of global and local conditions shifts across timesteps and generation states. Based on this hypothesis, we dynamically modulate their contributions rather than combining them at a fixed ratio. Accordingly, GlocalGate predicts global and local weights from the noisy motion <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> and timestep <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>t</mml:mi></mml:math></inline-formula>. This module is defined as:<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">w</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">w</mml:mi></mml:mrow><mml:mi>l</mml:mi></mml:msub><mml:mo stretchy="false">]</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mtext>Linear</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>SiLU</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>Linear</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mrow><mml:mtext>emb</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">]</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mo stretchy="false">[</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> denotes concatenation and <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">m</mml:mi><mml:mi mathvariant="normal">b</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> represents the timestep embedding. The output logits <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mrow><mml:mi mathvariant="bold">w</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mrow><mml:mi mathvariant="bold">w</mml:mi></mml:mrow><mml:mi>l</mml:mi></mml:msub></mml:math></inline-formula> are subsequently normalized via softmax to yield the final gating coefficients:<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mi mathvariant="bold-italic">&#x03B1;</mml:mi><mml:mi>g</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">w</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">w</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">w</mml:mi></mml:mrow><mml:mi>l</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msub><mml:mi mathvariant="bold-italic">&#x03B1;</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">w</mml:mi></mml:mrow><mml:mi>l</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">w</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">w</mml:mi></mml:mrow><mml:mi>l</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>These gating coefficients, <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mi mathvariant="bold-italic">&#x03B1;</mml:mi><mml:mi>g</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msub><mml:mi mathvariant="bold-italic">&#x03B1;</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, determine the relative contribution of global and local semantics at each timestep. Using these coefficients, we compute the part-specific condition <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msub><mml:mrow><mml:mi mathvariant="bold">c</mml:mi></mml:mrow><mml:mi>p</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> for part <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mi>p</mml:mi></mml:math></inline-formula> as the weighted sum of global and local embeddings:<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msub><mml:mrow><mml:mi mathvariant="bold">c</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">&#x03B1;</mml:mi><mml:mi>g</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">e</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">&#x03B1;</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">e</mml:mi></mml:mrow><mml:mrow><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>GlocalGate thus adaptively balances global and local influences based on its inputs. Specifically, the noisy motion <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> conveys the current geometric state and pose configuration, while the timestep <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi>t</mml:mi></mml:math></inline-formula> encodes the prevailing noise level. The learned coefficients <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msub><mml:mi mathvariant="bold-italic">&#x03B1;</mml:mi><mml:mi>g</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msub><mml:mi mathvariant="bold-italic">&#x03B1;</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:math></inline-formula> therefore shift the balance between global coherence and local specificity at each generation step. The EGN integrates both the transformation module and GlocalGate to produce the final part-specific conditioning embedding <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msub><mml:mrow><mml:mi mathvariant="bold">c</mml:mi></mml:mrow><mml:mi>p</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Conditional Generation with Part-Wise Pathways</title>
<p>Our approach maintains part-wise pathways within a single model: each part representation is generated under its global-local embedding, after which a lightweight fusion step exchanges inter-part information, as illustrated in <xref ref-type="fig" rid="fig-1">Fig. 1b</xref>. Accordingly, we separate the network into part-wise attention for conditional generation and <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mrow><mml:mi mathvariant="normal">P</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">r</mml:mi><mml:mi mathvariant="normal">t</mml:mi><mml:mi mathvariant="normal">F</mml:mi><mml:mi mathvariant="normal">u</mml:mi><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> for part-wise fusion. At each timestep <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mi>t</mml:mi></mml:math></inline-formula>, the generation process is divided into two specialized stages: semantic synthesis and structural coordination. Initially, part-wise attention utilizes the part-specific condition <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msub><mml:mrow><mml:mi mathvariant="bold">c</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> to update the latent representation of each part <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mi>p</mml:mi></mml:math></inline-formula> as follows:<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msubsup><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mtext>Attention</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">c</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:msubsup><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> represents the semantically updated state focusing on synthesizing the local action structure. Subsequently, <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mrow><mml:mi mathvariant="normal">P</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">r</mml:mi><mml:mi mathvariant="normal">t</mml:mi><mml:mi mathvariant="normal">F</mml:mi><mml:mi mathvariant="normal">u</mml:mi><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> enables controlled inter-part coordination while maintaining the separation of generation pathways. It is defined for each part <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mi>i</mml:mi></mml:math></inline-formula> as follows:<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mrow><mml:mtext>PartFuse</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mspace width="thinmathspace" /><mml:mo>;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>&#x2223;</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>&#x2223;</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mo stretchy="false">]</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x2260;</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext>LN</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mtext>MLP</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2218;</mml:mo><mml:mrow><mml:mtext>LN</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo>&#x2223;</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>&#x2223;</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>h</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mo stretchy="false">]</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x2260;</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msub><mml:mrow><mml:mi mathvariant="normal">M</mml:mi><mml:mi mathvariant="normal">L</mml:mi><mml:mi mathvariant="normal">P</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> functions as a part-specific alignment module that incorporates features from other parts into the current pathway.</p>
<p>A key advantage is the explicit separation, at the layer level, between part-wise attention and <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mrow><mml:mi mathvariant="normal">P</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">r</mml:mi><mml:mi mathvariant="normal">t</mml:mi><mml:mi mathvariant="normal">F</mml:mi><mml:mi mathvariant="normal">u</mml:mi><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. Decoupling semantic interpretation from inter-part coordination introduces an inductive bias that keeps part representations distinct. This architectural choice stabilizes training and provides consistent part-level control, even for complex textual conditions involving multiple parts.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiments</title>
<sec id="s4_1">
<label>4.1</label>
<title>Experiment Settings</title>
<p><bold>Dataset.</bold> We evaluate our method on HumanML3D [<xref ref-type="bibr" rid="ref-4">4</xref>], a widely used public benchmark for text-to-motion generation. HumanML3D consists of 3D human motion sequences paired with natural language annotations and has served as a standard dataset in prior T2M studies. It contains 14,616 motion sequences and 44,970 text descriptions. Since global text annotations in HumanML3D are often insufficient for part-level semantic alignment, we augment the dataset with local texts derived from the global annotations using the decomposition procedure described in <xref ref-type="sec" rid="s3_2">Section 3.2</xref>.</p>
<p><bold>Metrics.</bold> We evaluate motion quality and text&#x2013;motion alignment using five metrics: (1) R-Precision measures motion&#x2013;text retrieval accuracy in the feature space of a pretrained T2M evaluation network. For each motion sequence, we compute precision based on whether the ground-truth text is ranked within Top-1/2/3 among 32 candidate texts. (2) Fr&#x00E9;chet Inception Distance (FID) measures the distributional gap between generated and real motions by computing FID on motion features extracted by the T2M evaluator. (3) Multi-Modal Distance (MM-Dist) computes the average Euclidean distance between each text feature and the motion features generated from that text. (4) Diversity splits generated motions into two random subsets of equal size and computes the average Euclidean distance between their motion features [<xref ref-type="bibr" rid="ref-4">4</xref>,<xref ref-type="bibr" rid="ref-5">5</xref>,<xref ref-type="bibr" rid="ref-37">37</xref>]. (5) Part-level Multi-Modal Similarity (PMM Sim), adopted from Sun et al. [<xref ref-type="bibr" rid="ref-15">15</xref>], trains part-level text and motion encoders with contrastive learning following Petrovich et al. [<xref ref-type="bibr" rid="ref-38">38</xref>] and measures the correspondence between part-specific text descriptions and generated part motions.</p>
<p><bold>Baselines.</bold> We select three baselines to evaluate our part-wise diffusion generation framework, which jointly uses global and local text conditions: (1) MDM, a diffusion-based T2M model conditioned on global text, as a reference for full-body diffusion performance; (2) ParCo, a part-wise generation approach with separate body-part pathways, to assess the impact of part-wise generation on performance and consistency; and (3) LGTM, which decomposes global text into part-wise descriptions, to evaluate the effect of part-wise text conditioning.</p>
<p><bold>Implementation Details.</bold> We employ Qwen2.5-7B-Instruct [<xref ref-type="bibr" rid="ref-39">39</xref>] as the LLM for decomposition, prompting it to produce a local text for each of the six body parts given the global annotation. Our model consists of 155M total parameters, of which 21.2M are trainable. Training was conducted on a single NVIDIA RTX 4090 GPU for 83.3 h on the HumanML3D dataset. At inference, the model requires 0.45 s per sample (DDIM [<xref ref-type="bibr" rid="ref-40">40</xref>], 50 steps) and peaks at 1184 MB of GPU memory.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Qualitative Results</title>
<p><xref ref-type="fig" rid="fig-2">Fig. 2</xref> compares results on composite prompts in which a single sentence specifies a full-body action alongside multiple part-level instructions. MDM largely preserves natural full-body motion, but tends to converge to a similar global motion pattern, yielding limited changes when part instructions vary and occasionally omitting part-specific constraints. ParCo can reflect some parts&#x2019; instructions due to part-wise generation, but under composite prompts, inter-part interactions often become unstable, leading to motion collapse or suppression where constraints on one part inhibit the motions of other parts. LGTM improves responsiveness to part instructions by providing part-wise text conditions, but when part conditions conflict or when the fusion with global context is not sufficiently stable, some constraints are weakened, or the overall action consistency degrades. In contrast, our method injects part-wise signals progressively while preserving the global condition throughout denoising, producing motions that preserve full-body coherence while simultaneously satisfying multiple part instructions under composite prompts.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Qualitative comparison results with baseline models. Red dashed lines indicate missing instructions, and red solid lines indicate motion collapse artifacts. (<bold>a</bold>) MDM, (<bold>b</bold>) ParCo, and (<bold>c</bold>) LGTM exhibit missing or incorrectly reflected instructions along with body distortion artifacts, whereas (<bold>d</bold>) the proposed method reflects the given instructions while generating natural body motion.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_80992-fig-2.tif"/>
</fig>
<p><xref ref-type="fig" rid="fig-3">Fig. 3</xref> evaluates how well part-level text descriptions, decomposed from input sentences using an LLM, are preserved during full-body motion synthesis. For example, given the prompt <italic>&#x201C;A person sits down and stretches his legs straight.&#x201D;</italic>, LGTM generates part-wise motions from decomposed part texts and then applies strong global conditioning during the full-body motion optimization stage [<xref ref-type="bibr" rid="ref-15">15</xref>]. In this process, features induced by part conditions (e.g., stretches legs straight) are partially suppressed by the global context (e.g., sits down), leading to cases where the legs converge to a bent posture rather than being fully extended. In contrast, our method maintains the global condition as an anchor while explicitly injecting and aligning part-wise conditions during generation, ensuring full-body consistency while preserving part-level semantics. As a result, local semantics such as <italic>&#x201C;stretches legs straight&#x201D;</italic> are more consistently reflected in the final motion, even in composite sentences.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Generation results according to different embedding methods for part-wise instructions. Green lines indicate parts where instructions are well-reflected, and red lines indicate parts where instructions are missing. (<bold>a</bold>) LGTM fails to reflect the &#x201C;stretches legs straight&#x201D; instruction in both legs, whereas (<bold>b</bold>) the proposed method correctly captures both &#x201C;sits down&#x201D; and &#x201C;stretches legs straight&#x201D; across the corresponding body parts.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_80992-fig-3.tif"/>
</fig>
<p>Overall, our method reflects part-wise conditions more consistently when satisfying multiple part constraints simultaneously in composite prompts. Even in sentences with a strong global context, local prompts decomposed at the part level are stably preserved in the final motion, ensuring that key actions of specific parts appear without being weakened. This observation qualitatively supports the effectiveness of the global-local conditioning scheme, which stabilizes the integration of part-wise conditions throughout the diffusion process while preserving the global condition.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Quantitative Results</title>
<p><xref ref-type="table" rid="table-1">Table 1</xref> presents the quantitative evaluation results of the proposed method and baselines (MDM, ParCo, LGTM) on the HumanML3D test set. From the perspective of text-motion alignment, our method demonstrates consistent improvements across all metrics. Our method achieves substantial gains in R-Precision and MM-Dist compared with MDM and outperforms ParCo and LGTM in text-motion alignment. This quantitatively confirms that part-wise conditioning improves alignment over global-only conditioning, which tends to obscure fine-grained elements in complex sentences. In terms of generation quality (FID), our method shows slightly higher FID than ParCo, yet R-Precision and MM-Dist consistently show improvements. Meng et al. [<xref ref-type="bibr" rid="ref-24">24</xref>] have shown that the standard evaluation protocol can disproportionately favor VQ-based methods over diffusion-based methods. Accordingly, the FID gap between the proposed diffusion-based method and the VQ-based ParCo likely reflects intrinsic differences between the two generation paradigms. Among diffusion-based methods, our method achieves the best FID, suggesting that dynamically modulating global and local conditions throughout generation is effective for improving generation quality. Finally, for Diversity, our model maintains a level of diversity comparable to that of real motions, suggesting that it improves text-motion alignment without sacrificing variation in the generated results.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Quantitative comparison with baseline models on HumanML3D dataset. <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula> indicates higher is better, <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula> indicates lower is better, and <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> indicates closer to real motion is better. Bold indicates the best performance, and underline indicates the second-best performance.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th align="center" rowspan="2">Method</th>
<th align="center" colspan="3">R-Precision <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula></th>
<th align="center" rowspan="2">FID <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th align="center" rowspan="2">MM-Dist <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th align="center" rowspan="2">Div. <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula></th>
</tr>
<tr>
<th>Top-1</th>
<th>Top-2</th>
<th>Top-3</th>
</tr>
</thead>
<tbody>
<tr>
<td>Real</td>
<td>0.4616</td>
<td>0.6726</td>
<td>0.7746</td>
<td></td>
<td>3.2408</td>
<td>9.3808</td>
</tr>
<tr>
<td>MDM</td>
<td><inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msup><mml:mn>0.4078</mml:mn><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.007</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msup><mml:mn>0.6065</mml:mn><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.008</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:msup><mml:mn>0.7194</mml:mn><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.007</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msup><mml:mn>0.8482</mml:mn><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.081</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:msup><mml:mn>3.5088</mml:mn><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.024</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msup><mml:mn>9.1045</mml:mn><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.099</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
</tr>
<tr>
<td>ParCo</td>
<td><underline>0.4514</underline><sup>&#x00B1;0.003</sup></td>
<td><underline>0.6552</underline><sup>&#x00B1;0.003</sup></td>
<td><inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msup><mml:mi>0.7635</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><bold>0.1427</bold><sup>&#x00B1;0.006</sup></td>
<td><underline>3.2785</underline><sup>&#x00B1;0.010</sup></td>
<td><bold>9.4201</bold><sup>&#x00B1;0.068</sup></td>
</tr>
<tr>
<td>LGTM</td>
<td><inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:msup><mml:mi>0.4486</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:msup><mml:mi>0.6551</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><underline>0.7676</underline><sup>&#x00B1;0.003</sup></td>
<td><inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:msup><mml:mi>0.8015</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:msup><mml:mi>3.3048</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:msup><mml:mi>8.7743</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
</tr>
<tr>
<td>Ours</td>
<td><bold>0.4751</bold><sup>&#x00B1;0.003</sup></td>
<td><bold>0.6823</bold><sup>&#x00B1;0.003</sup></td>
<td><bold>0.7874</bold><sup>&#x00B1;0.003</sup></td>
<td><underline>0.1567</underline><sup>&#x00B1;0.005</sup></td>
<td><bold>3.1396</bold><sup>&#x00B1;0.009</sup></td>
<td><underline>9.5130</underline><sup>&#x00B1;0.072</sup></td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-2">Table 2</xref> reports the evaluation of part-level semantic alignment using PMM Sim. Our method achieves the highest scores on Left Arm and Right Arm, and remains competitive on Head, ranking second only to ParCo by a narrow margin. This is because arm movements are explicitly described in text annotations, yielding semantically rich local embeddings via LLM-based decomposition, which enables EGN&#x2019;s gating and dedicated generation pathways to operate effectively. In contrast, performance on Torso, Left Leg, and Right Leg is slightly below LGTM. These parts are less frequently described in text, often receiving null descriptors, which makes the local embeddings semantically sparse and causes EGN&#x2019;s gating to shift toward the global condition, reducing the benefit of part-specific conditioning. Nevertheless, our method consistently surpasses MDM and ParCo on these parts, maintaining balanced alignment performance across all body parts.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Quantitative comparison of part-level semantic alignment with baseline models on the HumanML3D dataset, evaluated using PMM Sim. Higher values indicate better performance. Bold indicates the best performance, and underline indicates the second-best performance.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Method</th>
<th>Head</th>
<th>Left Arm</th>
<th>Right Arm</th>
<th>Torso</th>
<th>Left Leg</th>
<th>Right Leg</th>
</tr>
</thead>
<tbody>
<tr>
<td>Real</td>
<td>0.8037</td>
<td>0.7160</td>
<td>0.7210</td>
<td>0.7583</td>
<td>0.7531</td>
<td>0.7578</td>
</tr>
<tr>
<td>MDM</td>
<td>0.7867</td>
<td>0.7000</td>
<td>0.6925</td>
<td>0.7425</td>
<td>0.7347</td>
<td>0.7231</td>
</tr>
<tr>
<td>ParCo</td>
<td><bold>0.8081</bold></td>
<td><underline>0.7206</underline></td>
<td><underline>0.7293</underline></td>
<td>0.7070</td>
<td>0.6380</td>
<td>0.6243</td>
</tr>
<tr>
<td>LGTM</td>
<td>0.7967</td>
<td>0.7198</td>
<td>0.7256</td>
<td><bold>0.7656</bold></td>
<td><bold>0.7575</bold></td>
<td><bold>0.7631</bold></td>
</tr>
<tr>
<td>Ours</td>
<td><underline>0.8069</underline></td>
<td><bold>0.7248</bold></td>
<td><bold>0.7349</bold></td>
<td><underline>0.7603</underline></td>
<td><underline>0.7516</underline></td>
<td><underline>0.7592</underline></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Overall, by preserving the global condition while adaptively injecting part-level conditions during generation, our method improves text-motion semantic alignment over existing baselines and remains stable in generation quality and diversity.</p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Ablation Study</title>
<p>This section analyzes the impact of the key components of the proposed framework on performance. We first examine the contributions of individual components within EGN, then investigate how framework-level design choices&#x2014;including global descriptions, local descriptions, the PartFuse module, and part-wise generation pathways&#x2014;affect overall performance.</p>
<p><xref ref-type="table" rid="table-3">Table 3</xref> reports the results of separately removing the two core components of EGN: the Transformation module and the Noisy Motion conditioning. <italic>&#x201C;w/o Transformation&#x201D;</italic> removes the residual feed-forward adapter <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:msub><mml:mi>f</mml:mi><mml:mi>g</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:msub><mml:mi>f</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:math></inline-formula> (<xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>), directly using the raw CLIP embeddings as conditioning inputs. <italic>&#x201C;w/o Noisy Motion&#x201D;</italic> removes <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> from the GlocalGate input (<xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>), so that the gating coefficients are determined solely by timestep information. Removing the Transformation module degrades R-Precision and MM-Dist, while FID rises sharply from 0.1567 to 0.4152. Without a learned transformation, the raw embeddings are insufficient for effective conditioning, destabilizing the generation distribution. Removing the Noisy Motion conditioning also results in a decline in R-Precision and MM-Dist, along with a modest increase in FID. This indicates that reflecting the current noisy motion state in the gating contributes to adaptive condition adjustment as generation progresses. Collectively, these results confirm that both the embedding transformation and the motion-state-based gating within EGN contribute to text-motion alignment and generation quality, and that neither component alone is sufficient.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Ablation results for EGN component variants on the HumanML3D dataset, including the removal of the Transformation module and Noisy Motion. <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula> indicates higher is better, <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula> indicates lower is better, and <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> indicates that values closer to the real-motion Diversity score (9.3808) are better. Bold indicates the best performance.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th align="center" rowspan="2">Method</th>
<th colspan="3">R-Precision <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula></th>
<th align="center" rowspan="2">FID <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th align="center" rowspan="2">MM-Dist <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th align="center" rowspan="2">Div. <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula></th>
</tr>
<tr>
<th>Top-1</th>
<th>Top-2</th>
<th>Top-3</th>
</tr>
</thead>
<tbody>
<tr>
<td>Full EGN</td>
<td><bold>0.4751</bold><sup>&#x00B1;0.003</sup></td>
<td><bold>0.6823</bold><sup>&#x00B1;0.003</sup></td>
<td><bold>0.7874</bold><sup>&#x00B1;0.003</sup></td>
<td><bold>0.1567</bold><sup>&#x00B1;0.005</sup></td>
<td><bold>3.1396</bold><sup>&#x00B1;0.009</sup></td>
<td><inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:msup><mml:mi>9.5130</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.105</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
</tr>
<tr>
<td>w/o Transformation</td>
<td><inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:msup><mml:mi>0.4670</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.002</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:msup><mml:mi>0.6665</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.002</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:msup><mml:mi>0.7735</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.002</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:msup><mml:mi>0.4152</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.008</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:msup><mml:mi>3.2141</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.009</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><bold>9.3378</bold><sup>&#x00B1;0.064</sup></td>
</tr>
<tr>
<td>w/o Noisy Motion</td>
<td><inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:msup><mml:mi>0.4655</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:msup><mml:mi>0.6700</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:msup><mml:mi>0.7800</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:msup><mml:mi>0.1835</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.008</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:msup><mml:mi>3.1893</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.009</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:msup><mml:mi>9.4955</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.085</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-4">Table 4</xref> analyzes the contributions of four framework-level components&#x2014;global descriptions, local descriptions, the PartFuse module, and part-wise generation pathways&#x2014;by evaluating their combinations. Removing PartFuse results in a slight decline in R-Precision and MM-Dist, while FID increases notably from 0.1567 to 0.2973. This suggests that simply combining independently generated part motions without a dedicated fusion mechanism degrades the quality of the full-body distribution, confirming that PartFuse helps ensure inter-part coherence. When the global description is removed and the generation is based solely on local descriptions, a sharp performance drop is observed across all metrics. This can be attributed to two factors. First, text annotations in the current dataset do not explicitly describe every body part, so relying on local descriptions alone leaves certain parts without conditioning information. Second, without the global context itself, there is no basis for capturing inter-part relationships and the overall semantic structure of the motion, making it difficult for part-wise generators to form coherent full-body motion. When local descriptions are removed, each part-wise generation pathway receives only the same global summary as its condition. Although the pathways are physically separated, the absence of differentiated semantic information prevents part-wise separation from being fully exploited, leading to degradation in both text-motion alignment and generation quality. When part-wise generation pathways are removed and generation relies solely on the global description, text-motion alignment metrics (MM-Dist, R-Precision) decline, and FID also increases. This indicates that relying solely on a global description through a single generation pathway limits the model&#x2019;s ability to capture fine-grained semantics, and that explicitly separating part-wise generation pathways is effective for improving both alignment and generation quality.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Ablation study of framework-level conditioning and architectural components on the HumanML3D dataset. &#x2713; and &#x2717; indicate the inclusion and exclusion of each component, respectively. <inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula> indicates higher is better, <inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula> indicates lower is better, and <inline-formula id="ieqn-100"><mml:math id="mml-ieqn-100"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> indicates that values closer to the real-motion Diversity score (9.3808) are better. Bold indicates the best performance.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th colspan="4">Method</th>
<th colspan="3">R-Precision <inline-formula id="ieqn-101"><mml:math id="mml-ieqn-101"><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula></th>
<th align="center" rowspan="2">FID <inline-formula id="ieqn-102"><mml:math id="mml-ieqn-102"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th align="center" rowspan="2">MM-Dist <inline-formula id="ieqn-103"><mml:math id="mml-ieqn-103"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th align="center" rowspan="2">Div. <inline-formula id="ieqn-104"><mml:math id="mml-ieqn-104"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula></th>
</tr>
<tr>
<th>Global Description</th>
<th>Local Description</th>
<th>Part<break/>Fuse</th>
<th>Part-wise Pathways</th>
<th>Top-1</th>
<th>Top-2</th>
<th>Top-3</th>
</tr>
</thead>
<tbody>
<tr>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td><bold>0.4751</bold><sup>&#x00B1;0.003</sup></td>
<td><bold>0.6823</bold><sup>&#x00B1;0.003</sup></td>
<td><bold>0.7874</bold><sup>&#x00B1;0.003</sup></td>
<td><bold>0.1567</bold><sup>&#x00B1;0.005</sup></td>
<td><bold>3.1396</bold><sup>&#x00B1;0.009</sup></td>
<td><inline-formula id="ieqn-110"><mml:math id="mml-ieqn-110"><mml:msup><mml:mi>9.5130</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.105</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
</tr>
<tr>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2717;</td>
<td>&#x2713;</td>
<td><inline-formula id="ieqn-111"><mml:math id="mml-ieqn-111"><mml:msup><mml:mi>0.4664</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-112"><mml:math id="mml-ieqn-112"><mml:msup><mml:mi>0.6717</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-113"><mml:math id="mml-ieqn-113"><mml:msup><mml:mi>0.7808</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-114"><mml:math id="mml-ieqn-114"><mml:msup><mml:mi>0.2973</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.009</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-115"><mml:math id="mml-ieqn-115"><mml:msup><mml:mi>3.1718</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.006</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><bold>9.3709</bold><sup>&#x00B1;0.107</sup></td>
</tr>
<tr>
<td>&#x2717;</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td><inline-formula id="ieqn-117"><mml:math id="mml-ieqn-117"><mml:msup><mml:mi>0.1762</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-118"><mml:math id="mml-ieqn-118"><mml:msup><mml:mi>0.2632</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.005</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-119"><mml:math id="mml-ieqn-119"><mml:msup><mml:mi>0.3234</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.005</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td>1<inline-formula id="ieqn-120"><mml:math id="mml-ieqn-120"><mml:msup><mml:mi>0.632</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.099</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-121"><mml:math id="mml-ieqn-121"><mml:msup><mml:mi>6.6868</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.017</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-122"><mml:math id="mml-ieqn-122"><mml:msup><mml:mi>7.6367</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.083</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
</tr>
<tr>
<td>&#x2713;</td>
<td>&#x2717;</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td><inline-formula id="ieqn-123"><mml:math id="mml-ieqn-123"><mml:msup><mml:mi>0.4581</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.006</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-124"><mml:math id="mml-ieqn-124"><mml:msup><mml:mi>0.6525</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.006</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-125"><mml:math id="mml-ieqn-125"><mml:msup><mml:mi>0.7640</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.006</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-126"><mml:math id="mml-ieqn-126"><mml:msup><mml:mi>0.6030</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.068</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-127"><mml:math id="mml-ieqn-127"><mml:msup><mml:mi>3.2864</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.021</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-128"><mml:math id="mml-ieqn-128"><mml:msup><mml:mi>9.3459</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.068</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
</tr>
<tr>
<td>&#x2713;</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
<td><inline-formula id="ieqn-129"><mml:math id="mml-ieqn-129"><mml:msup><mml:mi>0.4639</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-130"><mml:math id="mml-ieqn-130"><mml:msup><mml:mi>0.6670</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.002</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-131"><mml:math id="mml-ieqn-131"><mml:msup><mml:mi>0.7750</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.002</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-132"><mml:math id="mml-ieqn-132"><mml:msup><mml:mi>0.2274</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.009</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-133"><mml:math id="mml-ieqn-133"><mml:msup><mml:mi>3.2098</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.009</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-134"><mml:math id="mml-ieqn-134"><mml:msup><mml:mi>9.3021</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.070</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In summary, the global context serves as the structural foundation for overall motion, while local descriptions and part-wise pathways function complementarily to achieve fine-grained semantic alignment. PartFuse integrates independently generated part motions at the full-body level. When these components are combined, both text-motion alignment and generation quality improve jointly.</p>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Analysis of Timestep-Dependent Gating</title>
<p>To quantify the relative contribution of local conditioning with respect to global conditioning at diffusion timestep <inline-formula id="ieqn-135"><mml:math id="mml-ieqn-135"><mml:mi>t</mml:mi></mml:math></inline-formula>, we define <inline-formula id="ieqn-136"><mml:math id="mml-ieqn-136"><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> as the ratio of the local gating coefficient to the global gating coefficient, i.e., <inline-formula id="ieqn-137"><mml:math id="mml-ieqn-137"><mml:mi mathvariant="bold-italic">&#x03B1;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">&#x03B1;</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msub><mml:mi mathvariant="bold-italic">&#x03B1;</mml:mi><mml:mi>g</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. <xref ref-type="fig" rid="fig-4">Figs. 4</xref> and <xref ref-type="fig" rid="fig-5">5</xref> visualize the changes in <inline-formula id="ieqn-138"><mml:math id="mml-ieqn-138"><mml:mi mathvariant="bold-italic">&#x03B1;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> across diffusion timesteps. <inline-formula id="ieqn-139"><mml:math id="mml-ieqn-139"><mml:mi mathvariant="bold-italic">&#x03B1;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is high in the early generation stages and gradually decreases as denoising progresses, indicating that local conditions contribute strongly to part-level structure formation at the beginning and then shift toward refining full-body context under global conditions. This suggests that local information is not merely used as a late-stage detail refinement signal; rather, it serves as a structural constraint in the early steps to shape the outline and skeleton of part-specific actions. As sampling proceeds, the relative contribution of global conditions increases, strengthening overall motion context and full-body consistency. <inline-formula id="ieqn-140"><mml:math id="mml-ieqn-140"><mml:mi mathvariant="bold-italic">&#x03B1;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is not uniform across body parts: parts that are essential for establishing the global action tend to exhibit larger <inline-formula id="ieqn-141"><mml:math id="mml-ieqn-141"><mml:mi mathvariant="bold-italic">&#x03B1;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. This discrepancy is particularly pronounced in the early diffusion steps, implying that the gating network prioritizes injecting local signals into globally critical parts to stabilize the formation of the initial motion structure. As timesteps progress, local contributions decrease across all parts, while the early-stage prioritization for globally important parts remains consistent.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>The part-wise variation of the ratio of local gating coefficient to global gating coefficient across timesteps. Across all body parts, the local ratio is high in the early stages of generation and gradually decreases as denoising progresses, indicating that local conditions contribute primarily to initial structure formation and their influence diminishes in later stages.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_80992-fig-4.tif"/>
</fig><fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>The part-wise variation of the ratio of local gating coefficient to global gating coefficient across timesteps according to different prompts. (<bold>a</bold>) When the instructions for all parts are clear, the ratios show similar patterns across them. (<bold>b</bold>) When instructions for specific parts are critical, the ratios of those parts appear relatively higher.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_80992-fig-5.tif"/>
</fig>
<p><xref ref-type="table" rid="table-5">Table 5</xref> provides quantitative evidence supporting this interpretation. Removing local injection (<inline-formula id="ieqn-142"><mml:math id="mml-ieqn-142"><mml:mi mathvariant="bold-italic">&#x03B1;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>) degrades part-level semantic alignment, confirming that explicit local conditioning is necessary for part-level semantic alignment under composite prompts. Mirroring <inline-formula id="ieqn-143"><mml:math id="mml-ieqn-143"><mml:mi mathvariant="bold-italic">&#x03B1;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> over timesteps&#x2014;so that early local contributions shrink and late ones grow&#x2014;causes an overall performance drop, demonstrating that sufficiently strong local signals in the early stage are crucial. Furthermore, using a fixed <inline-formula id="ieqn-144"><mml:math id="mml-ieqn-144"><mml:mi mathvariant="bold-italic">&#x03B1;</mml:mi></mml:math></inline-formula> also underperforms the learned schedule, suggesting that a static global&#x2013;local mixture is insufficient and that timestep-dependent role transition is necessary.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Performance metrics according to local gating coefficients across timesteps. <inline-formula id="ieqn-145"><mml:math id="mml-ieqn-145"><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula> indicates higher is better, <inline-formula id="ieqn-146"><mml:math id="mml-ieqn-146"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula> indicates lower is better, and <inline-formula id="ieqn-147"><mml:math id="mml-ieqn-147"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> indicates that values closer to the real-motion Diversity score (9.3808) are better. Bold indicates the best performance.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th align="center" rowspan="2">Method</th>
<th colspan="3">R-Precision <inline-formula id="ieqn-148"><mml:math id="mml-ieqn-148"><mml:mo stretchy="false">&#x2191;</mml:mo></mml:math></inline-formula></th>
<th align="center" rowspan="2">FID <inline-formula id="ieqn-149"><mml:math id="mml-ieqn-149"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th align="center" rowspan="2">MM-Dist <inline-formula id="ieqn-150"><mml:math id="mml-ieqn-150"><mml:mo stretchy="false">&#x2193;</mml:mo></mml:math></inline-formula></th>
<th align="center" rowspan="2">Div. <inline-formula id="ieqn-151"><mml:math id="mml-ieqn-151"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula></th>
</tr>
<tr>
<th>Top-1</th>
<th>Top-2</th>
<th>Top-3</th>
</tr>
</thead>
<tbody>
<tr>
<td><inline-formula id="ieqn-152"><mml:math id="mml-ieqn-152"><mml:mi mathvariant="bold-italic">&#x03B1;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-153"><mml:math id="mml-ieqn-153"><mml:msup><mml:mi>0.4575</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.002</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-154"><mml:math id="mml-ieqn-154"><mml:msup><mml:mi>0.6631</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-155"><mml:math id="mml-ieqn-155"><mml:msup><mml:mi>0.7718</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-156"><mml:math id="mml-ieqn-156"><mml:msup><mml:mi>0.1738</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.005</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-157"><mml:math id="mml-ieqn-157"><mml:msup><mml:mi>3.2190</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.009</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><bold>9.3201</bold><sup>&#x00B1;0.072</sup></td>
</tr>
<tr>
<td><inline-formula id="ieqn-159"><mml:math id="mml-ieqn-159"><mml:mi mathvariant="bold-italic">&#x03B1;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mn>0.5</mml:mn></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-160"><mml:math id="mml-ieqn-160"><mml:msup><mml:mi>0.4295</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-161"><mml:math id="mml-ieqn-161"><mml:msup><mml:mi>0.6418</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-162"><mml:math id="mml-ieqn-162"><mml:msup><mml:mi>0.7470</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-163"><mml:math id="mml-ieqn-163"><mml:msup><mml:mi>0.6598</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.005</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-164"><mml:math id="mml-ieqn-164"><mml:msup><mml:mi>3.4003</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.009</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-165"><mml:math id="mml-ieqn-165"><mml:msup><mml:mi>9.1112</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.071</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
</tr>
<tr>
<td>Mirrored</td>
<td><inline-formula id="ieqn-166"><mml:math id="mml-ieqn-166"><mml:msup><mml:mi>0.3616</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.002</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-167"><mml:math id="mml-ieqn-167"><mml:msup><mml:mi>0.5355</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-168"><mml:math id="mml-ieqn-168"><mml:msup><mml:mi>0.6395</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.003</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-169"><mml:math id="mml-ieqn-169"><mml:msup><mml:mi>2.0028</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.036</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-170"><mml:math id="mml-ieqn-170"><mml:msup><mml:mi>4.2265</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.012</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-171"><mml:math id="mml-ieqn-171"><mml:msup><mml:mi>8.1536</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.071</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
</tr>
<tr>
<td>Learnable (Ours)</td>
<td><bold>0.4751</bold><sup>&#x00B1;0.003</sup></td>
<td><bold>0.6823</bold><sup>&#x00B1;0.003</sup></td>
<td><bold>0.7874</bold><sup>&#x00B1;0.003</sup></td>
<td><bold>0.1567</bold><sup>&#x00B1;0.005</sup></td>
<td><bold>3.1396</bold><sup>&#x00B1;0.009</sup></td>
<td><inline-formula id="ieqn-177"><mml:math id="mml-ieqn-177"><mml:msup><mml:mi>9.5130</mml:mi><mml:mrow><mml:mo>&#x00B1;</mml:mo><mml:mn>0.105</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In summary, the proposed EGN functions as a mechanism that uses local signals as early structural constraints to form part-wise outlines and progressively refines full-body consistency through global context.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>This study proposes an EGN that dynamically modulates the contributions of global and local conditions to enhance part-level semantic alignment in composite prompts. Additionally, we design a pathway-separated generation structure composed of part-wise attention and PartFuse modules, enabling each part to maintain its separated pathway while ensuring full-body consistency through inter-part coordination. Quantitative and qualitative evaluations on HumanML3D showed overall improvements in text-motion alignment and generation quality compared to existing full-body and part-based methods. Furthermore, modulating global-local fusion based on timestep and motion state produced more consistent alignment than either fixed-ratio or schedule-based alternatives. These results suggest that condition modulation reflecting generation progress can effectively improve fine-grained motion consistency and controllability in diffusion-based T2M.</p>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This work was supported by the Institute of Information &#x0026; Communications Technology Planning &#x0026; Evaluation (IITP) grant funded by the Korea government (MSIT) (No. RS-2025-02214780, Generative Haptics and Fine Response Inference for Flexible Tactile Interfaces) and the National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIT) (No. RS-2026-25479030).</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Conceptualization, Chanyoung Kim; methodology, Chanyoung Kim; software, Chanyoung Kim; validation, Chanyoung Kim; formal analysis, Chanyoung Kim; investigation, Chanyoung Kim; resources, Chanyoung Kim; data curation, Chanyoung Kim; writing&#x2014;original draft preparation, Chanyoung Kim; writing&#x2014;review and editing, Jion Kim, Byeong-Seok Shin; visualization, Chanyoung Kim; supervision, Byeong-Seok Shin; project administration, Byeong-Seok Shin; funding acquisition, Byeong-Seok Shin. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The HumanML3D dataset used in this study is publicly available at <ext-link ext-link-type="uri" xlink:href="https://github.com/EricGuo5513/HumanML3D">https://github.com/EricGuo5513/HumanML3D</ext-link>.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Plappert</surname> <given-names>M</given-names></string-name>, <string-name><surname>Mandery</surname> <given-names>C</given-names></string-name>, <string-name><surname>Asfour</surname> <given-names>T</given-names></string-name></person-group>. <article-title>The KIT motion-language dataset</article-title>. <source>Big Data</source>. <year>2016</year>;<volume>4</volume>(<issue>4</issue>):<fpage>236</fpage>&#x2013;<lpage>52</lpage>. doi:<pub-id pub-id-type="doi">10.1089/big.2016.0028</pub-id>; <pub-id pub-id-type="pmid">27992262</pub-id></mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ahn</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ha</surname> <given-names>T</given-names></string-name>, <string-name><surname>Choi</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yoo</surname> <given-names>H</given-names></string-name>, <string-name><surname>Oh</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Text2Action: generative adversarial synthesis from language to action</article-title>. In: <conf-name>Proceedings of the 2018 IEEE International Conference on Robotics and Automation (ICRA); 2018 May 21&#x2013;25</conf-name>; <publisher-loc>Brisbane, QLD, Australia</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>5</lpage>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ghosh</surname> <given-names>A</given-names></string-name>, <string-name><surname>Cheema</surname> <given-names>N</given-names></string-name>, <string-name><surname>Oguz</surname> <given-names>C</given-names></string-name>, <string-name><surname>Theobalt</surname> <given-names>C</given-names></string-name>, <string-name><surname>Slusallek</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Synthesis of compositional animations from textual descriptions</article-title>. In: <conf-name>Proceedings of the 2021 IEEE/CVF International Conference on Computer Vision (ICCV); 2021 Oct 10&#x2013;17</conf-name>; <publisher-loc>Montreal, QC, Canada</publisher-loc>. p. <fpage>1396</fpage>&#x2013;<lpage>406</lpage>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Guo</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zou</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zuo</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ji</surname> <given-names>W</given-names></string-name>, <string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Generating diverse and natural 3D human motions from text</article-title>. In: <conf-name>Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2022 Jun 18&#x2013;24</conf-name>; <publisher-loc>New Orleans, LA, USA</publisher-loc>. p. <fpage>5152</fpage>&#x2013;<lpage>61</lpage>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Petrovich</surname> <given-names>M</given-names></string-name>, <string-name><surname>Black</surname> <given-names>MJ</given-names></string-name>, <string-name><surname>Varol</surname> <given-names>G</given-names></string-name></person-group>. <chapter-title>TEMOS: generating diverse human motions from textual descriptions</chapter-title>. In: <source>Comput vision&#x2013;ECCV 2022</source>. Vol. <volume>13682</volume>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2022</year>. p. <fpage>480</fpage>&#x2013;<lpage>97</lpage>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ahuja</surname> <given-names>C</given-names></string-name>, <string-name><surname>Morency</surname> <given-names>LP</given-names></string-name></person-group>. <article-title>Language2Pose: natural language grounded pose forecasting</article-title>. In: <conf-name>Proceedings of the 2019 International Conference on 3D Vision (3DV); 2019 Sep 16&#x2013;19</conf-name>; <publisher-loc>Qu&#x00E9;bec City, QC, Canada</publisher-loc>. p. <fpage>719</fpage>&#x2013;<lpage>28</lpage>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Tevet</surname> <given-names>G</given-names></string-name>, <string-name><surname>Raab</surname> <given-names>S</given-names></string-name>, <string-name><surname>Gordon</surname> <given-names>B</given-names></string-name>, <string-name><surname>Shafir</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cohen-Or</surname> <given-names>D</given-names></string-name>, <string-name><surname>Bermano</surname> <given-names>AH</given-names></string-name></person-group>. <article-title>Human motion diffusion model</article-title>. <comment>arXiv:2209.14916. 2022</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2209.14916</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>L</given-names></string-name>, <string-name><surname>Hong</surname> <given-names>F</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>L</given-names></string-name></person-group>. <article-title>MotionDiffuse: text-driven human motion generation with diffusion model</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2024</year>;<volume>46</volume>(<issue>6</issue>):<fpage>4115</fpage>&#x2013;<lpage>28</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TPAMI.2024.3355414</pub-id>; <pub-id pub-id-type="pmid">38285589</pub-id></mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Guo</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zuo</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>L</given-names></string-name></person-group>. <chapter-title>TM2T: stochastic and tokenized modeling for the reciprocal generation of 3D human motions and texts</chapter-title>. In: <source>Comput vision&#x2013;ECCV 2022</source>. Vol. <volume>13695</volume>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2022</year>. p. <fpage>580</fpage>&#x2013;<lpage>97</lpage>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cun</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>H</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Generating human motion from textual descriptions with discrete representations (T2M-GPT)</article-title>. In: <conf-name>Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2023 Jun 17&#x2013;24</conf-name>; <publisher-loc>Vancouver, BC, Canada</publisher-loc>. p. <fpage>14730</fpage>&#x2013;<lpage>40</lpage>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Guo</surname> <given-names>C</given-names></string-name>, <string-name><surname>Mu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Javed</surname> <given-names>MG</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>L</given-names></string-name></person-group>. <article-title>MoMask: generative masked modeling of 3D human motions</article-title>. In: <conf-name>Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2024 Jun 16&#x2013;22</conf-name>; <publisher-loc>Seattle, WA, USA</publisher-loc>. p. <fpage>1900</fpage>&#x2013;<lpage>10</lpage>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Pinyoanuntapong</surname> <given-names>E</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>P</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>M</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>C</given-names></string-name></person-group>. <article-title>MMM: generative masked motion model</article-title>. In: <conf-name>Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2024 Jun 16&#x2013;22</conf-name>; <publisher-loc>Seattle, WA, USA</publisher-loc>. p. <fpage>1546</fpage>&#x2013;<lpage>55</lpage>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>T</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Executing your commands via motion diffusion in latent space</article-title>. In: <conf-name>Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2023 Jun 17&#x2013;24</conf-name>; <publisher-loc>Vancouver, BC, Canada</publisher-loc>. p. <fpage>18000</fpage>&#x2013;<lpage>10</lpage>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Zou</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Yuan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Du</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <chapter-title>ParCo: part-coordinating text-to-motion synthesis</chapter-title>. In: <source>Comput vision-ECCV 2024</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer Nature Switzerland</publisher-name>; <year>2025</year>. p. <fpage>126</fpage>&#x2013;<lpage>43</lpage>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Sun</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>R</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>C</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>R</given-names></string-name></person-group>. <article-title>LGTM: local-to-global text-driven human motion diffusion model</article-title>. In: <conf-name>SIGGRAPH &#x2019;24: Special Interest Group on Computer Graphics and Interactive Techniques Conference; 2024 Jul 27&#x2013;Aug 1</conf-name>; <publisher-loc>Denver, CO, USA</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3641519.3657422</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Athanasiou</surname> <given-names>N</given-names></string-name>, <string-name><surname>Petrovich</surname> <given-names>M</given-names></string-name>, <string-name><surname>Black</surname> <given-names>MJ</given-names></string-name>, <string-name><surname>Varol</surname> <given-names>G</given-names></string-name></person-group>. <article-title>SINC: spatial composition of 3D human motions for simultaneous action generation</article-title>. In: <conf-name>Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV); 2023 Oct 1&#x2013;6</conf-name>; <publisher-loc>Paris, France</publisher-loc>. p. <fpage>9984</fpage>&#x2013;<lpage>95</lpage>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Guo</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zuo</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zou</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Action2Motion: conditioned generation of 3D human motions</article-title>. In: <conf-name>Proceedings of the 28th ACM International Conference on Multimedia; 2020 Oct 12&#x2013;16</conf-name>; <publisher-loc>Seattle, WA, USA</publisher-loc>. p. <fpage>2021</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Petrovich</surname> <given-names>M</given-names></string-name>, <string-name><surname>Black</surname> <given-names>MJ</given-names></string-name>, <string-name><surname>Varol</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Action-conditioned 3D human motion synthesis with Transformer VAE</article-title>. In: <conf-name>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV); 2021 Oct 10&#x2013;17</conf-name>; <publisher-loc>Montreal, QC, Canada</publisher-loc>. p. <fpage>10985</fpage>&#x2013;<lpage>95</lpage>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Kingma</surname> <given-names>DP</given-names></string-name>, <string-name><surname>Welling</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Auto-encoding variational Bayes</article-title>. <comment>arXiv:1312.6114. 2013</comment>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Radford</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>JW</given-names></string-name>, <string-name><surname>Hallacy</surname> <given-names>C</given-names></string-name>, <string-name><surname>Ramesh</surname> <given-names>A</given-names></string-name>, <string-name><surname>Goh</surname> <given-names>G</given-names></string-name>, <string-name><surname>Agarwal</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Learning transferable visual models from natural language supervision</article-title>. In: <conf-name>Proceedings of the 38th International Conference on Machine Learning; 2021 Jul 18&#x2013;24</conf-name>; <publisher-loc>Virtual</publisher-loc>. p. <fpage>8748</fpage>&#x2013;<lpage>63</lpage>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ho</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jain</surname> <given-names>A</given-names></string-name>, <string-name><surname>Abbeel</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Denoising diffusion probabilistic models</article-title>. In: <conf-name>Proceedings of the 34th International Conference on Neural Information Processing Systems; 2020 Dec 6&#x2013;12</conf-name>; <publisher-loc>Vancouver, BC, Canada</publisher-loc>. p. <fpage>6840</fpage>&#x2013;<lpage>51</lpage>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Vaswani</surname> <given-names>A</given-names></string-name>, <string-name><surname>Shazeer</surname> <given-names>N</given-names></string-name>, <string-name><surname>Parmar</surname> <given-names>N</given-names></string-name>, <string-name><surname>Uszkoreit</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jones</surname> <given-names>L</given-names></string-name>, <string-name><surname>Gomez</surname> <given-names>AN</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Attention is all you need</article-title>. In: <conf-name>Proceedings of the 31st International Conference on Neural Information Processing Systems; 2017 Dec 4&#x2013;9</conf-name>; <publisher-loc>Long Beach, CA, USA</publisher-loc>. p. <fpage>6000</fpage>&#x2013;<lpage>10</lpage>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>van den Oord</surname> <given-names>A</given-names></string-name>, <string-name><surname>Vinyals</surname> <given-names>O</given-names></string-name>, <string-name><surname>Kavukcuoglu</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Neural discrete representation learning</article-title>. In: <conf-name>Proceedings of the 31st International Conference on Neural Information Processing Systems; 2017 Dec 4&#x2013;9</conf-name>; <publisher-loc>Long Beach, CA, USA</publisher-loc>. p. <fpage>6309</fpage>&#x2013;<lpage>18</lpage>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Meng</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Peng</surname> <given-names>X</given-names></string-name>, <string-name><surname>Han</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Rethinking diffusion for text-driven human motion generation: redundant representations, evaluation, and masked autoregression</article-title>. In: <conf-name>Proceedings of the 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2025 Jun 10&#x2013;17</conf-name>; <publisher-loc>Nashville, TN, USA</publisher-loc>. p. <fpage>27859</fpage>&#x2013;<lpage>71</lpage>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>M</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Leng</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Li</surname> <given-names>FWB</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Fg-T2M&#x002B;&#x002B;: LLMs-augmented fine-grained text driven human motion generation</article-title>. <source>Int J Comput Vis</source>. <year>2025</year>;<volume>133</volume>(<issue>7</issue>):<fpage>4277</fpage>&#x2013;<lpage>93</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11263-025-02392-9</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Fan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Du</surname> <given-names>B</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>X</given-names></string-name>, <string-name><surname>Peng</surname> <given-names>B</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>L</given-names></string-name></person-group>. <article-title>TextIM: part-aware interactive motion synthesis from text</article-title>. <comment>arXiv:2408.03302. 2024</comment>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Chang</surname> <given-names>CJ</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>QT</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>H</given-names></string-name>, <string-name><surname>Pavlovic</surname> <given-names>V</given-names></string-name>, <string-name><surname>Kapadia</surname> <given-names>M</given-names></string-name></person-group>. <article-title>CASIM: composite aware semantic injection for text to motion generation</article-title>. <comment>arXiv:2502.02063. 2025</comment>. doi:<pub-id pub-id-type="doi">10.48550/arxiv.2502.02063</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>K</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Motion generation from fine-grained textual descriptions</article-title>. In: <conf-name>LREC-COLING 2024-The 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation</conf-name>; <year>2024</year>. p. <fpage>11625</fpage>&#x2013;<lpage>41</lpage>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>N</given-names></string-name>, <string-name><surname>Li</surname> <given-names>S</given-names></string-name>, <string-name><surname>Du</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Torralba</surname> <given-names>A</given-names></string-name>, <string-name><surname>Tenenbaum</surname> <given-names>JB</given-names></string-name></person-group>. <chapter-title>Compositional visual generation with composable diffusion models</chapter-title>. In: <source>Comput vision-ECCV 2022</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer Nature Switzerland</publisher-name>; <year>2022</year>. p. <fpage>423</fpage>&#x2013;<lpage>39</lpage>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Feng</surname> <given-names>W</given-names></string-name>, <string-name><surname>He</surname> <given-names>X</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>TJ</given-names></string-name>, <string-name><surname>Jampani</surname> <given-names>V</given-names></string-name>, <string-name><surname>Akula</surname> <given-names>A</given-names></string-name>, <string-name><surname>Narayana</surname> <given-names>P</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Training-free structured diffusion guidance for compositional text-to-image synthesis</article-title>. <comment>arXiv:2212.05032. 2022</comment>. doi:<pub-id pub-id-type="doi">10.48550/arxiv.2212.05032</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zarei</surname> <given-names>A</given-names></string-name>, <string-name><surname>Rezaei</surname> <given-names>K</given-names></string-name>, <string-name><surname>Basu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Saberi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Moayeri</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kattakinda</surname> <given-names>P</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Improving compositional attribute binding in text-to-image generative models via enhanced text embeddings</article-title>. <comment>arXiv:2406.07844. 2024</comment>. doi:<pub-id pub-id-type="doi">10.48550/arxiv.2406.07844</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Song</surname> <given-names>W</given-names></string-name>, <string-name><surname>Ye</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>M</given-names></string-name>, <string-name><surname>Hou</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hao</surname> <given-names>A</given-names></string-name></person-group>. <article-title>AttriDiffuser: adversarially enhanced diffusion model for text-to-facial attribute image synthesis</article-title>. <source>Pattern Recognit</source>. <year>2025</year>;<volume>163</volume>(<issue>8</issue>):<fpage>111447</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.patcog.2025.111447</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Balaji</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Nah</surname> <given-names>S</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Vahdat</surname> <given-names>A</given-names></string-name>, <string-name><surname>Song</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Q</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>eDiff-I: text-to-image diffusion models with an ensemble of expert denoisers</article-title>. <comment>arXiv:2211.01324. 2022</comment>. doi:<pub-id pub-id-type="doi">10.48550/arxiv.2211.01324</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Cho</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ohana</surname> <given-names>R</given-names></string-name>, <string-name><surname>Jacobsen</surname> <given-names>C</given-names></string-name>, <string-name><surname>Jothi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>MH</given-names></string-name>, <string-name><surname>Mao</surname> <given-names>ZM</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>TC-LoRA: temporally modulated conditional LoRA for adaptive diffusion control</article-title>. <comment>arXiv:2510.09561. 2025</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2510.09561</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Rissanen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Heinonen</surname> <given-names>M</given-names></string-name>, <string-name><surname>Solin</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Generative modelling with inverse heat dissipation</article-title>. In: <conf-name>Proceedings of the Eleventh International Conference on Learning Representations; 2023 May 1&#x2013;5</conf-name>; <publisher-loc>Kigali, Rwanda</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>5</lpage>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Choi</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>J</given-names></string-name>, <string-name><surname>Shin</surname> <given-names>C</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yoon</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Perception prioritized training of diffusion models</article-title>. In: <conf-name>Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2022 Jun 18&#x2013;24</conf-name>; <publisher-loc>New Orleans, LA, USA</publisher-loc>. p. <fpage>11472</fpage>&#x2013;<lpage>81</lpage>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Heusel</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ramsauer</surname> <given-names>H</given-names></string-name>, <string-name><surname>Unterthiner</surname> <given-names>T</given-names></string-name>, <string-name><surname>Nessler</surname> <given-names>B</given-names></string-name>, <string-name><surname>Hochreiter</surname> <given-names>S</given-names></string-name></person-group>. <article-title>GANs trained by a two time-scale update rule converge to a local Nash equilibrium</article-title>. In: <conf-name>Proceedings of the 31st International Conference on Neural Information Processing Systems; 2017 Dec 4&#x2013;9</conf-name>; <publisher-loc>Long Beach, CA, USA</publisher-loc>. p. <fpage>6629</fpage>&#x2013;<lpage>40</lpage>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Petrovich</surname> <given-names>M</given-names></string-name>, <string-name><surname>Black</surname> <given-names>MJ</given-names></string-name>, <string-name><surname>Varol</surname> <given-names>G</given-names></string-name></person-group>. <article-title>TMR: text-to-motion retrieval using contrastive 3D human motion synthesis</article-title>. In: <conf-name>Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV); 2023 Oct 1&#x2013;6</conf-name>; <publisher-loc>Paris, France</publisher-loc>. p. <fpage>9488</fpage>&#x2013;<lpage>97</lpage>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>A</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Hui</surname> <given-names>B</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>B</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>B</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Qwen2.5 technical report</article-title>. <comment>arXiv:2412.15115. 2024</comment>. doi:<pub-id pub-id-type="doi">10.48550/arxiv.2412.15115</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Song</surname> <given-names>J</given-names></string-name>, <string-name><surname>Meng</surname> <given-names>C</given-names></string-name>, <string-name><surname>Ermon</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Denoising diffusion implicit models</article-title>. <comment>arXiv:2010.02502. 2020</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2010.02502</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>