<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">78743</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2026.078743</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>UniModal-LSR: A Unified Multimodal Framework for Joint Lip Reading and Sign Language Recognition in Video Sequences</article-title>
<alt-title alt-title-type="left-running-head">UniModal-LSR: A Unified Multimodal Framework for Joint Lip Reading and Sign Language Recognition in Video Sequences</alt-title>
<alt-title alt-title-type="right-running-head">UniModal-LSR: A Unified Multimodal Framework for Joint Lip Reading and Sign Language Recognition in Video Sequences</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Truong Hoang</surname><given-names>Vinh</given-names></name><xref rid="cor1" ref-type="corresp">&#x002A;</xref><email>vinh.th@ou.edu.vn</email></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Dinh</surname><given-names>Nghia</given-names></name></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Quang Phuong</surname><given-names>Luu</given-names></name></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Tran-Trung</surname><given-names>Kiet</given-names></name></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Duong Thi Hong</surname><given-names>Ha</given-names></name></contrib>
<contrib id="author-6" contrib-type="author">
<name name-style="western"><surname>Nguyen Van</surname><given-names>Bay</given-names></name></contrib>
<contrib id="author-7" contrib-type="author">
<name name-style="western"><surname>Nguyen Trung</surname><given-names>Hau</given-names></name></contrib>
<contrib id="author-8" contrib-type="author">
<name name-style="western"><surname>Ho Huong</surname><given-names>Thien</given-names></name></contrib>
<aff id="aff-1">
<institution>AI Lab, Faculty of Information Technology, Ho Chi Minh City Open University</institution>, <addr-line>35&#x2013;37 Ho Hao Hon Street, Co Giang Ward, District 1, Ho Chi Minh City</addr-line>, <country>Vietnam</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Vinh Truong Hoang. Email: <email>vinh.th@ou.edu.vn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>15</day><month>06</month><year>2026</year>
</pub-date>
<volume>88</volume>
<issue>2</issue>
<elocation-id>35</elocation-id>
<history>
<date date-type="received">
<day>07</day>
<month>01</month>
<year>2026</year>
</date>
<date date-type="accepted">
<day>16</day>
<month>03</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_78743.pdf"></self-uri>
<abstract>
<p>Visual speech recognition is a central problem in computer vision, encompassing both lip reading (visual speech recognition) and sign language recognition. Although substantial progress has been achieved independently on each task, their complementary characteristics have rarely been explored jointly. In this work we propose UniModal-LSR (Unified Multimodal Lip and Sign Recognition), a novel deep learning framework that jointly addresses lip reading and sign language recognition within a single multimodal architecture. By exploiting shared properties of visual communication channels, namely temporal dynamics, spatial articulation structure, and contextual dependencies, the proposed model enables bidirectional transfer of knowledge between modalities. The framework incorporates a Hierarchical Temporal-Spatial Encoder that captures multi-scale temporal patterns through the combination of local convolutions and global self-attention. It also includes a Cross-Modal Attention Fusion module that performs dynamic, context-aware information exchange via bidirectional cross-attention and adaptive gating. Additionally, a Contrastive Semantic Alignment loss enforces semantic consistency across modality-specific representations. Overall, the architecture integrates three-dimensional convolutional neural networks for spatiotemporal feature extraction with graph neural networks for explicit hand-pose modeling. Extensive experiments on several public benchmarks show that UniModal-LSR improves performance compared with recent methods. The model attains a Word Error Rate (WER) of 33.2% on LRS2-BBC, representing a 12.4% relative gain. On PHOENIX-2014, it achieves 18.3% WER, a 13.7% relative gain. Moreover, the unified model reduces parameter count by 25.9% relative to two separate task-specific systems. These results indicate that unified multimodal modeling can improve visual speech recognition performance and may support future communication technologies.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Multimodal learning</kwd>
<kwd>lip reading</kwd>
<kwd>visual speech recognition</kwd>
<kwd>deep learning</kwd>
<kwd>sign language recognition</kwd>
<kwd>cross-modal attention</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Ho Chi Minh City Open University (HCMCOU) and the Ministry of Education and Training (Vietnam)</funding-source>
<award-id>B2025-MBS-01</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Visual speech recognition is an important research area in human-computer interaction, encompassing the complementary tasks of lip reading and sign language recognition. Lip reading, often referred to as visual speech recognition (VSR), involves decoding spoken language from visible articulatory movements of the lips, teeth, and tongue. Sign language recognition (SLR), in contrast, interprets the semantically rich combination of manual gestures, non-manual markers, and spatial grammar that characterize natural sign languages [<xref ref-type="bibr" rid="ref-1">1</xref>]. Although these tasks appear methodologically distinct, they share core computational properties that motivate a unified treatment.</p>
<p>The alignment between lip reading and SLR arises from deep structural similarities in how visual language is conveyed. Both modalities operate at comparable frame rates, typically 25&#x2013;30 frames per second, and require models capable of handling variable-length sequences that may span hundreds of frames [<xref ref-type="bibr" rid="ref-2">2</xref>]. Both depend on fine-grained spatial control of biological articulators, whether the orofacial musculature or the hands and arms. Furthermore, both exhibit strong contextual dependence, as isolated visual patterns are often ambiguous without their temporal and linguistic surroundings, leading to viseme-level confusion in lip reading and coarticulation effects in SLR [<xref ref-type="bibr" rid="ref-3">3</xref>]. These similarities suggest that representations learned from one modality could benefit the other.</p>
<p>In many real-world contexts, both modalities co-occur. Proficient signers often mouth spoken words while signing, providing complementary linguistic cues [<xref ref-type="bibr" rid="ref-4">4</xref>]. Existing systems that process these channels independently disregard this inherent multimodal redundancy, leaving major gains untapped.</p>
<p>Current visual speech recognition techniques face several limitations that hinder practical deployment. Most studies treat lip reading and sign language recognition as separate tasks, resulting in distinct architectures, training pipelines, and evaluation protocols that limit cross-task knowledge transfer and increase computational overhead [<xref ref-type="bibr" rid="ref-5">5</xref>]. Many approaches rely on recurrent models such as LSTMs [<xref ref-type="bibr" rid="ref-6">6</xref>] to capture temporal dependencies, but these models struggle with the long-range dependencies present in continuous sign language sequences and restrict parallelization during training [<xref ref-type="bibr" rid="ref-7">7</xref>]. In addition, conventional CNN-based spatial representations often fail to adequately capture the structural relationships of hand configurations, while multimodal fusion strategies typically rely on simple operations such as concatenation or averaging, which cannot effectively model the dynamic, context-dependent interactions between lip and sign cues.</p>
<p>This work addresses these limitations through the following contributions: We introduce UniModal-LSR, a unified architecture that simultaneously handles lip reading and sign language recognition, enabling bidirectional knowledge transfer and shared representations across modalities. The unified architecture also reduces redundancy by 25.9% and improves generalization. To capture temporal dynamics, we design a Hierarchical Temporal-Spatial Encoder (HTSE) that combines 3D convolutions for local spatiotemporal feature extraction with multi-head self-attention for modeling long-range dependencies, enabling efficient processing of continuous visual speech sequences. We further propose a Cross-Modal Attention Fusion (CMAF) module with bidirectional cross-attention and adaptive gating to dynamically integrate lip and sign cues. A Contrastive Semantic Alignment (CSA) objective aligns modality-specific embeddings within a shared space to enhance semantic consistency and reduce modality-specific noise. Additionally, a Spatial-Temporal Graph Convolutional Network (ST-GCN) models hand skeletal structures to capture joint topology and better distinguish visually similar signs. Extensive experiments on multiple benchmarks demonstrate state-of-the-art performance, and comprehensive ablation studies quantify the contribution of each component; <xref ref-type="table" rid="table-1">Table 1</xref> further clarifies the novelty of the proposed framework.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Analysis of component novelty in UniModal-LSR.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Component</th>
<th>Status</th>
<th>Description</th>
</tr>
</thead>
<tbody>
<tr>
<td>3D ResNet Frontend</td>
<td>Adopted</td>
<td>Standard architecture from [<xref ref-type="bibr" rid="ref-8">8</xref>]; adapted for dual-stream</td>
</tr>
<tr>
<td>Transformer Layers</td>
<td>Adopted</td>
<td>Standard multi-head attention [<xref ref-type="bibr" rid="ref-7">7</xref>]</td>
</tr>
<tr>
<td>ST-GCN Module</td>
<td>Adopted</td>
<td>Architecture from [<xref ref-type="bibr" rid="ref-9">9</xref>]; integrated into unified pipeline</td>
</tr>
<tr>
<td>InfoNCE Loss</td>
<td>Adopted</td>
<td>Standard formulation [<xref ref-type="bibr" rid="ref-10">10</xref>]</td>
</tr>
<tr>
<td>Unified Lip-Sign Architecture</td>
<td>Novel</td>
<td>First framework jointly modeling both tasks</td>
</tr>
<tr>
<td>Hierarchical Multi-scale HTSE</td>
<td>Novel</td>
<td>Specific combination of local conv &#x002B; global attention</td>
</tr>
<tr>
<td>Bidirectional CMAF with Gating</td>
<td>Novel</td>
<td>Dynamic context-aware fusion mechanism for lip-sign modalities</td>
</tr>
<tr>
<td>Cross-Modal CSA for Lip-Sign</td>
<td>Novel</td>
<td>Application of contrastive alignment specifically</td>
</tr>
<tr>
<td>Joint Training Protocol</td>
<td>Novel</td>
<td>Multi-task learning scheme with aligned pair construction</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>To capture temporal patterns at multiple scales, we design a Hierarchical Temporal-Spatial Encoder (HTSE) that fuses three-dimensional convolutions for local spatiotemporal feature extraction with multi-head self-attention for modeling long-range temporal dependencies. Arranging these modules hierarchically facilitates the processing of long continuous visual-speech sequences.</p>
<p>The remainder of the paper is organized as follows. <xref ref-type="sec" rid="s2">Section 2</xref> situates our approach within the existing literature. <xref ref-type="sec" rid="s3">Section 3</xref> details the proposed framework. <xref ref-type="sec" rid="s4">Section 4</xref> presents an architectural and representational analysis. <xref ref-type="sec" rid="s5">Section 5</xref> describes the experimental setup. <xref ref-type="sec" rid="s6">Section 6</xref> reports empirical results. <xref ref-type="sec" rid="s7">Section 7</xref> discusses implications and limitations. <xref ref-type="sec" rid="s8">Section 8</xref> concludes the paper.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<sec id="s2_1">
<label>2.1</label>
<title>Lip Reading and Visual Speech Recognition</title>
<sec id="s2_1_1">
<label>2.1.1</label>
<title>Classical Approaches</title>
<p>Early lip reading methods relied on hand-crafted visual descriptors coupled with statistical sequence models. Potamianos et al. [<xref ref-type="bibr" rid="ref-11">11</xref>] employed discrete cosine transform (DCT) coefficients to encode mouth appearance. Matthews et al. [<xref ref-type="bibr" rid="ref-12">12</xref>] used active appearance models (AAMs) to jointly model shape and texture. These visual representations were typically paired with hidden Markov models (HMMs) for temporal modeling [<xref ref-type="bibr" rid="ref-13">13</xref>]. The recognition task was formally expressed as maximum a posteriori (MAP) inference:<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mrow><mml:mover><mml:mi>W</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mi>arg</mml:mi><mml:mo>&#x2061;</mml:mo><mml:munder><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mi>W</mml:mi></mml:munder><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>W</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>O</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mi>arg</mml:mi><mml:mo>&#x2061;</mml:mo><mml:munder><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mi>W</mml:mi></mml:munder><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>O</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>W</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>W</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>O</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>O</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>o</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>o</mml:mi><mml:mi>T</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denotes the visual frame sequence with each <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi>o</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, and <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>W</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mi>N</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is a hypothesized word sequence within vocabulary <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msup><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow></mml:msup></mml:math></inline-formula>. Although pioneering, these systems suffered from limited expressive power due to the restrictive nature of hand-crafted features.</p>
</sec>
<sec id="s2_1_2">
<label>2.1.2</label>
<title>Deep Learning Approaches</title>
<p>The advent of deep learning has dramatically advanced lip reading. LipNet [<xref ref-type="bibr" rid="ref-2">2</xref>] introduced an effective end-to-end architecture for sentence-level visual speech recognition, employing spatiotemporal convolutions together with connectionist temporal classification (CTC) to reach 93.4% accuracy on the GRID benchmark. The Watch, Listen, Attend and Spell (WLAS) model [<xref ref-type="bibr" rid="ref-14">14</xref>] incorporated attention mechanisms:<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msub><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>h</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:math></disp-formula>allowing the decoder to focus adaptively on relevant encoder states during generation.</p>
<p>Subsequent transformer-based approaches have set new performance benchmarks by exploiting self-attention to model long-range temporal relationships [<xref ref-type="bibr" rid="ref-15">15</xref>]. Ma et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] demonstrated that purely visual models can rival audio-visual systems when trained on sufficiently large data. Self-supervised pre-training strategies [<xref ref-type="bibr" rid="ref-17">17</xref>] have emerged as a means of reducing labeled data requirements. Ma et al. [<xref ref-type="bibr" rid="ref-18">18</xref>] introduced Auto-AVSR, which leverages automatic labels for audio-visual speech recognition, demonstrating the effectiveness of weakly supervised learning approaches. Shukla et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] investigated whether visual self-supervision improves speech representations for emotion recognition, providing insights into cross-task transfer learning. Recent work by Prajwal et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] achieved significant improvements through sub-word modeling.</p>
</sec>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Sign Language Recognition</title>
<p>Isolated sign language recognition (ISLR) focuses on classifying pre-segmented signs. The Inflated 3D ConvNet (I3D) [<xref ref-type="bibr" rid="ref-8">8</xref>] extended successful 2D CNN designs into the temporal dimension via filter inflation:<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mrow><mml:mn>3</mml:mn><mml:mi>D</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>h</mml:mi><mml:mo>,</mml:mo><mml:mi>w</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:mi>D</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>h</mml:mi><mml:mo>,</mml:mo><mml:mi>w</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:math></disp-formula>preserving spatial semantics while averaging over the temporal axis.</p>
<p>Skeleton-based methods exploit hand-joint coordinates to model structural relationships directly. Attention-driven approaches such as SignBERT [<xref ref-type="bibr" rid="ref-21">21</xref>] illustrate the benefits of large-scale pre-training on sign language corpora.</p>
<p>Continuous sign language recognition (CSLR) jointly performs segmentation and classification in unsegmented video streams. The Connectionist Temporal Classification (CTC) loss [<xref ref-type="bibr" rid="ref-22">22</xref>] is widely employed:<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>T</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>Y</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>X</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:munder><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>&#x03C0;</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi>&#x0212C;</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>Y</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:munder><mml:munderover><mml:mo movablelimits="false">&#x220F;</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:munderover><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>X</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>enabling alignment-free training.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Multimodal Learning and Fusion</title>
<p>Multimodal fusion is typically categorized into early (feature-level), late (decision-level), and intermediate (shared representation) strategies. Early fusion preserves detailed cross-modal interactions but yields high-dimensional feature spaces. Late fusion maintains modality-specific pipelines but limits interaction depth. Intermediate fusion balances both aspects.</p>
<p>Attention-based fusion has become prevalent for modeling dynamic relationships. Hierarchical co-attention [<xref ref-type="bibr" rid="ref-23">23</xref>] and multimodal transformers [<xref ref-type="bibr" rid="ref-24">24</xref>] exemplify recent progress. Contrastive learning techniques such as CLIP [<xref ref-type="bibr" rid="ref-25">25</xref>] further illustrate the efficacy of aligning multimodal representations.</p>
<p>Especially relevant to this study, Ge et al. [<xref ref-type="bibr" rid="ref-26">26</xref>] proposed an audio-text multimodal framework for speech recognition in air traffic control communications, showing that unified multimodal modeling can improve recognition accuracy through coordinated processing of audio and text. Similarly, Li et al. [<xref ref-type="bibr" rid="ref-27">27</xref>] developed an end-to-end audio-visual system for multi-channel speech separation, dereverberation, and recognition, while Wang et al. [<xref ref-type="bibr" rid="ref-28">28</xref>] introduced DCIM-AVSR, an efficient audio-visual speech recognition model using dual conformer interaction modules. However, these studies focus on audio-visual fusion. In contrast, our work extends the unified architecture paradigm to visual-visual modality fusion, addressing the integration of lip and sign information. The proposed CMAF module enables bidirectional cross-attention with adaptive gating, making it well suited for the temporal synchronization required in visual speech modalities.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Proposed Methodology</title>
<p>This section describes the UniModal-LSR architecture in detail. An overview is provided in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>System architecture of UniModal-LSR. The framework processes video through parallel lip and sign encoding streams, applies hierarchical temporal-spatial encoding, performs cross-modal attention fusion, and generates task-specific outputs via a shared transformer decoder.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_78743-fig-1.tif"/>
</fig>
<sec id="s3_1">
<label>3.1</label>
<title>Problem Formulation</title>
<p>Given an input video <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mrow><mml:mi mathvariant="bold">V</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">I</mml:mi></mml:mrow><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">I</mml:mi></mml:mrow><mml:mn>2</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">I</mml:mi></mml:mrow><mml:mi>T</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>, where <italic>T</italic> denotes the number of frames and each frame <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mrow><mml:mi mathvariant="bold">I</mml:mi></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, the objective includes lip reading, sign language recognition, and unified representation learning. The overall training loss combines task-specific objectives with cross-modal regularization:<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>total</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mtext>sign</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>sign</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mtext>CSA</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>CSA</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mtext>reg</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>reg</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mtext>sign</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> denote task-specific losses (CTC and cross-entropy), <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mtext>CSA</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> is the contrastive semantic alignment loss defined in <xref ref-type="disp-formula" rid="eqn-17">Eq. (17)</xref>, and <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mtext>reg</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> is an <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mi>&#x2113;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> weight-decay term. The non-negative scalars <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msub></mml:math></inline-formula> balance each term&#x2019;s contribution.</p>
<p>The CSA loss requires aligned lip-sign pairs that share semantic content. We construct such pairs from two sources. The first source is the How2Sign dataset, which provides natural co-occurrence of lip and sign modalities for the same utterances, enabling direct pairing without additional processing. The second source involves synthetic alignment of LRS2/LRS3 lip sequences with PHOENIX-2014 sign sequences when they share identical or semantically equivalent transcriptions, as determined by text matching with a minimum overlap threshold of 80%. For batches containing samples from only one modality, such as lip-only data from LRS2, the CSA loss is computed only over the subset of aligned pairs present in that batch. When no aligned pairs exist in a batch, the CSA term is set to zero for that iteration, and the model optimizes only the task-specific losses.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Preprocessing and Region Extraction</title>
<p>Accurate face detection is essential for effective lip reading. RetinaFace yields facial landmarks <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mrow><mml:mi mathvariant="bold">L</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mn>68</mml:mn></mml:mrow></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> under the standard 68-point annotation protocol. We obtain the lip patch <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">R</mml:mi></mml:mrow><mml:mi>t</mml:mi><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mo>&#x00D7;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> through affine alignment:<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">R</mml:mi></mml:mrow><mml:mi>t</mml:mi><mml:mrow><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D4AF;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>affine</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">I</mml:mi></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">L</mml:mi></mml:mrow><mml:mrow><mml:mn>48</mml:mn><mml:mo>:</mml:mo><mml:mn>67</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mrow><mml:mi mathvariant="bold">L</mml:mi></mml:mrow><mml:mrow><mml:mn>48</mml:mn><mml:mo>:</mml:mo><mml:mn>67</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> correspond to the outer and inner lip contour landmarks. The patch is normalized to <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mi>H</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn>88</mml:mn></mml:math></inline-formula> pixels.</p>
<p>For SLR, hand pose estimation provides structured motion cues. MediaPipe Hands [<xref ref-type="bibr" rid="ref-29">29</xref>] predicts 21 three-dimensional landmarks per hand. The skeleton is modeled as a spatio-temporal graph <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mrow><mml:mi mathvariant="bold">G</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mi>&#x2130;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, where <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mn>42</mml:mn></mml:mrow></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> comprises landmarks of both hands and <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mrow><mml:mi>&#x2130;</mml:mi></mml:mrow></mml:math></inline-formula> encodes anatomical connectivity. Each node <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> at time <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi>t</mml:mi></mml:math></inline-formula> has feature vector:<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">v</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>z</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:msup><mml:mo stretchy="false">]</mml:mo><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mn>4</mml:mn></mml:msup><mml:mo>,</mml:mo></mml:math></disp-formula>with <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>y</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>z</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denoting normalized coordinates and <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msubsup><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> the confidence score.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Hierarchical Temporal-Spatial Encoder (HTSE)</title>
<p>The HTSE captures multi-scale temporal dynamics through cascaded processing stages. An initial 3D ResNet-18 backbone extracts low-level spatiotemporal features:<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:msup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mtext>ResNet3D</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="bold">R</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mi>T</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:msup><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:mrow></mml:msup><mml:mo>,</mml:mo></mml:math></disp-formula>where <italic>T</italic><sup>&#x2032;</sup> accounts for temporal subsampling and <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mi>d</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn>512</mml:mn></mml:math></inline-formula> denotes feature dimensionality. Thereafter <italic>L</italic> hierarchical levels successively apply local temporal convolutions, global self-attention, feed-forward networks, and temporal pooling. This yields progressively coarser temporal resolutions while preserving fine-grained information. A Feature Pyramid Network style aggregation fuses multi-scale outputs:<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:msup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>agg</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mtext>Upsample</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi>T</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>where learned coefficients <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:math></inline-formula> weight each scale adaptively.</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Graph-Enhanced Sign Encoding</title>
<p>To complement appearance-based features, a spatial-temporal graph convolutional network (ST-GCN) models hand articulation. The spatial graph convolution aggregates information from anatomically connected joints:<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">f</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>out</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>v</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>u</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mi>&#x1D4A9;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>v</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:munder><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:msub><mml:mi>Z</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mi>u</mml:mi></mml:mrow></mml:msub></mml:mfrac></mml:mstyle><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">f</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>in</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>u</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x2113;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>u</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi>&#x2113;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>v</mml:mi><mml:mo>,</mml:mo><mml:mi>u</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> distinguishes edge types (centripetal, centrifugal, self-connections) and <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msub><mml:mi>Z</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mi>u</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is a normalization factor. Temporal graph convolutions then capture motion dynamics across frames. Resulting pose features <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mtext>pose</mml:mtext></mml:mrow></mml:msup></mml:math></inline-formula> are concatenated with appearance features and projected to a shared dimensionality before entering the HTSE.</p>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Cross-Modal Attention Fusion (CMAF)</title>
<p>CMAF proceeds in three stages. First, intra-modal self-attention refines each modality independently:<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>self</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mtext>LayerNorm</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:mrow><mml:mtext>MHSA</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:mi>m</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>sign</mml:mtext></mml:mrow><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula>where MHSA denotes multi-head self-attention. Second, bidirectional cross-attention allows each modality to query information from the other:<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mrow><mml:mtext>sign</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mtext>MHCA</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>self</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>self</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>sign</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>self</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>sign</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>sign</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mtext>MHCA</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>self</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>sign</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>self</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>self</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where MHCA denotes multi-head cross-attention with the first argument as query and the second/third as key/value. Third, adaptive gating computes element-wise weights <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mrow><mml:mi mathvariant="bold">g</mml:mi></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mi>d</mml:mi></mml:msup></mml:math></inline-formula>:<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mrow><mml:mi mathvariant="bold">g</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:msub><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msub><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>self</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>;</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>self</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>sign</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mrow><mml:mtext>sign</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo>;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>sign</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo stretchy="false">]</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">b</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msub><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msub><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>4</mml:mn><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msub><mml:mrow><mml:mi mathvariant="bold">b</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mi>d</mml:mi></mml:msup></mml:math></inline-formula>, and <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mi>&#x03C3;</mml:mi></mml:math></inline-formula> is the sigmoid function. The fused representation is:<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:msup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>fused</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mi mathvariant="bold">g</mml:mi></mml:mrow><mml:mo>&#x2299;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>self</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>sign</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mi mathvariant="bold">g</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2299;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>self</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>sign</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mrow><mml:mtext>sign</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo>.</mml:mo></mml:math></disp-formula></p>
</sec>
<sec id="s3_6">
<label>3.6</label>
<title>Contrastive Semantic Alignment (CSA) Loss</title>
<p>For a minibatch of <italic>B</italic> aligned lip-sign pairs, modality-specific embeddings are normalized:<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">z</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>proj</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mspace width="thinmathspace" /><mml:msubsup><mml:mrow><mml:mover><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow></mml:mrow></mml:msubsup></mml:mrow><mml:mrow><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>proj</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mspace width="thinmathspace" /><mml:msubsup><mml:mrow><mml:mover><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">z</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>sign</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>proj</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>sign</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mspace width="thinmathspace" /><mml:msubsup><mml:mrow><mml:mover><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>sign</mml:mtext></mml:mrow></mml:mrow></mml:msubsup></mml:mrow><mml:mrow><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>proj</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>sign</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mspace width="thinmathspace" /><mml:msubsup><mml:mrow><mml:mover><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>sign</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msubsup><mml:mrow><mml:mover><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> denotes the temporally pooled representation for sample <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mi>i</mml:mi></mml:math></inline-formula> and modality <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mi>m</mml:mi></mml:math></inline-formula>. A symmetric InfoNCE objective enforces alignment:<disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>CSA</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mn>2</mml:mn><mml:mi>B</mml:mi></mml:mrow></mml:mfrac><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:munderover><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="2.470em" minsize="2.470em">[</mml:mo></mml:mrow></mml:mstyle><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">z</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x22C5;</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">z</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>sign</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>&#x03C4;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msubsup><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">z</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x22C5;</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">z</mml:mi></mml:mrow><mml:mi>j</mml:mi><mml:mrow><mml:mrow><mml:mtext>sign</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>&#x03C4;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>+</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">z</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>sign</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x22C5;</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">z</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>&#x03C4;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msubsup><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">z</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mrow><mml:mtext>sign</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x22C5;</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">z</mml:mi></mml:mrow><mml:mi>j</mml:mi><mml:mrow><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>&#x03C4;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="2.470em" minsize="2.470em">]</mml:mo></mml:mrow></mml:mstyle><mml:mo>,</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> is a temperature hyperparameter.</p>
</sec>
<sec id="s3_7">
<label>3.7</label>
<title>Training Objective</title>
<p>The complete training loss integrates a CTC component with a cross-entropy term:<disp-formula id="eqn-18"><label>(18)</label><mml:math id="mml-eqn-18" display="block"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>task</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mspace width="thinmathspace" /><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>CTC</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>CE</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>The final optimization objective is:<disp-formula id="eqn-19"><label>(19)</label><mml:math id="mml-eqn-19" display="block"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>total</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>lip</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>task</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>sign</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>task</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mtext>CSA</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>CSA</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mtext>reg</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:msubsup><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn><mml:mn>2</mml:mn></mml:msubsup><mml:mo>.</mml:mo></mml:math></disp-formula></p><p>The overall training procedure is described in Algorithm 1.</p>
<fig id="fig-6">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_78743-fig-6.tif"/>
</fig>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Architectural and Representational Analysis</title>
<p>This section presents an analysis of the architectural properties and representational capacity of the proposed framework. Rather than restating general theoretical results, we focus on aspects directly relevant to the CMAF and HTSE designs that justify our specific architectural choices.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Representational Capacity of Cross-Modal Attention</title>
<p>The CMAF module&#x2019;s effectiveness derives from its ability to model complex interactions between lip and sign modalities. We formalize this capacity in terms of the function classes that the architecture can represent.</p>
<p><bold>Proposition 1</bold> (Expressive Capacity of Cross-Modal Attention). <italic>Let <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:msub><mml:mrow><mml:mi>&#x2131;</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denote the family of continuous functions mapping paired sequences <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mi mathvariant="bold">y</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:mrow></mml:msup><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula> to fused representations in <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>o</mml:mi></mml:msub></mml:mrow></mml:msup></mml:math></inline-formula>, where cross-modal dependencies are bounded by a Lipschitz constant <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:msub><mml:mi>L</mml:mi><mml:mi>f</mml:mi></mml:msub></mml:math></inline-formula>. For any <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mi>f</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x2131;</mml:mi></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mi>&#x03F5;</mml:mi><mml:mo>&#x003E;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>, the CMAF architecture with H attention heads, hidden dimension <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:msub><mml:mi>d</mml:mi><mml:mi>h</mml:mi></mml:msub></mml:math></inline-formula>, and N layers can approximate <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mi>f</mml:mi></mml:math></inline-formula> within error <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:mi>&#x03F5;</mml:mi></mml:math></inline-formula> when:</italic>
<disp-formula id="eqn-20"><label>(20)</label><mml:math id="mml-eqn-20" display="block"><mml:mi>H</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>h</mml:mi></mml:msub><mml:mo>&#x2265;</mml:mo><mml:mi>C</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mi>f</mml:mi></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>&#x03F5;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula></p>
<p><italic>where C is a constant depending on input dimensionality</italic>.</p>
<p><bold>Proof Sketch.</bold> The bidirectional cross-attention mechanism in CMAF can be decomposed into four components: intra-modal self-attention capturing within-modality dependencies, lip-to-sign cross-attention, sign-to-lip cross-attention, and adaptive gating for combination. Each cross-attention operation computes:
<disp-formula id="eqn-21"><label>(21)</label><mml:math id="mml-eqn-21" display="block"><mml:mrow><mml:mtext>CrossAttn</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="bold">Q</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mi mathvariant="bold">K</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mi mathvariant="bold">V</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mtext>softmax</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mrow><mml:mi mathvariant="bold">Q</mml:mi></mml:mrow><mml:msup><mml:mrow><mml:mi mathvariant="bold">K</mml:mi></mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msup></mml:mrow><mml:msqrt><mml:msub><mml:mi>d</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:msqrt></mml:mfrac></mml:mstyle><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi mathvariant="bold">V</mml:mi></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>By the results of Yun et al. [<xref ref-type="bibr" rid="ref-30">30</xref>], transformer architectures with softmax attention are universal approximators for sequence-to-sequence functions. The CMAF extends this by enabling queries from one modality to attend to keys and values from another, effectively doubling the functional space that can be represented. The adaptive gating mechanism <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mrow><mml:mi mathvariant="bold">g</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">W</mml:mi></mml:mrow><mml:mi>g</mml:mi></mml:msub><mml:mo stretchy="false">[</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">]</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> provides an additional multiplicative interaction that allows context-dependent weighting. Combined, these mechanisms can represent any continuous cross-modal function within the specified bounds.&#x2002;<inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:mi>&#x25FB;</mml:mi></mml:math></inline-formula></p>
<p>This result justifies our architectural choice of bidirectional cross-attention: the design ensures that both modalities can contribute information to the fused representation, while the adaptive gating allows the model to learn when each modality is most informative for a given context.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Effective Receptive Field of HTSE</title>
<p>The hierarchical structure of HTSE is designed to capture temporal dependencies at multiple scales efficiently. We analyze how the receptive field grows with network depth and how this relates to the temporal extent of visual speech phenomena.</p>
<p><bold>Lemma 1</bold> (Hierarchical Receptive Field Growth). <italic>An HTSE with L hierarchical levels, local convolution kernel size <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mi>k</mml:mi></mml:math></inline-formula>, and pooling factor <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn></mml:math></inline-formula> achieves an effective temporal receptive field of:</italic>
<disp-formula id="eqn-22"><label>(22)</label><mml:math id="mml-eqn-22" display="block"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mtext>eff</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>k</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:msup><mml:mi>p</mml:mi><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>=</mml:mo><mml:mi>k</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:math></disp-formula></p>
<p><italic>while the parameter count grows as <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>L</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mi>d</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, where <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mi>d</mml:mi></mml:math></inline-formula> is the hidden dimension</italic>.</p>
<p><bold>Proof.</bold> At level <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mi>l</mml:mi></mml:math></inline-formula>, the temporal resolution is reduced by factor <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:msup><mml:mi>p</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> relative to the input. A convolution with kernel <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mi>k</mml:mi></mml:math></inline-formula> at level <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mi>l</mml:mi></mml:math></inline-formula> therefore covers <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mi>k</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mi>p</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> frames of the original input. Summing over all levels:<disp-formula id="eqn-23"><label>(23)</label><mml:math id="mml-eqn-23" display="block"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mtext>eff</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:munderover><mml:mi>k</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mi>p</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>k</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>L</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:munderover><mml:msup><mml:mi>p</mml:mi><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mi>k</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mfrac><mml:mrow><mml:msup><mml:mi>p</mml:mi><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>p</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>For <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn></mml:math></inline-formula>, this simplifies to <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:mi>k</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msup><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. Each level contains convolution layers with <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mi>d</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> parameters and attention layers with <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>d</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> parameters, yielding total parameter growth of <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>L</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mi>d</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mi>&#x25FB;</mml:mi></mml:math></inline-formula></p>
<p>For our configuration with <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:mn>4</mml:mn></mml:math></inline-formula> levels and kernel size <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>3</mml:mn></mml:math></inline-formula>, the effective receptive field is <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mtext>eff</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>15</mml:mn><mml:mo>=</mml:mo><mml:mn>45</mml:mn></mml:math></inline-formula> frames. Combined with global self-attention at each level, the model can capture both local articulation patterns spanning a few frames and long-range dependencies extending across entire utterances. This design is motivated by the observation that lip movements exhibit local coarticulation effects while sign language requires understanding of phrase-level context.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Computational Complexity</title>
<p><xref ref-type="table" rid="table-2">Table 2</xref> summarizes the asymptotic complexities of each component. All symbols are defined as follows: <italic>T</italic> denotes input sequence length in frames, <italic>T<sup>&#x2032;</sup></italic> is the encoded sequence length after 3D CNN downsampling, <italic>H</italic> and <italic>W</italic> are spatial dimensions of input frames, <italic>C</italic> is the number of channels, <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:mi>k</mml:mi></mml:math></inline-formula> is the convolution kernel size, <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:mi>d</mml:mi></mml:math></inline-formula> is the hidden dimension, <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:msub><mml:mi>N</mml:mi><mml:mi>v</mml:mi></mml:msub></mml:math></inline-formula> is the number of graph vertices representing hand joints, <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:msub><mml:mi>N</mml:mi><mml:mi>d</mml:mi></mml:msub></mml:math></inline-formula> is the number of decoder layers, and <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:msub><mml:mi>T</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:math></inline-formula> is the output sequence length.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Computational complexity analysis.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Component</th>
<th>Time Complexity</th>
<th>Space Complexity</th>
</tr>
</thead>
<tbody>
<tr>
<td>3D CNN Frontend</td>
<td><inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>H</mml:mi><mml:mi>W</mml:mi><mml:msup><mml:mi>C</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mi>k</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>C</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mi>k</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td>HTSE (per level)</td>
<td><inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mi>d</mml:mi><mml:mo>+</mml:mo><mml:msup><mml:mi>T</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:msup><mml:msup><mml:mi>d</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mi>d</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td>ST-GCN Module</td>
<td><inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:msubsup><mml:mi>N</mml:mi><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mi>d</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>N</mml:mi><mml:mrow><mml:mi>v</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:mi>d</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td>CMAF</td>
<td><inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mi>d</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mi>d</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td>Transformer Decoder</td>
<td><inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:msubsup><mml:mi>T</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mi>d</mml:mi><mml:mo>+</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mi>d</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-100"><mml:math id="mml-ieqn-100"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mi>d</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
</tr>
<tr>
<td>Total</td>
<td><inline-formula id="ieqn-101"><mml:math id="mml-ieqn-101"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mi>d</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>H</mml:mi><mml:mi>W</mml:mi><mml:msup><mml:mi>C</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mi>k</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-102"><mml:math id="mml-ieqn-102"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mi>C</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mi>k</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The dominant term is the quadratic self-attention complexity <inline-formula id="ieqn-103"><mml:math id="mml-ieqn-103"><mml:mi>O</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>T</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mi>d</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. For typical video lengths where <inline-formula id="ieqn-104"><mml:math id="mml-ieqn-104"><mml:msup><mml:mi>T</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:msup><mml:mo>&#x003C;</mml:mo><mml:mn>500</mml:mn></mml:math></inline-formula> frames and our hidden dimension of <inline-formula id="ieqn-105"><mml:math id="mml-ieqn-105"><mml:mi>d</mml:mi><mml:mo>=</mml:mo><mml:mn>512</mml:mn></mml:math></inline-formula>, this remains computationally tractable. The unified architecture achieves computational savings relative to separate models by sharing the decoder and fusing features before decoding, rather than maintaining separate decoders for each task.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Experimental Setup</title>
<p><xref ref-type="table" rid="table-3">Table 3</xref> enumerates the benchmark datasets used in our evaluation. <xref ref-type="fig" rid="fig-2">Figs. 2</xref> and <xref ref-type="fig" rid="fig-3">3</xref> show example frames from the lip reading and sign language datasets, respectively. LRS2-BBC [<xref ref-type="bibr" rid="ref-14">14</xref>], LRS3-TED [<xref ref-type="bibr" rid="ref-31">31</xref>], and GRID [<xref ref-type="bibr" rid="ref-32">32</xref>] provide sentence- and word-level lip reading benchmarks collected from BBC broadcasts, TED talks, and controlled laboratory environments. For sign language recognition and translation, PHOENIX-2014 [<xref ref-type="bibr" rid="ref-33">33</xref>] and its translation counterpart PHOENIX-2014T [<xref ref-type="bibr" rid="ref-1">1</xref>] serve as standard continuous German Sign Language benchmarks, while CSL [<xref ref-type="bibr" rid="ref-34">34</xref>] and WLASL [<xref ref-type="bibr" rid="ref-35">35</xref>] provide isolated sign datasets for Chinese and American Sign Language, respectively. Finally, How2Sign [<xref ref-type="bibr" rid="ref-36">36</xref>] enables multimodal and cross-modal experiments as it contains synchronized lip and sign annotations for the same utterances. Performance for lip reading and continuous sign language recognition is measured using Word Error Rate (WER).</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Dataset statistics and characteristics.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Dataset</th>
<th>Task</th>
<th>Hours</th>
<th>Vocab</th>
<th>Samples</th>
<th>Signers</th>
</tr>
</thead>
<tbody>
<tr>
<td align="center" colspan="6"><italic>Lip Reading Benchmarks</italic></td>
</tr>
<tr>
<td>LRS2-BBC [<xref ref-type="bibr" rid="ref-14">14</xref>]</td>
<td>VSR</td>
<td>224</td>
<td>41K</td>
<td>144K</td>
<td>1000&#x002B;</td>
</tr>
<tr>
<td>LRS3-TED [<xref ref-type="bibr" rid="ref-31">31</xref>]</td>
<td>VSR</td>
<td>438</td>
<td>51K</td>
<td>151K</td>
<td>5000&#x002B;</td>
</tr>
<tr>
<td>GRID [<xref ref-type="bibr" rid="ref-32">32</xref>]</td>
<td>VSR</td>
<td>28</td>
<td>51</td>
<td>34K</td>
<td>34</td>
</tr>
<tr>
<td align="center" colspan="6"><italic>Sign Language Benchmarks</italic></td>
</tr>
<tr>
<td>PHOENIX-14 [<xref ref-type="bibr" rid="ref-33">33</xref>]</td>
<td>CSLR</td>
<td>11</td>
<td>1066</td>
<td>6,841</td>
<td>9</td>
</tr>
<tr>
<td>PHOENIX-14T [<xref ref-type="bibr" rid="ref-1">1</xref>]</td>
<td>SLT</td>
<td>11</td>
<td>1066</td>
<td>8,257</td>
<td>9</td>
</tr>
<tr>
<td>CSL [<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td>ISLR</td>
<td>100</td>
<td>500</td>
<td>25K</td>
<td>50</td>
</tr>
<tr>
<td>WLASL [<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>ISLR</td>
<td>&#x2013;</td>
<td>2000</td>
<td>21K</td>
<td>119</td>
</tr>
<tr>
<td align="center" colspan="6"><italic>Multimodal Benchmark</italic></td>
</tr>
<tr>
<td>How2Sign [<xref ref-type="bibr" rid="ref-36">36</xref>]</td>
<td>Multi</td>
<td>79</td>
<td>&#x2013;</td>
<td>35K</td>
<td>11</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Example frames from lip reading datasets showing extracted lip ROIs.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_78743-fig-2.tif"/>
</fig><fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Example frames from sign language datasets showing hand pose estimation.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_78743-fig-3.tif"/>
</fig>
<p>Face detection uses RetinaFace (confidence 0.9), with lip ROIs extracted from facial landmarks <inline-formula id="ieqn-106"><mml:math id="mml-ieqn-106"><mml:mn>48</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo></mml:mrow><mml:mn>67</mml:mn></mml:math></inline-formula> and resized to <inline-formula id="ieqn-107"><mml:math id="mml-ieqn-107"><mml:mn>88</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>88</mml:mn></mml:math></inline-formula> pixels via affine normalization. Hand pose estimation uses MediaPipe Hands (confidence 0.7), and the fewer than 2% of failed detections are filled by linear interpolation to maintain temporal continuity. Data augmentation is applied asymmetrically: horizontal flipping is used only for sign language due to handedness constraints, while both modalities use temporal jittering <inline-formula id="ieqn-108"><mml:math id="mml-ieqn-108"><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x00B1;</mml:mo><mml:mn>2</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> frames, color jittering <inline-formula id="ieqn-109"><mml:math id="mml-ieqn-109"><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x00B1;</mml:mo><mml:mn>0.2</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, and random erasing <inline-formula id="ieqn-110"><mml:math id="mml-ieqn-110"><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>0.1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. Decoding employs beam search (width 10, length normalization 0.6) without an external language model; tokenization uses character-level encoding for lip reading (28 tokens) and BPE with 1000 merges for sign language.</p>
<p>For reproducibility, all experiments use random seed 42, and results averaged over three runs show a standard deviation below 0.3% WER. Official dataset splits are used for all benchmarks, and the code along with pretrained models will be released upon publication to support replication and further research.</p>
</sec>
<sec id="s6">
<label>6</label>
<title>Results and Analysis</title>
<p><xref ref-type="table" rid="table-4">Table 4</xref> reports Word Error Rates on lip reading benchmarks, including recent baselines from 2022&#x2013;2023 for comprehensive comparison. All results are reported without external language models to ensure fair comparison. The unified model achieves the lowest WER across all evaluated datasets, with relative improvements of 5.7% on LRS2 and 5.1% on LRS3 compared to the strongest prior work.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Lip reading results (WER%; lower is better). All results without external language model.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Method</th>
<th>Year</th>
<th>LRS2</th>
<th>LRS3</th>
<th>GRID</th>
</tr>
</thead>
<tbody>
<tr>
<td>LipNet [<xref ref-type="bibr" rid="ref-2">2</xref>]</td>
<td>2016</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>4.8</td>
</tr>
<tr>
<td>TM-seq2seq [<xref ref-type="bibr" rid="ref-37">37</xref>]</td>
<td>2018</td>
<td>49.8</td>
<td>58.9</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>DC-TCN [<xref ref-type="bibr" rid="ref-38">38</xref>]</td>
<td>2020</td>
<td>44.3</td>
<td>47.1</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Ma et al. [<xref ref-type="bibr" rid="ref-15">15</xref>]</td>
<td>2021</td>
<td>37.9</td>
<td>43.3</td>
<td>1.2</td>
</tr>
<tr>
<td>Ma et al. [<xref ref-type="bibr" rid="ref-16">16</xref>]</td>
<td>2022</td>
<td>35.2</td>
<td>40.8</td>
<td>1.0</td>
</tr>
<tr>
<td>Prajwal et al. [<xref ref-type="bibr" rid="ref-20">20</xref>]</td>
<td>2022</td>
<td>34.8</td>
<td>39.5</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Kim et al. [<xref ref-type="bibr" rid="ref-39">39</xref>]</td>
<td>2022</td>
<td>34.5</td>
<td>39.2</td>
<td>0.9</td>
</tr>
<tr>
<td>UniModal-LSR (Lip only)</td>
<td>2025</td>
<td>35.6</td>
<td>41.2</td>
<td>0.9</td>
</tr>
<tr>
<td>UniModal-LSR (Full)</td>
<td>2025</td>
<td>33.2</td>
<td>38.7</td>
<td>0.8</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-5">Table 5</xref> shows results for continuous and isolated sign language recognition. The unified model achieves state-of-the-art performance on PHOENIX-14 with 18.3% WER, representing a 13.7% relative improvement over the previous best result.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Sign language recognition results.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Method</th>
<th>PHOENIX-14 (WER%)</th>
<th>CSL (Acc%)</th>
<th>WLASL (Acc%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>CNN &#x002B; LSTM &#x002B; HMM [<xref ref-type="bibr" rid="ref-40">40</xref>]</td>
<td>26.0</td>
<td>91.2</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>STMC [<xref ref-type="bibr" rid="ref-41">41</xref>]</td>
<td>21.1</td>
<td>94.6</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>FCN [<xref ref-type="bibr" rid="ref-42">42</xref>]</td>
<td>23.3</td>
<td>93.1</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>VAC [<xref ref-type="bibr" rid="ref-43">43</xref>]</td>
<td>21.2</td>
<td>95.2</td>
<td>65.8</td>
</tr>
<tr>
<td>UniModal-LSR (Sign only)</td>
<td>19.8</td>
<td>95.8</td>
<td>68.4</td>
</tr>
<tr>
<td>UniModal-LSR (Full)</td>
<td>18.3</td>
<td>96.5</td>
<td>70.2</td>
</tr>
<tr>
<td>Relative Improvement</td>
<td>13.7%</td>
<td>1.4%</td>
<td>6.7%</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-6">Table 6</xref> presents BLEU scores on PHOENIX-2014T. Our model achieves 24.89% BLEU-4, representing an improvement of 2.25 points over the previous state of the art.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Sign language translation results on PHOENIX-2014T.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Method</th>
<th>BLEU-1</th>
<th>BLEU-2</th>
<th>BLEU-3</th>
<th>BLEU-4</th>
</tr>
</thead>
<tbody>
<tr>
<td>Sign2Text [<xref ref-type="bibr" rid="ref-1">1</xref>]</td>
<td>44.13</td>
<td>31.47</td>
<td>23.89</td>
<td>19.26</td>
</tr>
<tr>
<td>Sign2(G&#x002B;T) [<xref ref-type="bibr" rid="ref-44">44</xref>]</td>
<td>47.26</td>
<td>34.40</td>
<td>26.31</td>
<td>21.32</td>
</tr>
<tr>
<td>TSPNet [<xref ref-type="bibr" rid="ref-45">45</xref>]</td>
<td>48.32</td>
<td>35.89</td>
<td>28.12</td>
<td>22.64</td>
</tr>
<tr>
<td>UniModal-LSR</td>
<td>51.24</td>
<td>38.67</td>
<td>30.45</td>
<td>24.89</td>
</tr>
<tr>
<td>Improvement</td>
<td>&#x002B;2.92</td>
<td>&#x002B;2.78</td>
<td>&#x002B;2.33</td>
<td>&#x002B;2.25</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-7">Table 7</xref> analyzes each component&#x2019;s contribution through cumulative addition, starting from a single-modality baseline and progressively adding each proposed component.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Ablation study: component contributions (WER%).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Configuration</th>
<th>LRS2</th>
<th>PH-14</th>
</tr>
</thead>
<tbody>
<tr>
<td>Single modality</td>
<td>37.9</td>
<td>21.2</td>
</tr>
<tr>
<td>&#x002B; HTSE</td>
<td>35.8 (&#x2212;2.1)</td>
<td>20.1 (&#x2212;1.1)</td>
</tr>
<tr>
<td>&#x002B; GNN</td>
<td>&#x2013;</td>
<td>19.4 (&#x2212;0.7)</td>
</tr>
<tr>
<td>&#x002B; CMAF</td>
<td>34.2 (&#x2212;1.6)</td>
<td>18.8 (&#x2212;0.6)</td>
</tr>
<tr>
<td>&#x002B; CSA loss</td>
<td>33.2 (&#x2212;1.0)</td>
<td>18.3 (&#x2212;0.5)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-8">Table 8</xref> provides detailed analysis of CMAF design choices, examining the impact of attention directionality and gating mechanisms.</p>
<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>CMAF structural ablation on LRS2.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>CMAF Variant</th>
<th>WER (%)</th>
<th><inline-formula id="ieqn-111"><mml:math id="mml-ieqn-111"><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:math></inline-formula></th>
</tr>
</thead>
<tbody>
<tr>
<td>No cross-modal fusion (concat only)</td>
<td>35.8</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Unidirectional: Lip<inline-formula id="ieqn-112"><mml:math id="mml-ieqn-112"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula>Sign only</td>
<td>34.8</td>
<td>&#x2212;1.0</td>
</tr>
<tr>
<td>Unidirectional: Sign<inline-formula id="ieqn-113"><mml:math id="mml-ieqn-113"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula>Lip only</td>
<td>34.6</td>
<td>&#x2212;1.2</td>
</tr>
<tr>
<td>Bidirectional (no gating)</td>
<td>34.1</td>
<td>&#x2212;1.7</td>
</tr>
<tr>
<td>Bidirectional &#x002B; static gating (<inline-formula id="ieqn-114"><mml:math id="mml-ieqn-114"><mml:mi>g</mml:mi><mml:mo>=</mml:mo><mml:mn>0.5</mml:mn></mml:math></inline-formula>)</td>
<td>33.8</td>
<td>&#x2212;2.0</td>
</tr>
<tr>
<td>Bidirectional &#x002B; learned scalar gating</td>
<td>33.5</td>
<td>&#x2212;2.3</td>
</tr>
<tr>
<td>Bidirectional &#x002B; adaptive element-wise gating</td>
<td>33.2</td>
<td>&#x2212;2.6</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The results confirm several design principles. Bidirectional attention outperforms unidirectional variants because both modalities contain complementary information that benefits the other. Sign-to-lip transfer provides slightly larger gains than lip-to-sign transfer, likely because sign language&#x2019;s richer spatial vocabulary provides additional discriminative cues for disambiguating visually similar lip movements. Adaptive element-wise gating achieves the best performance by allowing dimension-specific modality weighting, enabling the model to selectively combine different aspects of each modality&#x2019;s representation.</p>
<p><xref ref-type="table" rid="table-9">Table 9</xref> shows sensitivity to the CSA loss weight <inline-formula id="ieqn-115"><mml:math id="mml-ieqn-115"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mtext>CSA</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula>, and <xref ref-type="fig" rid="fig-4">Fig. 4</xref> visualizes the cross-modal alignment effect using t-SNE projections of the learned embeddings.</p>
<table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title>Sensitivity analysis of the CSA loss weight (<inline-formula id="ieqn-117"><mml:math id="mml-ieqn-117"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mtext>CSA</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula>). Word error rate (WER%) is reported for LRS2 and PH-14 datasets, while alignment score measures cross-modal alignment quality.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th><inline-formula id="ieqn-118"><mml:math id="mml-ieqn-118"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mtext>CSA</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula></th>
<th>LRS2</th>
<th>PH-14</th>
<th>Alignment Score</th>
</tr>
</thead>
<tbody>
<tr>
<td>0.00</td>
<td>34.8</td>
<td>19.2</td>
<td>0.62</td>
</tr>
<tr>
<td>0.05</td>
<td>33.9</td>
<td>18.7</td>
<td>0.71</td>
</tr>
<tr>
<td>0.10</td>
<td>33.2</td>
<td>18.3</td>
<td>0.78</td>
</tr>
<tr>
<td>0.20</td>
<td>33.5</td>
<td>18.5</td>
<td>0.81</td>
</tr>
<tr>
<td>0.50</td>
<td>34.1</td>
<td>19.0</td>
<td>0.84</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>t-SNE visualization of lip and sign embeddings for semantically matched pairs. Without CSA (left), modalities cluster separately in distinct regions. With CSA (right), semantically matched pairs align in the shared embedding space.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_78743-fig-4.tif"/>
</fig>
<p>The visualizations demonstrate that without CSA loss, lip and sign embeddings occupy distinct regions of the representation space, limiting cross-modal transfer. With the CSA loss at <inline-formula id="ieqn-116"><mml:math id="mml-ieqn-116"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mtext>CSA</mml:mtext></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0.1</mml:mn></mml:math></inline-formula>, semantically matched pairs from both modalities cluster together, enabling more effective information sharing. The optimal weight balances alignment strength against task-specific discriminability; larger weights improve alignment scores but begin to degrade recognition performance as the representations become overly constrained.</p>
<p><xref ref-type="table" rid="table-10">Table 10</xref> evaluates performance when one modality is degraded or missing during inference, addressing concerns about robustness to partial input availability.</p>
<table-wrap id="table-10">
<label>Table 10</label>
<caption>
<title>Robustness to modality degradation (inference only).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Condition</th>
<th>LRS2 (WER%)</th>
<th>PH-14 (WER%)</th>
<th>Relative Drop</th>
</tr>
</thead>
<tbody>
<tr>
<td>Full model (both modalities)</td>
<td>33.2</td>
<td>18.3</td>
<td>&#x2013;</td>
</tr>
<tr>
<td>Lip stream only (sign zeroed)</td>
<td>35.1</td>
<td>20.8</td>
<td>&#x002B;5.7%/&#x002B;13.7%</td>
</tr>
<tr>
<td>Sign stream only (lip zeroed)</td>
<td>36.8</td>
<td>18.9</td>
<td>&#x002B;10.8%/&#x002B;3.3%</td>
</tr>
<tr>
<td>Lip with 50% frame dropout</td>
<td>34.2</td>
<td>19.5</td>
<td>&#x002B;3.0%/&#x002B;6.6%</td>
</tr>
<tr>
<td>Sign with hand occlusion (30%)</td>
<td>34.8</td>
<td>19.2</td>
<td>&#x002B;4.8%/&#x002B;4.9%</td>
</tr>
<tr>
<td>Blurred mouth region (<inline-formula id="ieqn-119"><mml:math id="mml-ieqn-119"><mml:mi>&#x03C3;</mml:mi><mml:mo>=</mml:mo><mml:mn>5</mml:mn></mml:math></inline-formula>)</td>
<td>34.5</td>
<td>19.8</td>
<td>&#x002B;3.9%/&#x002B;8.2%</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The model exhibits graceful degradation when one modality is impaired. Performance drops range from 3% to 14% depending on the degradation type and target task, demonstrating that cross-modal training provides implicit robustness. As expected, lip reading performance depends more heavily on the lip stream, while sign recognition depends more on the sign stream. Partial degradations such as frame dropout or occlusion have smaller effects than complete modality removal, indicating that the model can leverage whatever partial information remains available. <xref ref-type="table" rid="table-11">Table 11</xref> compares different pre-training strategies, confirming that joint multimodal pre-training provides the strongest performance gains.</p>
<table-wrap id="table-11">
<label>Table 11</label>
<caption>
<title>Effect of cross-modal pre-training strategies (WER%).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Pre-training Strategy</th>
<th>LRS2</th>
<th>PH-14</th>
</tr>
</thead>
<tbody>
<tr>
<td>None (scratch)</td>
<td>37.9</td>
<td>21.2</td>
</tr>
<tr>
<td>Lip only</td>
<td>35.1</td>
<td>20.4</td>
</tr>
<tr>
<td>Sign only</td>
<td>36.8</td>
<td>19.6</td>
</tr>
<tr>
<td>Joint multimodal</td>
<td>33.2</td>
<td>18.3</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-12">Table 12</xref> evaluates generalization by training on one dataset and testing on another without fine-tuning.</p>
<table-wrap id="table-12">
<label>Table 12</label>
<caption>
<title>Cross-domain generalization (Train<inline-formula id="ieqn-120"><mml:math id="mml-ieqn-120"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula>Test).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Configuration</th>
<th>LRS2<inline-formula id="ieqn-121"><mml:math id="mml-ieqn-121"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula>LRS3</th>
<th>LRS3<inline-formula id="ieqn-122"><mml:math id="mml-ieqn-122"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula>LRS2</th>
<th>PH-14<inline-formula id="ieqn-123"><mml:math id="mml-ieqn-123"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula>CSL</th>
</tr>
</thead>
<tbody>
<tr>
<td>Single-task</td>
<td>52.3</td>
<td>48.7</td>
<td>82.4</td>
</tr>
<tr>
<td>UniModal-LSR</td>
<td>47.8</td>
<td>44.2</td>
<td>85.1</td>
</tr>
<tr>
<td>Improvement</td>
<td>&#x2212;8.6%</td>
<td>&#x2212;9.2%</td>
<td>&#x002B;3.3%</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The unified model shows improved cross-domain generalization, particularly for lip reading where the 8&#x2013;9% WER reduction suggests that cross-modal learning provides regularization benefits that transfer across dataset boundaries. The improvement on sign language transfer is smaller but still positive, indicating that the learned representations capture generalizable visual speech features rather than dataset-specific patterns.</p>
<p><xref ref-type="table" rid="table-13">Table 13</xref> demonstrates performance under reduced training data availability, showing relative improvements of up to 19% when training with only 10% of the data.</p>
<table-wrap id="table-13">
<label>Table 13</label>
<caption>
<title>Data efficiency: performance vs. training data fraction.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th rowspan="2">Data %</th>
<th colspan="2">LRS2 (WER%)</th>
<th colspan="2">PH-14 (WER%)</th>
</tr>
<tr>
<th>Baseline</th>
<th>UniModal-LSR</th>
<th>Baseline</th>
<th>UniModal-LSR</th>
</tr>
</thead>
<tbody>
<tr>
<td>10%</td>
<td>58.4</td>
<td>48.2 (&#x2212;17.5%)</td>
<td>35.6</td>
<td>28.9 (&#x2212;18.8%)</td>
</tr>
<tr>
<td>25%</td>
<td>48.7</td>
<td>41.3 (&#x2212;15.2%)</td>
<td>28.4</td>
<td>23.5 (&#x2212;17.3%)</td>
</tr>
<tr>
<td>50%</td>
<td>42.1</td>
<td>36.8 (&#x2212;12.6%)</td>
<td>24.2</td>
<td>20.4 (&#x2212;15.7%)</td>
</tr>
<tr>
<td>100%</td>
<td>37.9</td>
<td>33.2 (&#x2212;12.4%)</td>
<td>21.2</td>
<td>18.3 (&#x2212;13.7%)</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-14">Table 14</xref> compares computational requirements between separate single-task models and the unified architecture. Latency measurements are end-to-end, including all preprocessing steps, conducted on a single NVIDIA A100 GPU with batch size 1.</p>
<table-wrap id="table-14">
<label>Table 14</label>
<caption>
<title>Computational efficiency comparison.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Method</th>
<th>Params (M)</th>
<th>FLOPs (G)</th>
<th>Latency (ms)</th>
<th>Memory (GB)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Separate (Lip &#x002B; Sign)</td>
<td>120.6</td>
<td>42.9</td>
<td>98</td>
<td>10.0</td>
</tr>
<tr>
<td>UniModal-LSR (Unified)</td>
<td>89.4</td>
<td>31.2</td>
<td>68</td>
<td>7.2</td>
</tr>
<tr>
<td>Reduction</td>
<td>&#x2212;25.9%</td>
<td>&#x2212;27.3%</td>
<td>&#x2212;30.6%</td>
<td>&#x2212;28.0%</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>We address practical deployment scenarios beyond the full multimodal inference setting. When only one modality is available, such as lip-only or sign-only video, the model can operate with a single branch while still benefiting from joint training. <xref ref-type="table" rid="table-15">Table 15</xref> shows that single-branch variants extracted from the unified model outperform separately trained single-task models, demonstrating that the cross-modal training provides transferable improvements even when inference uses only one modality. <xref ref-type="table" rid="table-16">Table 16</xref> presents performance of reduced-capacity variants suitable for edge deployment with limited computational resources.</p>
<table-wrap id="table-15">
<label>Table 15</label>
<caption>
<title>Single-branch inference performance.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Configuration</th>
<th>LRS2 (WER%)</th>
<th>PH-14 (WER%)</th>
<th>Params (M)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Separate lip-only model</td>
<td>37.9</td>
<td>&#x2013;</td>
<td>58.2</td>
</tr>
<tr>
<td>UniModal-LSR lip branch only</td>
<td>35.6</td>
<td>&#x2013;</td>
<td>52.1</td>
</tr>
<tr>
<td>Separate sign-only model</td>
<td>&#x2013;</td>
<td>21.2</td>
<td>62.4</td>
</tr>
<tr>
<td>UniModal-LSR sign branch only</td>
<td>&#x2013;</td>
<td>19.8</td>
<td>56.3</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-16">
<label>Table 16</label>
<caption>
<title>Lightweight model variants.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Variant</th>
<th>LRS2 (WER%)</th>
<th>PH-14 (WER%)</th>
<th>Params (M)</th>
<th>Latency (ms)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Full model</td>
<td>33.2</td>
<td>18.3</td>
<td>89.4</td>
<td>68</td>
</tr>
<tr>
<td><inline-formula id="ieqn-124"><mml:math id="mml-ieqn-124"><mml:mi>d</mml:mi><mml:mo>=</mml:mo><mml:mn>256</mml:mn></mml:math></inline-formula> (half width)</td>
<td>35.8</td>
<td>19.9</td>
<td>45.2</td>
<td>42</td>
</tr>
<tr>
<td><inline-formula id="ieqn-125"><mml:math id="mml-ieqn-125"><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn></mml:math></inline-formula> (2 HTSE levels)</td>
<td>34.5</td>
<td>19.2</td>
<td>67.8</td>
<td>51</td>
</tr>
<tr>
<td>MobileNet3D backbone</td>
<td>36.2</td>
<td>20.4</td>
<td>32.1</td>
<td>28</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-17">Table 17</xref> evaluates performance under reduced temporal sampling, which is relevant for bandwidth-constrained applications where transmitting full frame-rate video is impractical.</p>
<table-wrap id="table-17">
<label>Table 17</label>
<caption>
<title>Performance vs. frame rate.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Frame Rate (fps)</th>
<th>LRS2 (WER%)</th>
<th>PH-14 (WER%)</th>
<th>Latency (ms)</th>
</tr>
</thead>
<tbody>
<tr>
<td>25 (full)</td>
<td>33.2</td>
<td>18.3</td>
<td>68</td>
</tr>
<tr>
<td>15</td>
<td>34.1</td>
<td>18.9</td>
<td>45</td>
</tr>
<tr>
<td>10</td>
<td>35.8</td>
<td>20.2</td>
<td>32</td>
</tr>
<tr>
<td>5</td>
<td>39.4</td>
<td>23.8</td>
<td>18</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="fig" rid="fig-5">Fig. 5</xref> decomposes errors by category. The unified model reduces errors across all categories, with notable improvements in viseme confusion (28% reduction) and boundary detection (33% reduction), which are areas where cross-modal information provides the greatest benefit.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Error analysis by category showing reductions across all error types.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_78743-fig-5.tif"/>
</fig>
</sec>
<sec id="s7">
<label>7</label>
<title>Discussion</title>
<p>Our work builds on the foundation established by multimodal architectures for speech-related tasks. The dual-tower framework proposed by Ge et al. [<xref ref-type="bibr" rid="ref-26">26</xref>] demonstrates the effectiveness of unified multimodal modeling for audio-text fusion in air traffic control communications, achieving improved recognition accuracy through coordinated processing of acoustic and textual information. While their approach improves multimodal speech recognition, it addresses a fundamentally different modality combination. Our CMAF module addresses the unique challenges of visual-visual fusion, where both modalities share temporal structure but differ in spatial semantics. The bidirectional cross-attention design enables more flexible information exchange than parallel tower processing, which is important when the usefulness of lip and sign cues varies across contexts. Furthermore, the adaptive gating mechanism provides context-dependent weighting that is especially suited to the dynamic relationship between visual speech modalities.</p>
<p>Despite encouraging results, several limitations remain. The benchmark coverage may not fully capture real-world variability in illumination, camera pose, and occlusion, and evaluation on more diverse conditions would strengthen practical applicability. The model is language-specific and must be retrained for each target spoken or signed language, highlighting the need for multilingual or language-agnostic extensions. Training is computationally expensive, requiring eight A100 GPUs during 72 h; although lightweight variants reduce inference cost, they do not address training efficiency. The framework also assumes both lip and sign modalities are available during inference, and while performance degrades gracefully when one modality is missing, explicit strategies such as modality dropout could improve robustness. Finally, the CSA loss relies on semantically aligned lip-sign pairs, and constructing such pairs across datasets may introduce noise, suggesting that future work could explore self-supervised alignment methods that enable training on unpaired data.</p>
</sec>
<sec id="s8">
<label>8</label>
<title>Conclusion</title>
<p>We introduced UniModal-LSR, a unified multimodal architecture for joint lip reading and sign language recognition. Through hierarchical temporal-spatial encoding, graph-enhanced hand modeling, cross-modal attention fusion with adaptive gating, and contrastive semantic alignment, the model achieves state-of-the-art performance while reducing parameters by 25.9% relative to separate task-specific models. Detailed ablations confirm the contribution of each component, and analysis demonstrates robustness to modality degradation and improved cross-domain generalization.</p>
<p>Future work will explore integration of additional modalities such as audio and full-body pose to capture a more complete picture of visual communication. Language-agnostic representation learning could enable a single model to serve multiple linguistic communities. Self-supervised pre-training methods may reduce labeled data requirements, making the technology more accessible for under-resourced languages. Integration with dialogue systems, such as the full-duplex speech dialogue schemes based on large language models, could enable more natural and responsive human-computer interaction.</p>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This work is funded by Ho Chi Minh City Open University (HCMCOU) and the Ministry of Education and Training (Vietnam) under grant number B2025-MBS-01.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors report that their specific contributions to this paper are as follows: the study was conceived and designed by Vinh Truong Hoang and Nghia Dinh; data were collected by Thien Ho Huong and Luu Quang Phuong; data analysis and interpretation of the results were undertaken by Kiet Tran-Trung and Ha Duong Thi Hong; and the initial manuscript draft was written by Bay Nguyen Van and Hau Nguyen Trung. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The authors confirm that the data supporting the findings of this study are openly available. LRS2 and LRS3 datasets are available at <ext-link ext-link-type="uri" xlink:href="https://www.robots.ox.ac.uk/vgg/data/lip_reading/">https://www.robots.ox.ac.uk/vgg/data/lip_reading/</ext-link>. PHOENIX dataset is available at <ext-link ext-link-type="uri" xlink:href="https://www-i6.informatik.rwth-aachen.de/koller/RWTH-PHOENIX/">https://www-i6.informatik.rwth-aachen.de/koller/RWTH-PHOENIX/</ext-link>. WLASL dataset is available at <ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/datasets/risangbaskoro/wlasl-processed">https://www.kaggle.com/datasets/risangbaskoro/wlasl-processed</ext-link>. CSL dataset is available at <ext-link ext-link-type="uri" xlink:href="https://ustc-slr.github.io/datasets/2015_csl">https://ustc-slr.github.io/datasets/2015_csl</ext-link>. How2Sign dataset is available at <ext-link ext-link-type="uri" xlink:href="https://how2sign.github.io/">https://how2sign.github.io/</ext-link>.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>This research uses publicly available benchmark datasets collected with appropriate consent and ethical approval by the original creators.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Camgoz</surname> <given-names>NC</given-names></string-name>, <string-name><surname>Hadfield</surname> <given-names>S</given-names></string-name>, <string-name><surname>Koller</surname> <given-names>O</given-names></string-name>, <string-name><surname>Ney</surname> <given-names>H</given-names></string-name>, <string-name><surname>Bowden</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Neural sign language translation</article-title>. In: <conf-name>2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2018 Jun 18&#x2013;23; Salt Lake City, UT, USA</conf-name>. p. <fpage>7784</fpage>&#x2013;<lpage>93</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR.2018.00812</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Assael</surname> <given-names>YM</given-names></string-name>, <string-name><surname>Shillingford</surname> <given-names>B</given-names></string-name>, <string-name><surname>Whiteson</surname> <given-names>S</given-names></string-name>, <string-name><surname>De Freitas</surname> <given-names>N</given-names></string-name></person-group>. <article-title>LipNet: end-to-end sentence-level lipreading</article-title>. <comment>arXiv:1611.01599. 2016</comment>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bear</surname> <given-names>HL</given-names></string-name>, <string-name><surname>Harvey</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Phoneme-to-viseme mappings: the good, the bad, and the ugly</article-title>. <source>Speech Commun</source>. <year>2017</year>;<volume>95</volume>(<issue>3</issue>):<fpage>40</fpage>&#x2013;<lpage>67</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.specom.2017.07.001</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Heracleous</surname> <given-names>P</given-names></string-name>, <string-name><surname>Beautemps</surname> <given-names>D</given-names></string-name>, <string-name><surname>Aboutabit</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Cued speech automatic recognition in normal-hearing and deaf subjects</article-title>. <source>Speech Commun</source>. <year>2010</year>;<volume>52</volume>(<issue>6</issue>):<fpage>504</fpage>&#x2013;<lpage>12</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.specom.2010.03.001</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Petridis</surname> <given-names>S</given-names></string-name>, <string-name><surname>Stafylakis</surname> <given-names>T</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>P</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>F</given-names></string-name>, <string-name><surname>Tzimiropoulos</surname> <given-names>G</given-names></string-name>, <string-name><surname>Pantic</surname> <given-names>M</given-names></string-name></person-group>. <article-title>End-to-end audiovisual speech recognition</article-title>. In: <conf-name>2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); 2018 Apr 15&#x2013;20; Calgary, AB, Canada</conf-name>. p. <fpage>6548</fpage>&#x2013;<lpage>52</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ICASSP.2018.8461326</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Subba Rao</surname> <given-names>MV</given-names></string-name>, <string-name><surname>Naga Amulya</surname> <given-names>T</given-names></string-name>, <string-name><surname>Aparna</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Pranavi</surname> <given-names>R</given-names></string-name>, <string-name><surname>Madhumitha</surname> <given-names>S</given-names></string-name>, <string-name><surname>Priya</surname> <given-names>SS</given-names></string-name></person-group>. <article-title>Speech reconstruction from silent lip movements using deep learning</article-title>. In: <conf-name>2025 2nd International Conference on Artificial Intelligence for Innovations in Healthcare Industries (ICAIIHI); 2025 Dec 4&#x2013;5; Raipur, India</conf-name>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.1109/icaiihi67124.2025.11403769</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Vaswani</surname> <given-names>A</given-names></string-name>, <string-name><surname>Shazeer</surname> <given-names>N</given-names></string-name>, <string-name><surname>Parmar</surname> <given-names>N</given-names></string-name>, <string-name><surname>Uszkoreit</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jones</surname> <given-names>L</given-names></string-name>, <string-name><surname>Gomez</surname> <given-names>AN</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Attention is all you need</article-title>. In: <conf-name>NIPS&#x2019;17: Proceedings of the 31st International Conference on Neural Information Processing Systems</conf-name>. <publisher-loc>Red Hook, NY, USA</publisher-loc>: <publisher-name>Curran Associates Inc.</publisher-name>; <year>2017</year>. p. <fpage>6000</fpage>&#x2013;<lpage>10</lpage>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Carreira</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zisserman</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Quo vadis, action recognition? A new model and the kinetics dataset</article-title>. In: <conf-name>2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2017 Jul 21&#x2013;26; Honolulu, HI, USA</conf-name>. p. <fpage>4724</fpage>&#x2013;<lpage>33</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR.2017.502</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Spatial temporal graph convolutional networks for skeleton-based action recognition</article-title>. <source>Proc AAAI Conf Artif Intell</source>. <year>2018</year>;<volume>32</volume>(<issue>1</issue>):<fpage>7444</fpage>&#x2013;<lpage>52</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v32i1.12328</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>van den Oord</surname> <given-names>A</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Vinyals</surname> <given-names>O</given-names></string-name></person-group>. <article-title>Representation learning with contrastive predictive coding</article-title>. <comment>arXiv:1807.03748. 2018</comment>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Potamianos</surname> <given-names>G</given-names></string-name>, <string-name><surname>Neti</surname> <given-names>C</given-names></string-name>, <string-name><surname>Gravier</surname> <given-names>G</given-names></string-name>, <string-name><surname>Garg</surname> <given-names>A</given-names></string-name>, <string-name><surname>Senior</surname> <given-names>AW</given-names></string-name></person-group>. <article-title>Recent advances in the automatic recognition of audiovisual speech</article-title>. <source>Proc IEEE</source>. <year>2003</year>;<volume>91</volume>(<issue>9</issue>):<fpage>1306</fpage>&#x2013;<lpage>26</lpage>. doi:<pub-id pub-id-type="doi">10.1109/JPROC.2003.817150</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Matthews</surname> <given-names>I</given-names></string-name>, <string-name><surname>Cootes</surname> <given-names>TF</given-names></string-name>, <string-name><surname>Bangham</surname> <given-names>JA</given-names></string-name>, <string-name><surname>Cox</surname> <given-names>S</given-names></string-name>, <string-name><surname>Harvey</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Extraction of visual features for lipreading</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2002</year>;<volume>24</volume>(<issue>2</issue>):<fpage>198</fpage>&#x2013;<lpage>213</lpage>. doi:<pub-id pub-id-type="doi">10.1109/34.982900</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gales</surname> <given-names>M</given-names></string-name>, <string-name><surname>Young</surname> <given-names>S</given-names></string-name></person-group>. <article-title>The application of hidden Markov models in speech recognition</article-title>. <source>Found Trends&#x00AE; Signal Process</source>. <year>2008</year>;<volume>1</volume>(<issue>3</issue>):<fpage>195</fpage>&#x2013;<lpage>304</lpage>. doi:<pub-id pub-id-type="doi">10.1561/2000000004</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chung</surname> <given-names>JS</given-names></string-name>, <string-name><surname>Senior</surname> <given-names>A</given-names></string-name>, <string-name><surname>Vinyals</surname> <given-names>O</given-names></string-name>, <string-name><surname>Zisserman</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Lip reading sentences in the wild</article-title>. In: <conf-name>2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2017 Jul 21&#x2013;26; Honolulu, HI, USA</conf-name>. p. <fpage>3444</fpage>&#x2013;<lpage>53</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR.2017.367</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ma</surname> <given-names>P</given-names></string-name>, <string-name><surname>Petridis</surname> <given-names>S</given-names></string-name>, <string-name><surname>Pantic</surname> <given-names>M</given-names></string-name></person-group>. <article-title>End-to-end audio-visual speech recognition with conformers</article-title>. In: <conf-name>ICASSP 2021&#x2013;2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); 2021 Jun 6&#x2013;11; Toronto, ON, Canada</conf-name>. p. <fpage>7613</fpage>&#x2013;<lpage>7</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ICASSP39728.2021.9414567</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ma</surname> <given-names>P</given-names></string-name>, <string-name><surname>Petridis</surname> <given-names>S</given-names></string-name>, <string-name><surname>Pantic</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Visual speech recognition for multiple languages in the wild</article-title>. <source>Nat Mach Intell</source>. <year>2022</year>;<volume>4</volume>(<issue>11</issue>):<fpage>930</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1038/s42256-022-00550-z</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Shi</surname> <given-names>B</given-names></string-name>, <string-name><surname>Hsu</surname> <given-names>WN</given-names></string-name>, <string-name><surname>Lakhotia</surname> <given-names>K</given-names></string-name>, <string-name><surname>Mohamed</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Learning audio-visual speech representation by masked multimodal cluster prediction</article-title>. <comment>arXiv:2201.02184. 2022</comment>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ma</surname> <given-names>P</given-names></string-name>, <string-name><surname>Haliassos</surname> <given-names>A</given-names></string-name>, <string-name><surname>Fernandez-Lopez</surname> <given-names>A</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>H</given-names></string-name>, <string-name><surname>Petridis</surname> <given-names>S</given-names></string-name>, <string-name><surname>Pantic</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Auto-AVSR: audio-visual speech recognition with automatic labels</article-title>. In: <conf-name>ICASSP 2023&#x2013;2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); 2023 Jun 4&#x2013;10; Rhodes Island, Greece</conf-name>. p. <fpage>1</fpage>&#x2013;<lpage>5</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ICASSP49357.2023.10096889</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shukla</surname> <given-names>A</given-names></string-name>, <string-name><surname>Petridis</surname> <given-names>S</given-names></string-name>, <string-name><surname>Pantic</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Does visual self-supervision improve learning of speech representations for emotion recognition?</article-title> <source>IEEE Trans Affect Comput</source>. <year>2023</year>;<volume>14</volume>(<issue>1</issue>):<fpage>406</fpage>&#x2013;<lpage>20</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TAFFC.2021.3062406</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Prajwal</surname> <given-names>KR</given-names></string-name>, <string-name><surname>Afouras</surname> <given-names>T</given-names></string-name>, <string-name><surname>Zisserman</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Sub-word level lip reading with visual attention</article-title>. In: <conf-name>2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2022 Jun 18&#x2013;24; New Orleans, LA, USA</conf-name>. p. <fpage>5152</fpage>&#x2013;<lpage>62</lpage>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Hu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>W</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name></person-group>. <article-title>SignBERT: pre-training of hand-model-aware representation for sign language recognition</article-title>. In: <conf-name>2021 IEEE/CVF International Conference on Computer Vision (ICCV); 2021 Oct 10&#x2013;17; Montreal, QC, Canada</conf-name>. p. <fpage>11067</fpage>&#x2013;<lpage>76</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ICCV48922.2021.01090</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Graves</surname> <given-names>A</given-names></string-name>, <string-name><surname>Fern&#x00E1;ndez</surname> <given-names>S</given-names></string-name>, <string-name><surname>Gomez</surname> <given-names>F</given-names></string-name>, <string-name><surname>Schmidhuber</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks</article-title>. In: <conf-name>Proceedings of the 23rd International Conference on Machine Learning&#x2014;ICML &#x2019;06; 2006 Jun 25&#x2013;29; Pittsburgh, PA, USA</conf-name>. p. <fpage>369</fpage>&#x2013;<lpage>76</lpage>. doi:<pub-id pub-id-type="doi">10.1145/1143844.1143891</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Lu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Batra</surname> <given-names>D</given-names></string-name>, <string-name><surname>Parikh</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Hierarchical question-image co-attention for visual question answering</article-title>. <comment>arXiv:1606.00061</comment>. <year>2016</year>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Tsai</surname> <given-names>YH</given-names></string-name>, <string-name><surname>Bai</surname> <given-names>S</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>PP</given-names></string-name>, <string-name><surname>Kolter</surname> <given-names>JZ</given-names></string-name>, <string-name><surname>Morency</surname> <given-names>LP</given-names></string-name>, <string-name><surname>Salakhutdinov</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Multimodal transformer for unaligned multimodal language sequences</article-title>. In: <conf-name>Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics</conf-name>. <publisher-loc>Stroudsburg, PA, USA</publisher-loc>: <publisher-name>ACL</publisher-name>; <year>2019</year>. p. <fpage>6558</fpage>&#x2013;<lpage>69</lpage>. doi:<pub-id pub-id-type="doi">10.18653/v1/p19-1656</pub-id>; <pub-id pub-id-type="pmid">32362720</pub-id></mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Radford</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>JW</given-names></string-name>, <string-name><surname>Hallacy</surname> <given-names>C</given-names></string-name>, <string-name><surname>Ramesh</surname> <given-names>A</given-names></string-name>, <string-name><surname>Goh</surname> <given-names>G</given-names></string-name>, <string-name><surname>Agarwal</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Learning transferable visual models from natural language supervision</article-title>. <comment>arXiv:2103.00020. 2021</comment>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ge</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>J</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Audio-text multimodal speech recognition via dual-tower architecture for mandarin air traffic control communications</article-title>. <source>Comput Mater Contin</source>. <year>2024</year>;<volume>78</volume>(<issue>3</issue>):<fpage>3215</fpage>&#x2013;<lpage>45</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmc.2023.046746</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>G</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Geng</surname> <given-names>M</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Audio-visual end-to-end multi-channel speech separation, dereverberation and recognition</article-title>. <source>IEEE/ACM Trans Audio Speech Lang Process</source>. <year>2023</year>;<volume>31</volume>:<fpage>2707</fpage>&#x2013;<lpage>23</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TASLP.2023.3294705</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Fang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>DCIM-AVSR: efficient audio-visual speech recognition via dual conformer interaction module</article-title>. In: <conf-name>ICASSP 2025&#x2013;2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); 2025 Apr 6&#x2013;11; Hyderabad, India</conf-name>. p. <fpage>1</fpage>&#x2013;<lpage>5</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ICASSP49660.2025.10890272</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>F</given-names></string-name>, <string-name><surname>Bazarevsky</surname> <given-names>V</given-names></string-name>, <string-name><surname>Vakunov</surname> <given-names>A</given-names></string-name>, <string-name><surname>Tkachenka</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sung</surname> <given-names>G</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>C-L</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>MediaPipe hands: on-device real-time hand tracking</article-title>. <comment>arXiv:2006.10214. 2020</comment>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Yun</surname> <given-names>C</given-names></string-name>, <string-name><surname>Bhojanapalli</surname> <given-names>S</given-names></string-name>, <string-name><surname>Rawat</surname> <given-names>AS</given-names></string-name>, <string-name><surname>Reddi</surname> <given-names>S</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Are transformers universal approximators of sequence-to-sequence functions?</article-title> <comment>arXiv:1912.10077. 2020</comment>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Afouras</surname> <given-names>T</given-names></string-name>, <string-name><surname>Chung</surname> <given-names>JS</given-names></string-name>, <string-name><surname>Zisserman</surname> <given-names>A</given-names></string-name></person-group>. <article-title>LRS3-TED: a large-scale dataset for visual speech recognition</article-title>. <comment>arXiv:1809.00496. 2018</comment>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cooke</surname> <given-names>M</given-names></string-name>, <string-name><surname>Barker</surname> <given-names>J</given-names></string-name>, <string-name><surname>Cunningham</surname> <given-names>S</given-names></string-name>, <string-name><surname>Shao</surname> <given-names>X</given-names></string-name></person-group>. <article-title>An audio-visual corpus for speech perception and automatic speech recognition</article-title>. <source>J Acoust Soc Am</source>. <year>2006</year>;<volume>120</volume>(<issue>5 Pt 1</issue>):<fpage>2421</fpage>&#x2013;<lpage>4</lpage>. doi:<pub-id pub-id-type="doi">10.1121/1.2229005</pub-id>; <pub-id pub-id-type="pmid">17139705</pub-id></mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Koller</surname> <given-names>O</given-names></string-name>, <string-name><surname>Forster</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ney</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Continuous sign language recognition: towards large vocabulary statistical recognition systems handling multiple signers</article-title>. <source>Comput Vis Image Underst</source>. <year>2015</year>;<volume>141</volume>(<issue>5</issue>):<fpage>108</fpage>&#x2013;<lpage>25</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.cviu.2015.09.013</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Video-based sign language recognition without temporal segmentation</article-title>. <source>Proc AAAI Conf Artif Intell</source>. <year>2018</year>;<volume>32</volume>(<issue>1</issue>):<fpage>2257</fpage>&#x2013;<lpage>64</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v32i1.11903</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>D</given-names></string-name>, <string-name><surname>Opazo</surname> <given-names>CR</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Word-level deep sign language recognition from video: a new large-scale dataset and methods comparison</article-title>. In: <conf-name>2020 IEEE Winter Conference on Applications of Computer Vision (WACV); 2020 Mar 1&#x2013;5; Snowmass Village, CO, USA</conf-name>. p. <fpage>1448</fpage>&#x2013;<lpage>58</lpage>. doi:<pub-id pub-id-type="doi">10.1109/wacv45572.2020.9093512</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Duarte</surname> <given-names>A</given-names></string-name>, <string-name><surname>Palaskar</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ventura</surname> <given-names>L</given-names></string-name>, <string-name><surname>Ghadiyaram</surname> <given-names>D</given-names></string-name>, <string-name><surname>DeHaan</surname> <given-names>K</given-names></string-name>, <string-name><surname>Metze</surname> <given-names>F</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>How2Sign: a large-scale multimodal dataset for continuous American sign language</article-title>. In: <conf-name> 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2021 Jun 20&#x2013;25; Nashville, TN, USA</conf-name>. p. <fpage>2734</fpage>&#x2013;<lpage>43</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cvpr46437.2021.00276</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Afouras</surname> <given-names>T</given-names></string-name>, <string-name><surname>Chung</surname> <given-names>JS</given-names></string-name>, <string-name><surname>Senior</surname> <given-names>A</given-names></string-name>, <string-name><surname>Vinyals</surname> <given-names>O</given-names></string-name>, <string-name><surname>Zisserman</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Deep audio-visual speech recognition</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2022</year>;<volume>44</volume>(<issue>12</issue>):<fpage>8717</fpage>&#x2013;<lpage>27</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TPAMI.2018.2889052</pub-id>; <pub-id pub-id-type="pmid">30582526</pub-id></mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Martinez</surname> <given-names>B</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>P</given-names></string-name>, <string-name><surname>Petridis</surname> <given-names>S</given-names></string-name>, <string-name><surname>Pantic</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Lipreading using temporal convolutional networks</article-title>. In: <conf-name>ICASSP 2020&#x2013;2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); 2020 May 4&#x2013;8; Barcelona, Spain</conf-name>. p. <fpage>6319</fpage>&#x2013;<lpage>23</lpage>. doi:<pub-id pub-id-type="doi">10.1109/icassp40776.2020.9053841</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kim</surname> <given-names>M</given-names></string-name>, <string-name><surname>Yeo</surname> <given-names>JH</given-names></string-name>, <string-name><surname>Ro</surname> <given-names>YM</given-names></string-name></person-group>. <article-title>Distinguishing homophenes using multi-head visual-audio memory for lip reading</article-title>. <source>Proc AAAI Conf Artif Intell</source>. <year>2022</year>;<volume>36</volume>(<issue>1</issue>):<fpage>1174</fpage>&#x2013;<lpage>82</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v36i1.20003</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Koller</surname> <given-names>O</given-names></string-name>, <string-name><surname>Camgoz</surname> <given-names>NC</given-names></string-name>, <string-name><surname>Ney</surname> <given-names>H</given-names></string-name>, <string-name><surname>Bowden</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Weakly supervised learning with multi-stream CNN-LSTM-HMMs to discover sequential parallelism in sign language videos</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2020</year>;<volume>42</volume>(<issue>9</issue>):<fpage>2306</fpage>&#x2013;<lpage>20</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TPAMI.2019.2911077</pub-id>; <pub-id pub-id-type="pmid">30990421</pub-id></mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Spatial-temporal multi-cue network for continuous sign language recognition</article-title>. <source>Proc AAAI Conf Artif Intell</source>. <year>2020</year>;<volume>34</volume>(<issue>7</issue>):<fpage>13009</fpage>&#x2013;<lpage>16</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v34i07.7001</pub-id>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Cheng</surname> <given-names>KL</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Tai</surname> <given-names>YW</given-names></string-name></person-group>. <chapter-title>Fully convolutional networks for continuous sign language recognition</chapter-title>. In: <source>Computer Vision&#x2014;ECCV 2020</source>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2020</year>. p. <fpage>697</fpage>&#x2013;<lpage>714</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-030-58586-0_41</pub-id>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Min</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Hao</surname> <given-names>A</given-names></string-name>, <string-name><surname>Chai</surname> <given-names>X</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Visual alignment constraint for continuous sign language recognition</article-title>. In: <conf-name>2021 IEEE/CVF International Conference on Computer Vision (ICCV); 2021 Oct 10&#x2013;17; Montreal, QC, Canada</conf-name>. p. <fpage>11522</fpage>&#x2013;<lpage>31</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ICCV48922.2021.01134</pub-id>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Cihan Camgoz</surname> <given-names>N</given-names></string-name>, <string-name><surname>Koller</surname> <given-names>O</given-names></string-name>, <string-name><surname>Hadfield</surname> <given-names>S</given-names></string-name>, <string-name><surname>Bowden</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Sign language transformers: joint end-to-end sign language recognition and translation</article-title>. In: <conf-name>2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2020 Jun 13&#x2013;19; Seattle, WA, USA</conf-name>. p. <fpage>10020</fpage>&#x2013;<lpage>30</lpage>. doi:<pub-id pub-id-type="doi">10.1109/cvpr42600.2020.01004</pub-id>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>D</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Swift</surname> <given-names>B</given-names></string-name>, <string-name><surname>Suominen</surname> <given-names>H</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>TSPNet: hierarchical feature learning via temporal semantic pyramid for sign language translation</article-title>. In: <conf-name>NIPS&#x2019;20: Proceedings of the 34th International Conference on Neural Information Processing Systems</conf-name>. <publisher-loc>Red Hook, NY, USA</publisher-loc>: <publisher-name>Curran Associates Inc.</publisher-name>; <year>2020</year>. p. <fpage>12034</fpage>&#x2013;<lpage>45</lpage>.</mixed-citation></ref>
</ref-list>
</back></article>