<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">81534</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2026.081534</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Intra-Video Temporal-Aware RAG: A Self-Contained Framework for Video-Based Question Answering</article-title>
<alt-title alt-title-type="left-running-head">Intra-Video Temporal-Aware RAG: A Self-Contained Framework for Video-Based Question Answering</alt-title>
<alt-title alt-title-type="right-running-head">Intra-Video Temporal-Aware RAG: A Self-Contained Framework for Video-Based Question Answering</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Shafiq</surname><given-names>Sumaira</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Ejaz</surname><given-names>Naveed</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Shah</surname><given-names>Munam Ali</given-names></name><xref ref-type="aff" rid="aff-3">3</xref><email>mashah@kfu.edu.sa</email></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Kamal</surname><given-names>Rashid</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Sohail</surname><given-names>Adnan</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-6" contrib-type="author">
<name name-style="western"><surname>Aslam</surname><given-names>Sheraz</given-names></name><xref ref-type="aff" rid="aff-4">4</xref><xref ref-type="aff" rid="aff-5">5</xref><xref ref-type="aff" rid="aff-6">6</xref></contrib>
<aff id="aff-1"><label>1</label><institution>Department of Computing and Technology, Islamabad Campus</institution>, <addr-line>Iqra University, Islamabad</addr-line>, <country>Pakistan</country></aff>
<aff id="aff-2"><label>2</label><institution>School of Computing, Ulster University</institution>, <addr-line>Belfast</addr-line>, <country>UK</country></aff>
<aff id="aff-3"><label>3</label><institution>Department of Computer Networks and Communication, College of Computer Science and Information Technology, King Faisal University</institution>, <addr-line>Al-Ahsa</addr-line>, <country>Saudi Arabia</country></aff>
<aff id="aff-4"><label>4</label><institution>Department of Computer Science, CTL Eurocollege</institution>, <addr-line>Limassol</addr-line>, <country>Cyprus</country></aff>
<aff id="aff-5"><label>5</label><institution>Department of Computer Science, American University of Cyprus</institution>, <addr-line>Larnaca</addr-line>, <country>Cyprus</country></aff>
<aff id="aff-6"><label>6</label><institution>International Digital Economy College, Minjiang University</institution>, <addr-line>Fuzhou</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Munam Ali Shah. Email: <email>mashah@kfu.edu.sa</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>15</day><month>06</month><year>2026</year>
</pub-date>
<volume>88</volume>
<issue>2</issue>
<elocation-id>96</elocation-id>
<history>
<date date-type="received">
<day>09</day>
<month>03</month>
<year>2026</year>
</date>
<date date-type="accepted">
<day>21</day>
<month>04</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_81534.pdf"></self-uri>
<abstract>
<p>Lecture videos are widely used in modern education, yet answering questions from them remains challenging. Relevant information is often distributed across time and expressed through multiple modalities, including speech, slides, and visual content. Existing VideoQA approaches, including recent retrieval-augmented generation (RAG) methods, typically rely on static text representations or global video features. Consequently, they may retrieve evidence that is semantically relevant but temporally misaligned, leading to inaccurate or weakly grounded responses. In addition, dependence on external knowledge sources can introduce hallucinations and reduce reliability in educational settings. To address these limitations, we propose a temporally aware, intra-video RAG framework tailored for lecture videos. The approach aligns automatic speech transcripts and visual captions into timestamped segments and performs retrieval constrained by temporal boundaries. Retrieved segments are further refined using a cross-encoder before answer generation, ensuring that responses are grounded in the correct portions of the video. We evaluate the proposed method on the LectQA-Vid dataset, consisting of 100 lecture videos and 3000 temporally annotated questions. Experimental results demonstrate improved factual alignment and robustness over non-temporal baselines, highlighting the importance of temporal grounding in lecture VideoQA.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Video question answering</kwd>
<kwd>retrieval-augmented generation</kwd>
<kwd>temporal grounding</kwd>
<kwd>multimodal retrieval</kwd>
<kwd>educational videos</kwd>
<kwd>whisper ASR</kwd>
<kwd>visual captioning</kwd>
<kwd>large language models</kwd>
<kwd>explainable AI</kwd>
<kwd>timestamped evidence</kwd>
</kwd-group><funding-group>
<award-group id="awg1">
<funding-source>King Faisal University</funding-source>
<award-id>KFU262069</award-id>
</award-group>
</funding-group></article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>In recent years, Video Question Answering (VideoQA) has become a popular research area in multimedia artificial intelligence. The goal of VideoQA is to generate responses to diverse questions from video content using visual, spoken, and temporal information [<xref ref-type="bibr" rid="ref-1">1</xref>]. The VideoQA problem differs from visual question answering (VQA), which focuses on a single image [<xref ref-type="bibr" rid="ref-2">2</xref>]. VideoQA requires reasoning over temporal sequences in which relevant cues may be partially observable, appear briefly, or emerge at different times across different modalities [<xref ref-type="bibr" rid="ref-3">3</xref>].</p>
<p>Video lectures have become a popular way to learn. In lecture videos, the required information to answer a particular question is usually distributed over time and may also rely on different communication modalities, including spoken explanations, slides, diagrams, and on-screen text [<xref ref-type="bibr" rid="ref-4">4</xref>]. Therefore, effective VideoQA in lecture videos demands robust multimodal understanding, fine-grained temporal grounding, and the retrieval of relevant evidence from specific moments in the video.</p>
<p>Retrieval-Augmented Generation (RAG) pipelines [<xref ref-type="bibr" rid="ref-5">5</xref>,<xref ref-type="bibr" rid="ref-6">6</xref>] have recently been used for open-domain VideoQA by combining dense retrieval with the generative reasoning capabilities of large language models (LLMs). Despite their success, conventional RAG approaches typically assume the availability of structured, self-contained textual passages and provide no mechanisms for temporal alignment or multimodal grounding [<xref ref-type="bibr" rid="ref-7">7</xref>]. Most of these methods do not incorporate temporal grounding and thus do not guarantee that the retrieved segments cover the appropriate moments in the video timeline. Moreover, VideoQA systems that use LLMs often rely heavily on external text corpora, which may lead to hallucinations. These issues demand a temporally aware, multimodal, and intra-video retrieval-and-reasoning framework for VideoQA in educational videos.</p>
<p>To address these challenges, we propose a temporal-aware retrieval-augmented generation (RAG) framework designed specifically for lecture videos. The main idea is to use the video&#x2019;s temporal structure to generate answers d from the correct parts of the lecture. To achieve this, we first combine transcripts and visual captions into meaningful segments that are aligned with the video timeline. When a question is asked, the system searches for relevant segments within the appropriate time ranges and then refines the results to retain only the most useful ones. Finally, a language model generates the answer using only this selected evidence. This ensures that the responses are accurate, based on the lecture, and easy to understand.</p>
<p>To evaluate the proposed framework, we developed a dataset called LectQA-Vid. This dataset consists of 100 computer science lecture videos collected from YouTube, each ranging from 2 to 5 min long. Three annotators created a set of 3000 temporally aligned question&#x2013;answer pairs. Each question was assigned one of three difficulty levels: &#x201C;Simple&#x201D;, &#x201C;Hard&#x201D;, or &#x201C;Very Hard&#x201D;. On the LectQA-Vid dataset, the model achieves semantic similarity of 0.71&#x2013;0.77 and F1 scores of 0.16&#x2013;0.29 for open-ended questions, and 50%&#x2013;57% accuracy for multiple-choice questions. Moreover, ablation results highlight the importance of temporal filtering, multimodal fusion, and timestamp grounding.</p>
<p>The major contributions of this work are as follows:<list list-type="bullet">
<list-item>
<p>We propose a <bold>temporally aware, intra-video</bold> retrieval-augmented generation framework for lecture VideoQA that enforces <bold>timestamp-constrained evidence selection</bold>, ensuring that generated answers are grounded exclusively in the source video.</p></list-item>
<list-item>
<p>We introduce a unified <bold>multimodal pre-processing pipeline</bold> that integrates Whisper ASR transcripts and Gemini-based visual captions into <bold>temporally aligned semantic units</bold> suitable for dense retrieval.</p></list-item>
<list-item>
<p>We develop a <bold>temporally aware retrieval strategy</bold> that combines FAISS-based dense retrieval with explicit timestamp filtering and cross-encoder re-ranking, enabling <bold>verifiable and interpretable grounding</bold>.</p></list-item>
</list></p>
<p>The remainder of this paper is arranged as follows. <xref ref-type="sec" rid="s2">Section 2</xref> provides a brief overview of the existing research in VideoQA. <xref ref-type="sec" rid="s3">Section 3</xref> discusses the LectQA-Vid dataset. <xref ref-type="sec" rid="s4">Section 4</xref> discusses the proposed Temporal-Aware RAG framework. <xref ref-type="sec" rid="s5">Section 5</xref> describes the experimental setup and detailed experimental results. Finally, <xref ref-type="sec" rid="s6">Section 6</xref> concludes the work.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>The goal of VideoQA is to answer natural-language questions using visual and audio cues from videos. VideoQA methods involve the usage of temporal dynamics, motion cues, and long-range semantic dependencies, which make the task extremely challenging [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-9">9</xref>]. This section provides a brief review of the existing methods in VideoQA. It also discusses methods for temporal reasoning and grounding. The section also briefly discusses RAG-based approaches for VideoQA.</p>
<p>Typical VideoQA pipelines involve steps including video feature extraction, question encoding, multimodal fusion, and answer prediction [<xref ref-type="bibr" rid="ref-1">1</xref>]. Spatio-temporal reasoning is a fundamental step of video question answering, as many questions require the joint capture of spatial relationships and temporal dependencies across frames. Early VideoQA approaches model spatio-temporal information using appearance and motion features extracted at the frame or clip level [<xref ref-type="bibr" rid="ref-10">10</xref>,<xref ref-type="bibr" rid="ref-11">11</xref>]. Graph-based methods model temporal, spatial, and visual-linguistic relations by representing video units as nodes and enabling interactions with question structures [<xref ref-type="bibr" rid="ref-12">12</xref>&#x2013;<xref ref-type="bibr" rid="ref-14">14</xref>]. However, their use of nodes with mixed semantics limits the explicit alignment of video and questions.</p>
<p>Transformer-based architectures can model complex dependencies across multimodal inputs. Transformer-based video&#x2013;language models [<xref ref-type="bibr" rid="ref-15">15</xref>&#x2013;<xref ref-type="bibr" rid="ref-17">17</xref>] adopt frame- or clip-level features as visual inputs. Yang et al. [<xref ref-type="bibr" rid="ref-18">18</xref>] used subtitles and visual concepts using BERT. Furthermore, Yang et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] conducted a comparative study of text-language Transformers, including BERT, XLNet, RoBERTa, and ALBERT. They demonstrated that Transformer-based architectures can better capture complex multimodal semantics than recurrent models. Garcia et al. proposed a knowledge-driven framework called KnowIT VQA [<xref ref-type="bibr" rid="ref-20">20</xref>], which incorporated external knowledge sources by retrieving and integrating knowledge representations with video and textual features. Wu et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] investigated knowledge-oriented transfer learning in VideoQA by distinguishing between domain-specific and domain-agnostic knowledge, and transferring the latter across datasets to improve generalization. However, existing Transformer-based, knowledge-driven, and transfer learning approaches primarily rely on global video representations, external knowledge, or dataset-level knowledge transfer. They do not explicitly enforce temporal grounding during retrieval or reasoning. As a result, retrieved or attended evidence can be semantically relevant but temporally misaligned. In contrast, our work explicitly models intra-video temporal structure via timestamp-constrained retrieval, ensuring that evidence aligns with the correct temporal segments.</p>
<p>Temporal moment localization has been incorporated in VideoQA, as answering many questions requires first localizing the relevant temporal moment before reasoning over its visual content. Related research on temporal moment localization aims to align a natural-language query with its corresponding temporal segment in a video, i.e., identifying the start and end timestamps at which the described event occurs, typically using cross-modal attention [<xref ref-type="bibr" rid="ref-22">22</xref>] and structured temporal modeling [<xref ref-type="bibr" rid="ref-23">23</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>]. Grounded VideoQA frameworks further extend this idea by highlighting supporting frames or regions during answer inference [<xref ref-type="bibr" rid="ref-25">25</xref>,<xref ref-type="bibr" rid="ref-26">26</xref>]. However, these approaches are typically embedded in end-to-end architectures that lack explicit retrieval mechanisms or timestamp constraints, leading models to attend to semantically relevant yet temporally misaligned moments. This is especially critical in lecture videos, where information is dense and highly structured over time.</p>
<p>RAG improves factual grounding by using retrieved evidence during generation. The original RAG framework retrieves text prior to LLM decoding [<xref ref-type="bibr" rid="ref-27">27</xref>], whereas later multimodal extensions incorporate images and cross-modal features [<xref ref-type="bibr" rid="ref-28">28</xref>]. However, most RAG systems are built for static data and do not model temporal structure. In VideoQA, existing approaches rely mainly on semantic similarity and do not enforce retrieval from the correct time segments. As a result, evidence may be relevant but temporally misaligned.</p>
<p>Across the above areas, three key gaps remain. First, existing VideoQA and grounding methods rarely incorporate explicit timestamp constraints, leading to semantically relevant but temporally misaligned evidence. Second, multimodal RAG systems lack mechanisms to constrain retrieval to a specific video or temporal window, which is essential for factual alignment in lecture settings. Third, current educational VideoQA datasets generally lack timestamped evidence and do not enforce Intra-video grounding, limiting transparency and verifiability. This work addresses these gaps through a temporally aware Intra-video RAG framework that retrieves multimodal, timestamped evidence directly from the lecture video. By combining Whisper transcriptions, Gemini keyframe captions, BGE embeddings, FAISS retrieval, temporal filtering, and cross-encoder re-ranking, the proposed method produces verifiable, grounded answers tailored to the structure of educational lecture videos.</p>
</sec>
<sec id="s3">
<label>3</label>
<title>Dataset: LectQA-Vid</title>
<p>For the evaluation of temporally grounded retrieval and reasoning in educational VideoQA, we constructed a benchmark dataset called &#x201C;LectQA-Vid&#x201D; with explicit temporal supervision. This section describes the LectQA-Vid dataset, including its source videos, annotation pipeline, question design, and scope, independent of the modeling and retrieval techniques introduced later.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Corpus Construction and Video Selection</title>
<p>Videos in LectQA-Vid were selected from YouTube via keyword searches, including queries such as &#x201C;short computer science lecture&#x201D; and topic-specific terms, including operating systems, algorithms, artificial intelligence, computer architecture, and machine learning. From the retrieved results, we sampled 100 videos subject to quality and format constraints.</p>
<p>We retained videos that (i) are in English, (ii) have clear and intelligible audio narration, (iii) are short (2&#x2013;5 min), (iv) feature a single speaker, and (v) contain minimal background music. We focused on slide-intensive, lecture-style videos, in which the instructor&#x2019;s narration accompanies on-screen slides. To encourage diversity and reduce redundancy, videos were drawn from different channels. The resulting corpus consists of <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:mn>100</mml:mn></mml:math></inline-formula> videos. Each video focuses on a computer science topic, including (but not limited to) operating systems, algorithms, databases, networking, and programming languages, etc.</p>
<p>LectQA-Vid was intentionally constructed from computer science lectures featuring a single speaker and a slide-based format to provide a controlled and consistent experimental setting. This design isolates the impact of temporal alignment from additional variability introduced by multi-speaker interactions, diverse domains, or heterogeneous presentation styles.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Modalities and Derived Annotations</title>
<p>For each video, we derived two complementary textual modalities:<list list-type="bullet">
<list-item>
<p><bold>Transcript (Whisper ASR).</bold> The audio track was transcribed using Whisper (large-v3), producing a time-aligned transcript that served as the primary textual source for question answering.</p></list-item>
<list-item>
<p><bold>Visual captions (Gemini).</bold> A set of keyframes was extracted from each video, and captions were generated using a Gemini-based vision&#x2013;language model. These captions summarize salient on-screen visual information. While OCR-based approaches can extract textual content from slides, they are limited in capturing semantic and contextual information beyond visible text. Therefore, we adopted a vision-language model to generate richer visual captions that better support downstream reasoning.</p></list-item>
</list></p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Data Representation</title>
<p>The dataset is defined as:<disp-formula id="ueqn-1"><mml:math id="mml-ueqn-1" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:mi>&#x1D49F;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D4AC;</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> denotes the <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msup><mml:mi>i</mml:mi><mml:mrow><mml:mtext>th</mml:mtext></mml:mrow></mml:msup></mml:math></inline-formula> video (represented by its YouTube link and metadata), and <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mrow><mml:mi>&#x1D4AC;</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> denotes the associated set of question&#x2013;answer pairs. For each video <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>, we also provide metadata such as title, topic, duration, and the derived transcript and visual captions.</p>
<p>For each video <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>, we define:<disp-formula id="ueqn-2"><mml:math id="mml-ueqn-2" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x1D4AC;</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>M</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:msubsup><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where:<list list-type="bullet">
<list-item>
<p><inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msup><mml:mi>j</mml:mi><mml:mrow><mml:mtext>th</mml:mtext></mml:mrow></mml:msup></mml:math></inline-formula> question associated with video <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>;</p></list-item>
<list-item>
<p><inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> is the corresponding <bold>ground-truth answer</bold>, manually created and verified from the video content;</p></list-item>
<list-item>
<p><inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mtext>start</mml:mtext></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mtext>end</mml:mtext></mml:mrow></mml:msubsup><mml:mo stretchy="false">]</mml:mo><mml:mo>&#x2286;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> is the annotated temporal interval within <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> supporting <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msubsup><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> (with <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> the video duration);</p></list-item>
<list-item>
<p><inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mtext>``Simple''</mml:mtext><mml:mo>,</mml:mo><mml:mtext>``Hard''</mml:mtext><mml:mo>,</mml:mo><mml:mtext>``Very Hard''</mml:mtext><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> denotes question difficulty;</p></list-item>
<list-item>
<p><inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mtext>MCQ</mml:mtext><mml:mo>,</mml:mo><mml:mtext>Open</mml:mtext><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> denotes question type (multiple-choice or open-ended).</p></list-item>
</list></p>
<p>The total number of question&#x2013;answer pairs across the corpus is:<disp-formula id="ueqn-3"><mml:math id="mml-ueqn-3" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>M</mml:mi><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>M</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Question Design and Grounding Constraint</title>
<p>From each video, 15 multiple-choice questions (MCQs) and 15 open-ended questions were designed. Each MCQ presents four possible answers, with only one correct option. The questions in the dataset were categorized into three levels: &#x201C;Simple&#x201D;, &#x201C;Hard&#x201D;, and &#x201C;Very Hard&#x201D;, based on their difficulty. The difficulty labels are inherently subjective; thus, annotators were provided with clear guidelines, and a consensus on difficulty labels was achieved during the annotation review stage.</p>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Annotation Pipeline and Quality Control</title>
<p>Three computer science graduates annotated the videos in the dataset. All annotators possessed subject matter expertise and adhered to standardized guidelines. They reviewed the complete video, Whisper-generated transcript, and Gemini-generated keyframe captions.</p>
<p>For every question&#x2013;answer pair, annotators labeled a ground-truth temporal interval as:<disp-formula id="ueqn-4"><mml:math id="mml-ueqn-4" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msup><mml:mi>t</mml:mi><mml:mrow><mml:mrow><mml:mtext>start</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi>t</mml:mi><mml:mrow><mml:mrow><mml:mtext>end</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo stretchy="false">]</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>representing a <italic>single contiguous supporting span</italic> that contains sufficient evidence to answer the question. Temporal boundaries were defined using transcript-aligned timestamps to ensure consistency between textual and temporal annotations. Annotators were instructed to mark a <italic>reasonable supporting span</italic> rather than the minimal possible interval; boundaries are intended to be approximate (i.e., looser but safe) rather than frame-precise. When the necessary evidence was distributed across multiple portions of the lecture (e.g., a concept introduced earlier and applied later, or an audio explanation paired with a slide containing a key formula), <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula> was chosen to cover the whole region encompassing all required supporting content.</p>
<p>All annotated questions were subsequently reviewed by the first author, who verified the correctness of answer pairs, the alignment between the annotated temporal interval and the supporting video content, and consistency with the assigned difficulty labels. Ambiguous or weakly grounded questions were revised or removed during this review. To assess annotation consistency, a random subset of the dataset was independently inspected by a second annotator. Agreement was evaluated qualitatively with respect to answer correctness and temporal-span alignment, indicating high consistency in question interpretation and the localization of supporting temporal intervals, with only minor boundary-level differences that did not affect answerability.</p>
</sec>
<sec id="s3_6">
<label>3.6</label>
<title>Data Splits</title>
<p>We construct an 80/10/10 train&#x2013;validation&#x2013;test split over the 100 videos, while maintaining a balanced distribution of difficulty levels across splits. Unless otherwise stated, experiments are reported on a held-out evaluation subset of 1000 QA pairs (500 MCQ and 500 open-ended) sampled from the complete set of 3000 QA pairs.</p>
</sec>
<sec id="s3_7">
<label>3.7</label>
<title>Illustrative Examples and Dataset Statistics</title>
<p><xref ref-type="table" rid="table-1">Table 1</xref> summarises the LectQA-Vid dataset. The dataset consists of 100 short computer science lecture videos (2&#x2013;5 min each). The dataset includes 3000 question&#x2013;answer pairs, which are evenly split between MCQs and open-ended questions. The questions are of three difficulty levels (&#x201C;Simple&#x201D;, &#x201C;Hard&#x201D;, &#x201C;Very Hard&#x201D;). Each video contributes 30 questions (15 MCQs and 15 open-ended). Additionally, the dataset includes time-aligned ASR transcripts (Whisper), Gemini-based visual captions, and precise temporal spans for each question. <xref ref-type="table" rid="table-2">Table 2</xref> shows examples of questions in the LectQA-Vid dataset along with the groundtruth answers.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Dataset statistics for the LectQA-Vid lecture VideoQA corpus.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Category</th>
<th>Value</th>
</tr>
</thead>
<tbody>
<tr>
<td>Total Number of Videos</td>
<td>100</td>
</tr>
<tr>
<td>Video Type</td>
<td>Computer science micro-lectures (YouTube)</td>
</tr>
<tr>
<td>Topics Covered</td>
<td>OS, Networking, Databases, Algorithms, AI/ML</td>
</tr>
<tr>
<td>Video Duration</td>
<td>2&#x2013;5 min</td>
</tr>
<tr>
<td>Total Number of QA Pairs</td>
<td>3000 (1500 MCQ &#x002B; 1500 open-ended)</td>
</tr>
<tr>
<td>Question Types</td>
<td>MCQ (4 options) &#x002B; open-ended</td>
</tr>
<tr>
<td>QA Difficulty Levels</td>
<td>Simple, Hard, Very Hard</td>
</tr>
<tr>
<td>Distribution Across Difficulty Levels</td>
<td>1000 Simple, 1000 Hard, 1000 Very Hard</td>
</tr>
<tr>
<td>Questions per Video</td>
<td>30 (15 MCQ &#x002B; 15 open-ended)</td>
</tr>
<tr>
<td>MCQ Difficulty Split (per video)</td>
<td>5 Simple, 5 Hard, 5 Very Hard</td>
</tr>
<tr>
<td>Open-ended Difficulty Split (per video)</td>
<td>5 Simple, 5 Hard, 5 Very Hard</td>
</tr>
<tr>
<td>Average Transcript Length</td>
<td>8.3k characters/video</td>
</tr>
<tr>
<td>ASR Transcript</td>
<td>Whisper (large-v3), time-aligned</td>
</tr>
<tr>
<td>Visual Captions</td>
<td>Gemini-generated keyframe captions (timestamped)</td>
</tr>
<tr>
<td>Temporal Supervision</td>
<td><inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mi>&#x03C4;</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msup><mml:mi>t</mml:mi><mml:mrow><mml:mtext>start</mml:mtext></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mi>t</mml:mi><mml:mrow><mml:mtext>end</mml:mtext></mml:mrow></mml:msup><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> per QA (timestamp-aligned)</td>
</tr>
<tr>
<td>Provided Artifacts</td>
<td>YouTube links &#x002B; transcripts &#x002B; keyframe captions &#x002B; QA pairs</td>
</tr>
<tr>
<td>Knowledge Constraint</td>
<td>Intra-video only; no external knowledge</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Examples from the LectQA-Vid dataset.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Question</th>
<th>Reference Answer</th>
<th>Difficulty</th>
<th>Primary Grounding</th>
</tr>
</thead>
<tbody>
<tr>
<td>What assumption is made when defining the empirical risk minimization objective in the lecture?</td>
<td>The objective assumes that training samples are independently and identically distributed and that the expected risk can be approximated by the empirical average shown on the slide.</td>
<td>Simple</td>
<td>Transcript &#x002B; Slide</td>
</tr>
<tr>
<td>Why does the instructor prefer stochastic gradient descent over full-batch gradient descent in the optimization discussion?</td>
<td>Because stochastic gradient descent enables scalable learning on large datasets and provides faster convergence in practice, as explained using the optimization slides.</td>
<td>Hard</td>
<td>Transcript</td>
</tr>
<tr>
<td>What limitation of the baseline algorithm is highlighted by the convergence plot shown in the lecture?</td>
<td>The convergence plot shows that the baseline algorithm converges slowly due to a fixed learning rate, motivating the adaptive method introduced later.</td>
<td>Very Hard</td>
<td>Visual (Plot)</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Materials and Methods</title>
<p>This section provides a detailed description of the proposed Temporal-Aware RAG framework. <xref ref-type="fig" rid="fig-1">Fig. 1</xref> provides an overview of the proposed framework.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Overview of the proposed temporal-aware, Intra-video RAG pipeline for lecture VideoQA.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81534-fig-1.tif"/>
</fig>
<sec id="s4_1">
<label>4.1</label>
<title>Data Extraction</title>
<p>Let a video dataset be represented as a collection <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, where each video <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> is characterized by its visual frame sequence and corresponding audio track, <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">I</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">a</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2223;</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">]</mml:mo><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>, with <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">I</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denoting the RGB frame at time <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mi>t</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">a</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> the corresponding audio signal, and <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mi>T</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> the total duration of <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>.</p>
<p>The goal of the data extraction phase is to convert each raw video <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> into two temporally aligned, semantically meaningful textual modalities: (1) a sequence of linguistic segments obtained from automatic speech recognition (ASR), and (2) a sequence of visual sentences obtained from keyframe captioning. Formally, this phase produces the intermediate representation:<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>&#x222A;</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>v</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></disp-formula>where <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msubsup><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">a</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msubsup><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">v</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> denote audio-derived and vision-derived textual items, respectively.</p>
<sec id="s4_1_1">
<label>4.1.1</label>
<title>Audio Transcription</title>
<p>The audio component <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">a</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is processed by an automatic speech recognition model <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">A</mml:mi><mml:mi mathvariant="normal">S</mml:mi><mml:mi mathvariant="normal">R</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, implemented using Whisper. The model generates a set of <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msubsup><mml:mi>N</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">a</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> speech segments:<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msubsup><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>start</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>end</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msubsup><mml:mi>N</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:msubsup></mml:math></disp-formula>where <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msubsup><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">a</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> denotes the transcribed text and <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">t</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">r</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">n</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:mrow></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> are the corresponding temporal boundaries. Equivalently, each segment is obtained via the mapping:<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mtext>ASR</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>:</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">a</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">&#x21A6;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>start</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>end</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>providing temporally localized linguistic evidence aligned with the spoken narrative of the video.</p>
</sec>
<sec id="s4_1_2">
<label>4.1.2</label>
<title>Keyframe Extraction and Captioning</title>
<p>The visual stream <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">I</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>t</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is sampled to obtain a set of representative keyframes:<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msub><mml:mrow><mml:mi>&#x1D4A6;</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">I</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>}</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msubsup><mml:mi>N</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>v</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:msubsup></mml:math></disp-formula>where the sampling instants <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> are determined either by scene-change detection or by uniform temporal intervals <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi></mml:math></inline-formula>. Each keyframe is converted into a natural-language description using a vision&#x2013;language captioning model <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">V</mml:mi><mml:mi mathvariant="normal">L</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, instantiated here as Gemini. Video transcripts are generated using Whisper, while visual captions are produced using Gemini. The resulting set of visual sentences is:<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msubsup><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>v</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>v</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msubsup><mml:mi>N</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>v</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:mrow></mml:msubsup></mml:math></disp-formula>where <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msubsup><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">v</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">V</mml:mi><mml:mi mathvariant="normal">L</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mspace width="negativethinmathspace" /><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">I</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> encodes the semantic content of frame <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">I</mml:mtext></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> at time <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>.</p>
</sec>
<sec id="s4_1_3">
<label>4.1.3</label>
<title>Temporal Unification</title>
<p>In the next step, the two modalities are temporally synchronized into a unified sequence ordered by time:<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msub><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>sort</mml:mtext></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:msubsup><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>&#x222A;</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>v</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>start</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>end</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:msubsup></mml:math></disp-formula>where sorting ensures chronological coherence. This unified representation <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msub><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> is then used for subsequent preprocessing, semantic chunking, and embedding steps.</p>
</sec>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Data Pre-Processing</title>
<p>The unified sequence <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msub><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">t</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">r</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>&#x03C4;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">n</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:mrow></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:msubsup></mml:math></inline-formula> obtained from the data extraction phase may be noisy and redundant. Therefore, preprocessing is performed to transform <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msub><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> into a temporally consistent set of semantically meaningful textual units. Formally, this transformation is represented as:<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:msub><mml:mi>g</mml:mi><mml:mrow><mml:mrow><mml:mtext>prep</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>:</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo stretchy="false">&#x27F6;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x1D49E;</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>start</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>end</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:msubsup></mml:math></disp-formula>where each <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes a normalized, semantically coherent chunk, temporally bounded by <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">t</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">r</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">n</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:mrow></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. The rest of the section describes the pre-processing steps.</p>
<sec id="s4_2_1">
<label>4.2.1</label>
<title>Text Normalization and Noise Removal</title>
<p>Each segment <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is normalized using the function <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:msub><mml:mrow><mml:mover><mml:mi>s</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">n</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">r</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, which removes non-linguistic tokens (e.g., [music], [applause]), repeated punctuation, and capitalization inconsistencies. Segments with <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>s</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mo>&#x003C;</mml:mo><mml:mtext>&#xA0;</mml:mtext><mml:msub><mml:mo>&#x03B5;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="normal">l</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">n</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> (shorter than a minimum token threshold) are discarded. Moreover, duplicate or near-duplicate sentences are eliminated using a cosine-similarity filter.</p>
</sec>
<sec id="s4_2_2">
<label>4.2.2</label>
<title>Sentence Boundary Refinement</title>
<p>Long ASR sentences are split into smaller parts using punctuation and grammar rules. These parts are then combined again when they are temporally close and syntactically relevant.</p>
<p>If <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mrow><mml:mi>&#x0212C;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>s</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>s</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>s</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>r</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> denote the boundary refinement operator producing <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:mi>r</mml:mi></mml:math></inline-formula> refined sub-sentences, the updated sequence after boundary refinement is
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:msubsup><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x22C3;</mml:mo><mml:mrow><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:munderover><mml:mrow><mml:mi>&#x0212C;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>s</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
</sec>
<sec id="s4_2_3">
<label>4.2.3</label>
<title>Temporal Alignment across Modalities</title>
<p>In <xref ref-type="disp-formula" rid="eqn-8">Eq. (8)</xref>, <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:msubsup><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> combines both audio and visual textual items. Next, a temporal alignment function <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">l</mml:mi><mml:mi mathvariant="normal">i</mml:mi><mml:mi mathvariant="normal">g</mml:mi><mml:mi mathvariant="normal">n</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is used to ensure consistent ordering and gap filling. The function is defined as:<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:msubsup><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mtext>align</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>start</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>end</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03B7;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:msubsup></mml:math></disp-formula>where <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msub><mml:mi>&#x03B7;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mspace width="negativethinmathspace" /><mml:mo>&#x2208;</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mrow><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">u</mml:mi><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="normal">i</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mi mathvariant="normal">v</mml:mi><mml:mi mathvariant="normal">i</mml:mi><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">u</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">l</mml:mi></mml:mrow><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> indicates the source modality.</p>
<p>In this step, audio transcripts and visual captions are fused at the text level by aligning modality-specific segments based on timestamps and combining them into a single sequence while preserving temporal order. We define a temporal overlap threshold <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> to resolve any potential conflicts between modalities. The temporal overlap between two segments is defined as:<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mrow><mml:mtext>overlap</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>q</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">(</mml:mo></mml:mrow></mml:mstyle><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>end</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>end</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>start</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>start</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">)</mml:mo></mml:mrow></mml:mstyle></mml:math></disp-formula></p>
<p>If the overlap between an audio and visual segment exceeds <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula>, the segments are considered temporally redundant. In such cases, the visual segment is either suppressed or merged into the corresponding audio segment, with priority given to the audio transcript.</p>
<p>Alternative fusion strategies, such as early feature-level fusion or late fusion of independently processed modalities, were considered. However, early fusion increases model complexity and requires cross-modal embedding alignment, while late fusion may lead to fragmented or inconsistent context across modalities. In contrast, the proposed temporally aligned text-level fusion provides a simple, interpretable, and computationally efficient mechanism that preserves both semantic coherence and temporal structure.</p>
</sec>
<sec id="s4_2_4">
<label>4.2.4</label>
<title>Semantic Chunking</title>
<p>Consecutive sentences in <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:msubsup><mml:mrow><mml:mi>&#x1D4AE;</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> that are semantically related and temporally contiguous are merged into higher-level units, termed <italic>semantic chunks</italic>. Let <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denote a semantic affinity function between two adjacent sentences:<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mrow><mml:mrow><mml:mtext>sem</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">h</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msubsup><mml:msub><mml:mrow><mml:mtext mathvariant="bold">h</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">h</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">h</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">h</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are contextual embeddings from a pretrained sentence encoder.</p>
<p>We introduce a semantic similarity threshold <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> to control chunk formation, and a temporal gap threshold <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">g</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">p</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> to enforce temporal continuity. The threshold <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">g</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">p</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> defines the maximum allowable time difference between two consecutive sentences for them to be considered temporally adjacent. Two sentences are merged if
<disp-formula id="ueqn-16"><mml:math id="mml-ueqn-16" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>&#x03D5;</mml:mi><mml:mrow><mml:mrow><mml:mtext>sem</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x003E;</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mrow><mml:mtext>sem</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mspace width="1em" /><mml:mrow><mml:mtext>and</mml:mtext></mml:mrow><mml:mspace width="1em" /><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:mtext>start</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>end</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x003C;</mml:mo><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mrow><mml:mtext>gap</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>ensuring both semantic coherence and temporal continuity.</p>
<p>The merged chunk <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is defined as
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2A01;</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>e</mml:mi></mml:msub></mml:mrow></mml:munderover><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="2em" /><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>start</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:msub><mml:mi>p</mml:mi><mml:mi>s</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:mtext>start</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mspace width="2em" /><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>end</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:msub><mml:mi>p</mml:mi><mml:mi>e</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:mtext>end</mml:mtext></mml:mrow></mml:mrow></mml:msubsup></mml:math></disp-formula>where <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mo>&#x2A01;</mml:mo></mml:math></inline-formula> denotes string concatenation preserving sentence order. The resulting set of chunks <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:msub><mml:mrow><mml:mi>&#x1D49E;</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> forms the temporally grounded, semantically consolidated representation of video <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>.</p>
</sec>
<sec id="s4_2_5">
<label>4.2.5</label>
<title>Output Representation</title>
<p>After pre-processing, each video <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> is represented by <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:msub><mml:mrow><mml:mi>&#x1D49E;</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">t</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">r</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">n</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:mrow></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:msubsup></mml:math></inline-formula>, where each <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> corresponds to a normalized, semantically coherent textual unit which is associated with a temporal interval within the video. This structured representation serves as input to the embedding and indexing stage, as discussed in <xref ref-type="sec" rid="s4_3">Section 4.3</xref>.</p>
</sec>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Representation and Indexing</title>
<p>The pre-processed video representation <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:msub><mml:mrow><mml:mi>&#x1D49E;</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">t</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">r</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">n</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:mrow></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:msubsup></mml:math></inline-formula> encapsulates semantically coherent textual units. Each textual unit is associated with a temporal interval in video <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:msub><mml:mi>v</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula>. Each chunk <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is then mapped to a dense vector representation in a continuous embedding space <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mi>d</mml:mi></mml:msup></mml:math></inline-formula>. This transformation is achieved by a sentence-level embedding model <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">m</mml:mi><mml:mi mathvariant="normal">b</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>:<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mtext>emb</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mi>d</mml:mi></mml:msup></mml:math></disp-formula>where <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is a <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mi>d</mml:mi></mml:math></inline-formula>-dimensional vector encoding the semantic and contextual meaning of chunk <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>.</p>
<sec id="s4_3_1">
<label>4.3.1</label>
<title>Embedding Function</title>
<p>The embedding model <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">m</mml:mi><mml:mi mathvariant="normal">b</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is instantiated using a pre-trained Bidirectional Generalized Embedding (BGE) encoder, denoted as <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:msub><mml:mtext>BGE</mml:mtext><mml:mrow><mml:mtext>base</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula>. For each textual input <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> composed of <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> tokens, the encoder produces token-level hidden states <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">H</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">h</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">h</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>, which are mean-pooled to obtain the chunk-level embedding:<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>&#x2113;</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munderover><mml:msubsup><mml:mrow><mml:mtext mathvariant="bold">h</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x2113;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></disp-formula></p>
<p>The embeddings are then <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:msub><mml:mi>&#x2113;</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula>-normalized:<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:msub><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>The embedding dimension <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mi>d</mml:mi></mml:math></inline-formula> is fixed (<inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:mi>d</mml:mi><mml:mo>=</mml:mo><mml:mn>768</mml:mn></mml:math></inline-formula> for <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:msub><mml:mtext>BGE</mml:mtext><mml:mrow><mml:mtext>base</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula>).</p>
</sec>
<sec id="s4_3_2">
<label>4.3.2</label>
<title>FAISS Index Construction</title>
<p>All normalized embeddings from the dataset are aggregated into a single matrix: <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:mrow><mml:mtext mathvariant="bold">Z</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>11</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>12</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:msup><mml:mo stretchy="false">]</mml:mo><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>M</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, where <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mi>M</mml:mi><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>N</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:math></inline-formula> is the total number of chunks across all videos. Each embedding <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is associated with metadata <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mtext>video_id</mml:mtext><mml:mo>=</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mtext>chunk_id</mml:mtext><mml:mo>=</mml:mo><mml:mi>n</mml:mi><mml:mo>,</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">t</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">r</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">n</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:mrow></mml:mrow></mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>, linking the vector to its source and temporal span.</p>
<p>For efficient approximate nearest-neighbor (ANN) search, we employ the FAISS (Facebook AI Similarity Search) library to construct an index <inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:mrow><mml:mi>&#x02110;</mml:mi></mml:mrow></mml:math></inline-formula>:<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:mrow><mml:mi>&#x02110;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext>FAISS.build</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext mathvariant="bold">Z</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>metric</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext>cosine</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula>where FAISS organizes the embeddings using an inverted-file index with product quantization (IVF-PQ) for sub-linear retrieval. This yields a mapping: <inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">i</mml:mi><mml:mi mathvariant="normal">n</mml:mi><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">x</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mspace width="negativethinmathspace" /><mml:mo>:</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">&#x21A6;</mml:mo><mml:mrow><mml:mi>&#x02110;</mml:mi></mml:mrow></mml:math></inline-formula>, allowing rapid lookup of semantically similar vectors.</p>
</sec>
<sec id="s4_3_3">
<label>4.3.3</label>
<title>Similarity Formulation</title>
<p>Given a query embedding <inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>q</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">m</mml:mi><mml:mi mathvariant="normal">b</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>q</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> derived from a textual question <inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:mi>q</mml:mi></mml:math></inline-formula>, semantic similarity between the query and a candidate chunk <inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is computed using cosine similarity:<disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:mrow><mml:mtext>sim</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>q</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>q</mml:mi><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msubsup><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p>We define a retrieval similarity threshold <inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>r</mml:mi></mml:msub></mml:math></inline-formula> to characterize relevance, and only those chunks with similarity greater than <inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>r</mml:mi></mml:msub></mml:math></inline-formula> are selected as semantically relevant candidates. For large-scale retrieval, FAISS efficiently estimates the top-<inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:mi>k</mml:mi></mml:math></inline-formula> nearest neighbors:<disp-formula id="ueqn-23"><mml:math id="mml-ueqn-23" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>TopK</mml:mtext></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtext>sim</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>q</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:mi>k</mml:mi></mml:math></inline-formula> denotes the number of retrieved candidates.</p>
<p>The resulting set is a ranked list of chunk identifiers and corresponding similarity scores:<disp-formula id="eqn-18"><label>(18)</label><mml:math id="mml-eqn-18" display="block"><mml:msubsup><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mo>{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mtext>sim</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mrow><mml:mtext>sim</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>q</mml:mi><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msubsup><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:mtext>top-</mml:mtext></mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula></p>
</sec>
<sec id="s4_3_4">
<label>4.3.4</label>
<title>Temporal Metadata Preservation</title>
<p>Each retrieved embedding in <inline-formula id="ieqn-100"><mml:math id="mml-ieqn-100"><mml:msubsup><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> retains its temporal and source metadata <inline-formula id="ieqn-101"><mml:math id="mml-ieqn-101"><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. This structure allows downstream modules to apply temporal constraints during the temporal-aware retrieval phase. The representation and indexing phase produces a searchable, semantically rich vector space, given by:<disp-formula id="ueqn-25"><label>(19)</label><mml:math id="mml-ueqn-25" display="block"><mml:msub><mml:mrow><mml:mi>&#x1D49F;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>index</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mi>N</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:msubsup></mml:math></disp-formula></p>
<p>This vector space <inline-formula id="ieqn-102"><mml:math id="mml-ieqn-102"><mml:msub><mml:mrow><mml:mi>&#x1D49F;</mml:mi></mml:mrow><mml:mrow><mml:mtext>index</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> serves as the knowledge base for efficient, timestamp-aware video question answering.</p>
</sec>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Temporal-Aware Retrieval and Reasoning</title>
<p>Given the indexed representation <inline-formula id="ieqn-103"><mml:math id="mml-ieqn-103"><mml:msub><mml:mrow><mml:mi>&#x1D49F;</mml:mi></mml:mrow><mml:mrow><mml:mtext>index</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula>, the goal of this stage is to find video segments that are both relevant to the query <inline-formula id="ieqn-104"><mml:math id="mml-ieqn-104"><mml:mi>q</mml:mi></mml:math></inline-formula> and correctly aligned in time. Unlike standard RAG systems that search over static text, this framework retrieves information directly from the video. It uses a temporally aware, intra-video retrieval process to ensure that all selected evidence comes from the correct parts of the same video. This design helps maintain factual accuracy and enables the system to produce answers with clear, timestamped evidence that is easy to interpret.</p>
<sec id="s4_4_1">
<label>4.4.1</label>
<title>Query Embedding and Initial Retrieval</title>
<p>Each user question <inline-formula id="ieqn-105"><mml:math id="mml-ieqn-105"><mml:mi>q</mml:mi></mml:math></inline-formula> is first encoded into a dense query vector, given by:<disp-formula id="eqn-20"><label>(20)</label><mml:math id="mml-eqn-20" display="block"><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mtext>emb</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>q</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mspace width="2em" /><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></disp-formula>using the same embedding function <inline-formula id="ieqn-106"><mml:math id="mml-ieqn-106"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">m</mml:mi><mml:mi mathvariant="normal">b</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> defined in <xref ref-type="sec" rid="s4_3">Section 4.3</xref>. An initial candidate set is obtained from the FAISS index by retrieving the top-<inline-formula id="ieqn-107"><mml:math id="mml-ieqn-107"><mml:mi>K</mml:mi></mml:math></inline-formula> semantically nearest neighbors:<disp-formula id="eqn-21"><label>(21)</label><mml:math id="mml-eqn-21" display="block"><mml:msubsup><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>K</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>TopK</mml:mtext></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mtext>sim</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:mrow><mml:mtext>sim</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>q</mml:mi></mml:mrow><mml:mi mathvariant="normal">&#x22A4;</mml:mi></mml:msubsup><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mtext mathvariant="bold">z</mml:mtext></mml:mrow><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p>The resulting retrieval set <inline-formula id="ieqn-108"><mml:math id="mml-ieqn-108"><mml:msubsup><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>K</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">i</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:msub><mml:mi>n</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> contains both the similarity scores and corresponding temporal metadata <inline-formula id="ieqn-109"><mml:math id="mml-ieqn-109"><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:msub><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:msub><mml:mi>n</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula>.</p>
</sec>
<sec id="s4_4_2">
<label>4.4.2</label>
<title>Temporal-Aware Filtering</title>
<p>To exploit temporal structure and maintain coherence, a timestamp-based filtering function <inline-formula id="ieqn-110"><mml:math id="mml-ieqn-110"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">t</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">m</mml:mi><mml:mi mathvariant="normal">p</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is applied to <inline-formula id="ieqn-111"><mml:math id="mml-ieqn-111"><mml:msubsup><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>K</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>. The parameter <inline-formula id="ieqn-112"><mml:math id="mml-ieqn-112"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi></mml:math></inline-formula> represents a temporal tolerance window that expands the query-specific interval to account for minor misalignment between annotated timestamps and retrieved segments.
<disp-formula id="eqn-22"><label>(22)</label><mml:math id="mml-eqn-22" display="block"><mml:msubsup><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mrow><mml:mtext>temp</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>K</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mspace width="thinmathspace" /><mml:mrow></mml:mrow><mml:mo>|</mml:mo><mml:mrow></mml:mrow><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:msub><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:msub><mml:mi>n</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:mtext>start</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:msub><mml:mi>i</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:msub><mml:mi>n</mml:mi><mml:mi>k</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:mtext>end</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo stretchy="false">]</mml:mo><mml:mo>&#x2229;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mi>t</mml:mi><mml:mi>q</mml:mi><mml:mrow><mml:mrow><mml:mtext>start</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msubsup><mml:mi>t</mml:mi><mml:mi>q</mml:mi><mml:mrow><mml:mrow><mml:mtext>end</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi><mml:mo stretchy="false">]</mml:mo><mml:mo>&#x2260;</mml:mo><mml:mi mathvariant="normal">&#x2205;</mml:mi><mml:mo>}</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>The system then applies local temporal smoothing to group adjacent chunks belonging to the same video:<disp-formula id="eqn-23"><label>(23)</label><mml:math id="mml-eqn-23" display="block"><mml:msubsup><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mrow><mml:mtext>smooth</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mtext>merge</mml:mtext></mml:mrow><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.623em" minsize="1.623em">(</mml:mo></mml:mrow></mml:mstyle><mml:msubsup><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>K</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup><mml:mo>;</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.623em" minsize="1.623em">)</mml:mo></mml:mrow></mml:mstyle></mml:math></disp-formula>where <inline-formula id="ieqn-113"><mml:math id="mml-ieqn-113"><mml:mrow><mml:mi mathvariant="normal">m</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">r</mml:mi><mml:mi mathvariant="normal">g</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is used to aggregate consecutive results whose time gaps are below <inline-formula id="ieqn-114"><mml:math id="mml-ieqn-114"><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula>. This temporal-aware filtering ensures that retrieved evidence is not only semantically relevant but also temporally contiguous and contextually stable.</p>
</sec>
<sec id="s4_4_3">
<label>4.4.3</label>
<title>Cross-Encoder Re-Ranking</title>
<p>We refine the top-<inline-formula id="ieqn-115"><mml:math id="mml-ieqn-115"><mml:mi>M</mml:mi></mml:math></inline-formula> candidates (<inline-formula id="ieqn-116"><mml:math id="mml-ieqn-116"><mml:mi>M</mml:mi><mml:mo>&#x003C;</mml:mo><mml:mi>K</mml:mi></mml:math></inline-formula>) of the retrieval using a cross-encoder scoring model <inline-formula id="ieqn-117"><mml:math id="mml-ieqn-117"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">c</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. This cross-encoder scoring model evaluates the pairwise relevance between the query <inline-formula id="ieqn-118"><mml:math id="mml-ieqn-118"><mml:mi>q</mml:mi></mml:math></inline-formula> and each candidate chunk <inline-formula id="ieqn-119"><mml:math id="mml-ieqn-119"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>:<disp-formula id="eqn-24"><label>(24)</label><mml:math id="mml-eqn-24" display="block"><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mtext>ce</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>q</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-120"><mml:math id="mml-ieqn-120"><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is a scalar relevance score in <inline-formula id="ieqn-121"><mml:math id="mml-ieqn-121"><mml:mo stretchy="false">[</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>. The final ranked set is:<disp-formula id="eqn-25"><label>(25)</label><mml:math id="mml-eqn-25" display="block"><mml:msubsup><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>sort</mml:mtext></mml:mrow><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2223;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mrow><mml:mtext>temp</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>This set is ordered by descending <inline-formula id="ieqn-122"><mml:math id="mml-ieqn-122"><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. <inline-formula id="ieqn-123"><mml:math id="mml-ieqn-123"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">c</mml:mi><mml:mi mathvariant="normal">e</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is implemented using a transformer-based cross-encoder fine-tuned on QA relevance pairs. This re-ranking step yields high-precision evidence selection, improving over embedding-only retrieval.</p>
</sec>
<sec id="s4_4_4">
<label>4.4.4</label>
<title>Context Assembly and Reasoning</title>
<p>From the top-<inline-formula id="ieqn-124"><mml:math id="mml-ieqn-124"><mml:mi>L</mml:mi></mml:math></inline-formula> elements of <inline-formula id="ieqn-125"><mml:math id="mml-ieqn-125"><mml:msubsup><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>, a compact context window <inline-formula id="ieqn-126"><mml:math id="mml-ieqn-126"><mml:msub><mml:mrow><mml:mi>&#x2130;</mml:mi></mml:mrow><mml:mi>q</mml:mi></mml:msub></mml:math></inline-formula> is constructed:<disp-formula id="eqn-26"><label>(26)</label><mml:math id="mml-eqn-26" display="block"><mml:msub><mml:mrow><mml:mi>&#x2130;</mml:mi></mml:mrow><mml:mi>q</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">{</mml:mo></mml:mrow></mml:mstyle><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>c</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x2113;</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>t</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x2113;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>start</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>t</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x2113;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>end</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:msubsup><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.2em" minsize="1.2em">}</mml:mo></mml:mrow></mml:mstyle><mml:mrow><mml:mi>&#x2113;</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>c</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x2113;</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>t</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x2113;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>start</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>t</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x2113;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>end</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2208;</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow><mml:mi>q</mml:mi><mml:mrow><mml:mo>&#x2217;</mml:mo></mml:mrow></mml:msubsup></mml:math></disp-formula></p>
<p>The corresponding text snippets are concatenated into a temporally ordered context:<disp-formula id="eqn-27"><label>(27)</label><mml:math id="mml-eqn-27" display="block"><mml:msub><mml:mi>C</mml:mi><mml:mi>q</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2A01;</mml:mo><mml:mrow><mml:mi>&#x2113;</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:munderover><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.623em" minsize="1.623em">[</mml:mo></mml:mrow></mml:mstyle><mml:mo stretchy="false">[</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>t</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x2113;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>start</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mover><mml:mi>t</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x2113;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>end</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo stretchy="false">]</mml:mo><mml:mo>:</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>c</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>&#x2113;</mml:mi></mml:mrow></mml:msub><mml:mstyle scriptlevel="0"><mml:mrow><mml:mo maxsize="1.623em" minsize="1.623em">]</mml:mo></mml:mrow></mml:mstyle></mml:math></disp-formula></p>
<p>An instruction-tuned LLM <inline-formula id="ieqn-127"><mml:math id="mml-ieqn-127"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">L</mml:mi><mml:mi mathvariant="normal">L</mml:mi><mml:mi mathvariant="normal">M</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> then generates a grounded natural-language answer:<disp-formula id="eqn-28"><label>(28)</label><mml:math id="mml-eqn-28" display="block"><mml:msub><mml:mi>a</mml:mi><mml:mi>q</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:mtext>LLM</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>q</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mi>q</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula>subject to the grounding constraint that all information used in <inline-formula id="ieqn-128"><mml:math id="mml-ieqn-128"><mml:msub><mml:mi>a</mml:mi><mml:mi>q</mml:mi></mml:msub></mml:math></inline-formula> must originate from <inline-formula id="ieqn-129"><mml:math id="mml-ieqn-129"><mml:msub><mml:mi>C</mml:mi><mml:mi>q</mml:mi></mml:msub></mml:math></inline-formula>.</p>
<p>This constraint instructs the LLM to act primarily as a reasoning engine by restricting answer generation to the retrieved video-derived context.</p>
<p><xref ref-type="fig" rid="fig-2">Fig. 2</xref> <bold>summarizes the structured input formulation used in the proposed framework.</bold></p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Preparation of structured, temporally grounded input for the proposed VideoQA framework.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81534-fig-2.tif"/>
</fig>
</sec>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Experiments and Results</title>
<p>This section presents the experimental evaluation of the proposed framework on the LectQA-Vid dataset. We first describe the evaluation setup, implementation details, and assessment metrics. We then report both quantitative and qualitative results. Finally, we discuss performance over different difficulty levels and the approach&#x2019;s strengths and limitations.</p>
<p>All experiments were conducted using <bold>Google Colab Pro&#x002B;</bold>, configured with an NVIDIA Tesla T4 GPU (16 GB VRAM), 32 GB RAM, and 4 vCPUs. The software environment included Python 3.10, PyTorch 2.2.0, HuggingFace Transformers 4.39.3, FAISS-gpu 1.7.4, and SentenceTransformers 2.6.0. All runs were made deterministic by setting the random seed to 42 and enabling deterministic CUDA kernels.</p>
<p>The experiments were performed using the LectQA-Vid dataset (<xref ref-type="sec" rid="s3">Section 3</xref>). The dataset is partitioned at the video level into 80% training, 10% validation, and 10% test splits. For computational efficiency, evaluation is conducted on a stratified subset of 1000 question&#x2013;answer pairs sampled from the test videos (500 MCQs and 500 open-ended questions).</p>
<sec id="s5_1">
<label>5.1</label>
<title>Threshold Selection</title>
<p>The proposed framework uses several thresholds to control semantic chunking, temporal alignment, and retrieval relevance. These thresholds include the semantic similarity threshold <inline-formula id="ieqn-130"><mml:math id="mml-ieqn-130"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, temporal overlap threshold <inline-formula id="ieqn-131"><mml:math id="mml-ieqn-131"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula>, temporal gap threshold <inline-formula id="ieqn-132"><mml:math id="mml-ieqn-132"><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">g</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">p</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, duplicate similarity threshold <inline-formula id="ieqn-133"><mml:math id="mml-ieqn-133"><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="normal">u</mml:mi><mml:mi mathvariant="normal">p</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, retrieval similarity threshold <inline-formula id="ieqn-134"><mml:math id="mml-ieqn-134"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>r</mml:mi></mml:msub></mml:math></inline-formula>, and the temporal tolerance parameter <inline-formula id="ieqn-135"><mml:math id="mml-ieqn-135"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi></mml:math></inline-formula> used in timestamp-based filtering, along with the retrieval parameter <inline-formula id="ieqn-136"><mml:math id="mml-ieqn-136"><mml:mi>k</mml:mi></mml:math></inline-formula>. All similarity-based thresholds are derived from cosine similarity. Temporal thresholds are defined in seconds.</p>
<p>We used validation data to find the right values for these parameters that balance semantic coherence, temporal consistency, and retrieval precision. <xref ref-type="table" rid="table-3">Table 3</xref> contains the selected thresholds.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Thresholds used in the proposed framework.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Parameter</th>
<th>Symbol</th>
<th>Value</th>
<th>Range</th>
</tr>
</thead>
<tbody>
<tr>
<td>Semantic similarity threshold</td>
<td><inline-formula id="ieqn-137"><mml:math id="mml-ieqn-137"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>0.75</td>
<td>[0, 1]</td>
</tr>
<tr>
<td>Temporal overlap threshold</td>
<td><inline-formula id="ieqn-138"><mml:math id="mml-ieqn-138"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula></td>
<td>2 s</td>
<td>[0, <inline-formula id="ieqn-139"><mml:math id="mml-ieqn-139"><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:math></inline-formula>)</td>
</tr>
<tr>
<td>Temporal gap threshold</td>
<td><inline-formula id="ieqn-140"><mml:math id="mml-ieqn-140"><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">g</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">p</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>5 s</td>
<td>[0, <inline-formula id="ieqn-141"><mml:math id="mml-ieqn-141"><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:math></inline-formula>)</td>
</tr>
<tr>
<td>Duplicate similarity threshold</td>
<td><inline-formula id="ieqn-142"><mml:math id="mml-ieqn-142"><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">d</mml:mi><mml:mi mathvariant="normal">u</mml:mi><mml:mi mathvariant="normal">p</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>0.90</td>
<td>[0, 1]</td>
</tr>
<tr>
<td>Retrieval similarity threshold</td>
<td><inline-formula id="ieqn-143"><mml:math id="mml-ieqn-143"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>r</mml:mi></mml:msub></mml:math></inline-formula></td>
<td>0.60</td>
<td>[0, 1]</td>
</tr>
<tr>
<td>Temporal tolerance window</td>
<td><inline-formula id="ieqn-144"><mml:math id="mml-ieqn-144"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi></mml:math></inline-formula></td>
<td>3 s</td>
<td>[0, <inline-formula id="ieqn-145"><mml:math id="mml-ieqn-145"><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:math></inline-formula>)</td>
</tr>
<tr>
<td>Top-<inline-formula id="ieqn-146"><mml:math id="mml-ieqn-146"><mml:mi>k</mml:mi></mml:math></inline-formula> retrieval</td>
<td><inline-formula id="ieqn-147"><mml:math id="mml-ieqn-147"><mml:mi>k</mml:mi></mml:math></inline-formula></td>
<td>10</td>
<td><inline-formula id="ieqn-148"><mml:math id="mml-ieqn-148"><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">Z</mml:mi></mml:mrow><mml:mo>+</mml:mo></mml:msup></mml:math></inline-formula></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Validation across different threshold values shows that the system remains stable near the selected configuration. Lowering <inline-formula id="ieqn-149"><mml:math id="mml-ieqn-149"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">m</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> yields overly long representations, whereas increasing it yields fragmented outputs and reduced contextual continuity. Increasing <inline-formula id="ieqn-150"><mml:math id="mml-ieqn-150"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>r</mml:mi></mml:msub></mml:math></inline-formula> improves precision but reduces recall by admitting less relevant evidence into the retrieved context.</p>
<p>Temporal thresholds (<inline-formula id="ieqn-151"><mml:math id="mml-ieqn-151"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-152"><mml:math id="mml-ieqn-152"><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">g</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">p</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, and <inline-formula id="ieqn-153"><mml:math id="mml-ieqn-153"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi></mml:math></inline-formula>) balance alignment and coverage. The tolerance parameter <inline-formula id="ieqn-154"><mml:math id="mml-ieqn-154"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi></mml:math></inline-formula> provides flexibility when matching retrieved segments to the query&#x2019;s temporal window. Smaller <inline-formula id="ieqn-155"><mml:math id="mml-ieqn-155"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>t</mml:mi></mml:math></inline-formula> values enforce tighter alignment but may miss relevant evidence, while larger values increase recall but can include temporally distant segments. Similarly, very small <inline-formula id="ieqn-156"><mml:math id="mml-ieqn-156"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-157"><mml:math id="mml-ieqn-157"><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">g</mml:mi><mml:mi mathvariant="normal">a</mml:mi><mml:mi mathvariant="normal">p</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> values may exclude useful multimodal information, while larger values can introduce temporal drift.</p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Evaluation Metrics</title>
<p>This section provides details of the evaluation metrics used to assess the proposed framework. We assume that <inline-formula id="ieqn-158"><mml:math id="mml-ieqn-158"><mml:mi>y</mml:mi></mml:math></inline-formula> denotes the ground-truth answer and <inline-formula id="ieqn-159"><mml:math id="mml-ieqn-159"><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> denotes the predicted answer.</p>
<sec id="s5_2_1">
<label>5.2.1</label>
<title>Metrics for Open-Ended Questions</title>
<p>For Open-ended questions, we used standard text-based and semantic evaluation metrics. Token-level precision, recall, and F1 are computed as:<disp-formula id="eqn-29"><label>(29)</label><mml:math id="mml-eqn-29" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:mtext>Precision</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>y</mml:mi></mml:msub><mml:mo>&#x2229;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>,</mml:mo><mml:mspace width="2em" /><mml:mrow><mml:mtext>Recall</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>y</mml:mi></mml:msub><mml:mo>&#x2229;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>y</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-30"><label>(30)</label><mml:math id="mml-eqn-30" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mrow><mml:mtext>F1</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mi>P</mml:mi><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>R</mml:mi></mml:mrow></mml:mfrac></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-160"><mml:math id="mml-ieqn-160"><mml:msub><mml:mi>T</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-161"><mml:math id="mml-ieqn-161"><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> denote the sets of tokens in the ground-truth and generated answers, respectively.</p>
<p>The next used metric is BLEU score which measures <inline-formula id="ieqn-162"><mml:math id="mml-ieqn-162"><mml:mi>n</mml:mi></mml:math></inline-formula>-gram overlap between <inline-formula id="ieqn-163"><mml:math id="mml-ieqn-163"><mml:mi>y</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-164"><mml:math id="mml-ieqn-164"><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> and is defined as:<disp-formula id="eqn-31"><label>(31)</label><mml:math id="mml-eqn-31" display="block"><mml:mrow><mml:mtext>BLEU</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext>BP</mml:mtext></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:mi>exp</mml:mi><mml:mspace width="negativethinmathspace" /><mml:mrow><mml:mo>(</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mn>4</mml:mn></mml:munderover><mml:msub><mml:mi>w</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msub><mml:mi>w</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>4</mml:mn></mml:mfrac></mml:math></disp-formula>where BP is the brevity penalty and <inline-formula id="ieqn-165"><mml:math id="mml-ieqn-165"><mml:msub><mml:mi>p</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:math></inline-formula> denotes modified <inline-formula id="ieqn-166"><mml:math id="mml-ieqn-166"><mml:mi>n</mml:mi></mml:math></inline-formula>-gram precision.</p>
<p>The metric METEOR emphasizes recall and incorporates a penalty for fragmented matches:<disp-formula id="eqn-32"><label>(32)</label><mml:math id="mml-eqn-32" display="block"><mml:mrow><mml:mtext>METEOR</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:mtext>pen</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mspace width="2em" /><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>&#x03B1;</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>10</mml:mn><mml:mi>P</mml:mi><mml:mi>R</mml:mi></mml:mrow><mml:mrow><mml:mi>R</mml:mi><mml:mo>+</mml:mo><mml:mn>9</mml:mn><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-167"><mml:math id="mml-ieqn-167"><mml:mi>P</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-168"><mml:math id="mml-ieqn-168"><mml:mi>R</mml:mi></mml:math></inline-formula> denote precision and recall, respectively. The penalty term <inline-formula id="ieqn-169"><mml:math id="mml-ieqn-169"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mtext>pen</mml:mtext></mml:mrow></mml:msub></mml:math></inline-formula> accounts for fragmented alignments and is defined as:<disp-formula id="eqn-33"><label>(33)</label><mml:math id="mml-eqn-33" display="block"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mrow><mml:mtext>pen</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03B3;</mml:mi><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mi>c</mml:mi><mml:mi>h</mml:mi></mml:mrow><mml:mi>m</mml:mi></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>&#x03B2;</mml:mi></mml:mrow></mml:msup></mml:math></disp-formula>where <inline-formula id="ieqn-170"><mml:math id="mml-ieqn-170"><mml:mi>m</mml:mi></mml:math></inline-formula> is the number of matched unigrams, <inline-formula id="ieqn-171"><mml:math id="mml-ieqn-171"><mml:mi>c</mml:mi><mml:mi>h</mml:mi></mml:math></inline-formula> is the number of contiguous chunks in the alignment, and <inline-formula id="ieqn-172"><mml:math id="mml-ieqn-172"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-173"><mml:math id="mml-ieqn-173"><mml:mi>&#x03B2;</mml:mi></mml:math></inline-formula> are tunable parameters (<inline-formula id="ieqn-174"><mml:math id="mml-ieqn-174"><mml:mi>&#x03B3;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.5</mml:mn></mml:math></inline-formula>, <inline-formula id="ieqn-175"><mml:math id="mml-ieqn-175"><mml:mi>&#x03B2;</mml:mi><mml:mo>=</mml:mo><mml:mn>3</mml:mn></mml:math></inline-formula>).</p>
<p>Next, ROUGE-1 is used, which measures unigram recall:<disp-formula id="eqn-34"><label>(34)</label><mml:math id="mml-eqn-34" display="block"><mml:mrow><mml:mtext>ROUGE-1</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>y</mml:mi></mml:msub><mml:mo>&#x2229;</mml:mo><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>T</mml:mi><mml:mi>y</mml:mi></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-176"><mml:math id="mml-ieqn-176"><mml:msub><mml:mi>T</mml:mi><mml:mi>y</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-177"><mml:math id="mml-ieqn-177"><mml:msub><mml:mi>T</mml:mi><mml:mrow><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> denote the reference and generated unigram sets, respectively.</p>
<p>To account for paraphrasing and semantic equivalence beyond surface-form overlap, we compute cosine similarity between sentence embeddings:<disp-formula id="eqn-35"><label>(35)</label><mml:math id="mml-eqn-35" display="block"><mml:mrow><mml:mtext>Sim</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mtext mathvariant="bold">e</mml:mtext></mml:mrow><mml:mi>y</mml:mi></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">e</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">e</mml:mtext></mml:mrow><mml:mi>y</mml:mi></mml:msub><mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mrow><mml:mtext mathvariant="bold">e</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow></mml:msub><mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mfrac></mml:math></disp-formula>where embeddings <inline-formula id="ieqn-178"><mml:math id="mml-ieqn-178"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">e</mml:mtext></mml:mrow><mml:mi>y</mml:mi></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-179"><mml:math id="mml-ieqn-179"><mml:msub><mml:mrow><mml:mtext mathvariant="bold">e</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> are generated using all-MiniLM-L6-v2.</p>
</sec>
<sec id="s5_2_2">
<label>5.2.2</label>
<title>Metrics for Multiple-Choice Questions</title>
<p>Multiple-choice questions (MCQs) are treated as a <italic>classification task</italic>, since each question contains four answer options with exactly one correct choice. Accordingly, we report objective classification metrics. The accuracy measures the proportion of correctly selected options as:<disp-formula id="eqn-36"><label>(36)</label><mml:math id="mml-eqn-36" display="block"><mml:mrow><mml:mtext>Accuracy</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mrow><mml:mtext>correct</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:msub><mml:mi>N</mml:mi><mml:mrow><mml:mrow><mml:mtext>total</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mfrac></mml:math></disp-formula></p>
<p>For MCQs, precision, recall, and F1 are computed over the predicted and ground-truth option labels, providing a balanced view of classification performance, particularly under varying difficulty levels.</p>
<p>In summary, text-generation and semantic metrics (F1, BLEU, METEOR, ROUGE, and semantic similarity) are used exclusively for <italic>open-ended</italic> questions. In contrast, classification metrics (Accuracy, Precision, Recall, and F1) are used exclusively for <italic>multiple-choice</italic> questions.</p>
</sec>
</sec>
<sec id="s5_3">
<label>5.3</label>
<title>Experimental Results</title>
<p>This section presents a quantitative evaluation of the proposed framework on both open-ended and multiple-choice (MCQ) questions. We compare the proposed Temporal-Aware RAG framework with LLaVA-1.6, a strong open-source multimodal foundation model. This comparison assesses whether a large vision-language model, when provided with both visual and textual context, can implicitly perform the temporal localization and evidence grounding required for lecture-style VideoQA without explicit retrieval or alignment mechanisms.</p>
<p>LLaVA-1.6 combines a CLIP-based visual encoder with a large language model trained on a broad multimodal instruction-following corpus. The model has demonstrated strong performance on general image reasoning, visual question answering, and OCR-related tasks. As such, it serves as a representative and competitive baseline for evaluating whether general-purpose multimodal reasoning alone is sufficient for answering questions grounded in long-form instructional videos. However, LLaVA-1.6 does not incorporate explicit mechanisms for temporal modeling, retrieval, or evidence re-ranking across extended video content.</p>
<p>In our evaluation setting, LLaVA-1.6 is provided with 4&#x2013;6 uniformly sampled keyframes from each lecture video together with the full Whisper ASR transcript. The model must answer questions using only this information. The following prompt is used:</p>
<p><bold>Prompt:</bold> You are given (i) a set of keyframes extracted from a lecture video and (ii) the full ASR transcript of the same video. Answer the question using <italic>only</italic> the provided keyframes and transcript. Do not use any external knowledge. If the evidence is insufficient, respond with &#x201C;Insufficient evidence from the provided transcript/keyframes.&#x201D;</p>
<p><bold>Question:</bold> {QUESTION}</p>
<p><bold>Transcript:</bold> {WHISPER_TRANSCRIPT}</p>
<p><bold>Keyframes:</bold> (images provided above)</p>
<p><bold>Output format:</bold> Provide a short answer only. For MCQs, output exactly one letter from {A,B,C,D}.</p>
<p>For multiple-choice questions, the four answer options (A&#x2013;D) are appended to the question text. For open-ended questions, the model is instructed to generate concise answers grounded strictly in the provided transcript and keyframes.</p>
<sec id="s5_3_1">
<label>5.3.1</label>
<title>Open-Ended Question Performance</title>
<p><xref ref-type="table" rid="table-4">Table 4</xref> shows the results of open-ended QA for LLaVA-1.6 and the proposed framework. It can be seen that the performance of both systems decreases as the difficulty level of the questions increases. However, the proposed framework consistently outperforms LLaVA-1.6 on all metrics and difficulty levels. The proposed Temporal-Aware RAG improves F1 from 15.30% to 23.52% and ROUGE-1 from 21.70% to 29.76%, indicating more faithful overlap with the reference answers and better coverage of the key facts required by the questions.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Open-ended QA performance comparison (in%).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Model</th>
<th>Difficulty</th>
<th>F1</th>
<th>Sim</th>
<th>BLEU</th>
<th>METEOR</th>
<th>R1</th>
</tr>
</thead>
<tbody>
<tr>
<td>LLaVA-1.6 (Keyframes &#x002B; Transcript)</td>
<td>Simple</td>
<td>21.29</td>
<td>70.34</td>
<td>5.18</td>
<td>30.73</td>
<td>28.41</td>
</tr>
<tr>
<td></td>
<td>Hard</td>
<td>15.62</td>
<td>66.11</td>
<td>3.49</td>
<td>24.28</td>
<td>21.95</td>
</tr>
<tr>
<td></td>
<td>Very Hard</td>
<td>10.47</td>
<td>62.85</td>
<td>2.67</td>
<td>18.39</td>
<td>16.12</td>
</tr>
<tr>
<td></td>
<td>Overall</td>
<td>15.30</td>
<td>66.88</td>
<td>3.30</td>
<td>24.57</td>
<td>21.70</td>
</tr>
<tr>
<td>Temporal-Aware RAG (Ours)</td>
<td>Simple</td>
<td>29.47</td>
<td>77.23</td>
<td>8.61</td>
<td>38.15</td>
<td>36.82</td>
</tr>
<tr>
<td></td>
<td>Hard</td>
<td>24.38</td>
<td>74.56</td>
<td>5.29</td>
<td>34.71</td>
<td>30.94</td>
</tr>
<tr>
<td></td>
<td>Very Hard</td>
<td>16.72</td>
<td>71.48</td>
<td>2.83</td>
<td>24.36</td>
<td>22.51</td>
</tr>
<tr>
<td></td>
<td>Overall</td>
<td>23.52</td>
<td>74.42</td>
<td>5.58</td>
<td>32.41</td>
<td>29.76</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>LLaVA-1.6 exhibits moderate semantic similarity and a relatively low F1 and ROUGE-1. This suggests that LLaVA-1.6 can often infer the question&#x2019;s general topic and produce broadly plausible answers. However, it frequently fails to respond to the specific lecture segment where the correct information is stated. This behavior is consistent with the absence of an explicit evidence-selection mechanism. The model implicitly searches over long transcripts and sparsely sampled visual context, which increases the likelihood of generating generic, underspecified, or partially correct answers.</p>
<p>In contrast, the proposed framework demonstrates higher similarity and significantly improved lexical and coverage-based scores. These results suggest that temporally filtered retrieval and re-ranking not only identify conceptually relevant evidence but also enable the generator to produce answers that more closely align with the reference content.</p>
</sec>
<sec id="s5_3_2">
<label>5.3.2</label>
<title>Multiple-Choice Question Performance</title>
<p><xref ref-type="table" rid="table-5">Table 5</xref> reports MCQ performance. Since each MCQ contains four answer options with exactly one correct choice, we treat MCQs as a classification task and report Accuracy (ACC), Precision, Recall, and F1-score.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>MCQ performance comparison (in%).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Model</th>
<th>Difficulty</th>
<th>ACC</th>
<th>Precision</th>
<th>Recall</th>
<th>F1</th>
</tr>
</thead>
<tbody>
<tr>
<td>LLaVA-1.6 (Keyframes &#x002B; Transcript)</td>
<td>Simple</td>
<td>45.37</td>
<td>46.58</td>
<td>46.21</td>
<td>46.84</td>
</tr>
<tr>
<td></td>
<td>Hard</td>
<td>40.62</td>
<td>41.13</td>
<td>41.79</td>
<td>41.45</td>
</tr>
<tr>
<td></td>
<td>Very Hard</td>
<td>35.48</td>
<td>36.92</td>
<td>36.27</td>
<td>36.64</td>
</tr>
<tr>
<td></td>
<td>Overall</td>
<td>40.00</td>
<td>41.56</td>
<td>41.34</td>
<td>41.89</td>
</tr>
<tr>
<td>Temporal-Aware RAG (Ours)</td>
<td>Simple</td>
<td>57.29</td>
<td>62.15</td>
<td>62.08</td>
<td>62.11</td>
</tr>
<tr>
<td></td>
<td>Hard</td>
<td>52.34</td>
<td>53.18</td>
<td>53.25</td>
<td>53.21</td>
</tr>
<tr>
<td></td>
<td>Very Hard</td>
<td>50.67</td>
<td>51.42</td>
<td>51.39</td>
<td>51.40</td>
</tr>
<tr>
<td></td>
<td>Overall</td>
<td>53.43</td>
<td>55.58</td>
<td>55.57</td>
<td>55.57</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>For MCQs, the proposed Temporal-Aware RAG framework achieves consistently higher performance. The accuracy, precision, recall, and F1 values were higher for our method than for LLaVA-1.6. The performance of LLaVA-1.6 decreases steadily from &#x201C;Simple&#x201D; to &#x201C;Very Hard&#x201D; questions. This shows that it struggles when questions require careful distinction between similar options or when information must be combined from different parts of the lecture. This happens because LLaVA-1.6 uses only a few keyframes and a full, unstructured transcript. As a result, it is easily affected by irrelevant information and by details that appear far apart in time.</p>
<p>On the other hand, the proposed framework shows a gradual decline as difficulty rises, but it still performs better than LLaVA-1.6. These results show that explicit retrieval, temporal filtering, and re-ranking make systems more robust, even when questions require detailed reasoning or distributed evidence.</p>
<p>Overall, the results across both open-ended and MCQ settings demonstrate that explicit, timestamp-aware evidence selection is essential for lecture-style VideoQA. While LLaVA-1.6 provides a strong general-purpose multimodal baseline, the proposed Temporal-Aware RAG framework yields more accurate and reliable answers by combining multimodal cues with structured retrieval, temporal constraints, and re-ranking.</p>
</sec>
<sec id="s5_3_3">
<label>5.3.3</label>
<title>Comparison with Standard RAG Baselines</title>
<p>We next compare the proposed Temporal-Aware RAG framework against standard retrieval-based baselines that progressively introduce multimodal evidence while explicitly excluding temporal reasoning. This controlled comparison isolates the contribution of temporal grounding beyond retrieval and multimodality alone.</p>
<p><bold>Text-only RAG (no temporal awareness):</bold> A FAISS index is constructed over semantically chunked Whisper ASR transcript segments, and the top-<inline-formula id="ieqn-180"><mml:math id="mml-ieqn-180"><mml:mi>k</mml:mi></mml:math></inline-formula> chunks are retrieved using embedding similarity. The generator produces answers using only the retrieved transcript evidence, without visual captions or timestamp constraints.</p>
<p><bold>Multimodal RAG (no timestamps):</bold> ASR transcript chunks and Gemini-generated visual caption chunks are jointly indexed in a single retrieval space. Top-<inline-formula id="ieqn-181"><mml:math id="mml-ieqn-181"><mml:mi>k</mml:mi></mml:math></inline-formula> chunks are retrieved by similarity without enforcing any timestamp constraints, allowing retrieved evidence to be semantically relevant but temporally misaligned.</p>
<p><xref ref-type="table" rid="table-6">Table 6</xref> reports results for open-ended question answering. The results show that retrieval based on ASR transcripts alone is not robust as question difficulty increases. While &#x201C;Simple&#x201D; questions achieve 24.37% F1, performance drops sharply for &#x201C;Hard&#x201D; (9.01%) and &#x201C;Very Hard&#x201D; (3.89%) questions, with a corresponding decline in semantic similarity. This indicates that similarity-based retrieval over transcripts struggles to retrieve precise, question-specific evidence for complex queries.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Standard RAG baselines for open-ended QA without temporal awareness (in%).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Difficulty</th>
<th>F1</th>
<th>Semantic Similarity</th>
<th>BLEU</th>
<th>METEOR</th>
</tr>
</thead>
<tbody>
<tr>
<td align="center" colspan="5"><bold>Text-only RAG (ASR only)</bold></td>
</tr>
<tr>
<td>Simple</td>
<td>24.37</td>
<td>60.06</td>
<td>8.05</td>
<td>41.05</td>
</tr>
<tr>
<td>Hard</td>
<td>9.01</td>
<td>50.59</td>
<td>1.23</td>
<td>18.18</td>
</tr>
<tr>
<td>Very Hard</td>
<td>3.89</td>
<td>34.06</td>
<td>1.58</td>
<td>10.53</td>
</tr>
<tr>
<td>Overall</td>
<td>12.42</td>
<td>48.24</td>
<td>3.62</td>
<td>23.25</td>
</tr>
<tr>
<td align="center" colspan="5"><bold>Multimodal RAG (ASR &#x002B; captions, no timestamps)</bold></td>
</tr>
<tr>
<td>Simple</td>
<td>27.90</td>
<td>71.00</td>
<td>6.12</td>
<td>38.89</td>
</tr>
<tr>
<td>Hard</td>
<td>11.30</td>
<td>62.00</td>
<td>1.21</td>
<td>16.34</td>
</tr>
<tr>
<td>Very Hard</td>
<td>5.30</td>
<td>52.00</td>
<td>1.51</td>
<td>9.73</td>
</tr>
<tr>
<td>Overall</td>
<td>14.80</td>
<td>61.00</td>
<td>2.95</td>
<td>21.65</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Incorporating visual captions substantially improves semantic alignment across all difficulty levels, increasing overall similarity from 48.24% to 61.00%. This highlights the complementary value of slide content and on-screen information. However, without timestamp constraints, multimodal RAG remains vulnerable to temporal drift and can retrieve conceptually related but contextually incorrect segments.</p>
</sec>
<sec id="s5_3_4">
<label>5.3.4</label>
<title>Ablation Study</title>
<p>To assess the contribution of each individual component in the proposed framework, an ablation study was conducted on open-ended questions. The ablation study evaluates four model variants, each created by removing one key component: visual captions, temporal filtering, the cross-encoder re-ranker, or normalization.</p>
<p><xref ref-type="table" rid="table-7">Table 7</xref> presents the ablation results. It can be seen that removing the temporal filter from the framework had the greatest impact on performance, as the overall F1 score drops to 11.40 and the semantic similarity to 36.67. This performance decline is especially more evident for &#x201C;Very Hard&#x201D; questions. The decline in performance is due to the retriever becoming less effective at localizing relevant segments without temporal constraints.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Ablation study for open-ended QA (in%). Each block removes one component from the full temporal-aware RAG pipeline.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Difficulty</th>
<th>F1</th>
<th>Semantic Similarity</th>
<th>BLEU</th>
<th>METEOR</th>
</tr>
</thead>
<tbody>
<tr>
<td align="center" colspan="5"><bold>No Visual Captions</bold></td>
</tr>
<tr>
<td>Simple</td>
<td>26.23</td>
<td>63.56</td>
<td>7.50</td>
<td>35.62</td>
</tr>
<tr>
<td>Hard</td>
<td>14.50</td>
<td>55.21</td>
<td>3.50</td>
<td>24.12</td>
</tr>
<tr>
<td>Very Hard</td>
<td>7.80</td>
<td>42.13</td>
<td>1.20</td>
<td>12.32</td>
</tr>
<tr>
<td>Overall</td>
<td>16.40</td>
<td>55.80</td>
<td>4.10</td>
<td>24.00</td>
</tr>
<tr>
<td align="center" colspan="5"><bold>No Temporal Filter</bold></td>
</tr>
<tr>
<td>Simple</td>
<td>21.80</td>
<td>44.23</td>
<td>9.05</td>
<td>31.05</td>
</tr>
<tr>
<td>Hard</td>
<td>8.70</td>
<td>37.57</td>
<td>2.23</td>
<td>17.18</td>
</tr>
<tr>
<td>Very Hard</td>
<td>3.80</td>
<td>28.23</td>
<td>2.58</td>
<td>9.53</td>
</tr>
<tr>
<td>Overall</td>
<td>11.40</td>
<td>36.67</td>
<td>4.52</td>
<td>22.85</td>
</tr>
<tr>
<td align="center" colspan="5"><bold>No Cross-Encoder Re-Rank</bold></td>
</tr>
<tr>
<td>Simple</td>
<td>31.15</td>
<td>60.31</td>
<td>8.19</td>
<td>33.57</td>
</tr>
<tr>
<td>Hard</td>
<td>24.07</td>
<td>48.72</td>
<td>4.79</td>
<td>26.19</td>
</tr>
<tr>
<td>Very Hard</td>
<td>11.30</td>
<td>42.48</td>
<td>0.56</td>
<td>7.31</td>
</tr>
<tr>
<td>Overall</td>
<td>22.21</td>
<td>52.58</td>
<td>4.53</td>
<td>26.20</td>
</tr>
<tr>
<td align="center" colspan="5"><bold>No Normalization</bold></td>
</tr>
<tr>
<td>Simple</td>
<td>30.15</td>
<td>58.31</td>
<td>8.19</td>
<td>33.57</td>
</tr>
<tr>
<td>Hard</td>
<td>23.07</td>
<td>47.72</td>
<td>4.79</td>
<td>26.19</td>
</tr>
<tr>
<td>Very Hard</td>
<td>10.30</td>
<td>28.48</td>
<td>0.56</td>
<td>7.31</td>
</tr>
<tr>
<td>Overall</td>
<td>20.21</td>
<td>44.83</td>
<td>4.53</td>
<td>26.20</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The ablation study after removing the cross-encoder re-ranker shows that the performance in the simple category remains relatively stable (F1 &#x003D; 31.15). However, for the &#x201C;Hard&#x201D; and &#x201C;Very Hard&#x201D; questions, a relatively sharp decline is observed. This indicates that, without re-ranking, the framework struggles with hard questions due to its inability to prioritize the most relevant evidence.</p>
<p>When visual captions were removed, a moderate decline in performance was observed as the overall F1 score decreased to 16.40, and semantic similarity fell to 55.80. Again, the decline is more evident in the more difficult categories. This indicates that visual captions are an important clue to answer complex questions.</p>
<p>Finally, the ablation study after removing the normalization component also resulted in a moderate performance decrease. The &#x201C;Very Hard&#x201D; category shows a more substantial decrease in semantic similarity (to 28.48). These results suggest that normalization supports consistency in more challenging scenarios.</p>
</sec>
</sec>
<sec id="s5_4">
<label>5.4</label>
<title>Qualitative Analysis of Success and Failure Cases</title>
<p>This section shows examples of both failure and success from the LectQA-Vid dataset. <xref ref-type="fig" rid="fig-3">Fig. 3</xref> shows a case of failure with a multiple choice question about when the narrator talks about off-page SEO. The ground truth shows a section where the speaker discusses topics such as backlinks, guest blogging, and influencer marketing. The model, on the other hand, retrieves content from an earlier section on on-page SEO, indicating that the system chooses an option that is relevant to the topic but not to the time.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>An example MCQ where the proposed method failed.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81534-fig-3.tif"/>
</fig>
<p><xref ref-type="fig" rid="fig-4">Fig. 4</xref> shows an example of correct detection. The question asks about the hierarchical structure of Internet Service Providers and their roles in connectivity. The corresponding video segment contains clear slide diagrams and explicit verbal references to each tier, which resulted in highly distinctive lexical cues (e.g., &#x201C;Tier 1 backbone&#x201D;, &#x201C;Tier 2 regional providers&#x201D;, &#x201C;Tier 3 access networks&#x201D;). In this case, both embedding-based retrieval and cross-encoder re-ranking were able to discriminate the relevant chunk from the rest of the lecture.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>An example MCQ where the proposed method gave the correct result.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81534-fig-4.tif"/>
</fig>
<p>The example in <xref ref-type="fig" rid="fig-5">Fig. 5</xref> shows a failure case of an open-ended question. The question asks which algorithms are discussed in the lecture, with the correct answer being linear and logistic regression, decision trees, and convolutional neural networks. Instead, the model outputs an answer related to Prim&#x2019;s algorithm. This lecture introduces different algorithmic examples in different segments. This error was because of the incomplete coverage of the slide text by visual captioning. The retrieval step picks up only a limited number of top-ranked chunks, and thus those chunks may emphasize &#x201C;algorithmic thinking&#x201D; in general and thus the LLM generates an answer that is thematically compatible yet factually incorrect.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>An example open-ended question where the proposed method failed.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81534-fig-5.tif"/>
</fig>
<p>Lastly, <xref ref-type="fig" rid="fig-6">Fig. 6</xref> shows an open-ended question which was answered correctly. The question asks when the instructor highlights a particular selling point of the program. The correct answer requires identifying a segment where small class sizes and close student&#x2013;faculty relationships are discussed. In this case, the ASR transcript contains unambiguous phrases (e.g., &#x201C;small class sizes&#x201D;, &#x201C;personalized attention&#x201D;) that appear within a narrow time window, and the corresponding frames show slides that emphasize the same message. Temporal-aware retrieval, therefore, focuses on the initial seconds of the lecture, and the cross-encoder re-ranker promotes the segment that contains the correct emphasis.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>An example open-ended question where the proposed method succeeded.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81534-fig-6.tif"/>
</fig>
<p>These visual examples support our quantitative results. The model works well when key concepts are clear, occur within a short time window, and are consistently captured in both ASR and visual captions. In such cases, the temporal-aware retriever and re-ranker can select a small set of highly relevant chunks. Failures usually occur due to temporal drift between similar segments when the question requires information from multiple distant parts of the video.</p>
</sec>
<sec id="s5_5">
<label>5.5</label>
<title>Efficiency Analysis</title>
<p>Next, we evaluated the computational efficiency of the proposed framework. We measure index construction time, retrieval latency, temporal filtering cost, re-ranking time, and answer generation time under varying video lengths and retrieval scales.</p>
<p>The results in <xref ref-type="table" rid="table-8">Table 8</xref> show that retrieval latency increases gradually with the number of indexed chunks. Temporal filtering incurs negligible overhead, and the cost of re-ranking increases with larger top-<inline-formula id="ieqn-182"><mml:math id="mml-ieqn-182"><mml:mi>k</mml:mi></mml:math></inline-formula> values because more candidate segments are evaluated. Across all settings, answer generation is the dominant component of the total response time. Importantly, total query time remains stable (around 2 s) even as video length and retrieval scale increase.</p>
<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Efficiency analysis under varying video lengths and retrieval scales.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Setting</th>
<th>Video Length</th>
<th>Chunks</th>
<th>Index (s)</th>
<th>Retr. (ms)</th>
<th>Temp. (ms)</th>
<th>Re-Rank (ms)</th>
<th>Gen. (s)</th>
<th>Total (s)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Short (small)</td>
<td>2 min</td>
<td>36</td>
<td>0.3</td>
<td>11</td>
<td>2</td>
<td>34</td>
<td>1.86</td>
<td>1.91</td>
</tr>
<tr>
<td>Medium (medium)</td>
<td>3 min</td>
<td>54</td>
<td>0.5</td>
<td>14</td>
<td>3</td>
<td>47</td>
<td>2.01</td>
<td>2.07</td>
</tr>
<tr>
<td>Long (large)</td>
<td>4 min</td>
<td>72</td>
<td>0.7</td>
<td>17</td>
<td>4</td>
<td>63</td>
<td>2.18</td>
<td>2.26</td>
</tr>
<tr>
<td>Medium, <inline-formula id="ieqn-183"><mml:math id="mml-ieqn-183"><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>20</mml:mn></mml:math></inline-formula></td>
<td>3 min</td>
<td>54</td>
<td>0.5</td>
<td>20</td>
<td>4</td>
<td>86</td>
<td>2.09</td>
<td>2.20</td>
</tr>
<tr>
<td>Medium, <inline-formula id="ieqn-184"><mml:math id="mml-ieqn-184"><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>30</mml:mn></mml:math></inline-formula></td>
<td>3 min</td>
<td>54</td>
<td>0.5</td>
<td>26</td>
<td>5</td>
<td>121</td>
<td>2.19</td>
<td>2.34</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5_6">
<label>5.6</label>
<title>Discussions</title>
<p>Experimental results demonstrate that a range of answers related to lecture videos can be answered by the proposed framework. Lecture videos are characterized by a high degree of structure, conceptual density, and sequential organization. Therefore, accurate answers frequently require grounding evidence in the specific temporal segment in which a concept is introduced or further developed. The proposed framework uses timestamp-based filtering to reduce temporal drift by restricting retrieval to the relevant time window, thereby enhancing semantic faithfulness.</p>
<p>The experiments also show that lecture VideoQA is not purely a speech-understanding problem. The inclusion of Gemini-generated visual captions improves semantic alignment, indicating that on-screen slide content provides complementary evidence that is usually essential for reliable question answering. The ablation study further confirms this dependency: removing visual captions reduces robustness as question difficulty increases.</p>
<p>There are, however, certain limitations of the proposed method:<list list-type="bullet">
<list-item>
<p>The evaluation is conducted in a micro-lecture setting in which the videos are 2&#x2013;5 min long. Thus, the proposed method has not been tested on long lecture videos where long-range dependencies and increased topic drift are likely to make the task of video question answering more challenging.</p></list-item>
<list-item>
<p>The proposed pipeline relies on Whisper ASR transcripts and Gemini-generated visual captions. Any errors in transcription and captioning (e.g., due to pronunciation, domain-specific terminology, or noise) can propagate to retrieval and temporal grounding.</p></list-item>
<list-item>
<p>The proposed retrieval design uses fixed thresholds for granularity. The evidence in videos is extracted using fixed-size semantic chunks and by using a single-shot top-<inline-formula id="ieqn-185"><mml:math id="mml-ieqn-185"><mml:mi>k</mml:mi></mml:math></inline-formula> strategy. Therefore, the retrieval phase may miss fine-grained evidence in the videos or fail to assemble the distributed evidence needed for multi-hop instructional questions.</p></list-item>
<list-item>
<p>The proposed framework was designed to use only intra-video evidence. However, the framework lacks an explicit mechanism to detect hallucinations or verify factual consistency.</p></list-item>
<list-item>
<p>The proposed framework deliberately focuses on intra-video reasoning only and does not incorporate external knowledge, so the answers to the questions do not include any additional explanations which may be helpful to some students.</p></list-item>
<list-item>
<p>Although the proposed method has been compared against a strong multimodal model and controlled RAG baselines, we do not include recent VideoQA models designed for long-video understanding. Many of these approaches operate end-to-end, without explicit retrieval or timestamp-constrained evidence selection. In contrast, our work focuses on intra-video, retrieval-based reasoning with strict temporal grounding. Due to this difference in task formulation, direct comparison is not fully aligned. Extending evaluation to recent long-video VideoQA is an important direction for future work.</p></list-item>
</list></p>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusions</title>
<p>This work introduced a Temporal-Aware, Intra-video Retrieval-Augmented Generation (RAG) framework for Video Question Answering (VideoQA) over short lecture videos. Unlike conventional LLM-based or non-temporal RAG approaches, the proposed system retrieves multimodal, timestamped evidence from the target video alone before generating an answer. By integrating Whisper ASR, Gemini visual captioning, semantic chunking, BGE embeddings, FAISS retrieval, cross-encoder re-ranking, and temporal filtering, the framework ensures that all predictions are explicitly grounded in verifiable video segments. The experiments on the LectVidQA dataset show that the proposed framework outperforms text-only RAG and multimodal non-temporal baselines. The extensive ablation studies also confirm the importance of multimodal cues and temporal constraints. The framework is simple and efficient as it does not rely on external knowledge and works well across different lecture topics.</p>
<p>In future work, we aim to improve the alignment between video frames and text by developing stronger temporal models and expanding the system to support longer instructional videos. We also plan to add uncertainty estimation to make the system more explainable. Currently, audio is converted to text using ASR, which makes it easier to combine different types of data but omits important acoustic details such as pauses and emphasis. A valuable next step would be to use raw audio features along with transcripts and visual data.</p>
</sec>
</body>
<back>
<ack>
<p>Not Applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This work was supported by the Deanship of Scientific Research, Vice Presidency for Graduate Studies and Scientific Research, King Faisal University, Saudi Arabia, under Grant KFU262069.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: conceptualization, Sumaira Shafiq and Munam Ali Shah; methodology, Sumaira Shafiq; software, Sumaira Shafiq; validation, Sumaira Shafiq, Naveed Ejaz and Munam Ali Shah; formal analysis, Sumaira Shafiq; investigation, Sumaira Shafiq; resources, Naveed Ejaz, Rashid Kamal and Sheraz Aslam; data curation, Sumaira Shafiq and Adnan Sohail; writing&#x2014;original draft preparation, Sumaira Shafiq; writing&#x2014;review and editing, Naveed Ejaz, Munam Ali Shah, Rashid Kamal, Adnan Sohail and Sheraz Aslam; visualization, Sumaira Shafiq; supervision, Munam Ali Shah; project administration, Munam Ali Shah; funding acquisition, Munam Ali Shah. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The dataset developed in this study is publicly available at <ext-link ext-link-type="uri" xlink:href="https://data.mendeley.com/datasets/yt4nmz9mcv/1">https://data.mendeley.com/datasets/yt4nmz9mcv/1</ext-link>.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<glossary content-type="abbreviations" id="glossary-1">
<title>Abbreviations </title>
<def-list>
<def-item>
<term>The following abbreviations are used in this manuscript:</term> 
</def-item>
<def-item>
<term>ASR</term>
<def>
<p>Automatic Speech Recognition</p>
</def>
</def-item>
<def-item>
<term>CMC</term>
<def>
<p>Computers, Materials &#x0026; Continua</p>
</def>
</def-item>
<def-item>
<term>FAISS</term>
<def>
<p>Facebook AI Similarity Search</p>
</def>
</def-item>
<def-item>
<term>GPT</term>
<def>
<p>Generative Pre-trained Transformer</p>
</def>
</def-item>
<def-item>
<term>LLM</term>
<def>
<p>Large Language Model</p>
</def>
</def-item>
<def-item>
<term>MCQ</term>
<def>
<p>Multiple-Choice Question</p>
</def>
</def-item>
<def-item>
<term>QA</term>
<def>
<p>Question Answering</p>
</def>
</def-item>
<def-item>
<term>RAG</term>
<def>
<p>Retrieval-Augmented Generation</p>
</def>
</def-item>
<def-item>
<term>VLM</term>
<def>
<p>Vision-Language Model</p>
</def>
</def-item>
<def-item>
<term>VideoQA</term>
<def>
<p>Video Question Answering</p>
</def>
</def-item>
</def-list>
</glossary>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jeshmol</surname> <given-names>PJ</given-names></string-name>, <string-name><surname>Kovoor</surname> <given-names>BC</given-names></string-name></person-group>. <article-title>Video question answering: a survey of the state-of-the-art</article-title>. <source>J Vis Commun Image Represent</source>. <year>2024</year>;<volume>105</volume>(<issue>3</issue>):<fpage>104320</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jvcir.2024.104320</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ishmam</surname> <given-names>MF</given-names></string-name>, <string-name><surname>Shovon</surname> <given-names>MSH</given-names></string-name>, <string-name><surname>Mridha</surname> <given-names>MF</given-names></string-name>, <string-name><surname>Dey</surname> <given-names>N</given-names></string-name></person-group>. <article-title>From image to language: a critical analysis of visual question answering (VQA) approaches, challenges, and opportunities</article-title>. <source>Inf Fusion</source>. <year>2024</year>;<volume>106</volume>(<issue>6</issue>):<fpage>102270</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.inffus.2024.102270</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Islam</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Nagarajan</surname> <given-names>T</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Bertasius</surname> <given-names>G</given-names></string-name>, <string-name><surname>Torresani</surname> <given-names>L</given-names></string-name></person-group>. <article-title>BIMBA: selective-scan compression for long-range video question answering</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2025 Jun 10&#x2013;17</conf-name>; <publisher-loc>Nashville, TN, USA</publisher-loc>. p. <fpage>29096</fpage>&#x2013;<lpage>107</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR52734.2025.02709</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Alemdag</surname> <given-names>E</given-names></string-name></person-group>. <article-title>A scoping review of the literature on embodied instructional videos</article-title>. <source>Res Pract Technol Enhanc Learn</source>. <year>2023</year>;<volume>18</volume>:29. doi:<pub-id pub-id-type="doi">10.58459/rptel.2023.18029</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Oche</surname> <given-names>AJ</given-names></string-name>, <string-name><surname>Folashade</surname> <given-names>AG</given-names></string-name>, <string-name><surname>Ghosal</surname> <given-names>T</given-names></string-name>, <string-name><surname>Biswas</surname> <given-names>A</given-names></string-name></person-group>. <article-title>A systematic review of key retrieval-augmented generation (RAG) systems: progress, gaps, and future directions</article-title>. <comment>arXiv:2507.18910. 2025</comment>. doi:<pub-id pub-id-type="doi">10.48550/arxiv.2507.18910</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Han</surname> <given-names>B</given-names></string-name>, <string-name><surname>Susnjak</surname> <given-names>T</given-names></string-name>, <string-name><surname>Mathrani</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Automating systematic literature reviews with retrieval-augmented generation: a comprehensive overview</article-title>. <source>Appl Sci</source>. <year>2024</year>;<volume>14</volume>(<issue>19</issue>):<fpage>9103</fpage>. doi:<pub-id pub-id-type="doi">10.3390/app14199103</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Raja</surname> <given-names>R</given-names></string-name>, <string-name><surname>Vats</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Multimedia-aware question answering: a review of retrieval and cross-modal reasoning architectures</article-title>. In: <conf-name>Proceedings of the 2nd ACM Workshop on AI-Powered Question &#x0026; Answering Systems</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>The Association for Computing Machinery (ACM)</publisher-name>; <year>2025</year>. p. <fpage>28</fpage>&#x2013;<lpage>35</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3746274.3760393</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sun</surname> <given-names>G</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Li</surname> <given-names>T</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Video question answering: a survey of models and datasets</article-title>. <source>Mob Netw Appl</source>. <year>2021</year>;<volume>26</volume>(<issue>5</issue>):<fpage>1904</fpage>&#x2013;<lpage>37</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11036-020-01730-0</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zhong</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ji</surname> <given-names>W</given-names></string-name>, <string-name><surname>Xiao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>W</given-names></string-name>, <string-name><surname>Chua</surname> <given-names>TS</given-names></string-name></person-group>. <article-title>Video question answering: datasets, algorithms and challenges</article-title>. In: <conf-name>Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing</conf-name>. <publisher-loc>Abu Dhabi, United Arab Emirates</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>; <year>2022</year>. p. <fpage>6439</fpage>&#x2013;<lpage>55</lpage>. doi:<pub-id pub-id-type="doi">10.18653/v1/2022.emnlp-main.432</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Gao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ge</surname> <given-names>R</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>K</given-names></string-name>, <string-name><surname>Nevatia</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Motion-appearance co-memory networks for video question answering</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2018 Jun 18&#x2013;23</conf-name>; <publisher-loc>Salt Lake City, UT, USA</publisher-loc>. p. <fpage>6576</fpage>&#x2013;<lpage>85</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR.2018.00688</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zeng</surname> <given-names>P</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>L</given-names></string-name>, <string-name><surname>Song</surname> <given-names>J</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>HT</given-names></string-name></person-group>. <article-title>Video question answering with prior knowledge and object-sensitive learning</article-title>. <source>IEEE Trans Image Process</source>. <year>2022</year>;<volume>31</volume>:<fpage>5936</fpage>&#x2013;<lpage>48</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TIP.2022.3205212</pub-id>; <pub-id pub-id-type="pmid">36083958</pub-id></mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Park</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>J</given-names></string-name>, <string-name><surname>Sohn</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Bridge to answer: structure-aware graph interaction network for video question answering</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2021 Jun 20&#x2013;25</conf-name>; <publisher-loc>Nashville, TN, USA</publisher-loc>. p. <fpage>15521</fpage>&#x2013;<lpage>30</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR46437.2021.01527</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>F</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>H</given-names></string-name></person-group>. <article-title>HAIR: hierarchical visual-semantic relational reasoning for video question answering</article-title>. In: <conf-name>Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV); 2021 Oct 10&#x2013;17</conf-name>; <publisher-loc>Montreal, QC, Canada</publisher-loc>. p. <fpage>1678</fpage>&#x2013;<lpage>87</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ICCV48922.2021.00172</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cherian</surname> <given-names>A</given-names></string-name>, <string-name><surname>Hori</surname> <given-names>C</given-names></string-name>, <string-name><surname>Marks</surname> <given-names>TK</given-names></string-name>, <string-name><surname>Le Roux</surname> <given-names>J</given-names></string-name></person-group>. <article-title>(2.5 &#x002B; 1)D spatio-temporal scene graphs for video question answering</article-title>. <source>Proc AAAI Conf Artif Intell</source>. <year>2022</year>;<volume>36</volume>(<issue>1</issue>):<fpage>444</fpage>&#x2013;<lpage>53</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v36i1.19922</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lei</surname> <given-names>J</given-names></string-name>, <string-name><surname>Berg</surname> <given-names>T</given-names></string-name>, <string-name><surname>Bansal</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Revealing single frame bias for video-and-language learning</article-title>. In: <conf-name>Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL)</conf-name>. <publisher-loc>Toronto, ON, Canada</publisher-loc>: <publisher-name>Association for Computational Linguistics</publisher-name>; <year>2023</year>. p. <fpage>487</fpage>&#x2013;<lpage>507</lpage>. doi:<pub-id pub-id-type="doi">10.18653/v1/2023.acl-long.29</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Li</surname> <given-names>W</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>G</given-names></string-name></person-group>. <article-title>SHMamba: structured hyperbolic state space model for audio-visual question answering</article-title>. <source>IEEE Trans Audio Speech Lang Process</source>. <year>2025</year>;<volume>33</volume>:<fpage>3582</fpage>&#x2013;<lpage>93</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TASLPRO.2025.3597461</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zong</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wan</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Video-context aligned transformer for video question answering</article-title>. <source>Proc AAAI Conf Artif Intell</source>. <year>2024</year>;<volume>38</volume>(<issue>17</issue>):<fpage>19795</fpage>&#x2013;<lpage>803</lpage>. doi:<pub-id pub-id-type="doi">10.1609/aaai.v38i17.29954</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Garcia</surname> <given-names>N</given-names></string-name>, <string-name><surname>Chu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Otani</surname> <given-names>M</given-names></string-name>, <string-name><surname>Nakashima</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Takemura</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Bert representations for video question answering</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision; 2020 Mar 1&#x2013;5</conf-name>; <publisher-loc>Snowmass Village, CO, USA</publisher-loc>. p. <fpage>1556</fpage>&#x2013;<lpage>65</lpage>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Garcia</surname> <given-names>N</given-names></string-name>, <string-name><surname>Chu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Otani</surname> <given-names>M</given-names></string-name>, <string-name><surname>Nakashima</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Takemura</surname> <given-names>H</given-names></string-name></person-group>. <article-title>A comparative study of language transformers for video question answering</article-title>. <source>Neurocomputing</source>. <year>2021</year>;<volume>445</volume>:<fpage>121</fpage>&#x2013;<lpage>33</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.neucom.2021.02.092</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Garcia</surname> <given-names>N</given-names></string-name>, <string-name><surname>Otani</surname> <given-names>M</given-names></string-name>, <string-name><surname>Chu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Nakashima</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>KnowIT VQA: answering knowledge-based questions about videos</article-title>. <source>Proc AAAI Conf Artif Intell</source>. <year>2020</year>;<volume>34</volume>:<fpage>10826</fpage>&#x2013;<lpage>34</lpage>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Garcia</surname> <given-names>N</given-names></string-name>, <string-name><surname>Otani</surname> <given-names>M</given-names></string-name>, <string-name><surname>Chu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Nakashima</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Takemura</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Transferring domain-agnostic knowledge in video question answering</article-title>. <comment>arXiv:211013395. 2021</comment>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Nie</surname> <given-names>L</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>B</given-names></string-name>, <string-name><surname>Chua</surname> <given-names>TS</given-names></string-name></person-group>. <article-title>Cross-modal moment localization in videos</article-title>. In: <conf-name>Proceedings of the 26th ACM International Conference on MuLtimedia</conf-name>. <publisher-loc>New York, NY, USA</publisher-loc>: <publisher-name>The Association for Computing Machinery (ACM)</publisher-name>; <year>2018</year>. p. <fpage>843</fpage>&#x2013;<lpage>51</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3240508.3240549</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Paul</surname> <given-names>S</given-names></string-name>, <string-name><surname>Mithun</surname> <given-names>NC</given-names></string-name>, <string-name><surname>Roy-Chowdhury</surname> <given-names>AK</given-names></string-name></person-group>. <article-title>Text-based temporal localization of novel events</article-title>. In: <conf-name>Proceedings of the European Conference on Computer Vision (ECCV)</conf-name>. <publisher-loc>Berlin/Heidelberg, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2022</year>. p. <fpage>567</fpage>&#x2013;<lpage>87</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-031-19781-9_33</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Nie</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Rui</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>A survey on video moment localization</article-title>. <source>ACM Comput Surv</source>. <year>2023</year>;<volume>55</volume>(<issue>9</issue>):<fpage>1</fpage>&#x2013;<lpage>37</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3556537</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Xiao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yao</surname> <given-names>A</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chua</surname> <given-names>TS</given-names></string-name></person-group>. <article-title>Can I trust your answer? Visually grounded video question answering</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2024 Jun 16&#x2013;22</conf-name>; <publisher-loc>Seattle, WA, USA</publisher-loc>. p. <fpage>13204</fpage>&#x2013;<lpage>14</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR52733.2024.01254</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Nie</surname> <given-names>L</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A survey on video temporal grounding with multimodal large language model</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2026</year>;<volume>48</volume>(<issue>2</issue>):<fpage>1521</fpage>&#x2013;<lpage>41</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TPAMI.2025.3615586</pub-id>; <pub-id pub-id-type="pmid">41021939</pub-id></mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Lewis</surname> <given-names>P</given-names></string-name>, <string-name><surname>Perez</surname> <given-names>E</given-names></string-name>, <string-name><surname>Piktus</surname> <given-names>A</given-names></string-name>, <string-name><surname>Petroni</surname> <given-names>F</given-names></string-name>, <string-name><surname>Karpukhin</surname> <given-names>V</given-names></string-name>, <string-name><surname>Goyal</surname> <given-names>N</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Retrieval-augmented generation for knowledge-intensive NLP tasks</article-title>. <comment>arXiv:2005.11401. 2020</comment>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zheng</surname> <given-names>X</given-names></string-name>, <string-name><surname>Weng</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Lyu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Xue</surname> <given-names>H</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>B</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Retrieval augmented generation and understanding in vision: a survey and new outlook</article-title>. <comment>arXiv:2503.18016. 2025</comment>.</mixed-citation></ref>
</ref-list>
</back></article>