<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">58193</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2024.058193</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Improving Badminton Action Recognition Using Spatio-Temporal Analysis and a Weighted Ensemble Learning Model</article-title>
<alt-title alt-title-type="left-running-head">Improving Badminton Action Recognition Using Spatio-Temporal Analysis and a Weighted Ensemble Learning Model</alt-title>
<alt-title alt-title-type="right-running-head">Improving Badminton Action Recognition Using Spatio-Temporal Analysis and a Weighted Ensemble Learning Model</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Asriani</surname><given-names>Farida</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Azhari</surname><given-names>Azhari</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><email>arisn@ugm.ac.id</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Wahyono</surname><given-names>Wahyono</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<aff id="aff-1"><label>1</label><institution>Department of Computer Science and Electronics, Universitas Gadjah Mada</institution>, <addr-line>Yogyakarta, 55281</addr-line>, <country>Indonesia</country></aff>
<aff id="aff-2"><label>2</label><institution>Electrical Engineering Department, Universitas Jenderal Soedirman</institution>, <addr-line>Purbalingga, 53371</addr-line>, <country>Indonesia</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Azhari Azhari. Email: <email>arisn@ugm.ac.id</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2024</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>18</day><month>11</month><year>2024</year>
</pub-date>
<volume>81</volume>
<issue>2</issue>
<fpage>3079</fpage>
<lpage>3096</lpage>
<history>
<date date-type="received">
<day>06</day>
<month>9</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>18</day>
<month>10</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2024 The Authors.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_58193.pdf"></self-uri>
<abstract>
<p>Incredible progress has been made in human action recognition (HAR), significantly impacting computer vision applications in sports analytics. However, identifying dynamic and complex movements in sports like badminton remains challenging due to the need for precise recognition accuracy and better management of complex motion patterns. Deep learning techniques like convolutional neural networks (CNNs), long short-term memory (LSTM), and graph convolutional networks (GCNs) improve recognition in large datasets, while the traditional machine learning methods like SVM (support vector machines), RF (random forest), and LR (logistic regression), combined with handcrafted features and ensemble approaches, perform well but struggle with the complexity of fast-paced sports like badminton. We proposed an ensemble learning model combining support vector machines (SVM), logistic regression (LR), random forest (RF), and adaptive boosting (AdaBoost) for badminton action recognition. The data in this study consist of video recordings of badminton stroke techniques, which have been extracted into spatiotemporal data. The three-dimensional distance between each skeleton point and the right hip represents the spatial features. The temporal features are the results of Fast Dynamic Time Warping (FDTW) calculations applied to 15 frames of each video sequence. The weighted ensemble model employs soft voting classifiers from SVM, LR, RF, and AdaBoost to enhance the accuracy of badminton action recognition. The E2 ensemble model, which combines SVM, LR, and AdaBoost, achieves the highest accuracy of 95.38%.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Weighted ensemble learning</kwd>
<kwd>badminton action</kwd>
<kwd>soft voting classifier</kwd>
<kwd>joint skeleton</kwd>
<kwd>fast dynamic time warping</kwd>
<kwd>spatiotemporal</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Higher Education Funding (BPPT)</funding-source>
</award-group>
<award-group id="awg2">
<funding-source>Indonesia Endowment Fund for Education (LPDP)</funding-source>
<award-id>02092/J5.2.3/BPI.06/9/2022</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Incredible progress has been made in detecting human action recognition (HAR), which significantly impacts computer vision applications in sports analytics. Despite these developments, it is still challenging to identify dynamic and complex movements like those in badminton because precise recognition accuracy and improved management of complex motion patterns are required [<xref ref-type="bibr" rid="ref-1">1</xref>&#x2013;<xref ref-type="bibr" rid="ref-3">3</xref>].</p>
<p>Spatiotemporal dynamics in video and skeleton data generated with deep learning techniques like convolutional neural networks (CNNs), long short-term memory (LSTM), and graph convolutional networks (GCNs) are highly effective [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-4">4</xref>]. These techniques improve recognition performance in large datasets by automatically learning hierarchical features. However, they can be resource-intensive and might not translate well to smaller, more specialized datasets, like badminton action recognition datasets.</p>
<p>On the other hand, when paired with handcrafted feature extraction, conventional machine learning techniques such as support vector machines (SVM), random forests (RF), and logistic regression (LR) show good performance [<xref ref-type="bibr" rid="ref-5">5</xref>&#x2013;<xref ref-type="bibr" rid="ref-7">7</xref>]. Ensemble approaches improve accuracy even more by utilizing the advantages of each model [<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-9">9</xref>]. Even with their efficacy, these models cannot fully represent the intricacy of human movement, especially in high-speed sports like badminton.</p>
<p>Feature extraction methods based on deep learning and handcrafted techniques have been used in recent studies. While handcrafted methods like motion trajectories and histograms of oriented gradients (HOG) are practical and straightforward, they frequently lack the sophistication required to capture the subtleties of intricate actions. Deep learning techniques utilizing 3D skeleton data provide more comprehensive spatiotemporal representations and have been extensively demonstrated in previous research [<xref ref-type="bibr" rid="ref-1">1</xref>&#x2013;<xref ref-type="bibr" rid="ref-3">3</xref>]. Tasnim et al. [<xref ref-type="bibr" rid="ref-10">10</xref>] employed MobileNetV2, DenseNet121, and ResNet18 models in conjunction with transfer learning and fusion techniques, resulting in high accuracy on benchmark dataset.</p>
<p>Action recognition could be significantly improved by integrating 3D skeleton data, which records joint movements&#x2019; temporal and spatial progression [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-12">12</xref>]. While skeleton-based methods have been used for human action recognition, deep learning implementations customized to badminton are still in their infancy and could be improved [<xref ref-type="bibr" rid="ref-13">13</xref>]. Furthermore, although 3D skeleton data is robust against common problems such as lighting and camera angles, it is still unclear how to extract meaningful features from this data to create scalable and resilient models.</p>
<p>Combining complementary color and texture features using the Choquet fuzzy integral to model complex, multi-modal distributions is one of the critical multi-view learning strategies, particularly in complex environments. Atanassov&#x2019;s intuitive 3D fuzzy Histon roughness index and other methods for encoding spatial information capture intricate spatial relationships, while wavelet transform image scaling improves the differentiation of spatial dynamics. Better feature extraction is achieved through content-aware feature weighting, which determines significance based on context. Additionally, the multi-view Bag of Words (BoW) model enhances scene categorization by integrating spatial and semantic information from multiple perspectives. Lastly, combining color and spatial features solves insufficient shape-based recognition [<xref ref-type="bibr" rid="ref-14">14</xref>,<xref ref-type="bibr" rid="ref-15">15</xref>].</p>
<p>This paper proposes a weighted ensemble learning technique and a spatiotemporal analysis of 3D skeletal data to address these issues. We suggest applying Fast Dynamic Time Warping (FDTW) for the temporal feature and using the right hip&#x2019;s 3D Euclidean distance as the spatial feature. The contributions of this paper are:
<list list-type="order">
<list-item>
<p>A new 3D spatiotemporal feature extraction method, in which the temporal feature is extracted using FDTW, and the spatial feature is the separation between the right hip and other skeleton coordinates.</p></list-item>
<list-item>
<p>A weighted ensemble learning model for badminton action recognition that combines LR, SVM, and AdaBoost.</p></list-item>
<list-item>
<p>A performance comparison indicates that the proposed weighted ensemble model surpasses deep learning methods, such as CNN and LSTM, on small datasets like badminton action detection, delivering competitive accuracy with reduced computational expense for real-time applications.</p></list-item>
</list></p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Works</title>
<sec id="s2_1">
<label>2.1</label>
<title>Three-Dimensional (3D) Human Skeleton Detection</title>
<p>The real-time posture estimation system, based on a deep neural network, has been researched extensively. For example, the system used convolutional neural networks (CNNs) and recurrent neural networks (RNNs) to determine a person&#x2019;s posture from a single RGB camera. The system exhibited a high accuracy in estimating an individual human&#x2019;s three-dimensional stance and utilizing a publicly accessible dataset [<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-16">16</xref>&#x2013;<xref ref-type="bibr" rid="ref-19">19</xref>]. Li et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] proposed a posture estimation system using a similar methodology based on an RGB-D camera. The convolutional neural network estimated the human joint positions using a depth map and the pose. The system could calculate the pose accurately and in real-time on a publicly available dataset. A novel posture estimation system was created, employing convolutional neural networks and recurrent neural network techniques. The technology was designed to precisely forecast the human body&#x2019;s position in real time with a sliding window mechanism. The suggested method successfully executed real-time posture estimation and achieved remarkable accuracy on a publicly available dataset [<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-16">16</xref>,<xref ref-type="bibr" rid="ref-20">20</xref>&#x2013;<xref ref-type="bibr" rid="ref-22">22</xref>]. The study uses media pipes and convolutional neural networks to measure the human body&#x2019;s position in three dimensions. This multi-step method accurately assesses human posture in publicly available datasets [<xref ref-type="bibr" rid="ref-16">16</xref>,<xref ref-type="bibr" rid="ref-20">20</xref>,<xref ref-type="bibr" rid="ref-22">22</xref>&#x2013;<xref ref-type="bibr" rid="ref-24">24</xref>]. GCN can be used to effectively model human joints and improve performance in pose estimation and action recognition tasks. Using GCN with spatiotemporal information enhances accuracy in complex scenes [<xref ref-type="bibr" rid="ref-25">25</xref>]. Action recognition improvement is also achieved by combining behavioral dependencies and contextual cues. This method allows the model to distinguish between similar poses [<xref ref-type="bibr" rid="ref-26">26</xref>].</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Spatiotemporal Analysis in Action Recognition</title>
<p>Researchers have analyzed human actions using a spatiotemporal feature analysis method based on deep learning. In action recognition, temporal sequence recordings are essential. Shahroudy et al. [<xref ref-type="bibr" rid="ref-27">27</xref>] established the NTU RGB&#x002B;D dataset, which has now become one of the largest and most diversified datasets for 3D action recognition. People extensively use it to evaluate the effectiveness of spatiotemporal models. Liu et al. [<xref ref-type="bibr" rid="ref-28">28</xref>] used global context-aware attention LSTM networks on 3D frame data, which helped the model focus on important joints and time segments. Graph Convolutional Networks (GCNs) have demonstrated efficacy as a powerful instrument for spatiotemporal modeling. Spatial-Temporal Graph Convolutional Networks (ST-GCN) serve as a technique for skeleton-based action recognition. The spatial interaction between joints and the temporal dynamics along the sequence are wrapped in this network [<xref ref-type="bibr" rid="ref-29">29</xref>]. The adaptive two-stream graph convolutional network (2s-AGCN) was proposed by Shi et al. [<xref ref-type="bibr" rid="ref-30">30</xref>] to support this idea of two-stream adaptive graph convolutional networks (2s-AGCN). These networks enhance accuracy by adaptively learning the graph topology from the supplied input. Moreover, Temporal Convolutional Networks (TCNs) exhibit considerable potential in this domain. Long Short-Term Memory (LSTM) models aren&#x2019;t as good at sequence modeling as Temporal Convolutional Networks (TCN) models because they can&#x2019;t capture long-range relationships as well as TCN models [<xref ref-type="bibr" rid="ref-31">31</xref>].</p>
<p>Besides deep learning methods, various methodologies have contributed to spatiotemporal analysis. Action posture is encoded using a bag of 3D points from depth data, and action graph dynamics are modeled for each action. Koniusz et al. [<xref ref-type="bibr" rid="ref-32">32</xref>] portrayed that space-time action sequences can efficiently be captured in a tensor space. Such methods have already been used to recognize badminton actions in RGB videos [<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>]. Feature extraction methods that have been applied are bounding box and histogram of oriented gradient (HOG). However, these methods produce misclassification for some classes involving the same pose during stroke [<xref ref-type="bibr" rid="ref-33">33</xref>]. The sliding window and Haar-like methods have also been applied. However, the recognition rate for different badminton players is not ideal [<xref ref-type="bibr" rid="ref-34">34</xref>]. In addition, detection of shot badminton players converts video data into sequential frames, applying sliding windows for feature extraction to detect shot timing and position with an accuracy of 95.9%. Similarly, image recognition of badminton swing motion based on a single inertial sensor transforms motion data into sequences, using sliding and action windows for accurate segmentation. This enables the Deep Residual LSTM model to recognize six swing types with over 90% accuracy, automating the extraction and recognition process for performance analysis in badminton [<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-36">36</xref>]. Atanassov&#x2019;s intuitive 3D fuzzy Histon roughness index method has been discovered to detect moving objects with colored data. (IA-IA3DFHRI) [<xref ref-type="bibr" rid="ref-14">14</xref>]. The proposed method demonstrates resilience to dynamic backgrounds.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Ensemble Learning in Action Recognition</title>
<p>Ensemble learning has been applied to several other fields, as pointed out by prior investigations. For example, Karim et al. [<xref ref-type="bibr" rid="ref-37">37</xref>] used ensemble learning combining logistic regression, decision trees, and SVM in healthcare. Regarding handwritten recognition, Karray et al. [<xref ref-type="bibr" rid="ref-38">38</xref>] proved improved accuracy by ensemble deep learning, RF, and SVM models. The accuracy of HAR can be enhanced by integrating SVM and HMM [<xref ref-type="bibr" rid="ref-39">39</xref>]. Combining SVM, KNN, and LR with a soft voting classifier can improve accuracy to 92.78% [<xref ref-type="bibr" rid="ref-40">40</xref>]. Das et al. [<xref ref-type="bibr" rid="ref-41">41</xref>] combined ensemble models (DT, RF, SVM, and KNN) with soft voting techniques and weighted optimization using FOX optimization. Research in human action recognition has also integrated deep learning models. The suggested ensemble learning algorithm includes DNN and CNN stacked at the gated recurrent unit (GRU) to recognize human actions [<xref ref-type="bibr" rid="ref-42">42</xref>]. Kaur et al. [<xref ref-type="bibr" rid="ref-43">43</xref>] combined ResNet50 and a custom CNN in human action recognition.</p>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Badminton Action Recognition</title>
<p>Badminton action recognition methods are categorized into two groups: machine learning-based classification and deep learning-based classification. Ting et al. [<xref ref-type="bibr" rid="ref-44">44</xref>] applied SVM for badminton recognition, classifying badminton strokes into ten distinct classes. Later, researchers further proved SVM&#x2019;s efficacy in badminton action recognition. Anik et al. [<xref ref-type="bibr" rid="ref-45">45</xref>] identified badminton stroke techniques. In this study, the accuracy of SVM reached 88%. Ghazali et al. [<xref ref-type="bibr" rid="ref-46">46</xref>] also confirmed that SVM provides superior classification results compared to decision trees, KNN, and SVM, with SVM achieving an accuracy of 83.4%. He concluded that future research should enhance recognition accuracy by expanding and refining feature extraction techniques. Another conventional machine learning method applied to BAR is AdaBoost. The recognition rates of AdaBoost are higher than those of HMM [<xref ref-type="bibr" rid="ref-34">34</xref>].</p>
<p>Deep learning has been applied in human action recognition. Wang et al. [<xref ref-type="bibr" rid="ref-47">47</xref>] applied the Adaptive Feature Extraction Block (AFEB) AlexNet. Rahmad et al. [<xref ref-type="bibr" rid="ref-4">4</xref>] have proven that GoogleNet has the highest classification performance compared to AlexNet, VGGNet-16, and VGGNet-19. GoogleNet provided the best accuracy for differentiating between smash and non-smash techniques [<xref ref-type="bibr" rid="ref-48">48</xref>,<xref ref-type="bibr" rid="ref-49">49</xref>]. Steels et al. [<xref ref-type="bibr" rid="ref-50">50</xref>] confirmed the accuracy of CNN in classifying nine activities using accelerometer data.</p>
<p>Analyzing badminton movements using skeleton key point data is a mostly unexplored area of research. Using the AlphaPose framework, Liang et al. [<xref ref-type="bibr" rid="ref-1">1</xref>] analyzed key point data from certain badminton films. Comparative investigation shows that the Long Short-Term Memory (LSTM) model surpasses the CNN model. Another model implemented is PDDRNet for 3D human pose estimation and XGBoost for classification [<xref ref-type="bibr" rid="ref-2">2</xref>]. Liu et al. [<xref ref-type="bibr" rid="ref-3">3</xref>] applied GCNN with skeletal data for badminton action recognition.</p>
<p>This research contributes to the spatiotemporal analysis of skeleton data and weighted ensemble learning for badminton action recognition. The use of video data processed into 3D skeleton coordinates. The data skeleton was chosen because, according to Sun et al. [<xref ref-type="bibr" rid="ref-51">51</xref>], the data skeleton has the advantage of providing 3D pose information of an object. It&#x2019;s simple yet informative and not sensitive to viewpoints, background, or light intensity changes. On the other hand, if the data is processed in RGB format, it will be susceptible to changes in viewpoint, light intensity, and changes in the background. The spatiotemporal feature captures the changes in the position of objects or body parts in space and time (temporal). This feature allows the model to understand the movement sequence, which is critical for recognizing actions that involve dynamic changes. In previous research, DTW was applied in human action recognition as a classifier by seeking the similarity between two sequences. In this paper, fast DTW is applied as a method for temporal feature extraction. One of the advantages of using fast DTW as a feature extraction method is its ability to reduce the number of features to a smaller amount. For example, the temporal feature in the right hip distance with dimensions 15 &#x00D7; 33, when processed with FDTW, will become a feature with dimensions 1 &#x00D7; 33.</p>
<p>Another contribution is the weighted ensemble of SVM, LR, RF, and AdaBoost. Machine learning classics were more efficient when compared with deep learning models. With appropriate feature extraction, machine learning models can achieve high accuracy. Ensemble learning methods that combine several classifier models have been proven to enhance accuracy.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>The Proposed Method</title>
<p><xref ref-type="fig" rid="fig-1">Fig. 1</xref> presents the proposed method. The first stage of video data acquisition for badminton action. The second stage is data preprocessing. The third stage is spatiotemporal data extraction. The fourth stage develops an ensemble learning model for badminton action recognition.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Proposed method</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_58193-fig-1.tif"/>
</fig>
<sec id="s3_1">
<label>3.1</label>
<title>Data Acquisition</title>
<p>Data was collected using a 15MP DSLR camera with the installation as presented in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>. The video data resolution is 30 fps. An indoor badminton court served as the location for data collection. We used the existing lighting in the field without any additions. The video records a single badminton stroke technique. The execution of the badminton stroke begins from the athlete&#x2019;s ready position at the center of the court and returns to the center after performing the badminton stroke. The subjects included badminton athletes registered with PBSI, aged 15 to 22 years. Badminton techniques are presented in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>. The dataset consists of 1333 videos for training and 433 videos for validation.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Camera installation</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_58193-fig-2.tif"/>
</fig><fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Badminton stroke techniques: (a) Drive backhand; (b) Drive forehand; (c) Overhead backhand; (d) Overhead forehand; (e) Underhand backhand; (f) Underhand forehand</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_58193-fig-3.tif"/>
</fig>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Data Preprocessing</title>
<p>The data preprocessing stage segmented the <italic>T</italic>-frame video data into <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>l</mml:mi></mml:math></inline-formula>-frame sequences. We perform the frame segmentation length of <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>l</mml:mi></mml:math></inline-formula> &#x003D; 15 by taking frames from 1 to <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>T</mml:mi></mml:math></inline-formula>, using the frame interval as presented in <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>. To extract the 3D coordinates for 33 pose landmarks, we apply media pipe to each frame. <xref ref-type="fig" rid="fig-4">Fig. 4</xref> presents the 33 pose landmarks. <xref ref-type="fig" rid="fig-5">Fig. 5</xref> presents the 15-frame pose landmarks.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>33 Pose landmarks</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_58193-fig-4.tif"/>
</fig><fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>15-Frame landmark pose</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_58193-fig-5.tif"/>
</fig>
<p><disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>v</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mi>T</mml:mi><mml:mi>l</mml:mi></mml:mfrac><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Feature Extraction</title>
<sec id="s3_3_1">
<label>3.3.1</label>
<title>Spatial Feature Extraction</title>
<p>Preprocessing data produces an array with dimensions l &#x00D7; m &#x00D7; 3, where l represents the number of frames (l &#x003D; 0, 1, 2, ..., 14), m represents the number of skeleton points (m &#x003D; 0, 1, 2, ..., 32), and 3 represents the coordinates <italic>x</italic>, <italic>y</italic>, and <italic>z</italic>. The spatial feature is extracted by measuring the right hip distance or the three-dimensional distance between the right hip and other skeletal points. (d). spatial data is the right hip distance in frame 7. Based on the skeleton number in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>, the right hip coordinate from frame seven was written by <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>7</mml:mn><mml:mo>,</mml:mo><mml:mn>24</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>7</mml:mn><mml:mo>,</mml:mo><mml:mn>24</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>x</mml:mi><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>7</mml:mn><mml:mo>,</mml:mo><mml:mn>24</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref> represents the calculation of the right hip distance, while <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref> represents the spatial features.
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mi>d</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msqrt><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>7</mml:mn><mml:mo>,</mml:mo><mml:mn>24</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mn>7</mml:mn><mml:mo>,</mml:mo><mml:mn>24</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mi>m</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mn>7</mml:mn><mml:mo>,</mml:mo><mml:mn>24</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:msqrt></mml:math></disp-formula>
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mi>S</mml:mi><mml:mi>p</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>f</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mo>=</mml:mo><mml:mi>d</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mn>7</mml:mn><mml:mo>,</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mn>0</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mi>d</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mn>7</mml:mn><mml:mo>,</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>d</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mn>7</mml:mn><mml:mo>,</mml:mo><mml:mspace width="negativethinmathspace" /><mml:mn>32</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
</sec>
<sec id="s3_3_2">
<label>3.3.2</label>
<title>Temporal Feature Extraction</title>
<p>Temporal features are generated with FDTW applied to each column of right hip distance. In each data sample, d (l,m) is calculated. A sequence is defined as the order d (l,m) for each m, arranged according to the order of frames. For example, if m &#x003D; 12 (right shoulder), the series is: (d (0,12), d (1,12), &#x2026;, d (14,12)). The calculation of FDTW requires a series as a reference. This paper uses data 0 (n &#x003D; 0) as the reference. The temporal feature, which is the result of FDTW calculation in the form of a total distance (TD) sequence, is represented in <xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref>. Algorithm 1 describes the FDTW algorithm.
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mi>T</mml:mi><mml:mi>e</mml:mi><mml:mi>m</mml:mi><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>f</mml:mi><mml:mi>e</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mo>=</mml:mo><mml:mi>T</mml:mi><mml:mi>D</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>T</mml:mi><mml:mi>D</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>T</mml:mi><mml:mi>D</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mn>2</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2026;</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:mi>T</mml:mi><mml:mi>D</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mn>32</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<fig id="fig-9">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_58193-fig-9.tif"/>
</fig>
<p>The calculation of DTW, illustrated in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>, displays the results from a single joint skeleton of the right hip and the right hand in one video sequence, showing the calculation of DTW. The spatial data has 33 features, and the temporal data has 33 features, so the spatiotemporal data has 66 features. This integration of spatial and temporal information provides a comprehensive representation of the athlete&#x2019;s movements, encompassing both the static and dynamic aspects of the badminton stroke techniques.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Calculation of DTW for skeleton points on right-hand</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_58193-fig-6.tif"/>
</fig>
</sec>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Machine Learning &#x0026; Ensemble Model</title>
<p>We applied SVM, LR, RF, and AdaBoost in this research. Deep learning models (CNN and LSTM) are applied as a comparison in the machine learning model analysis. Each model was trained and evaluated individually. Five-fold cross-validation was used in validation.</p>
<p>The ensemble learning models in this study are arranged based on several variations of machine learning methods, as shown in <xref ref-type="table" rid="table-1">Table 1</xref>, such as Ensemble-1 (E1) is configured from SVM, LR, RF, and AdaBoost, and configuration individual models for Ensemble-2 (E2), Ensemble-3 (E3), and Ensemble-4 (E4). The weighted soft voting classifier method (E1, E2, E3, and E4) can contribute to the final decision. The voting process will choose the more excellent score for more robust models.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Ensemble model</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Ensemble model</th>
<th>Model configuration</th>
</tr>
</thead>
<tbody>
<tr>
<td>Ensemble-1 (E1)</td>
<td>SVM, LR, RF and AdaBoost</td>
</tr>
<tr>
<td>Ensemble-2 (E2)</td>
<td>SVM, LR and AdaBoost</td>
</tr>
<tr>
<td>Ensemble-3 (E3)</td>
<td>SVM, RF and AdaBoost</td>
</tr>
<tr>
<td>Ensembel-4 (E4)</td>
<td>SVM and AdaBoost</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Result and Discussion</title>
<sec id="s4_1">
<label>4.1</label>
<title>Machine Learning Models</title>
<p>We performed optimization by systematically testing several hyperparameter configurations for each model to determine the most effective combination that produces the maximum degree of performance. This method uses five-fold cross-validation to analyze the hyperparameters thoroughly. The training model that delivers the best performance metrics, such as accuracy, precision, recall, and F1 score, is chosen as the determinant of the best hyperparameters. The best hyperparameter model is shown in <xref ref-type="table" rid="table-2">Table 2</xref>.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Best hyperparameter</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>Best hyperparameters</th>
</tr>
</thead>
<tbody>
<tr>
<td>SVM</td>
<td>C: 10, gamma: 0.01, kernel: rbf</td>
</tr>
<tr>
<td>LR</td>
<td>C: 10, solver: liblinear</td>
</tr>
<tr>
<td>RF</td>
<td>criterion: gini, max_depth: 30, max_features: log2, min_samples_leaf: 1, min_samples_ split: 5, n_estimators: 300</td>
</tr>
<tr>
<td>AdaBoost</td>
<td>estimator__max_depth: 5, learning_rate: 1, n_estimators: 300</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>We conduct measurements of complexity, training time, and memory usage during model training. <xref ref-type="table" rid="table-3">Table 3</xref> presents the results. <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref> to <xref ref-type="disp-formula" rid="eqn-10">Eq. (10)</xref> allow for the calculation of the training time complexity (Big O). Compared to CNN and LSTM, classical machine learning (SVM, LR, RF, and AdaBoost) has lower complexity, training time, and memory usage during training.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Algorithm&#x2019;s complexity and training cost model</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Algorithm</th>
<th>Training time complexity</th>
<th>Training time (s)</th>
<th>Training memory (MB)</th>
</tr>
</thead>
<tbody>
<tr>
<td>SVM</td>
<td>O (2.04 &#x00D7; 10<sup>8</sup>) until O (2.98 &#x00D7; 10<sup>11</sup>)</td>
<td>0.47</td>
<td>0.32</td>
</tr>
<tr>
<td>LR</td>
<td>O (8.80 &#x00D7; 10<sup>4</sup>)</td>
<td>0.36</td>
<td>0.10</td>
</tr>
<tr>
<td>RF</td>
<td>O (8.25 &#x00D7; 10<sup>7</sup>)</td>
<td>4.86</td>
<td>0.26</td>
</tr>
<tr>
<td>AdaBoost</td>
<td>O (2.64 &#x00D7; 10<sup>7</sup>)</td>
<td>17.04</td>
<td>3.12</td>
</tr>
<tr>
<td>CNN</td>
<td>O (2.15 &#x00D7; 10<sup>9</sup>)</td>
<td>14.09</td>
<td>689.77</td>
</tr>
<tr>
<td>LSTM</td>
<td>O (4.97 &#x00D7; 10<sup>8</sup>)</td>
<td>19.75</td>
<td>189.93</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mi>B</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>S</mml:mi><mml:mi>V</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>O</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>n</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>.</mml:mo><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>u</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>l</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>O</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>n</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msup><mml:mo>.</mml:mo><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mi>B</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>L</mml:mi><mml:mi>R</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>O</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>.</mml:mo><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mi>B</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>R</mml:mi><mml:mi>F</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>O</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>m</mml:mi><mml:mo>.</mml:mo><mml:mi>n</mml:mi><mml:mo>.</mml:mo><mml:mi>d</mml:mi><mml:mspace width="thinmathspace" /><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mi>B</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>A</mml:mi><mml:mi>d</mml:mi><mml:mi>a</mml:mi><mml:mi>B</mml:mi><mml:mi>o</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>O</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>m</mml:mi><mml:mo>.</mml:mo><mml:mi>n</mml:mi><mml:mo>.</mml:mo><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mi>B</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mi>N</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>O</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>.</mml:mo><mml:msup><mml:mi>d</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>.</mml:mo><mml:msup><mml:mi>k</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>.</mml:mo><mml:mi>f</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mi>B</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>L</mml:mi><mml:mi>S</mml:mi><mml:mi>T</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mo>.</mml:mo><mml:mi>t</mml:mi><mml:mo>.</mml:mo><mml:mi>h</mml:mi><mml:mo>.</mml:mo><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>With <italic>n</italic>: number of samples, <italic>d</italic>: number of features, <italic>k</italic>: number of kernels, <italic>m</italic>: number of trees, <italic>t</italic>: sequence length.</p>
<p>We performed the validation test using 433 distinct data points that differed from the training data. <xref ref-type="table" rid="table-4">Table 4</xref> shows that the performance of each deep learning method is greater than that of machine learning methods. For example, the accuracy of deep learning (CNN, 93.06%; LSTM, 94.06%). The precision of deep learning (CNN, 94.66%; LSTM, 96.31%), The recall of deep learning (CNN, 94.61; LSTM, 95.26), et cetera., the accuracy of deep learning (CNN, 93.06%; LSTM, 94.06%). The precision of deep learning (CNN, 94.66%; LSTM, 96.31%). The recall of deep learning (CNN, 94.61; LSTM, 95.26), etc.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Evaluation models</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Model (Individually)</th>
<th>Accuracy</th>
<th>Precision</th>
<th>Recall</th>
<th>F1 score</th>
</tr>
</thead>
<tbody>
<tr>
<td>SVM</td>
<td>91.22</td>
<td>91.67</td>
<td>91.22</td>
<td>91.15</td>
</tr>
<tr>
<td>LR</td>
<td>90.53</td>
<td>91.14</td>
<td>90.53</td>
<td>90.48</td>
</tr>
<tr>
<td>RF</td>
<td>90.76</td>
<td>91.05</td>
<td>90.76</td>
<td>90.75</td>
</tr>
<tr>
<td>AdaBoost</td>
<td>92.15</td>
<td>91.68</td>
<td>91.28</td>
<td>91.15</td>
</tr>
<tr>
<td>CNN</td>
<td>93.06</td>
<td>94.66</td>
<td>95.10</td>
<td>94.61</td>
</tr>
<tr>
<td>LSTM</td>
<td>94.06</td>
<td>96.31</td>
<td>95.26</td>
<td>95.54</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Weighted Ensemble Model</title>
<p>A soft voting classifier method in an ensemble model enhances accuracy and generalization in badminton action recognition. The weight value of each model SVM, LR, RF, and AdaBoost is assigned based on accuracy. We determine the weights by normalizing each model&#x2019;s accuracy values. <xref ref-type="table" rid="table-5">Table 5</xref> shows the comparison evaluation matrix of ensemble models. E2 achieved the highest performance with an accuracy of 95.38%, precision of 96.01%, recall of 94.89%, and F1 score of 95, 25%.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>The evaluation of ensemble model validation</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Ensemble model</th>
<th>Accuracy (%)</th>
<th>Precision (%)</th>
<th>Recall (%)</th>
<th>F1 score (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>SVM-LR-RF- AdaBoost (E1)</td>
<td>94.69</td>
<td>95.25</td>
<td>94.19</td>
<td>94.53</td>
</tr>
<tr>
<td>SVM-LR- AdaBoost (E2)</td>
<td><bold>95</bold>.<bold>38</bold></td>
<td><bold>96</bold>.<bold>01</bold></td>
<td><bold>94</bold>.<bold>89</bold></td>
<td><bold>95</bold>.<bold>25</bold></td>
</tr>
<tr>
<td>SVM-RF- AdaBoost (E3)</td>
<td>93.30</td>
<td>93.59</td>
<td>92.86</td>
<td>93.03</td>
</tr>
<tr>
<td>SVM- AdaBoost (E4)</td>
<td>93.53</td>
<td>93.92</td>
<td>93.06</td>
<td>93.32</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>According to the validation test results, ensemble model E2, which combines SVM, LR, and AB, outperforms ensemble model E1, which combines SVM, LR, RF, and AB. The superior performance of E2 indicates that adding RF to E1 does not contribute to performance improvement and may even introduce noise that leads to a decrease in model efficiency. It also emphasizes that the selection of the right individual models is necessary to obtain an optimal ensemble model. E2 successfully leverages the strengths of SVM, LR, and AB more synergistically without the additional complexity that could weaken the overall generalization power of the model, as seen in E1. This finding shows that the number of models in an ensemble does not always correlate with the final performance; instead, the quality and suitability of the combined models play a crucial role in determining the outcome.</p>
<p>Refer to <xref ref-type="fig" rid="fig-7">Fig. 7</xref> for the evaluation results in the confusion matrix. An analysis of the confusion matrix from the top two ensemble models, E1 and E2, reveals a notable disparity in performance, especially in the categories of drive backhand and overhead backhand. Group E2 had the highest accuracy. Model E2 has superior sensitivity in detecting drive backhand patterns. Moreover, in the overhead backhand class, model E2 again demonstrates its superiority by achieving greater accuracy in differentiating this movement from other classes. That indicates a superior capability to identify patterns of overhead backhand action.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Confusion matrix: (a) Confusion matrix ensemble E1 (SVM-LR-RF-AdaBoost); (b) Confusion matrix ensemble E2 (SVM-LR-AdaBoost)</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_58193-fig-7.tif"/>
</fig>
<p>The classification levels for each class are presented in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>. Based on that figure, model E2 has an accuracy level of 100% for the drive backhand class, and four classes have an accuracy above 95%. In comparison, there are two classes with an accuracy below 95%, namely the overhead backhand class at 90.2% and the underhand forehand class at 87.2%. The accuracy value for the underhand backhand suggests that the model struggles to differentiate this movement from other movements with similar characteristics in this classification study.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>The accuracy of each class for E2</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_58193-fig-8.tif"/>
</fig>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Analysis Result</title>
<p>To evaluate the results of the proposed ensemble learning model, we created a comparison table of the results of recognizing badminton stroke techniques using skeleton data as presented in <xref ref-type="table" rid="table-6">Table 6</xref>. Each model in this table uses different data for its recognition process. Ensemble learning with 3D spatiotemporal skeleton features has higher accuracy compared to other models from the state of the art.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Comparison of results and state-of-the-art</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>Preprocessing data</th>
<th>Akurasi (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>E2</bold> (<bold>weighted ensemble SVM-LR-AdaBoost</bold>)</td>
<td><bold>Spatiotemporal 3D skeleton</bold></td>
<td><bold>95</bold>.<bold>38</bold></td>
</tr>
<tr>
<td>LSTM [<xref ref-type="bibr" rid="ref-1">1</xref>]</td>
<td>Skeleton data extracted by AlphaPose</td>
<td>80</td>
</tr>
<tr>
<td>CNN [<xref ref-type="bibr" rid="ref-1">1</xref>]</td>
<td>Skeleton data extracted by AlphaPose</td>
<td>60</td>
</tr>
<tr>
<td>PDDRNet-XGBOOST [<xref ref-type="bibr" rid="ref-2">2</xref>]</td>
<td>Human pose joint skeleton</td>
<td>93,5</td>
</tr>
<tr>
<td>GCN [<xref ref-type="bibr" rid="ref-3">3</xref>]</td>
<td>Human Skeleton Sequence</td>
<td>92</td>
</tr>
<tr>
<td>SVM [<xref ref-type="bibr" rid="ref-44">44</xref>]</td>
<td>RGB-D sensors</td>
<td>92</td>
</tr>
<tr>
<td>SVM [<xref ref-type="bibr" rid="ref-45">45</xref>]</td>
<td>Accelerometer and gyroscope</td>
<td>88.89</td>
</tr>
<tr>
<td>SVM [<xref ref-type="bibr" rid="ref-46">46</xref>]</td>
<td>Inertial senor</td>
<td>83.4</td>
</tr>
<tr>
<td>HMM [<xref ref-type="bibr" rid="ref-33">33</xref>]</td>
<td>Bounding box, HOG</td>
<td>83.33</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>Based on the experiments, spatiotemporal features of 3D skeleton data and ensemble models have been proven to provide high accuracy in recognizing badminton stroke techniques. The selection of machine learning models is significant in obtaining the best accuracy of the ensemble model. Based on the experiments, spatiotemporal features of 3D skeleton data and ensemble models have been proven to provide high accuracy in recognizing badminton stroke techniques. The selection of machine learning models is crucial to obtain the best accuracy of the ensemble model. E2 achieved the best performance with an accuracy rate of 95.38%.</p>
</sec>
</body>
<back>
<ack><p>We are truly thankful for the financial assistance from the Center for Higher Education Funding (BPPT) and the Indonesia Endowment Fund for Education (LPDP). Also, thank you to badminton athletes from Banyumas as the object of data acquisition. Thank you very much to the reviewers and editors, which helps improve the overall quality of ideas, concepts, and papers.</p>
</ack>
<sec><title>Funding Statement</title>
<p>This work is supported by the Center for Higher Education Funding (BPPT) and the Indonesia Endowment Fund for Education (LPDP), as acknowledged in decree number 02092/J5.2.3/BPI.06/9/2022.</p>
</sec>
<sec><title>Author Contributions</title>
<p>The contributions of the papers in this study are very diverse. Farida Asriani is responsible for data collection, research methodology design, experimental design, data analysis, and research reports. Azhari Azhari provides research concepts and planning, supervision, and correspondence of authors. Wahyono Wahyono plays a role in guiding and supervising research. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability"><title>Availability of Data and Materials</title>
<p>The datasets used and/or analyzed during the current study are available from the corresponding author upon reasonable request.</p>
</sec>
<sec><title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Liang</surname></string-name> and <string-name><given-names>T. E.</given-names> <surname>Nyamasvisva</surname></string-name></person-group>, &#x201C;<article-title>Badminton action classification based on human skeleton data extracted by AlphaPose</article-title>,&#x201D; in <conf-name>Int. Conf. Sens. Meas. Data Anal. Era Artif. Intell.</conf-name>, <year>2023</year>. doi: <pub-id pub-id-type="doi">10.1109/ICSMD60522.2023.10490491</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X. -W.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Ruan</surname></string-name>, <string-name><given-names>S. -S.</given-names> <surname>Yu</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Lai</surname></string-name>, <string-name><given-names>Z. -F.</given-names> <surname>Li</surname></string-name> and <string-name><given-names>W. -T.</given-names> <surname>Chen</surname></string-name></person-group>, &#x201C;<article-title>Badminton action classification based on PDDRNe</article-title>,&#x201D; in <source>3rd Int. Conf. Internet, Educ. Inf. Technol.</source>, vol. <volume>10</volume>, no. <issue>17</issue>, pp. <fpage>980</fpage>&#x2013;<lpage>987</lpage>, <year>2023</year>. doi: <pub-id pub-id-type="doi">10.2991/978-94-6463-230-9_118</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Liu</surname></string-name> and <string-name><given-names>B.</given-names> <surname>Liang</surname></string-name></person-group>, &#x201C;<article-title>An action recognition technology for badminton players using deep learning</article-title>,&#x201D; <source>Mobile Inf. Syst.</source>, vol. <volume>2022</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>10</lpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.1155/2022/3413584</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N. A.</given-names> <surname>Rahmad</surname></string-name>, <string-name><given-names>N. A. J.</given-names> <surname>Sufri</surname></string-name>, <string-name><given-names>M. A.</given-names> <surname>As&#x2019;Ari</surname></string-name>, and <string-name><given-names>A.</given-names> <surname>Azaman</surname></string-name></person-group>, &#x201C;<article-title>Recognition of badminton action using convolutional neural network</article-title>,&#x201D; <source>Indonesian J. Elect. Eng. Inf.</source>, vol. <volume>7</volume>, no. <issue>4</issue>, pp. <fpage>750</fpage>&#x2013;<lpage>756</lpage>, <year>Dec. 2019</year>. doi: <pub-id pub-id-type="doi">10.11591/ijeei.v7i4.968</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Ramasinghe</surname></string-name>, <string-name><given-names>K. G. M.</given-names> <surname>Chathuramali</surname></string-name>, and <string-name><given-names>R.</given-names> <surname>Rodrigo</surname></string-name></person-group>, &#x201C;<article-title>Recognition of badminton strokes using dense trajectories</article-title>,&#x201D; in <conf-name>Int. Conf. Inf. Automat. Sustain.</conf-name>, <publisher-name>IEEE</publisher-name>, <year>Mar. 2014</year>, pp. <fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi: <pub-id pub-id-type="doi">10.1109/ICIAFS.2014.7069620</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Feng</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Shi</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Zhou</surname></string-name>, and <string-name><given-names>L.</given-names> <surname>Mo</surname></string-name></person-group>, &#x201C;<article-title>Research on human activity recognition based on random forest classifier</article-title>,&#x201D; in <conf-name>Int. Conf. Control, Electron. Comput. Technol.</conf-name>, <publisher-name>IEEE</publisher-name>, <year>2023</year>, pp. <fpage>1507</fpage>&#x2013;<lpage>1513</lpage>. doi: <pub-id pub-id-type="doi">10.1109/ICCECT57938.2023.10140545</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Zaki</surname></string-name></person-group>, &#x201C;<article-title>Logistic regression based human activities recognition</article-title>,&#x201D; <source>J. Mech. Contin. Math. Sci.</source>, vol. <volume>15</volume>, no. <issue>4</issue>, pp. <fpage>228</fpage>&#x2013;<lpage>246</lpage>, <year>Apr. 2020</year>. doi: <pub-id pub-id-type="doi">10.26782/jmcms.2020.04.00018</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Wang</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Deep 3D human pose estimation: A review</article-title>,&#x201D; <source>Comput. Vis. Imag. Underst.</source>, vol. <volume>210</volume>, no. <issue>2</issue>, <year>Sep. 2021, Art. no. 103225</year>. doi: <pub-id pub-id-type="doi">10.1016/j.cviu.2021.103225</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Cui</surname></string-name> and <string-name><given-names>N.</given-names> <surname>Dahnoun</surname></string-name></person-group>, &#x201C;<article-title>Real-Time short-range human posture estimation using mmwave radars and neural networks</article-title>,&#x201D; <source>IEEE Sens. J.</source>, vol. <volume>22</volume>, no. <issue>1</issue>, pp. <fpage>535</fpage>&#x2013;<lpage>543</lpage>, <year>Jan. 2022</year>. doi: <pub-id pub-id-type="doi">10.1109/JSEN.2021.3127937</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Tasnim</surname></string-name>, <string-name><given-names>M. K.</given-names> <surname>Islam</surname></string-name>, and <string-name><given-names>J. H.</given-names> <surname>Baek</surname></string-name></person-group>, &#x201C;<article-title>Deep learning based human activity recognition using spatio-temporal image formation of skeleton joints</article-title>,&#x201D; <source>Appl. Sci.</source>, vol. <volume>11</volume>, no. <issue>6</issue>, <year>Mar. 2021, Art. no. 2675</year>. doi: <pub-id pub-id-type="doi">10.3390/app11062675</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Ramirez</surname></string-name>, <string-name><given-names>S. A.</given-names> <surname>Velastin</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Aguayo</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Fabregas</surname></string-name>, and <string-name><given-names>G.</given-names> <surname>Farias</surname></string-name></person-group>, &#x201C;<article-title>Human activity recognition by sequences of skeleton features</article-title>,&#x201D; <source>J. Sens.</source>, vol. <volume>22</volume>, no. <issue>11</issue>, <year>Jun. 2022, Art. no. 3991</year>. doi: <pub-id pub-id-type="doi">10.3390/s22113991</pub-id>; <pub-id pub-id-type="pmid">35684613</pub-id></mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Morais</surname></string-name>, <string-name><given-names>V.</given-names> <surname>Le</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Tran</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Saha</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Mansour</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Venkatesh</surname></string-name></person-group>, &#x201C;<article-title>Learning regularity in skeleton trajectories for anomaly detection in videos</article-title>,&#x201D; in <conf-name>Proc. IEEE Comput. Soci. Conf. Comput. Vis. Pattern Recognit.</conf-name>, <year>2019</year>, pp. <fpage>11988</fpage>&#x2013;<lpage>11996</lpage>. doi: <pub-id pub-id-type="doi">10.1109/CVPR.2019.01227</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. S.</given-names> <surname>Alsawadi</surname></string-name>, <string-name><given-names>E. S. M.</given-names> <surname>El-Kenawy</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Rio</surname></string-name></person-group>, &#x201C;<article-title>Using blazepose on spatial temporal graph convolutional networks for action recognition</article-title>,&#x201D; <source>Comput. Mater. Contin.</source>, vol. <volume>74</volume>, no. <issue>1</issue>, pp. <fpage>19</fpage>&#x2013;<lpage>36</lpage>, <year>2023</year>. doi: <pub-id pub-id-type="doi">10.32604/cmc.2023.032499</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Giveki</surname></string-name></person-group>, &#x201C;<article-title>Robust moving object detection based on fusing Atanassov&#x2019;s Intuitionistic 3D Fuzzy Histon Roughness Index and texture features</article-title>,&#x201D; <source>Int. J. Approx. Reasoning</source>, vol. <volume>135</volume>, no. <issue>18</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>20</lpage>, <year>Aug. 2021</year>. doi: <pub-id pub-id-type="doi">10.1016/j.ijar.2021.04.007</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Giveki</surname></string-name></person-group>, &#x201C;<article-title>Scale-space multi-view bag of words for scene categorization</article-title>,&#x201D; <source>Multimed. Tools Appl.</source>, vol. <volume>80</volume>, no. <issue>1</issue>, pp. <fpage>1223</fpage>&#x2013;<lpage>1245</lpage>, <year>Jan. 2021</year>. doi: <pub-id pub-id-type="doi">10.1007/s11042-020-09759-9</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Chen</surname></string-name>, and <string-name><given-names>Y.</given-names> <surname>Liu</surname></string-name></person-group>, &#x201C;<article-title>Human deep squat detection method based on MediaPipe combined with Yolov5 network</article-title>,&#x201D; in <conf-name>Proc. 41st Chin. Control Conf.</conf-name>, <publisher-loc>Hefei, China</publisher-loc>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.23919/CCC55666.2022.9902631</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Ma</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Ma</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Ruan</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Chen</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Feng</surname></string-name></person-group>, &#x201C;<article-title>A wushu posture recognition system based on MediaPipe</article-title>,&#x201D; in <conf-name>Int. Conf. Inf. Technol. Contemp. Sports</conf-name>, <publisher-name>IEEE</publisher-name>, <year>2022</year>, pp. <fpage>10</fpage>&#x2013;<lpage>13</lpage>. doi: <pub-id pub-id-type="doi">10.1109/TCS56119.2022.9918744</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. H.</given-names> <surname>Kim</surname></string-name> and <string-name><given-names>J. Y.</given-names> <surname>Chang</surname></string-name></person-group>, &#x201C;<article-title>Single-shot 3D multi-person shape reconstruction from a single RGB image</article-title>,&#x201D; <source>Entropy</source>, vol. <volume>22</volume>, no. <issue>8</issue>, <year>Aug. 2020, Art. no. 806</year>. doi: <pub-id pub-id-type="doi">10.3390/e22080806</pub-id>; <pub-id pub-id-type="pmid">33286577</pub-id></mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Han</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Jiang</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Jia</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Lu</surname></string-name></person-group>, &#x201C;<article-title>HEMlets PoSh: Learning part-centric heatmap triplets for 3D human pose and shape estimation</article-title>,&#x201D; <source>IEEE Trans. Pattern Anal. Mach. Intell.</source>, vol. <volume>44</volume>, no. <issue>6</issue>, pp. <fpage>3000</fpage>&#x2013;<lpage>3014</lpage>, <year>Jun. 2022</year>. doi: <pub-id pub-id-type="doi">10.1109/TPAMI.2021.3051173</pub-id>; <pub-id pub-id-type="pmid">33434125</pub-id></mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Gu</surname></string-name>, and <string-name><given-names>Z.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>Fitness action counting based on MediaPipe</article-title>,&#x201D; in <conf-name>Proc. Int. Congr. Imag. Signal Process., Biomed. Eng. Inf.</conf-name>, <publisher-name>IEEE</publisher-name>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.1109/CISP-BMEI56279.2022.9980337</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Bai</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Zhao</surname></string-name>, and <string-name><given-names>X.</given-names> <surname>Xiu</surname></string-name></person-group>, &#x201C;<article-title>Exploration of computer vision and image processing technology based on OpenCV</article-title>,&#x201D; in <conf-name>Proc. Int. Conf. Comput. Sci. Eng. Tech.</conf-name>, <publisher-name>IEEE</publisher-name>, <year>2022</year>, pp. <fpage>145</fpage>&#x2013;<lpage>147</lpage>. doi: <pub-id pub-id-type="doi">10.1109/SCSET55041.2022.00042</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Palani</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Panigrahi</surname></string-name>, <string-name><given-names>S. A.</given-names> <surname>Jammi</surname></string-name>, and <string-name><given-names>A.</given-names> <surname>Thondiyath</surname></string-name></person-group>, &#x201C;<article-title>Real-time joint angle estimation using Mediapipe framework and inertial sensors</article-title>,&#x201D; in <conf-name>Proc. Int. Conf. Bioinf. Bioeng.</conf-name>, <publisher-name>IEEE</publisher-name>, <year>2022</year>, pp. <fpage>128</fpage>&#x2013;<lpage>133</lpage>. doi: <pub-id pub-id-type="doi">10.1109/BIBE55377.2022.00035</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Adhikary</surname></string-name>, <string-name><given-names>A. K.</given-names> <surname>Talukdar</surname></string-name>, and <string-name><given-names>K.</given-names> <surname>Kumar Sarma</surname></string-name></person-group>, &#x201C;<article-title>A vision-based system for recognition of words used in indian sign language using MediaPipe</article-title>,&#x201D; in <conf-name>Int. Conf. Imag. Inf. Process.</conf-name>, <publisher-name>IEEE</publisher-name>, <year>2021</year>, pp. <fpage>390</fpage>&#x2013;<lpage>394</lpage>. doi: <pub-id pub-id-type="doi">10.1109/ICIIP53038.2021.9702551</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Padhi</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Das</surname></string-name></person-group>, &#x201C;<article-title>Hand gesture recognition using DenseNet201-Mediapipe hybrid modelling</article-title>,&#x201D; in <conf-name>Int. Conf. Automat. Comput. Renew. Syst.</conf-name>, <year>2022</year>, pp. <fpage>995</fpage>&#x2013;<lpage>999</lpage>. doi: <pub-id pub-id-type="doi">10.1109/ICACRS55517.2022.10029038</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Fu</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Spatiotemporal correlation based self-adaptive pose estimation in complex scenes</article-title>,&#x201D; <source>Digit. Commun. Netw.</source>, <year>2024</year>. doi: <pub-id pub-id-type="doi">10.1016/j.dcan.2024.03.007</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Xia</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Liu</surname></string-name></person-group>, &#x201C;<article-title>BCCLR: A skeleton-based action recognition with graph convolutional network combining behavior dependence and context clues</article-title>,&#x201D; <source>Comput. Mater. Contin.</source>, vol. <volume>78</volume>, no. <issue>3</issue>, pp. <fpage>4489</fpage>&#x2013;<lpage>4507</lpage>, <year>2024</year>. doi: <pub-id pub-id-type="doi">10.32604/cmc.2024.048813</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Shahroudy</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>T. T.</given-names> <surname>Ng</surname></string-name>, and <string-name><given-names>G.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>NTU RGB&#x002B;D: A large scale dataset for 3D human activity analysis</article-title>,&#x201D; in <conf-name>Proc. IEEE Comput. Soc. Conf. Comput. Vis. Pattern Recognit.</conf-name>, <year>Dec. 2016</year>, pp. <fpage>1010</fpage>&#x2013;<lpage>1019</lpage>. doi: <pub-id pub-id-type="doi">10.1109/CVPR.2016.115</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Shahroudy</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>A. C.</given-names> <surname>Kot</surname></string-name>, and <string-name><given-names>G.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Skeleton-based action recognition using spatio-temporal LSTM network with trust gates</article-title>,&#x201D; <source>IEEE Trans. Pattern Anal. Mach. Intell.</source>, vol. <volume>40</volume>, no. <issue>12</issue>, pp. <fpage>3007</fpage>&#x2013;<lpage>3021</lpage>, <year>Dec. 2018</year>. doi: <pub-id pub-id-type="doi">10.1109/TPAMI.2017.2771306</pub-id>; <pub-id pub-id-type="pmid">29990167</pub-id></mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Yan</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Xiong</surname></string-name>, and <string-name><given-names>D.</given-names> <surname>Lin</surname></string-name></person-group>, &#x201C;<article-title>Spatial temporal graph convolutional networks for skeleton-based action recognition</article-title>,&#x201D; in <conf-name>Proc. AAAI Conf. Artif. Intell.</conf-name>, <year>2018</year>. doi: <pub-id pub-id-type="doi">10.1609/aaai.v32i1.12328</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Shi</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Cheng</surname></string-name>, and <string-name><given-names>H.</given-names> <surname>Lu</surname></string-name></person-group>, &#x201C;<article-title>Two-stream adaptive graph convolutional networks for skeleton-based action recognition</article-title>,&#x201D; in <conf-name>Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR)</conf-name>, <year>2019</year>, pp. <fpage>12026</fpage>&#x2013;<lpage>12035</lpage>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Bai</surname></string-name>, <string-name><given-names>J. Z.</given-names> <surname>Kolter</surname></string-name>, and <string-name><given-names>V.</given-names> <surname>Koltun</surname></string-name></person-group>, &#x201C;<article-title>An empirical evaluation of generic convolutional and recurrent networks for sequence modeling</article-title>,&#x201D; <comment>Mar. 2018, <italic>arXiv:1803.01271</italic></comment>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Koniusz</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Wang</surname></string-name>, and <string-name><given-names>A.</given-names> <surname>Cherian</surname></string-name></person-group>, &#x201C;<article-title>Tensor representations for action recognition</article-title>,&#x201D; <source>IEEE Trans. Pattern Anal. Mach. Intell.</source>, vol. <volume>44</volume>, no. <issue>2</issue>, pp. <fpage>648</fpage>&#x2013;<lpage>665</lpage>, <year>Feb. 2022</year>. doi: <pub-id pub-id-type="doi">10.1109/TPAMI.2021.3107160</pub-id>; <pub-id pub-id-type="pmid">34428136</pub-id></mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>W. T.</given-names> <surname>Chu</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Situmeang</surname></string-name></person-group>, &#x201C;<article-title>Badminton video analysis based on spatiotemporal and stroke features</article-title>,&#x201D; in <conf-name>Proc. Int. Conf. Multimed. Retr.</conf-name>, <year>2017</year>, pp. <fpage>448</fpage>&#x2013;<lpage>451</lpage>. doi: <pub-id pub-id-type="doi">10.1145/3078971.3079032</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Long</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Rong</surname></string-name></person-group>, &#x201C;<article-title>Application of machine learning to badminton action decomposition teaching</article-title>,&#x201D; <source>Wirel. Commun. Mobile Comput.</source>, vol. <volume>2022</volume>, no. <issue>9</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>10</lpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.1155/2022/3707407</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N.</given-names> <surname>Tanaka</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Shishido</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Suita</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Nishijima</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Kamed</surname></string-name> and <string-name><given-names>I.</given-names> <surname>Kitahara</surname></string-name></person-group>, &#x201C;<article-title>Detection of shot information using footwork trajectory and skeletal information of badminton players</article-title>,&#x201D; <source>Int. Conf. Sport Sci. Res. Technol. Support</source>, pp. <fpage>112</fpage>&#x2013;<lpage>119</lpage>, <year>2023</year>. doi: <pub-id pub-id-type="doi">10.5220/0012162700003587</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Chu</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Li</surname></string-name></person-group>, &#x201C;<article-title>Image recognition of badminton swing motion based on single inertial Sensor</article-title>,&#x201D; <source>J. Sens.</source>, vol. <volume>2021</volume>, no. <issue>1</issue>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.1155/2021/3736923</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Karim</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Azhari</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Shahroz</surname></string-name>, <string-name><given-names>S. B.</given-names> <surname>Belhaouri</surname></string-name>, and <string-name><given-names>K.</given-names> <surname>Mustofa</surname></string-name></person-group>, &#x201C;<article-title>LDSVM: Leukemia cancer classification using machine learning</article-title>,&#x201D; <source>Comput. Mater. Contin.</source>, vol. <volume>71</volume>, no. <issue>2</issue>, pp. <fpage>3887</fpage>&#x2013;<lpage>3903</lpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.32604/cmc.2022.021218</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Karray</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Triki</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Ksantini</surname></string-name></person-group>, &#x201C;<article-title>A new speed limit recognition methodology based on ensemble learning: Hardware validation</article-title>,&#x201D; <source>Comput. Mater. Contin.</source>, vol. <volume>80</volume>, no. <issue>1</issue>, pp. <fpage>119</fpage>&#x2013;<lpage>138</lpage>, <year>2024</year>. doi: <pub-id pub-id-type="doi">10.32604/cmc.2024.051562</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Pan</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Xiong</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Xu</surname></string-name></person-group>, &#x201C;<article-title>Human activity recognition based on the combined SVM&#x0026;HMM</article-title>,&#x201D; in <conf-name>Int. Conf. Inf. Automat.</conf-name>, <publisher-loc>Hailar, China</publisher-loc>, <year>2014</year>. doi: <pub-id pub-id-type="doi">10.1109/ICInfA.2014.6932656</pub-id>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Jindal</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Sachdeva</surname></string-name>, <string-name><given-names>A. K. S.</given-names> <surname>Kushwaha</surname></string-name>, and <string-name><given-names>I. K.</given-names> <surname>Gujral</surname></string-name></person-group>, &#x201C;<article-title>Performance evaluation of machine learning based voting classifier system for human activity recognition</article-title>,&#x201D; <source>Kuwait J. Sci.</source>, pp. <fpage>1</fpage>&#x2013;<lpage>12</lpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.48129/kjs.splml.19189</pub-id>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Das</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Maity</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Jana</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Biswas</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Biswas</surname></string-name> and <string-name><given-names>P. K.</given-names> <surname>Samanta</surname></string-name></person-group>, &#x201C;<article-title>Automated improved human activity recognition using ensemble modeling</article-title>,&#x201D; in <conf-name>Int. Conf. Recent Adv. Elect. Electron. Ubiquitous Commun. Comput. Intell.</conf-name>, <publisher-name>IEEE</publisher-name>, <year>2024</year>, pp. <fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi: <pub-id pub-id-type="doi">10.1109/RAEEUCCI61380.2024.10547875</pub-id>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T. H.</given-names> <surname>Tan</surname></string-name>, <string-name><given-names>J. Y.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>S. H.</given-names> <surname>Liu</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Gochoo</surname></string-name></person-group>, &#x201C;<article-title>Human activity recognition using an ensemble learning algorithm with smartphone sensor data</article-title>,&#x201D; <source>Electronics</source>, vol. <volume>11</volume>, no. <issue>3</issue>, <year>Feb. 2022</year>. doi: <pub-id pub-id-type="doi">10.3390/electronics11030322</pub-id>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Kaur</surname></string-name> and <string-name><given-names>D.</given-names> <surname>Veer Sharma</surname></string-name></person-group>, &#x201C;<article-title>Human action recognition using an ensemble deep learning model for video datasets</article-title>,&#x201D; <source>J. Harbin Eng. Univ.</source>, vol. <volume>44</volume>, no. <issue>7</issue>, pp. <fpage>1006</fpage>&#x2013;<lpage>7043</lpage>, <year>2023</year>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H. Y.</given-names> <surname>Ting</surname></string-name>, <string-name><given-names>K. S.</given-names> <surname>Sim</surname></string-name>, and <string-name><given-names>F. S.</given-names> <surname>Abas</surname></string-name></person-group>, &#x201C;<article-title>Automatic badminton action recognition using RGB-D sensor</article-title>,&#x201D; <source>Adv. Mater. Res.</source>, vol. <volume>1042</volume>, pp. <fpage>89</fpage>&#x2013;<lpage>93</lpage>, <year>2014</year>. doi: <pub-id pub-id-type="doi">10.4028/www.scientific.net/AMR.1042.89</pub-id>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M. A. I.</given-names> <surname>Anik</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Hassan</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Mahmud</surname></string-name>, and <string-name><given-names>M. K.</given-names> <surname>Hasan</surname></string-name></person-group>, &#x201C;<article-title>Activity recognition of a badminton game through accelerometer and gyroscope</article-title>,&#x201D; in <conf-name>Int. Conf. Comput. Inf. Technol.</conf-name>, <publisher-name>IEEE</publisher-name>, <year>Feb. 2017</year>, pp. <fpage>213</fpage>&#x2013;<lpage>217</lpage>. doi: <pub-id pub-id-type="doi">10.1109/ICCITECHN.2016.7860197</pub-id>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>N. F.</given-names> <surname>Ghazali</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Shahar</surname></string-name>, and <string-name><given-names>M. A.</given-names> <surname>As&#x2019;ari</surname></string-name></person-group>, &#x201C;<article-title>Badminton strokes recognition using inertial sensor and machine learning approach</article-title>,&#x201D; in <conf-name>Int. Conf. Intell. Cybern. Technol. Appl.</conf-name>, <publisher-name>IEEE</publisher-name>, <year>2022</year>, pp. <fpage>1</fpage>&#x2013;<lpage>5</lpage>. doi: <pub-id pub-id-type="doi">10.1109/ICICyTA57421.2022.10037897</pub-id>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Guo</surname></string-name>, and <string-name><given-names>C.</given-names> <surname>Zhao</surname></string-name></person-group>, &#x201C;<article-title>Badminton stroke recognition based on body sensor networks</article-title>,&#x201D; <source>IEEE Trans. Hum. Mach. Syst.</source>, vol. <volume>46</volume>, no. <issue>5</issue>, pp. <fpage>769</fpage>&#x2013;<lpage>775</lpage>, <year>Oct. 2016</year>. doi: <pub-id pub-id-type="doi">10.1109/THMS.2016.2571265</pub-id>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N. A.</given-names> <surname>Rahmad</surname></string-name> and <string-name><given-names>M. A.</given-names> <surname>As&#x2019;ari</surname></string-name></person-group>, &#x201C;<article-title>The new Convolutional Neural Network (CNN) local feature extractor for automated badminton action recognition on vision-based data</article-title>,&#x201D; <source>Int. Conf. Emerg. Comput. Technol. Sport.</source>, vol. <volume>1529</volume>, no. <issue>2</issue>, <year>Jun. 2020</year>. doi: <pub-id pub-id-type="doi">10.1088/1742-6596/1529/2/022021</pub-id>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N. A.</given-names> <surname>Rahmad</surname></string-name>, <string-name><given-names>M. A.</given-names> <surname>As&#x2019;Ari</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Soeed</surname></string-name>, and <string-name><given-names>I.</given-names> <surname>Zulkapri</surname></string-name></person-group>, &#x201C;<article-title>Automated badminton smash recognition using convolutional neural network on the vision-based data</article-title>,&#x201D; <source>Mater. Sci. Eng.</source>, vol. <volume>884</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>7</lpage>, <year>2020</year>. doi: <pub-id pub-id-type="doi">10.1088/1757-899X/884/1/012009</pub-id>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Steels</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Van Herbruggen</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Fontaine</surname></string-name>, <string-name><given-names>T.</given-names> <surname>De Pessemier</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Plets</surname></string-name> and <string-name><given-names>E.</given-names> <surname>De Poorter</surname></string-name></person-group>, &#x201C;<article-title>Badminton activity recognition using accelerometer data</article-title>,&#x201D; <source>J. Sens.</source>, vol. <volume>20</volume>, no. <issue>17</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>16</lpage>, <year>Sep. 2020</year>. doi: <pub-id pub-id-type="doi">10.3390/s20174685</pub-id>; <pub-id pub-id-type="pmid">32825134</pub-id></mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Ke</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Rahmani</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Bennamoun</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Wang</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Liu</surname></string-name></person-group>, &#x201C;<article-title>Human action recognition from various data modalities: A review</article-title>,&#x201D; <source>IEEE Trans. Pattern Anal. Mach. Intell.</source>, vol. <volume>45</volume>, pp. <fpage>3200</fpage>&#x2013;<lpage>3225</lpage>, <year>Mar. 2023</year>. doi: <pub-id pub-id-type="doi">10.1109/TPAMI.2022.3183112</pub-id>; <pub-id pub-id-type="pmid">35700242</pub-id></mixed-citation></ref>
</ref-list>
</back></article>