<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">SDHM</journal-id>
<journal-id journal-id-type="nlm-ta">SDHM</journal-id>
<journal-id journal-id-type="publisher-id">SDHM</journal-id>
<journal-title-group>
<journal-title>Structural Durability &#x0026; Health Monitoring</journal-title>
</journal-title-group>
<issn pub-type="epub">1930-2991</issn>
<issn pub-type="ppub">1930-2983</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">69611</article-id>
<article-id pub-id-type="doi">10.32604/sdhm.2025.069611</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Automatic Potential Safety Hazard Detection for High-Speed Railroad Surrounding Environment Using Lightweight Hybrid Dual Tasks Architecture</article-title>
<alt-title alt-title-type="left-running-head">Automatic Potential Safety Hazard Detection for High-Speed Railroad Surrounding Environment Using Lightweight Hybrid Dual Tasks Architecture</alt-title>
<alt-title alt-title-type="right-running-head">Automatic Potential Safety Hazard Detection for High-Speed Railroad Surrounding Environment Using Lightweight Hybrid Dual Tasks Architecture</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Zhao</surname><given-names>Zheda</given-names></name></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Xu</surname><given-names>Tao</given-names></name></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Yang</surname><given-names>Tong</given-names></name></contrib>
<contrib id="author-4" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Wu</surname><given-names>Yunpeng</given-names></name><email>wuyunpeng@bjtu.edu.cn</email></contrib>
<contrib id="author-5" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Guo</surname><given-names>Fengxiang</given-names></name><email>guofengxiang@kust.edu.cn</email></contrib>
<aff id="aff-1">
<institution>The Faculty of Transportation Engineering, Kunming University of Science and Technology</institution>, <addr-line>Kunming, 650031</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Authors: Yunpeng Wu. Email: <email>wuyunpeng@bjtu.edu.cn</email>; Fengxiang Guo. Email: <email>guofengxiang@kust.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2025</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>17</day><month>11</month><year>2025</year>
</pub-date>
<volume>19</volume>
<issue>6</issue>
<fpage>1457</fpage>
<lpage>1472</lpage>
<history>
<date date-type="received">
<day>26</day>
<month>6</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>13</day>
<month>8</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 The Authors.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="_SDHM_69611.pdf"></self-uri>
<abstract>
<p>Utilizing unmanned aerial vehicle (UAV) photography to timely detect and evaluate potential safety hazards (PSHs) around high-speed rail has great potential to complement and reform the existing manual inspections by providing better overhead views and mitigating safety issues. However, UAV inspections based on manual interpretation, which heavily rely on the experience, attention, and judgment of human inspectors, still inevitably suffer from subjectivity and inaccuracy. To address this issue, this study proposes a lightweight hybrid learning algorithm named HDTA (hybrid dual tasks architecture) to automatically and efficiently detect the PSHs of UAV imagery. First, this HDTA architecture seamlessly integrates both detection and segmentation branches within a unified framework. This design enables the model to simultaneously perform PSH detection and railroad parsing, thereby providing comprehensive scene understanding. Such joint learning also lays the foundation for PSH assessment tasks. Second, an innovative lightweight backbone based on the shuffle selective state space model (S<sup>4</sup>M) is incorporated into HDTA. The state space model approach allows for global contextual information extraction while maintaining linear computational complexity. Furthermore, the incorporation of shuffle operation facilitates more efficient information flow across feature dimensions, enhancing both feature representation and fusion capabilities. Finally, extensive experiments conducted on a railroad environment dataset constructed from UAV imagery demonstrate that the proposed method achieves high detection accuracy while maintaining efficiency and practicality.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Railroad inspection</kwd>
<kwd>hybrid architecture</kwd>
<kwd>drone image</kwd>
<kwd>pixel-level segmentation</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>National Natural Science Foundation of China</funding-source>
<award-id>52362048</award-id>
</award-group>
<award-group id="awg2">
<funding-source>Yunnan Fundamental Research Projects</funding-source>
<award-id>202301BE070001-042</award-id>
<award-id>202401AT070409</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>With the rapid expansion of high-speed railways, potential safety hazards (PSHs) along the tracks&#x2014;such as foreign objects&#x2014;can be blown or displaced onto the railway due to adverse weather conditions or strong winds, posing serious threats to safe railway operations. For instance, in 2021, over ten incidents involving foreign objects like plastic tarps and covered nets entering high-speed rail tracks led to numerous train delays and cancellations in China, causing significant economic losses [<xref ref-type="bibr" rid="ref-1">1</xref>]. Therefore, the timely and accurate detection of the surrounding environment of high-speed railways is essential for ensuring the safe operation of railway systems. At present, PSH detection primarily relies on manual inspection, which is labor-intensive and inefficient. Fortunately, aerial inspection using unmanned aerial vehicles (UAVs) has emerged as a promising solution. However, manually interpreting UAV imagery to identify PSH near railway tracks remains time-consuming and subjectivity. Consequently, there is an urgent need to develop an automated visual inspection (VI) method capable of accurately detecting PSH and interpreting both tracks and surrounding infrastructure in UAV imagery.</p>
<p>In recent years, convolutional neural networks (CNNs), known for their ability to extract deep features, have been extensively utilized across a range of vision tasks. Representative object detection frameworks include two stage models such as RCNN series [<xref ref-type="bibr" rid="ref-2">2</xref>,<xref ref-type="bibr" rid="ref-3">3</xref>], as well as prominent one stage detectors like SSD [<xref ref-type="bibr" rid="ref-4">4</xref>], YOLOv4 [<xref ref-type="bibr" rid="ref-5">5</xref>], YOLOv5 [<xref ref-type="bibr" rid="ref-6">6</xref>], YOLOv7 [<xref ref-type="bibr" rid="ref-7">7</xref>], YOLOv8 [<xref ref-type="bibr" rid="ref-8">8</xref>]. Compared to two-stage approaches, the YOLO series incorporates more sophisticated designs in terms of network structure, anchor strategies, loss formulations, and data augmentation techniques, leading to superior detection accuracy and real-time performance. Notably, YOLOv10 integrates an NMS-free (Non-Maximum Suppression) and end-to-end processing mechanism, achieving competitive performance and high inference speed. With regard to saliency detection, classical networks include DeepLabv3&#x002B; [<xref ref-type="bibr" rid="ref-9">9</xref>], MobileNetv3 [<xref ref-type="bibr" rid="ref-10">10</xref>], GCnet [<xref ref-type="bibr" rid="ref-11">11</xref>], SegFormer [<xref ref-type="bibr" rid="ref-12">12</xref>], Mask2former [<xref ref-type="bibr" rid="ref-13">13</xref>] and other methods [<xref ref-type="bibr" rid="ref-14">14</xref>&#x2013;<xref ref-type="bibr" rid="ref-16">16</xref>]. Although the aforementioned models perform well in individual detection or segmentation tasks, their specialized design for single tasks and lack of distance information make them unsuitable for distance-based evaluation of PSHs in high-speed railroad environments.</p>
<p>To date, numerous CNN-based approaches have been proposed for railroad safety and infrastructure monitoring. Reference [<xref ref-type="bibr" rid="ref-17">17</xref>] proposed a defect detection method called ELA-YOLO based on YOLOv8, which integrates linear attention for better representation with lower complexity, a selective feature pyramid network for improved multi-level feature fusion, and a lightweight detection head for efficient output. Reference [<xref ref-type="bibr" rid="ref-18">18</xref>] developed a spatial&#x2013;temporal neural network model based on ResNet-Transformer architecture to predict the occurrence of broken rails one year in advance. Furthermore, research on engineering applications based on UAV imagery has also been widely developed, reference [<xref ref-type="bibr" rid="ref-19">19</xref>] presented a UAV and Mixed Reality (MR)-based inspection method to enhance efficiency and personnel involvement. It builds an MR digital environment for semi-automatic UAV control, defect data collection, and visual management. Experiments show the method effectively supports exterior wall defect inspection and provides a foundation for advancing MMESE. Reference [<xref ref-type="bibr" rid="ref-20">20</xref>] proposed CDCR-ISeg, an interactive segmentation model supported by a new 1500-image dataset (UAV-CrackX4, X8, X16) covering diverse zoom levels and domains. The model integrates super-resolution and domain adaptation to boost generalization and reduce annotation effort, and uses a vector map with reversed directions to enhance boundary detection. Reference [<xref ref-type="bibr" rid="ref-21">21</xref>] proposed PDIS-Net, an instance segmentation framework for UAV-based pavement distress detection. It uses dynamic convolution by predicting kernel positions and weights, refines them through metric learning and fusion, and applies them to feature maps to generate accurate instance masks.</p>
<p>In the field of PSH detection around railways, reference [<xref ref-type="bibr" rid="ref-22">22</xref>] developed an automatic real-time instance segmentation framework called YOLARC (You Only Look at Railroad Coefficients), which utilizes UAV imagery for high-speed railway monitoring. Building upon [<xref ref-type="bibr" rid="ref-22">22</xref>], reference [<xref ref-type="bibr" rid="ref-23">23</xref>] proposed the YOLARC&#x002B;&#x002B; framework for PSH detection in UAV images. This framework incorporates a Large Selective Kernel (LSK) module that adaptively adjusts the receptive field and mitigates the impact of scale variation. However, applying instance segmentation to object detection introduces unnecessary computational overhead. To address this issue, reference [<xref ref-type="bibr" rid="ref-24">24</xref>] proposed an all-in-one YOLO (AOYOLO) framework for multi-task railway component detection. AOYOLO employs a ConvNeXt-based backbone to extract super features and integrates a U-shaped salient object segmentation branch to enhance railway surface defect (RSD) segmentation and component detection. Furthermore, reference [<xref ref-type="bibr" rid="ref-25">25</xref>] introduced YOLORS, a hybrid learning architecture designed to segment railroad from drone images and assess PSHs along the railway. It features a stripe-pooling-based segment branch for pixel-level tasks and introduces a novel object-guided mosaic data augmentation method to improve small-object detection performance. However, these models generally require a large number of parameters and incur high computational costs for railway environment detection, which limits their usability and practical deployment.</p>
<p>Specifically, applying deep learning techniques for the timely and accurate detection of potential safety hazards (PSH) surrounding high-speed railways presents the following challenges: 1) Balancing computational cost and accuracy. Existing models often require substantial computational resources to achieve high detection accuracy. However, their large parameter sizes and high computational complexity significantly hinder deployment on edge devices, limiting their practical engineering utility. 2) Lack of global modeling capability. Traditional CNN-based methods struggle to capture long-range contextual dependencies. While Transformer-based approaches offer improved global modeling, their inherent quadratic computational complexity severely reduces efficiency and practical applicability.</p>
<p>To address these limitations, we propose a novel hybrid learning framework named HDTA (Hybrid Dual-Task Architecture), designed for accurate parsing of UAV imagery to detect railway lanes and tracks, while simultaneously identifying PSH through a unified architecture that integrates a pixel-level sparse segmentation branch and an object detection branch. Specifically, HDTA is equipped with a lightweight backbone built upon S<sup>4</sup>M (Shuffled Selective State Space Module) blocks. The S<sup>4</sup>M module first leverages the global modeling capability of the Selective State Space Model (SSM) to extract long-range contextual features from UAV images, and then applies a channel-wise shuffle operation to reduce computational complexity and enhance multi-source feature fusion. This design enables the S<sup>4</sup>M-Backbone to achieve an optimal trade-off between robust global feature extraction and computational efficiency.</p>
<p>Extensive experiments demonstrate that HDTA delivers outstanding performance in both pixel-level parsing and object detection tasks. The primary contributions of this research are as follows:
<list list-type="simple">
<list-item><label>1)</label><p>This study introduces a hybrid learning architecture, named HDTA (Hybrid Dual Tasks Architecture), to simultaneously segment railroad and detect PSHs along high-speed railways in UAV imagery, laying the foundation for PSH evaluation task.</p></list-item>
<list-item><label>2)</label><p>An innovative lightweight backbone network based on the shuffle selective state space model(S<sup>4</sup>M) has been integrated into HDTA to achieve efficient and high-accuracy PSH detection and track segmentation.</p></list-item>
<list-item><label>3)</label><p>Extensive experiments on the railroad environment dataset created using drone images demonstrate that the proposed system achieves high detection rates while maintaining efficiency and convenience.</p></list-item>
</list></p>
<p>The remaining content of the article is organized as follows: The proposed method is elaborated in <xref ref-type="sec" rid="s2">Section 2</xref>. The related experiments and results are presented in <xref ref-type="sec" rid="s3">Section 3</xref>. The conclusion is described at the end of the paper.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Methods</title>
<p>Recently, networks based on selective state spaces [<xref ref-type="bibr" rid="ref-26">26</xref>&#x2013;<xref ref-type="bibr" rid="ref-28">28</xref>] have emerged as strong competitors to conventional CNN and Transformer architectures, owing to their ability to model long-range dependencies with linear computational complexity inherent to the state space model (SSM). Inspired by these advances, this study introduces a PSH detection system built upon a novel &#x201C;SSM-with-shuffle&#x201D; framework, which enables simultaneous detection of PSHs along railroad tracks and parsing of both tracks and roads.</p>
<sec id="s2_1">
<label>2.1</label>
<title>HDTA Network</title>
<p><xref ref-type="fig" rid="fig-1">Fig. 1</xref> illustrates the overall architecture of the proposed HDTA. HDTA integrates a backbone network built upon Shuffle Selective State Space Model (S<sup>4</sup>M) blocks with detection and segmentation branches based on the C3Ghost module. This design enables the extraction of two-scale high-level features for object detection and facilitates pixel-level segmentation of railroad.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>The architecture of HDTA. Note the C3Ghost enclosed by the dashed line serves as the starting point of parsing branch. <bold>(a)</bold> The Shuffle Selective State Space Model (S<sup>4</sup>M) block. <bold>(b)</bold> SSM block. <bold>(c)</bold> C3Ghost block</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_69611-fig-1.tif"/>
</fig>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Backbone Using Shuffle Selective State Space Model (S<sup><bold><italic>4</italic></bold></sup>M)</title>
<p>The backbone section in <xref ref-type="fig" rid="fig-1">Fig. 1</xref> presents the detailed composition of the S<sup>4</sup>M-based backbone network. This backbone consists of one downsampling layer followed by three S<sup>4</sup>M blocks, each of which incorporates a built-in downsampling function. The input image is progressively downsampled through the backbone network. <xref ref-type="fig" rid="fig-1">Fig. 1a</xref> illustrates the overall structure of the S<sup>4</sup>M (Shuffle Selective State Space Model)-block. Each S<sup>4</sup>M-block consists of a SSM block, DWConv, Batch Normalization layer, and an activation function. The information flow within the S<sup>4</sup>M-block is determined by the downsampling ratio. The extracted feature space is then processed with a shuffle operation to promote feature fusion and enable lightweight computation. Specifically, for the S<sup>4</sup>M block (stride &#x003D; 2), given an input <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msup><mml:mi>M</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, where B, H, W, and C denote the batch size, height, width and number of channels, its output can be expressed as:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msup><mml:mi>M</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>L</mml:mi><mml:mi>U</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>B</mml:mi><mml:mi>N</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>D</mml:mi><mml:mi>W</mml:mi><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>S</mml:mi><mml:mi>S</mml:mi><mml:mi>M</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>D</mml:mi><mml:mi>W</mml:mi><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msup><mml:mi>M</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msup><mml:mi>M</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi>H</mml:mi><mml:mn>2</mml:mn></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi>W</mml:mi><mml:mn>2</mml:mn></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mn>2</mml:mn><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. Subsequently, <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msup><mml:mi>M</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is passed through a shuffle operation:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msup><mml:mi>M</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mi>h</mml:mi><mml:mi>u</mml:mi><mml:mi>f</mml:mi><mml:mi>f</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msup><mml:mi>M</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where the number of groups for the shuffle operation is set to 4. <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mi>S</mml:mi><mml:mrow><mml:mn>4</mml:mn></mml:mrow></mml:msup><mml:mi>M</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi>H</mml:mi><mml:mn>2</mml:mn></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi>W</mml:mi><mml:mn>2</mml:mn></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mn>2</mml:mn><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. <xref ref-type="fig" rid="fig-1">Fig. 1b</xref> illustrates the detailed structure of the SSM block. The SSM block integrates a normalization layer and a core SSM process unit, and employs extensive residual connections to enhance long-range contextual feature extraction. Specifically, for the SSM block, given an input <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msubsup><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>in</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>SSM</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>B</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>H</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>C</mml:mtext></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula>, where B, H, W, and C denote the batch size, height, width and number of channels, its output can be expressed as:
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>S</mml:mi><mml:mi>S</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>S</mml:mi><mml:mi>S</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:mi>S</mml:mi><mml:mi>S</mml:mi><mml:msup><mml:mi>M</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>N</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow><mml:mrow><mml:mi>S</mml:mi><mml:mi>S</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi>S</mml:mi><mml:mi>S</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>S</mml:mi><mml:mi>S</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:mi>S</mml:mi><mml:mi>S</mml:mi><mml:msup><mml:mi>M</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>N</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>m</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>S</mml:mi><mml:mi>S</mml:mi><mml:mi>M</mml:mi></mml:mrow></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>The <italic>SSM</italic> process module adopts 2D Selective Scanning [<xref ref-type="bibr" rid="ref-28">28</xref>] to bridge the gap between the sequential na-ture of 1D selective scanning and the non-sequential structure of 2D visual data.</p>
<p>The SPPF module can also be incorporated into backbone to enhance the network&#x2019;s representational capacity by aggregating multi-scale pooled features. This module has been validated in various versions of YOLO (v8&#x2013;v12), where it is integrated into the backbone to collect such features. In summary, the proposed lightweight backbone leverages selective state spaces, offering strong feature extraction capabilities and effective long-range contextual modeling.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Detection Branch</title>
<p>Two detection heads of different scales were used in the detection branch to detect objects in receptive fields of varying sizes, thereby improving detection accuracy. To reduce the computational cost of the model, DWConv and C3Ghost were also utilized in the detection branch. PSH in the railroad environment primarily include garbage, plastic protective sheets, sunshades, covering nets, and temporary shelters. According to the Median Relative Area (MRA) [<xref ref-type="bibr" rid="ref-25">25</xref>], the detection categories in the dataset are categorized into small objects and ordinary objects based on the typical size of the majority of samples in each category, as shown in <xref ref-type="table" rid="table-1">Table 1</xref>. <xref ref-type="table" rid="table-1">Table 1</xref> also lists the number of samples and the range of pixel numbers about each category, the pixel count for each category is calculated within images of size 640 &#x00D7; 640. During the training process of HDTA, the input images are first resized to 640 &#x00D7; 640 to improve inference efficiency, and then fed into the model, followed by four downsampling operations. Accordingly, inputting a 640 &#x00D7; 640 image yields dual prediction grids at resolutions of 80 &#x00D7; 80 and 40 &#x00D7; 40.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Size classification, sample numbers and range of pixel numbers of categories</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Category</th>
<th>Object size</th>
<th>number of objects</th>
<th>Range of pixel numbers</th>
</tr>
</thead>
<tbody>
<tr>
<td>Vehicle</td>
<td>Small</td>
<td>6676</td>
<td>6&#x2013;14,665</td>
</tr>
<tr>
<td>Temporary dwelling</td>
<td>Ordinary</td>
<td>3561</td>
<td>18&#x2013;103,902</td>
</tr>
<tr>
<td>Building</td>
<td>Ordinary</td>
<td>889</td>
<td>69&#x2013;26,710</td>
</tr>
<tr>
<td>Rubbish</td>
<td>Small</td>
<td>613</td>
<td>27&#x2013;38,398</td>
</tr>
<tr>
<td>Covered-mash</td>
<td>Ordinary</td>
<td>140</td>
<td>40&#x2013;164,422</td>
</tr>
<tr>
<td>Black plastic</td>
<td>Small</td>
<td>241</td>
<td>10&#x2013;52,670</td>
</tr>
<tr>
<td>White plastic</td>
<td>Small</td>
<td>256</td>
<td>26&#x2013;10,815</td>
</tr>
<tr>
<td>Shade net</td>
<td>Ordinary</td>
<td>467</td>
<td>6&#x2013;60,832</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Segmentation Branch</title>
<p>To perform fine-grained analysis of tracks and roads in UAV imagery with minimal extra computational burden, DWConv and C3Ghost were also applied in the segmentation branch of HDTA. This approach maintains the model&#x2019;s lightweight nature and accuracy. Attributed to the feature fusion in multi scale and long-distance modeling capabilities provided by the S<sup>4</sup>M-Backbone, the performance of track analysis is further enhanced. As shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, the 14th-layer of HDTA is selected as the starting point for image reconstruction in semantic parsing tasks.</p>
</sec>
<sec id="s2_5">
<label>2.5</label>
<title>Hybrid Multi-Task Loss Function</title>
<p>Since HDTA performs both detection and segmentation tasks, a hybrid loss function is proposed for both tasks and simultaneously constrains the ESH detection and railroad pixel-level parsing tasks. The hybrid loss can be expressed as:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>h</mml:mi><mml:mi>y</mml:mi><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03BE;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mi>D</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BE;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula>where <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>D</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is for object detection and <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is for railroad and road parsing. <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>&#x03BE;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi>&#x03BE;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> represent the weighted coefficients for object detection and parsing, respectively. <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>D</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is used to supervise the bounding box regression in ESH detection, it can be defined as:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>D</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mi>o</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula>where <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mi>o</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> leverages the CIoU loss, <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> represent the classification loss and confidence loss, respectively. The <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> loss term employs an enhanced cross-entropy (CE) function to reduce pixel-wise classification errors between the ground truth and predicted mask in track and road parsing [<xref ref-type="bibr" rid="ref-25">25</xref>].</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Experiments</title>
<p>In this section, ablation experiments are performed with backbone of various models(EfficientNet [<xref ref-type="bibr" rid="ref-29">29</xref>], StarNet [<xref ref-type="bibr" rid="ref-30">30</xref>] and Swin-TransFormer [<xref ref-type="bibr" rid="ref-31">31</xref>]), including the newly introduced S<sup>4</sup>M-Backbone within the same HDTA framework, to assess the impact of backbone choice on the architecture&#x2019;s effectiveness. Experiments between HDTA and other SOTA models were also conducted. These included detection comparisons with (DINO [<xref ref-type="bibr" rid="ref-32">32</xref>], Faster RCNN [<xref ref-type="bibr" rid="ref-2">2</xref>], YOLOv8 [<xref ref-type="bibr" rid="ref-33">33</xref>], YOLOv11 [<xref ref-type="bibr" rid="ref-34">34</xref>] and YOLOv12 [<xref ref-type="bibr" rid="ref-35">35</xref>]) and segmentation comparisons with (MoblieNetv3, GCnet, SegFormer, Mask2Former and other models) to validate the superiority of the proposed HDTA.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Dataset Description and Training Settings</title>
<p>The railroad environment drone images are taken using the DJI Mavic 3 Pro, which captures images with a focal length range from 24 to 166 mm and the maximum resolution of captured images is 8064 &#x00D7; 6048. A cruising velocity of 10&#x2013;15 m/s is maintained during drone operation. More than 3000 images are taken for the railroad environment dataset. To enable clearer visualization of object classification, <xref ref-type="fig" rid="fig-2">Fig. 2</xref> presents a visual representation of PSH categories classification. Experiments are conducted on a Linux server configured with a CPU: i7-13700, GPU: RTX 4070, and 32 GB of RAM. The training process spans 300 epochs, starting with an initial learning rate of 3.84 &#x00D7; 10<sup>&#x2212;3</sup>, which is reduced by a factor of 10 after 243 epochs. The dataset is randomly split into training, validation, and testing sets in a 7:2:1 ratio. Mosaic methods are used to data augmentation. The hyperparameters &#x03B7;1&#x2013;&#x03B7;3 are set to 0.06, 0.3, 0.95, respectively. The PR curve, mIoU, mAP0.5, mAP0.5&#x2212;0.95 and recall are the performance indicators used in the experiments.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Visualization of PSH for each category</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_69611-fig-2.tif"/>
</fig>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Ablation Experiments</title>
<p>The PR curves and mAP50 reflects hazard detection performance are employed to validate the effectiveness of the S<sup>4</sup>M-Backbone in HDTA. Additionally, the mIoU is also used to reflect the accuracy of pixel-level parsing tasks. Ablation experiments are performed with a backbone of various models (EfficientNet, StarNet and Swin-TransFormer). The computational complexity of different backbones was kept within a similar range.</p>
<p><bold>Ablation study on key modules:</bold> The ablation study uses the performance of HDTA with a backbone composed of SSM blocks only as the baseline. <xref ref-type="fig" rid="fig-3">Fig. 3</xref> illustrates the impact of key modules&#x2014;namely the shuffle operation and the SPPF module&#x2014;on the performance of the S<sup>4</sup>M-based backbone. It can be observed that as the shuffle operation and SPPF module are progressively added, HDTA demonstrates consistent improvements across various performance metrics. In terms of mIoU, the inclusion of the shuffle operation and SPPF brings segmentation performance gains of 0.53% and 0.34%, respectively. For mAP50, the progressive addition of the shuffle operation and SPPF results in detection performance improvements of 1.86% and 2.22%, respectively. The shuffle operation facilitates more effective inter-channel information exchange, while the SPPF module enhances the model&#x2019;s ability to capture multi-scale contextual information. These advantages strengthen the backbone&#x2019;s feature extraction and fusion capabilities, thereby improving the overall recognition performance of HDTA.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>MIoU, mAP50, mAP50&#x2013;95 and recall of HDTA gradually integrates base blocks on S<sup>4</sup>M backbone. <bold>(a)</bold> The impact of progressively adding basic blocks on mIoU performance for railroad segmentation. <bold>(b&#x2013;d)</bold> Effects of progressively adding basic blocks on detection performance in terms of mAP50, mAP50&#x2013;95, and Recall, respectively</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_69611-fig-3a.tif"/>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_69611-fig-3b.tif"/>
</fig>
<p><bold>Ablation study on backbone:</bold> Ablation experiments were conducted using various backbone networks (EfficientNet, StarNet and Swin Transformer), particularly focusing on the S<sup>4</sup>M backbone network proposed in HDTA. All backbone variants are trained and validated within the same HDTA framework, with all other training parameters kept identical to ensure consistency. As shown in <xref ref-type="fig" rid="fig-4">Fig. 4a</xref>, HDTA with the S<sup>4</sup>M backbone achieved the best parsing performance in railroad and road segmentation, with an mIoU of 86.5%, surpassing the EfficientNetv2, StarNet, and Swin-Transformer backbone networks by 8.38%, 1.11%, and 1.79%, respectively. Thanks to the long-sequence contextual modeling capability empowered by the SSM theory, the proposed S<sup>4</sup>M backbone network demonstrated significant advantages in pixel-level parsing tasks on image scales. In PSH detection, as shown in <xref ref-type="fig" rid="fig-4">Fig. 4b</xref>,<xref ref-type="fig" rid="fig-4">d</xref>, the proposed S<sup>4</sup>M backbone still achieved the best detection performance compared to other backbone methods within the same HDTA framework. In terms of mAP@0.5, S<sup>4</sup>M reached the highest accuracy of 79.7%, surpassing the EfficientNetv2, StarNet, and Swin-Transformer backbone by 5.0%, 4.3%, and 8.8%, respectively. In terms of mAP@0.5:0.95, S<sup>4</sup>M showed an even greater performance advantage, achieving an accuracy of 44.9%, which is higher by 4.6%, 3.72%, and 6.82% compared to the EfficientNetv2, StarNet, and Swin-Transformer backbone networks. <xref ref-type="fig" rid="fig-4">Fig. 4e</xref> presents the detection recall performance of different backbone networks. As indicated by the red line in the figure, S<sup>4</sup>M also achieved the best recall performance.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Comparison of different backbones in terms of performance, number of parameters, and computational cost. <bold>(a)</bold> Railroad and road parsing performance comparisons on mIoU. <bold>(b,d,e)</bold> PSH detect performance comparisons on mAP0.5, mAP0.5:0.95 and recall. <bold>(c,f)</bold> Comparison of computational cost, number of parameters, and performance</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_69611-fig-4.tif"/>
</fig>
<p><xref ref-type="fig" rid="fig-4">Fig. 4c</xref>,<xref ref-type="fig" rid="fig-4">f</xref> illustrates the mAP@0.5 performance of different backbone networks under varying parameter counts and computational loads. As shown in <xref ref-type="fig" rid="fig-4">Fig. 4c</xref>, S<sup>4</sup>M achieves the highest performance (79.7% mAP@0.5) with only 1.52 GFLOPs of computational cost, while the best-performing alternative, StarNet, requires 6.16 GFLOPs to reach a lower mAP@0.5 of 75.4%. In terms of parameter count, S<sup>4</sup>M delivers the best performance among all backbone networks with only 0.44 MB of parameters, achieving the highest energy efficiency across all evaluated backbones.</p>
<p>Furthermore, <xref ref-type="fig" rid="fig-5">Fig. 5</xref> highlights the advantages of HDTA equipped with the S<sup>4</sup>M backbone in terms of mAP@50 and AP@50. As shown in <xref ref-type="fig" rid="fig-5">Fig. 5a</xref>, HDTA with the S<sup>4</sup>M backbone achieves the highest mAP of 79.7% across all categories, outperforming HDTA variants using EfficientNetV2, StarNet, and Swin Transformer backbones by 5.0%, 4.3%, and 8.8%, respectively. Additionally, as shown by the red curves in <xref ref-type="fig" rid="fig-5">Fig. 5b</xref>&#x2013;<xref ref-type="fig" rid="fig-5">h</xref>, HDTA based on the S<sup>4</sup>M backbone consistently achieves the best precision-recall (PR) performance across nearly all categories when compared to other backbone methods. Moreover, HDTA with the S<sup>4</sup>M backbone shows notable performance advantages in detecting small objects such as vehicles, exceeding the performance of backbones based on EfficientNetV2, StarNet, and Swin Transformer by 3.2%, 9%, and 17%, respectively. This demonstrates that the global feature extraction capabilities enabled by state space model remain highly effective for small object detection. Furthermore, supported by state space model, HDTA also slightly outperforms other backbone networks in dense object categories such as temporary stops and buildings.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Comparison of different backbones in terms of mAP and AP. <bold>(a)</bold> The mAP across all categories. <bold>(b-h)</bold> are APs of vehicle, temporary dwelling, covered mash, black plastic, white plastic, building, shade net</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_69611-fig-5.tif"/>
</fig>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Comparative Experiments between HDTA and SOTA Models</title>
<p><bold>Comparisons on object detection:</bold> <xref ref-type="fig" rid="fig-6">Fig. 6</xref> presents the experimental results of various models on the PSH detection task, including mAP and recall. As shown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>, DINO, Faster R-CNN, Cascade R-CNN, and FSAF generally underperform compared to the YOLO series in terms of accuracy. However, HDTA remains the best-performing model among all, achieving approximately 80% mAP and 75% recall. Specifically, the mAP of HDTA surpasses that of YOLOv8s, YOLOv8m, YOLOv11s, YOLOv11m, YOLOv12s, and YOLOv12m by 6.8%, 4.3%, 9%, 4.9%, 11.1%, and 5.2%, respectively. In terms of recall, HDTA exceeds these models by 7.2%, 3.3%, 6.9%, 4.9%, 11.4%, and 4.6%, respectively.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Comparison of PSH detection performance in terms of mAP and recall</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_69611-fig-6.tif"/>
</fig>
<p><xref ref-type="table" rid="table-2">Table 2</xref> presents a performance comparison between HDTA and other models on specific categories within the environmental dataset. As shown in the table, DINO, Cascade R-CNN, and FSAF perform worse than the YOLO series across all categories. However, our proposed HDTA demonstrates significant performance advantages over the YOLO series in categories such as Vehicle, Rubbish, Black Plastic, and Shade Net. For instance, in the vehicle category, HDTA achieves AP50 scores that are 27.7%, 21.4%, 40.1%, 34.5%, 30.9%, 21%, 41%, 34.6%, and 23.3% higher than those of YOLOv8s, YOLOv8m, YOLOv10n, YOLOv11n, YOLOv11s, YOLOv11m, YOLOv12n, YOLOv12s, and YOLOv12m, respectively. This demonstrates that even small categories such as Vehicle can be accurately detected by HDTA, which may be attributed to HDTA&#x2019;s long-sequence global modeling capability, allowing it to effectively extract features of the Vehicle category across vast regions. HDTA also shows strong competitiveness in the remaining categories. In summary, the advanced network architecture and the S<sup>4</sup>M backbone empower HDTA with highly effective PSH target detection capabilities in high-speed railroad scenarios.</p>
<table-wrap id="table-2">
<label>Table 2 </label>
<caption>
<title>APs of models on different classes on railroad environmental dataset</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Category models</th>
<th>Vehcicle</th>
<th>Rubbish</th>
<th>Covered mash</th>
<th>Black plastic</th>
<th>Shade net</th>
</tr>
</thead>
<tbody>
<tr>
<td>DINO</td>
<td>0.488</td>
<td>0.182</td>
<td>0.341</td>
<td>0.362</td>
<td>0.467</td>
</tr>
<tr>
<td>Casacde RCNN</td>
<td>0.112</td>
<td>0.400</td>
<td>0.424</td>
<td>0.125</td>
<td>0.329</td>
</tr>
<tr>
<td>FSAF</td>
<td>0.243</td>
<td>0.202</td>
<td>0.383</td>
<td>0.083</td>
<td>0.256</td>
</tr>
<tr>
<td>YOLOv8s</td>
<td>0.431</td>
<td>0.638</td>
<td>0.959</td>
<td>0.638</td>
<td>0.605</td>
</tr>
<tr>
<td>YOLOv8m</td>
<td>0.494</td>
<td>0.681</td>
<td>0.963</td>
<td>0.655</td>
<td>0.614</td>
</tr>
<tr>
<td>YOLOv10n</td>
<td>0.303</td>
<td>0.620</td>
<td>0.883</td>
<td>0.588</td>
<td>0.497</td>
</tr>
<tr>
<td>YOLOv11n</td>
<td>0.363</td>
<td>0.357</td>
<td>0.863</td>
<td>0.689</td>
<td>0.595</td>
</tr>
<tr>
<td>YOLOv11s</td>
<td>0.399</td>
<td>0.562</td>
<td>0.966</td>
<td>0.636</td>
<td>0.573</td>
</tr>
<tr>
<td>YOLOv11m</td>
<td>0.498</td>
<td>0.653</td>
<td>0.984</td>
<td>0.672</td>
<td>0.600</td>
</tr>
<tr>
<td>YOLOv12n</td>
<td>0.298</td>
<td>0.690</td>
<td>0.943</td>
<td>0.548</td>
<td>0.479</td>
</tr>
<tr>
<td>YOLOv12s</td>
<td>0.362</td>
<td>0.532</td>
<td>0.977</td>
<td>0.576</td>
<td>0.566</td>
</tr>
<tr>
<td>YOLOv12m</td>
<td>0.475</td>
<td>0.656</td>
<td>0.982</td>
<td>0.656</td>
<td>0.595</td>
</tr>
<tr>
<td>HDTA(ours)</td>
<td>0.708</td>
<td>0.729</td>
<td>0.96</td>
<td>0.849</td>
<td>0.697</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><bold>Comparisons on segmentation:</bold> <xref ref-type="fig" rid="fig-7">Fig. 7</xref> illustrates the experimental results of various models on railroad and roadway parsing tasks in terms of mean Accuracy (mACC) and mean Intersection over Union (mIoU). As shown in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>, the proposed HDTA ranks first in mACC, outperforming SegFormer, GCNet, MobileNetv3, BiSeNetv2, DeepLabv3&#x002B;, EMANet, DDRNet, and Mask2Former by 2.5%, 13.7%, 1.3%, 5.1%, 4.6%, 5.5%, 0.9%, and 3.5%, respectively. In terms of mIoU, HDTA also achieves the highest value of 86.5%, exceeding the performance of SegFormer, GCNet, MobileNetv3, BiSeNetv2, DeepLabv3&#x002B;, EMANet, DDRNet, and Mask2Former by 2.6%, 14.8%, 7.1%, 14.4%, 5.1%, 10%, 4.4%, and 3.3%, respectively. These results convincingly demonstrate the superior capability of HDTA in pixel-level parsing. Notably, unlike conventional models that focus solely on semantic segmentation, the HDTA architecture integrates both object detection and parsing branches, enabling not only precise segmentation of railroad and road regions but also effective PSH detection.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Comparison of parsing performance in terms of mIoU and mACC</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_69611-fig-7.tif"/>
</fig>
<p><bold>Comparisons on inference speed, GFLOPs and params:</bold> <xref ref-type="table" rid="table-3">Table 3</xref> presents a comparison of inference speed, GFLOPs, and parameter counts, with all models evaluated on an RTX 4070. It can be observed that DINO, Faster R-CNN, Cascade R-CNN, and FSAF are inferior to the YOLO series in terms of inference speed. Compared to the YOLO series, our proposed HDTA achieves the lowest GFLOPs and parameter count&#x2014;4.5 GFLOPs and 0.89 MB, respectively. In terms of inference speed, HDTA outperforms most YOLO models, reaching 140.8 FPS. While maintaining a lightweight design and high inference speed, HDTA also demonstrates superior detection accuracy over all compared methods Therefore, we believe that HDTA possesses the capability for fast detection of PSH in the high-speed railway surrounding environment.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Comparison of inference speed, GFLOPs and Params on different methods</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>FPS</th>
<th>Times</th>
<th>GFLOPs</th>
<th>Params (MB)</th>
</tr>
</thead>
<tbody>
<tr>
<td>DINO</td>
<td>17.8</td>
<td>56.2</td>
<td>87.9</td>
<td>47.55</td>
</tr>
<tr>
<td>Faster-RCNN</td>
<td>68.6</td>
<td>14.6</td>
<td>69.5</td>
<td>41.39</td>
</tr>
<tr>
<td>Casacde-RCNN</td>
<td>46.1</td>
<td>21.7</td>
<td>97.5</td>
<td>69.39</td>
</tr>
<tr>
<td>FSAF</td>
<td>83.5</td>
<td>12.0</td>
<td>59.4</td>
<td>36.42</td>
</tr>
<tr>
<td>YOLOv8s</td>
<td>166.7</td>
<td>6.0</td>
<td>27.9</td>
<td>10.01</td>
</tr>
<tr>
<td>YOLOv8m</td>
<td>126.6</td>
<td>7.9</td>
<td>78.1</td>
<td>23.88</td>
</tr>
<tr>
<td>YOLOv10n</td>
<td>136.9</td>
<td>7.3</td>
<td>8.1</td>
<td>2.43</td>
</tr>
<tr>
<td>YOLOv11n</td>
<td>128.2</td>
<td>7.8</td>
<td>10.2</td>
<td>2.40</td>
</tr>
<tr>
<td>YOLOv11s</td>
<td>121.9</td>
<td>8.2</td>
<td>28.5</td>
<td>8.55</td>
</tr>
<tr>
<td>YOLOv11m</td>
<td>101.1</td>
<td>9.9</td>
<td>88.8</td>
<td>19.04</td>
</tr>
<tr>
<td>YOLOv12n</td>
<td>136.9</td>
<td>7.3</td>
<td>6.2</td>
<td>2.43</td>
</tr>
<tr>
<td>YOLOv12s</td>
<td>125.0</td>
<td>8.0</td>
<td>19.7</td>
<td>8.70</td>
</tr>
<tr>
<td>YOLOv12m</td>
<td>100.0</td>
<td>10.0</td>
<td>60.4</td>
<td>18.76</td>
</tr>
<tr>
<td>Ours</td>
<td>140.8</td>
<td>7.1</td>
<td>4.5</td>
<td>0.89</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><bold>Visualization:</bold> As shown in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>, the proposed system is capable of accurately parsing high-speed railroad tracks and roads, while successfully detecting all PSHs along the railroad. Furthermore, as illustrated in <xref ref-type="fig" rid="fig-8">Fig. 8b</xref> and <xref ref-type="fig" rid="fig-8">d</xref>, the system effectively detects small object categories such as vehicles and rubbish, as well as dense object categories like building and temporary dwelling. These results demonstrate that the proposed HDTA maintains strong adaptability and robustness, even in complex high-speed railroad environments. However, some limitations still exist in current methods, for example, there are misclassification issues between buildings and temporary dwellings, as seen in the first and fourth rows of column (f); and cases of vehicles being misclassified as temporary dwellings or entirely missed, as shown in the third and fourth rows of column (f). To address these issues, future work can focus on increasing the quantity and diversity of samples from easily confused categories such as Temporary Dwelling and vehicle to enhance the model&#x2019;s discriminative ability. Additionally, incorporating more powerful context-aware mechanisms could further strengthen the model&#x2019;s feature extraction and classification capabilities.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Visualization examples of the prediction results by HDTA using the railroad environment dataset. <bold>(a,c,e)</bold> are ground truth, <bold>(b,d,f)</bold> are corresponding predicted results. <bold>(e,f)</bold> display examples of false positives and false negatives</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_69611-fig-8.tif"/>
</fig>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Conclusions</title>
<p>To better parse tracks in drone images, detect potential safety hazards (PSHs) along the tracks, and leverage the advantages of drone imagery, this paper proposes a hybrid learning architecture called HDTA (Hybrid Dual Tasks Architecture). HDTA is designed for parsing railroad and road in drone images while detecting PSHs around the tracks simultaneously. The HDTA assembles a newly lightweight backbone based on the S<sup>4</sup>M-block. The S<sup>4</sup>M-block could facilitate feature fusion and extract global contextual information while maintaining linear complexity. This design enables the S<sup>4</sup>M-Backbone to achieve both strong global feature extraction capability and low computational cost. Compared to advanced segmentation models like MaskFormer and SegFormer, the proposed HDTA achieves superior performance with 90.8% mACC and 86.5% mIoU. It also outperforms leading object detection models, including the YOLO series, with a recall of 74.8% and a mAP of 79.7%. Additionally, among all the compared models, HDTA has the smallest computational cost of 4.5 GFLOPs and the smallest parameters of only 0.89MB. However, there is still room for improvement in HDTA in the following aspects:
<list list-type="simple">
<list-item><label>(1)</label><p>The algorithm currently lacks sufficient capability for edge deployment. A UAV-based onboard inspection system with high accuracy and low equipment requirement is wait for development.</p></list-item>
<list-item><label>(2)</label><p>Adverse weather conditions present major obstacles for UAV-based PSH detection in railway settings. Elements like rain, snow, or thick fog can impair visibility, leading UAVs to capture images that are blurred, distorted, or poorly lit, ultimately reducing the effectiveness and accuracy of exist models.</p></list-item>
</list></p>
<p>Accordingly, this research outlines the insights and development directions for future work:
<list list-type="simple">
<list-item><label>(1)</label><p>Designing an integrated onboard inspection system with tailored software, optimized for real-time UAV-based PSH inspection. The system aims to achieve high accuracy while maintaining low computational load, cost, and power consumption.</p></list-item>
<list-item><label>(2)</label><p>Expanding on existing research, create a robust and efficient PSH detection algorithm that can perform reliably in harsh weather conditions, thereby improving the model&#x2019;s resilience and adaptability to challenging environments.</p></list-item>
</list></p>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This work was supported in part by the National Natural Science Foundation of China (grant No. 52362048), and in part by Yunnan Fundamental Research Projects (grant No. 202301BE070001-042 and grant No. 202401AT070409).</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Conceptualization, Zheda Zhao, Tao Xu, Yunpeng Wu; methodology, Zheda Zhao, Tao Xu; software, Zheda Zhao, Tao Xu, Tong Yang; validation, Zheda Zhao, Tao Xu, Tong Yang, Yunpeng Wu; formal analysis, Zheda Zhao, Tao Xu, Tong Yang; investigation, Yunpeng Wu, Fengxiang Guo; resources, Yunpeng Wu, Fengxiang Guo; data curation, Yunpeng Wu, Fengxiang Guo; writing&#x2014;original draft preparation, Zheda Zhao, Tao Xu, Tong Yang; writing&#x2014;review and editing, Zheda Zhao, Tao Xu, Yunpeng Wu; visualization, Zheda Zhao; supervision, Yunpeng Wu, Fengxiang Guo; project administration, Yunpeng Wu, Fengxiang Guo; funding acquisition, Yunpeng Wu, Fengxiang Guo. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>Data available on request from the authors. The data that support the findings of this study are available from the Corresponding Author, [Yunpeng Wu], upon reasonable request.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Qin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Qin</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Automatic potential safety hazard evaluation system for environment around high-speed railroad using hybrid u-shape learning architecture</article-title>; <year>2025</year>;<volume>26</volume>(<issue>1</issue>):<fpage>1071</fpage>&#x2013;<lpage>87</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TITS.2024.3487592</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ren</surname> <given-names>S</given-names></string-name>, <string-name><surname>He</surname> <given-names>K</given-names></string-name>, <string-name><surname>Girshick</surname> <given-names>R</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Faster r-cnn: towards real-time object detection with region proposal networks</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2017</year>;<volume>39</volume>(<issue>6</issue>):<fpage>1137</fpage>&#x2013;<lpage>49</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tpami.2016.2577031</pub-id>; <pub-id pub-id-type="pmid">27295650</pub-id></mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Cai</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Vasconcelos</surname> <given-names>N</given-names>
</string-name></person-group>, editors. <article-title>Cascade r-cnn: delving into high quality object detection</article-title>. In: <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>; <year>2018 Jun 18&#x2013;23</year>; <publisher-loc>Salt Lake City, UT, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>. p. <fpage>6154</fpage>&#x2013;<lpage>62</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR.2018.00644</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Anguelov</surname> <given-names>D</given-names></string-name>, <string-name><surname>Erhan</surname> <given-names>D</given-names></string-name>, <string-name><surname>Szegedy</surname> <given-names>C</given-names></string-name>, <string-name><surname>Reed</surname> <given-names>S</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>C-Y</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Ssd: single shot multibox detector</article-title>. In: <conf-name>Computer Vision&#x2013;ECCV 2016: 14th European Conference</conf-name>; <year>2016 Oct 11&#x2013;14</year>; <publisher-loc>Amsterdam, The Netherlands</publisher-loc>: <publisher-name>Springer</publisher-name>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.1512.02325</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Bochkovskiy</surname> <given-names>A</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>C-Y</given-names></string-name>, <string-name><surname>Liao</surname> <given-names>H-YM</given-names></string-name></person-group>. <article-title>Yolov4: optimal speed and accuracy of object detection</article-title>. <comment>arXiv:2004.10934. 2020</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2004.10934</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Jocher</surname> <given-names>G</given-names></string-name>, <string-name><surname>Mineeva</surname> <given-names>T</given-names></string-name>, <string-name><surname>Vilari&#x00F1;o</surname> <given-names>R</given-names></string-name></person-group>. <article-title>YOLOv5 [Online]. [cited 2025 Aug 12]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://github.com/ultralytics/yolov5">https://github.com/ultralytics/yolov5</ext-link>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>C-Y</given-names></string-name>, <string-name><surname>Bochkovskiy</surname> <given-names>A</given-names></string-name>, <string-name><surname>Liao</surname> <given-names>H-YM,</given-names>
<suffix>editors</suffix></string-name></person-group>. <article-title>YOLOv7: trainable bag-of-freebies sets new state-of-the-art for real-time object detectors</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>. <publisher-loc>Vancouver, BC, Canada</publisher-loc>; <year>2023</year>. p. <fpage>7464</fpage>&#x2013;<lpage>75</lpage>. doi:<pub-id pub-id-type="doi">10.1109/CVPR52729.2023.00721</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Jocher</surname> <given-names>G</given-names></string-name></person-group>. <article-title>YOLOv8 [Online]. [cited 2025 Aug 12]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://github.com/ultralytics/ultralytics">https://github.com/ultralytics/ultralytics</ext-link>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>L-C</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Papandreou</surname> <given-names>G</given-names></string-name>, <string-name><surname>Schroff</surname> <given-names>F</given-names></string-name>, <string-name><surname>Adam</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Encoder-decoder with atrous separable convolution for semantic image segmentation</article-title>. In: <conf-name>Proceedings of the European Conference on Computer Vision (ECCV)</conf-name>. <publisher-loc>Munich, Germany</publisher-loc>; <year>2018</year>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.1802.02611</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Howard</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sandler</surname> <given-names>M</given-names></string-name>, <string-name><surname>Chu</surname> <given-names>G</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>L-C</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>B</given-names></string-name>, <string-name><surname>Tan</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Searching for mobilenetv3</article-title>. In: <conf-name>Proceedings of the IEEE/CVF International Conference on Computer Vision</conf-name>; <year>2019</year>. p. <fpage>1314</fpage>&#x2013;<lpage>24</lpage>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.1905.02244</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Cao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>F</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Gcnet: non-local networks meet squeeze-excitation networks and beyond</article-title>. In: <conf-name>Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops</conf-name>. <publisher-loc>Seoul, Republic of Korea</publisher-loc>; <year>2019</year>. p. <fpage>1971</fpage>&#x2013;<lpage>80</lpage>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.1904.11492</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xie</surname> <given-names>E</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Anandkumar</surname> <given-names>A</given-names></string-name>, <string-name><surname>Alvarez</surname> <given-names>JM</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>PJ</given-names></string-name></person-group>. <article-title>SegFormer: simple and efficient design for semantic segmentation with transformers</article-title>. <year>2021</year>;<volume>34</volume>:<fpage>12077</fpage>&#x2013;<lpage>90</lpage>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2105.15203</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Cheng</surname> <given-names>B</given-names></string-name>, <string-name><surname>Misra</surname> <given-names>I</given-names></string-name>, <string-name><surname>Schwing</surname> <given-names>AG</given-names></string-name>, <string-name><surname>Kirillov</surname> <given-names>A</given-names></string-name>, <string-name><surname>Girdhar</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Masked-attention mask transformer for universal image segmentation</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</conf-name>. <publisher-loc>New Orleans, LA, USA</publisher-loc>; <year>2022</year>. p. <fpage>1280</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2112.01527</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>G</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>C</given-names></string-name>, <string-name><surname>Sang</surname> <given-names>N</given-names></string-name></person-group>. <article-title>Bisenet v2: bilateral network with guided aggregation for real-time semantic segmentation</article-title>. <source>Int J Comput Vis</source>. <year>2021</year>;<volume>129</volume>(<issue>11</issue>):<fpage>3051</fpage>&#x2013;<lpage>68</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11263-021-01515-2</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhong</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Expectation-maximization attention networks for semantic segmentation</article-title>. In: <conf-name>Proceedings of the IEEE/CVF International Conference on Computer Vision</conf-name>. <publisher-loc>Seoul, Republic of Korea</publisher-loc>; <year>2019</year>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.1907.13426</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Hong</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>H</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>W</given-names></string-name>, <string-name><surname>Jia</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Deep dual-resolution networks for real-time and accurate semantic segmentation of road scenes</article-title>. <comment>arXiv:2101.06085. 2021</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2101.06085</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ma</surname> <given-names>R</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>J</given-names></string-name></person-group>. <article-title>ELA-YOLO: an efficient method with linear attention for steel surface defect detection during manufacturing</article-title>. <source>Adv Eng Inform</source>. <year>2025</year>;<volume>65</volume>(<issue>2</issue>):<fpage>103377</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.aei.2025.103377</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Dai</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>X</given-names></string-name></person-group>. <article-title>A spatial-temporal neural network based on ResNet-Transformer for predicting railroad broken rails</article-title>. <source>Adv Eng Inform</source>. <year>2025</year>;<volume>65</volume>:<fpage>103126</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.aei.2025.103126</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yi</surname> <given-names>W</given-names></string-name>, <string-name><surname>Li</surname> <given-names>S</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>P</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>R</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Intelligent inspection of building exterior walls using UAV and mixed reality based on man-machine-environment system engineering</article-title>. <source>Autom Constr</source>. <year>2025</year>;<volume>177</volume>(<issue>3</issue>):<fpage>106344</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.autcon.2025.106344</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shan</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Bridging cross-domain and cross-resolution gaps for UAV-based pavement crack segmentation</article-title>. <source>Autom Constr</source>. <year>2025</year>;<volume>174</volume>:<fpage>106141</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.autcon.2025.106141</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Efficient instance segmentation framework for UAV-based pavement distress detection</article-title>. <source>Autom Constr</source>. <year>2025</year>;<volume>175</volume>(<issue>1</issue>):<fpage>106195</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.autcon.2025.106195</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Meng</surname> <given-names>F</given-names></string-name>, <string-name><surname>Qin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>F</given-names></string-name>, <string-name><surname>Jia</surname> <given-names>L</given-names></string-name></person-group>. <article-title>UAV imagery based potential safety hazard evaluation for high-speed railroad using Real-time instance segmentation</article-title>. <source>Adv Eng Inform</source>. <year>2023</year>;<volume>55</volume>(<issue>6</issue>):<fpage>101819</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.aei.2022.101819</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Meng</surname> <given-names>F</given-names></string-name>, <string-name><surname>Qin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Shao</surname> <given-names>C</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Jia</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Automatic risk level evaluation system for potential environmental hazards along high-speed railroad using UAV aerial photograph</article-title>. <source>Expert Syst Appl</source>. <year>2025</year>;<volume>277</volume>(<issue>5</issue>):<fpage>127257</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2025.127257</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>P</given-names></string-name>, <string-name><surname>Qin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>F</given-names></string-name>, <string-name><surname>Jia</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Automatic railroad track components inspection using hybrid deep learning framework</article-title>. <source>IEEE Trans Instrum Meas</source>. <year>2023</year>;<volume>72</volume>:<fpage>5011415</fpage>&#x2013;<lpage>15</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tim.2023.3265636</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>P</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>F</given-names></string-name>, <string-name><surname>Qin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Long</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Hybrid learning architecture for high-speed railroad scene parsing and potential safety hazard evaluation of UAV images</article-title>. <source>Measurement</source>. <year>2025</year>;<volume>239</volume>:<fpage>115504</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.measurement.2024.115504</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Liao</surname> <given-names>B</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Vision mamba: efficient visual representation learning with bidirectional state space model</article-title>. <year>2024</year>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2401.09417</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ma</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>F</given-names></string-name>, <string-name><surname>B</surname> <given-names>Wang</given-names></string-name></person-group>. <article-title>U-mamba: enhancing long-range dependency for biomedical image segmentation</article-title>. <comment>arXiv:2401.09417. 2024</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2401.04722</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Vmamba: visual state space model</article-title>. <comment>arXiv:2401.10166. 2024</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2401.10166</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Le</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>Efficientnet: rethinking model scaling for convolutional neural networks</article-title>. In: <source>Proceedings of the 36th International Conference on Machine Learning</source>; <year>2019</year>. Vol. <volume>67</volume>, p. <fpage>6105</fpage>&#x2013;<lpage>14</lpage>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.1905.11946</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ma</surname> <given-names>X</given-names></string-name>, <string-name><surname>Dai</surname> <given-names>X</given-names></string-name>, <string-name><surname>Bai</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Rewrite the Stars</article-title>. <comment>arXiv:2403.19967. 2024</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2403.19967</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Swin transformer: Hierarchical vision transformer using shifted windows</article-title>. In: <conf-name>Proceedings of the IEEE/CVF International Conference on Computer Vision</conf-name>; <year>2021</year>; <publisher-loc>Montreal, QC, Canada</publisher-loc>. p. <fpage>10012</fpage>&#x2013;<lpage>22</lpage>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2103.14030</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Caron</surname> <given-names>M</given-names></string-name>, <string-name><surname>Touvron</surname> <given-names>H</given-names></string-name>, <string-name><surname>Misra</surname> <given-names>I</given-names></string-name>, <string-name><surname>J&#x00E9;gou</surname> <given-names>H</given-names></string-name>, <string-name><surname>Mairal</surname> <given-names>J</given-names></string-name>, <string-name><surname>Bojanowski</surname> <given-names>P</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Emerging properties in self-supervised vision transformers</article-title>. In: <conf-name>Proceedings of the IEEE/CVF International Conference on Computer Vision</conf-name>. <publisher-loc>Montreal, QC, Canada</publisher-loc>; <year>2021</year>. p. <fpage>9630</fpage>&#x2013;<lpage>40</lpage>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2104.14294</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Varghese</surname> <given-names>R</given-names></string-name>, <string-name><surname>Sambath</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Yolov8: A novel object detection algorithm with enhanced performance and robustness</article-title>. In: <conf-name>International Conference on Advances in Data Engineering and Intelligent Computing Systems (ADICS); 2024 Apr 18&#x2013;19</conf-name>. <publisher-loc>Chennai, India</publisher-loc>: <publisher-name>IEEE</publisher-name>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ADICS58448.2024.10533619</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Khanam</surname> <given-names>R</given-names></string-name>, <string-name><surname>Hussain</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Yolov11: an overview of the key architectural enhancements</article-title>. <comment>arXiv:2410.17725. 2024</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2410.17725</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Tian</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ye</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Doermann</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Yolov12: attention-centric real-time object detectors</article-title>. <comment>arXiv:2502.12524. 2025</comment>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2502.12524</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>









