<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">54525</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2024.054525</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Rail-PillarNet: A 3D Detection Network for Railway Foreign Object Based on LiDAR</article-title>
<alt-title alt-title-type="left-running-head">Rail-PillarNet: A 3D Detection Network for Railway Foreign Object Based on LiDAR</alt-title>
<alt-title alt-title-type="right-running-head">Rail-PillarNet: A 3D Detection Network for Railway Foreign Object Based on LiDAR</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Li</surname><given-names>Fan</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Zhang</surname><given-names>Shuyao</given-names></name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<contrib id="author-3" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Yang</surname><given-names>Jie</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref><email>yangjie@jxust.edu.cn</email></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Feng</surname><given-names>Zhicheng</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Chen</surname><given-names>Zhichao</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<aff id="aff-1"><label>1</label><institution>School of Electrical Engineering and Automation, Jiangxi University of Science and Technology</institution>, <addr-line>Ganzhou, 341000</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>Jiangxi Provincial Key Laboratory of Maglev Technology</institution>, <addr-line>Ganzhou, 341000</addr-line>, <country>China</country></aff>
<aff id="aff-3"><label>3</label><institution>School of Information Engineering, Jiangxi University of Science and Technology</institution>, <addr-line>Ganzhou, 341000</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Jie Yang. Email: <email>yangjie@jxust.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2024</year></pub-date>
<pub-date date-type="pub" publication-format="electronic"><day>12</day><month>9</month><year>2024</year></pub-date>
<volume>80</volume>
<issue>3</issue>
<fpage>3819</fpage>
<lpage>3833</lpage>
<history>
<date date-type="received">
<day>30</day>
<month>5</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>01</day>
<month>8</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2024 The Authors.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_54525.pdf"></self-uri>
<abstract>
<p>Aiming at the limitations of the existing railway foreign object detection methods based on two-dimensional (2D) images, such as short detection distance, strong influence of environment and lack of distance information, we propose Rail-PillarNet, a three-dimensional (3D) LIDAR (Light Detection and Ranging) railway foreign object detection method based on the improvement of PointPillars. Firstly, the parallel attention pillar encoder (PAPE) is designed to fully extract the features of the pillars and alleviate the problem of local fine-grained information loss in PointPillars pillars encoder. Secondly, a fine backbone network is designed to improve the feature extraction capability of the network by combining the coding characteristics of LIDAR point cloud feature and residual structure. Finally, the initial weight parameters of the model were optimised by the transfer learning training method to further improve accuracy. The experimental results on the OSDaR23 dataset show that the average accuracy of Rail-PillarNet reaches 58.51%, which is higher than most mainstream models, and the number of parameters is 5.49 M. Compared with PointPillars, the accuracy of each target is improved by 10.94%, 3.53%, 16.96% and 19.90%, respectively, and the number of parameters only increases by 0.64 M, which achieves a balance between the number of parameters and accuracy.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Railway foreign object</kwd>
<kwd>light detection and ranging (LiDAR)</kwd>
<kwd>3D object detection</kwd>
<kwd>PointPillars</kwd>
<kwd>parallel attention mechanism</kwd>
<kwd>transfer learning</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>National Key Research and Development Project</funding-source>
<award-id>2023YFB4302100</award-id>
</award-group>
<award-group id="awg2">
<funding-source>Research and Development Project of Jiangxi Province</funding-source>
<award-id>20232ACE01011</award-id>
</award-group>
<award-group id="awg3">
<funding-source>Independent Deployment Project of Ganjiang Innovation Research Institute</funding-source>
<award-id>E255J001</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>With the rapid development of railway transport and railway network systems, the railway-driving environment has become increasingly complex. In the process of train travel, the foreign objects into the railway boundaries can cause serious traffic accidents, seriously threatening the safety of people&#x2019;s lives and property, so there is an urgent need to conduct research on railway foreign object detection to ensure the safe operation of the train. However, the existing railway foreign object detection research is basically based on 2D images, and in the case of low light, bad weather, and long-distance small objects, the foreign object features in the image are not obvious, and there will be different degrees of omission and misdetection [<xref ref-type="bibr" rid="ref-1">1</xref>&#x2013;<xref ref-type="bibr" rid="ref-3">3</xref>]. In addition, 2D images lack distance information, and accurate distance perception is important for improving the level of train automation [<xref ref-type="bibr" rid="ref-4">4</xref>,<xref ref-type="bibr" rid="ref-5">5</xref>]. Unlike optical cameras, LiDAR has the advantages of all-weather operation, wide application environment, and long detection distance, etc. Its 3D point cloud data contains the real size, shape, and distance information of the foreign object, which is more suitable for railway operation characteristics.</p>
<p>Currently, LiDAR-based railway obstacle detection is still mainly based on traditional point cloud filtering and clustering methods. Vatavu et al. [<xref ref-type="bibr" rid="ref-6">6</xref>] designed a grid map estimation method based on particle filtering, which can estimate the speed of the objects, but it is difficult to track the objects when they are occluded. Xie et al. [<xref ref-type="bibr" rid="ref-7">7</xref>] clustered foreign objects by rasterizing and using the eight-neighbour cell clustering method, but it is still difficult to distinguish between neighbouring objects.</p>
<p>In recent years, deep learning-based lidar point cloud 3D object detection methods have achieved great success in the field of automated vehicle driving. The 3D object detection methods can be classified into multi-view-based, voxels or pillars-based [<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-9">9</xref>] and point-based [<xref ref-type="bibr" rid="ref-10">10</xref>,<xref ref-type="bibr" rid="ref-11">11</xref>]. Yuan et al. [<xref ref-type="bibr" rid="ref-12">12</xref>] proposed a car detection method, which maps the point cloud onto a 2D image from a bird&#x2019;s eye view perspective but only achieves the detection of the car. Liu et al. [<xref ref-type="bibr" rid="ref-13">13</xref>] proposed a combined image and point cloud method for Bus Rapid Transit vehicle detection with good detection results. Wang et al. [<xref ref-type="bibr" rid="ref-14">14</xref>] fused cameras and LiDAR to detect targets on railway tracks, but the detection effect depends on the semantic segmentation results of 2D images. Wen et al. [<xref ref-type="bibr" rid="ref-15">15</xref>] used a semantic segmentation network based on Double Spiral Transformer module to achieve the detection of multiple types of obstacles under complex weather conditions by using an obstacle anomaly sensing cross-modal discrimination strategy. Neri et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] produced a virtual railway environment to generate railway point cloud datasets and proposed a VoxelNet-based method for railway 3D object detection. Wisultschew et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] designed a lightweight object detection and embedded detection platform, which realized the detection and tracking of car and pedestrian targets at railway level crossings.</p>
<p>The above methods are only for a single object, a single scene, or rely on the detection results of 2D images, which have certain limitations in terms of generality. In addition, some methods use the computationally heavy Transformer and VoxelNet structures. Recently, the first generic multi-sensor dataset for the railway domain, OSDaR23 [<xref ref-type="bibr" rid="ref-18">18</xref>], has been published by a research group led by German Centre for Rail Traffic Research. Recorded in Hamburg, Germany, the dataset contains data from various sensors, and provides fine-grained data labels that offer a new solution idea for detecting foreign objects in railways.</p>
<p>This paper adopts LiDAR data from the universal multi-sensor dataset OSDaR23 [<xref ref-type="bibr" rid="ref-18">18</xref>] for the study of railway foreign object detection, and proposes a deep learning-based LiDAR railway foreign object 3D detection network, Rail-PillarNet. The main contributions of this paper are as follows:
<list list-type="simple">
<list-item><label>(1)</label><p>Based on PointPillars, for the long-range small objects in the railway scenario, we propose Parallel Attention Pillar Encoding (PAPE), which reduces the loss of fine-grained information in the pillars.</p></list-item>
<list-item><label>(2)</label><p>Combined the LIDAR point cloud encoding characteristics, the fusion of information at different scales is achieved by a finely designed 2D backbone network.</p></list-item>
<list-item><label>(3)</label><p>The model performance is further improved by pre-training on a similar traffic scene dataset, KITTI [<xref ref-type="bibr" rid="ref-19">19</xref>], while using transfer learning for training.</p></list-item>
</list></p>
<p>The general structure of this paper as follows. <xref ref-type="sec" rid="s2">Section 2</xref> describes the general framework of PointPillars. <xref ref-type="sec" rid="s3">Section 3</xref> describes the structure of Rail-PillarNet. <xref ref-type="sec" rid="s4">Section 4</xref> conducts the experiments. <xref ref-type="sec" rid="s5">Section 5</xref> discusses the conclusion of the experimental results. <xref ref-type="sec" rid="s6">Section 6</xref> gives the conclusion.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Structure of PointPillars</title>
<p>PointPillars [<xref ref-type="bibr" rid="ref-9">9</xref>] represents the original point cloud as pillars while extracting the point cloud features using a pillar encoder. Next, the pillars are converted into a sparse pseudo image, feature extraction is performed using a 2D convolutional backbone network, and finally the detection results are output through the detection head as shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. PointPillars [<xref ref-type="bibr" rid="ref-9">9</xref>] greatly reduces the amount of data to be processed by converting the 3D point cloud into a 2D pseudo image, and detects the object on the pseudo image using the 2D object detection algorithm, avoiding the use of computationally expensive 3D convolutions, making the algorithm lightweight and easy to use.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>PointPillars network structure</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_54525-fig-1.tif"/>
</fig>
<p>PointPillars [<xref ref-type="bibr" rid="ref-9">9</xref>] uses a simplified PointNet [<xref ref-type="bibr" rid="ref-20">20</xref>] in the pillars encoder to aggregate features in each pillars. However, this results in the loss of local fine-grained information, which is critical for 3D detection (small object at long distances) [<xref ref-type="bibr" rid="ref-21">21</xref>]. In addition, the backbone network of PointPillars uses a 2D convolutional network with the structure of Vgg [<xref ref-type="bibr" rid="ref-22">22</xref>,<xref ref-type="bibr" rid="ref-23">23</xref>] for feature extraction, which ignores the exchange of local features and input-output information, and feature extraction is insufficient [<xref ref-type="bibr" rid="ref-24">24</xref>,<xref ref-type="bibr" rid="ref-25">25</xref>].</p>
</sec>
<sec id="s3">
<label>3</label>
<title>Structure of Rail-PillarNet</title>
<p>This paper proposes Rail-PillarNet to address the above issues. Rail-PillarNet takes the LIDAR point cloud as input and first processes the point cloud by pillar division, then extracts the pillars features by Parallel Attention Pillar Encoding (PAPE), reduces the loss of local information, and obtains finer pillars features, which are then converted into 2D pseudo image. The extraction and fusion of features at different scales is achieved by a finely designed 2D backbone. Finally, object classification and box regression are performed using the detection head to generate prediction results as shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Rail-PillarNet network structure</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_54525-fig-2.tif"/>
</fig>
<sec id="s3_1">
<label>3.1</label>
<title>Parallel Attention Pillar Encoding (PAPE)</title>
<p>PointPillars [<xref ref-type="bibr" rid="ref-9">9</xref>] simply uses fully connected and max pooling to extract pillars features, which leads to the loss of fine-grained information in the pillars, and is prone to problems such as omission and misdetection of long-distance small objects in railway scenes. To improve the model&#x2019;s ability to detect small objects at long distances, the Parallel Attention Pillar Encoding (PAPE) module is embedded in the point cloud pillar encoding to mitigate the problem of fine-grained information loss in the point cloud encoding. As shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>, the PAPE module mainly consists of three units: (1) <italic>point-coding</italic>, (2) <italic>point-attention coding</italic>, (3) <italic>channel-attention coding</italic>.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Parallel attention pillar encoding</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_54525-fig-3.tif"/>
</fig>
<sec id="s3_1_1">
<label>3.1.1</label>
<title>Point-Coding</title>
<p>Assuming that the point cloud extends <italic>L</italic>, <italic>W</italic> and <italic>H</italic> along the <italic>X</italic>, <italic>Y</italic> and <italic>Z</italic>-axes in 3D space, the point cloud is uniformly divided into specific pillars of size <italic>l</italic>, <italic>w</italic> and <italic>H</italic>. As with PointPillars, only the point cloud in the <italic>X</italic>-<italic>Y</italic> plane is divided into pillars. Let <italic>P</italic> &#x003D; {<italic>p</italic><sub><italic>i</italic></sub> &#x003D; [<italic>x</italic><sub><italic>i</italic></sub>, <italic>y</italic><sub><italic>i</italic></sub>, <italic>z</italic><sub><italic>i</italic></sub>, <italic>r</italic><sub><italic>i</italic></sub>] &#x2208; <italic>R</italic><sup><italic>Nv</italic> &#x00D7; 4</sup>]}, where <italic>P</italic> is a non-empty pillars consisting of <italic>N</italic> points, <italic>p</italic><sub><italic>i</italic></sub> is the <italic>i</italic>-th point in the pillars, and each <italic>p</italic><sub><italic>i</italic></sub> has feature dimension D &#x003D; 4, <italic>i</italic> &#x2208; {1,...,<italic>N</italic><sub><italic>v</italic></sub>}, and <italic>N</italic><sub><italic>v</italic></sub> is the number of points in pillars <italic>v</italic>.</p>
<p>During the point encoding process, the points in each pillar is expanded as <italic>p</italic><sub><italic>i</italic></sub> &#x003D; {[<italic>x</italic><sub><italic>i</italic></sub>, <italic>y</italic><sub><italic>i</italic></sub>, <italic>z</italic><sub><italic>i</italic></sub>, <italic>r</italic><sub><italic>i</italic></sub>, <italic>xc</italic><sub><italic>i</italic></sub>, <italic>yc</italic><sub><italic>i</italic></sub>, <italic>zc</italic><sub><italic>i</italic></sub>, <italic>xp</italic><sub><italic>i</italic></sub>, <italic>yp</italic><sub><italic>i</italic></sub>, <italic>zp</italic><sub><italic>i</italic></sub>] &#x2208; <italic>R</italic><sup><italic>Nv</italic> &#x00D7; 10</sup>], where [<italic>x</italic><sub><italic>i</italic></sub>, <italic>y</italic><sub><italic>i</italic></sub>, <italic>z</italic><sub><italic>i</italic></sub>, <italic>r</italic><sub><italic>i</italic></sub>] are the coordinates and reflectance intensities of each point in the pillars, [<italic>xc</italic><sub><italic>i</italic></sub>, <italic>yc</italic><sub><italic>i</italic></sub>, <italic>zc</italic><sub><italic>i</italic></sub>,] are the offset of each point from the mean of all point clouds in that pillars, and [<italic>xp</italic><sub><italic>i</italic></sub>, <italic>yp</italic><sub><italic>i</italic></sub>, <italic>zp</italic><sub><italic>i</italic></sub>] are the offset of each point from the centre of the coordinates of that pillars.</p>
</sec>
<sec id="s3_1_2">
<label>3.1.2</label>
<title>Point-Attention Coding</title>
<p>After obtaining the augmented non-empty pillars <italic>P</italic><sup><italic>k</italic></sup>, in order to capture the features in the pointwise dimension within the pillars, we use the point-attention coding to aggregate the pointwise features of the input pillars, as shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>. First, two different kinds of pointwise information are generated in the pointwise dimension using the max pooling and average pooling: <italic>F</italic><sub><italic>p</italic></sub><sup>mean</sup> &#x2208; <italic>R</italic><sup><italic>P</italic> &#x00D7; <italic>N</italic> &#x00D7; 1</sup> and <italic>F</italic><sub><italic>p</italic></sub><sup>max</sup> &#x2208; <italic>R</italic><sup><italic>P</italic> &#x00D7; <italic>N</italic> &#x00D7; 1</sup>. Second, the two kinds of information are fed into a shared network consisting of a fully connected layer and a nonlinear activation function. Finally, the two types of information are summed and normalised weights are generated by the sigmoid activation function to describe the importance of each point within the pillars. The formulas are as follows:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msup><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:mtext>mean</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mtext>avgpool</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>P</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msup><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mtext>maxpool</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>P</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:mtext>mean</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <italic>A</italic><sub><italic>p</italic></sub> &#x2208; <italic>R</italic><sup><italic>P</italic> &#x00D7; <italic>N</italic> &#x00D7; 1</sup> is the attention score of each point, <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> is the Relu activation function, <italic>w</italic><sub><italic>0</italic></sub> and <italic>w</italic><sub><italic>1</italic></sub> are the weight parameters of the two fully connected layers, <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>&#x03C3;</mml:mi></mml:math></inline-formula> is the Sigmoid function.</p>
</sec>
<sec id="s3_1_3">
<label>3.1.3</label>
<title>Channel-Attention Coding</title>
<p>To capture the channel characteristics of the augmented pillars more comprehensively. Similarly, we extract features in the channel direction using max pooling and average pooling. Then, the importance of each feature channel is computed using the fully connected layer and the activation function. The corresponding equations are as follows:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msup><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:mtext>mean</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mtext>avgpool</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>P</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msup><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mtext>maxpool</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>P</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>mean</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>w</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>F</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <italic>A</italic><sub><italic>c</italic></sub> &#x2208; <italic>R</italic><sup><italic>P</italic> &#x00D7; 1 &#x00D7; C</sup> is the attention score of each channel, <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> is the Relu activation function, <italic>w</italic>&#x2032;<sub><italic>0</italic></sub> and <italic>w</italic>&#x2032;<sub><italic>1</italic></sub> are the weight parameters of the two fully connected layers, <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>&#x03C3;</mml:mi></mml:math></inline-formula> is the sigmoid function.</p>
<p>The parallel mechanism is used to combine point-attention coding and channel-attention coding to form a parallel attention pillars encoder. By multiplying the point attention <italic>A</italic><sub><italic>p</italic></sub> and the channel attention <italic>A</italic><sub><italic>c</italic></sub> with the original pillars <italic>P</italic>, respectively, the attention-weighted features of the pillars in the channel direction and in the point direction can be obtained. The output features are obtained by adding both with the original pillars <italic>P</italic>. Finally, the output is processed by fully connected layers and max pooling to obtain refined pillars features. The definition is as follows:
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mrow><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mi>f</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext>maxpool</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>w</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <italic>A</italic><sub><italic>p</italic></sub> &#x2208; <italic>R</italic><sup><italic>P</italic> &#x00D7; <italic>N</italic> &#x00D7; 1</sup> is the attention score per point, <italic>A</italic><sub><italic>c</italic></sub> &#x2208; <italic>R</italic><sup><italic>P</italic> &#x00D7; 1 &#x00D7; <italic>C</italic></sup> is the attention score per channel, <italic>P</italic> is the original pillars feature, <italic>w</italic> is the weight parameter of the fully connected layer.</p>
</sec>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Backbone Networks</title>
<p>The Vgg [<xref ref-type="bibr" rid="ref-22">22</xref>] structure of the backbone network is used in PointPillars, which ignores the exchange of local features and input-output information. In this paper, the residual structure block [<xref ref-type="bibr" rid="ref-23">23</xref>] is used instead of the ordinary convolution block to improve the feature extraction capability of the backbone network. In addition, unlike RGB images, features such as spatial distances and shapes of objects are explicitly encoded in the LIDAR point cloud, which does not require too much computational resources for later geometric modelling [<xref ref-type="bibr" rid="ref-25">25</xref>,<xref ref-type="bibr" rid="ref-26">26</xref>]. By adjusting the number of iterations in each stage of the backbone networks, more computations are allocated to the early stage to better integrate the geometric information contained in the point cloud.</p>
<p>The structure of backbone network is shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>. For the feature extraction module with stride of 1, two 3 &#x00D7; 3 convolutions are used to extract features, while the skip connection fuses the inputs and outputs of the modul. For the feature extraction module with stride of 2, a two-branch structure is used. In the main branch, a 3 &#x00D7; 3 convolution with stride of 2 is used for feature extraction and downsampling, followed by information fusion via another 3 &#x00D7; 3 convolution. In the other branch, feature mapping is performed on the input using a 1 &#x00D7; 1 convolution with stride of 2. In addition, the outputs of both branches are fused. Finally, the pseudo-image is downsampled using the backbone network to obtain feature maps of different sizes, and the downsampled multiple feature maps are upsampled to the same size for stitching to generate the final feature map.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Backbone network structure</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_54525-fig-4.tif"/>
</fig>
<p><xref ref-type="table" rid="table-1">Table 1</xref> shows the proposed backbone network structure. The size of the pseudo-image is assumed to be 640 &#x00D7; 320 &#x00D7; 64. In the table, Stage is the three stages of feature extraction, Residual is the residual block, Stride is the size of the step in the operation, Repeat is the number of repetitions.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Network structure of backbone network</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Feature map</th>
<th>Input size</th>
<th>Operate</th>
<th>Output channels</th>
<th>Stride</th>
<th>Repeat</th>
</tr>
</thead>
<tbody>
<tr>
<td>Pseudo-image</td>
<td>640 &#x00D7; 320 &#x00D7; 64</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
</tr>
<tr>
<td rowspan="2">Stage1</td>
<td>640 &#x00D7; 320 &#x00D7; 64</td>
<td>Residual</td>
<td>64</td>
<td>2</td>
<td>1</td>
</tr>
<tr>
<td>320 &#x00D7; 160 &#x00D7; 64</td>
<td>Residual</td>
<td>64</td>
<td>1</td>
<td>4</td>
</tr>
<tr>
<td rowspan="2">Stage2</td>
<td>320 &#x00D7; 160 &#x00D7; 64</td>
<td>Residual</td>
<td>128</td>
<td>2</td>
<td>1</td>
</tr>
<tr>
<td>160 &#x00D7; 80 &#x00D7; 128</td>
<td>Residual</td>
<td>128</td>
<td>1</td>
<td>2</td>
</tr>
<tr>
<td rowspan="2">Stage3</td>
<td>160 &#x00D7; 80 &#x00D7; 128</td>
<td>Residual</td>
<td>256</td>
<td>2</td>
<td>1</td>
</tr>
<tr>
<td>80 &#x00D7; 40 &#x00D7; 256</td>
<td>Residual</td>
<td>256</td>
<td>1</td>
<td>1</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Detection Head and Loss Function</title>
<p>In this paper, 3D object detection is performed by the single shot multibox detector (SSD) detection head [<xref ref-type="bibr" rid="ref-27">27</xref>], which uses 3D intersection over union (IoU) to match the anchors boxes with the ground-truth boxes. The network detection head finally outputs a 3D prediction frame with seven parameters, denoted as (<italic>x</italic>, <italic>y</italic>, <italic>z</italic>, <italic>w</italic>, <italic>l</italic>, <italic>h</italic>, <italic>&#x03B8;</italic>), where <italic>x</italic>, <italic>y</italic> and <italic>z</italic> are the centroids of the prediction boxes, <italic>w</italic>, <italic>l</italic> and <italic>h</italic> are the width, length, and height of the prediction boxes, and <italic>&#x03B8;</italic> is the angle of the orientation of the prediction boxes. The ground-truth boxes are defined as (<italic>x</italic><sub>gt</sub>, <italic>y</italic><sub>gt</sub>, <italic>z</italic><sub>gt</sub>, <italic>w</italic><sub>gt</sub>, <italic>l</italic><sub>gt</sub>, <italic>h</italic><sub>gt</sub>, <italic>&#x03B8;</italic><sub>gt</sub>). The position errors between the ground-truth boxes and the prediction boxes are (&#x2206;<italic>x</italic>, &#x2206;<italic>y</italic>, &#x2206;<italic>z</italic>, &#x2206;<italic>w</italic>, &#x2206;<italic>l</italic>, &#x2206;<italic>h</italic>, &#x2206;<italic>&#x03B8;</italic>):
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mi>x</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mrow><mml:mtext>gt</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mi>d</mml:mi></mml:mfrac><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mrow><mml:mtext>gt</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mi>d</mml:mi></mml:mfrac><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mi>z</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mrow><mml:mtext>gt</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mi>z</mml:mi></mml:mrow><mml:mi>h</mml:mi></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mi>w</mml:mi><mml:mo>=</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mfrac><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mrow><mml:mtext>gt</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mi>w</mml:mi></mml:mfrac><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mfrac><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mrow><mml:mtext>gt</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mi>l</mml:mi></mml:mfrac><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mi>h</mml:mi><mml:mo>=</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mfrac><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mrow><mml:mtext>gt</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mi>h</mml:mi></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mi>&#x03B8;</mml:mi><mml:mo>=</mml:mo><mml:mi>sin</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mrow><mml:mtext>gt</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where the subscript gt denotes the parameters of the ground-truth boxes, using the SmoothL1 loss as the 3D prediction boxes localisation loss:
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>loc</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>b</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#xA0;</mml:mtext><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#xA0;</mml:mtext><mml:mi>z</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#xA0;</mml:mtext><mml:mi>w</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#xA0;</mml:mtext><mml:mi>h</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#xA0;</mml:mtext><mml:mi>l</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#xA0;</mml:mtext><mml:mi>&#x03B8;</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:munder><mml:mrow><mml:mtext>SmoothL</mml:mtext></mml:mrow><mml:mn>1</mml:mn><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>Using Focus Loss to alleviate the problem of inefficient training and difficulty in model convergence, the object classification loss can be expressed as:</p>
<p><disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>cls</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>&#x03B3;</mml:mi></mml:mrow></mml:msup><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:msup><mml:mi>p</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msup></mml:math></disp-formula>where <italic>p</italic><sup><italic>t</italic></sup> is the category probability of the object in the anchor boxes, <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>&#x03B1;</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula> are set to 0.25 and 2.0, respectively.</p>
<p>In addition, the cross-entropy function is used to perform directional regression on the anchor boxes. Then the total loss function is defined as:</p>
<p><disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>loc</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>cls</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>3</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mrow><mml:mtext>dir</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:math></disp-formula>where <italic>&#x03B2;</italic><sub>1</sub>, <italic>&#x03B2;</italic><sub>2</sub> and <italic>&#x03B2;</italic><sub>3</sub> are the weighting parameters for the different loss components, set to 2.0, 1.0 and 0.2, respectively.</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Transfer Learning</title>
<p>Transfer learning uses the knowledge already learned in one domain and applies it to another related but different domain, thus avoiding the need to learn the model from scratch and providing the model with better initialised weight parameters. Rail-PillarNet is pre-trained on the KITTI dataset [<xref ref-type="bibr" rid="ref-19">19</xref>], which contains a rich variety of traffic scene objects (e.g., cars and pedestrians, etc.) in different environments such as city streets and highways. By pre-training on the KITTI dataset, the model can learn rich scene features and objects behaviours to better adapt to similar traffic scene objects in railway scenarios. The specific steps are as follows: following the training strategy of PointPillars, Rail-PillarNet is pre-trained on the KITTI dataset and the weights file is saved. Next, when training on the OSDaR23 [<xref ref-type="bibr" rid="ref-18">18</xref>] dataset, the originally saved pre-training weights are loaded and the mismatched operational layers are skipped.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experimental Datasets and Environment</title>
<sec id="s4_1">
<label>4.1</label>
<title>Experimental Datasets</title>
<p>In this paper, we use the LiDAR dataset from OSDaR23 [<xref ref-type="bibr" rid="ref-18">18</xref>], released by German Centre for Rail Traffic Research in 2023, for model training. As shown in <xref ref-type="fig" rid="fig-5">Fig. 5a</xref>, the acquisition equipment for the dataset consists of multiple calibrated and synchronised cameras, sensors such as LiDAR and millimetre-wave radar. The dataset includes 45 sequences totalling 1534 frames of data, and contains 20 categories of moving and static objects such as passengers, workers, vehicles, trains and buffer stop, as shown in <xref ref-type="fig" rid="fig-5">Fig. 5b</xref>. However, the number of samples for some of the categories in the dataset is extremely limited. Therefore, in this paper, four common and more numerous target categories: pedestrians, vehicles, trains and buffer stop are selected, and the original point cloud is filtered and sifted, resulting in a total of 1356 frames of LIDAR point cloud, divided into 1084 frames of the training set and 272 frames of the validation set with a ratio of 8:2.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>OSDaR23 acquisition equipment and part of the data. (a) Dataset acquisition equipment, (b) Example of dataset</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_54525-fig-5.tif"/>
</fig>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Experimental Environment and Setup</title>
<p>The experiments were conducted using the OpenPCDet object detection framework, the operating system was Centos 7, and the network was trained using the Nvidia TITAN RTX GPU (24G) platform with a batch size of 2. Training was performed using the Adam optimiser with an learning rate of 0.03, a weight decay value of 0.01, and a momentum value of 0.9. The <italic>X</italic>-<italic>Y</italic>-<italic>Z</italic> dimensions of each voxel are set to [0.16, 0.16, 12] m. The maximum number of pillar is 16,000, and each pillar contains 32 points.</p>
<p>In addition, due to the very long probing distance of the LIDAR in the dataset, the detection range of the setup point cloud is limited to [&#x2212;30, 354] m along the <italic>X</italic>-axis, [&#x2212;25, 39] m along the <italic>Y</italic>-axis and [&#x2212;4, 8] m along the <italic>Z</italic>-axis, and the points beyond this range are excluded from consideration. The anchor boxes dimensions for vehicles, pedestrians, buffer stop and trains were set to the average dimensions of the corresponding ground-truth boxes, the length, width and height ([l, w, h]) were [4.29, 3.06, 2.78] m, [0.88, 0.85, 1.89] m, [1.78, 2.85, 2.50] m and [62.95, 3.99, 4.28] m, respectively.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Experimental Results and Analysis</title>
<sec id="s5_1">
<label>5.1</label>
<title>Ablation Experiments</title>
<p>We perform ablation experiments on OSDaR23 to validate the contribution of each module in Rail-PillarNet to model performance improvement. In the ablation experiments, the accuracy of the model based on 40 recall thresholds under the IoU of 0.5 is mainly used as the evaluation indicator, and the ablation experiments include PAPE, fine-designed backbone network and transfer learning. <xref ref-type="table" rid="table-2">Table 2</xref> shows the experimental results.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Comparison of ablation study results</title>
</caption>
<table frame="hsides">
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead valign="top">
<tr>
<th rowspan="2">Model</th>
<th rowspan="2">PA</th>
<th rowspan="2">CA</th>
<th rowspan="2">Backbone</th>
<th rowspan="2">TL</th>
<th align="center" colspan="2">Vehicles (AP/%)</th>
<th align="center" colspan="2">Pedestrians (AP/%)</th>
<th align="center" colspan="2">Buffer stop (AP/%)</th>
<th align="center" colspan="2">Trains (AP/%)</th>
</tr>
<tr>
<th>0.7</th>
<th>0.5</th>
<th>0.5</th>
<th>0.25</th>
<th>0.7</th>
<th>0.5</th>
<th>0.7</th>
<th>0.5</th>
</tr>
</thead>
<tbody>
<tr>
<td>Model A</td>
<td></td>
<td></td>
<td></td>
<td></td>
<td>30.56</td>
<td>61.58</td>
<td>16.42</td>
<td>20.47</td>
<td>43.88</td>
<td>65.64</td>
<td>20.77</td>
<td>39.07</td>
</tr>
<tr>
<td>Model B</td>
<td>&#x221A;</td>
<td></td>
<td></td>
<td></td>
<td>32.39</td>
<td>63.37</td>
<td>14.68</td>
<td>18.95</td>
<td>45.48</td>
<td>59.31</td>
<td>23.78</td>
<td>45.18</td>
</tr>
<tr>
<td>Model C</td>
<td></td>
<td>&#x221A;</td>
<td></td>
<td></td>
<td>31.23</td>
<td>62.51</td>
<td>18.92</td>
<td>23.57</td>
<td>54.95</td>
<td>65.88</td>
<td>4.26</td>
<td>32.42</td>
</tr>
<tr>
<td>Model D</td>
<td>&#x221A;</td>
<td>&#x221A;</td>
<td></td>
<td></td>
<td>39.85</td>
<td>64.38</td>
<td>16.53</td>
<td>20.45</td>
<td>59.93</td>
<td>77.45</td>
<td>23.79</td>
<td>49.64</td>
</tr>
<tr>
<td>Model E</td>
<td></td>
<td></td>
<td>&#x221A;</td>
<td></td>
<td>34.06</td>
<td>63.02</td>
<td>14.54</td>
<td>18.81</td>
<td>61.88</td>
<td>73.69</td>
<td>24.10</td>
<td>41.70</td>
</tr>
<tr>
<td>Model F</td>
<td>&#x221A;</td>
<td>&#x221A;</td>
<td>&#x221A;</td>
<td></td>
<td>34.82</td>
<td>67.17</td>
<td>16.03</td>
<td>19.80</td>
<td>57.67</td>
<td>74.49</td>
<td>25.98</td>
<td>48.92</td>
</tr>
<tr>
<td><bold>Model G</bold></td>
<td><bold>&#x221A;</bold></td>
<td><bold>&#x221A;</bold></td>
<td><bold>&#x221A;</bold></td>
<td><bold>&#x221A;</bold></td>
<td><bold>45.42</bold></td>
<td><bold>72.52</bold></td>
<td><bold>19.95</bold></td>
<td><bold>24.30</bold></td>
<td><bold>67.60</bold></td>
<td><bold>82.60</bold></td>
<td><bold>26.01</bold></td>
<td><bold>58.97</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Using PointPillars as a baseline (Model A), point-attention (PA) and channel-attention (CA) are first individually integrated into the baseline model to form Model B and Model C. Model B has improved accuracy by 1.79% and 6.11% for the vehicle and train categories, respectively, i.e., the point-attention module is more effective for objects with larger dimensions and is better able to aggregate features in the point dimension of large objects. There is a small improvement in accuracy for small objects using the channel-attention. However, a more significant improvement was obtained by fusing the two.</p>
<p>In addition, the proposed backbone network takes into account the explicit encoding of the object in the point cloud and at the same time improves the feature extraction capability of the network through the residual structure, which increases the detection accuracy. The accuracy of vehicles, buffer stop and trains in the experiment improved by 1.44%, 8.05% and 2.63%, respectively. Finally, the effectiveness of the transfer learning (TL) training method is further analysed and the model obtains better initial weight parameters based on the results of the experiments and is able to adapt better to traffic scene targets similar to those in the KITTI dataset. Compared to Model F without transfer learning, its accuracies for vehicles and pedestrians are improved by 5.35% and 4.50%, respectively, and the detection accuracies for the remaining objects are also improved.</p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Comparison Experiments</title>
<p>Rail-PillarNet is compared with common 3D object detection algorithms to verify its effectiveness. The comparison algorithms include SECOND [<xref ref-type="bibr" rid="ref-8">8</xref>], PointPillars [<xref ref-type="bibr" rid="ref-9">9</xref>], PartA2 [<xref ref-type="bibr" rid="ref-10">10</xref>], PV-RCNN [<xref ref-type="bibr" rid="ref-11">11</xref>], PillarNet [<xref ref-type="bibr" rid="ref-24">24</xref>], Voxel RCNN [<xref ref-type="bibr" rid="ref-26">26</xref>], and Centerpoint [<xref ref-type="bibr" rid="ref-28">28</xref>].</p>
<p><xref ref-type="table" rid="table-3">Table 3</xref> shows the experimental results, where <italic>P</italic> is the number of model parameter values. Rail-PillarNet achieves satisfactory performance, where the average accuracy (mAP@0.5) reaches 58.51%, and the accuracies for each objective are 72.52%, 24.30%, 82.60% and 58.97%, respectively. Compared to PointPillars, the proposed method achieves an improvement of 12.83% in the average accuracy and about 10.94%, 3.53%, 16.96% and 19.90% for each category. SECOND consists of the 3D convolution, which achieves an average accuracy of 58.91%. This is attributed to the power of 3D convolution, but the high amount of 3D convolutional computation has a negative impact on the real-time performance of the model. However, Rail-PillarNet is able to achieve a significant increase in detection performance with a small increase in the number of parameters.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Comparison of the results of each model in the OSDaR23 dataset</title>
</caption>
<table frame="hsides">
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th rowspan="2">Model</th>
<th align="center" colspan="2">Vehicles (AP/%)</th>
<th align="center" colspan="2">Pedestrians (AP/%)</th>
<th align="center" colspan="2">Buffer stop (AP/%)</th>
<th align="center" colspan="2">Trains (AP/%)</th>
<th rowspan="2">mAP@<break/>0.5/%</th>
<th rowspan="2"><italic>P</italic>/M</th>
<th rowspan="2"><italic>T</italic><sub><italic>avg</italic></sub><italic>/</italic><break/>ms</th>
</tr>
<tr>
<th>0.7</th>
<th>0.5</th>
<th>0.5</th>
<th>0.25</th>
<th>0.7</th>
<th>0.5</th>
<th>0.7</th>
<th>0.5</th>
</tr>
</thead>
<tbody>
<tr>
<td>PointPillars</td>
<td>30.56</td>
<td>61.58</td>
<td>16.42</td>
<td>20.47</td>
<td>43.88</td>
<td>65.64</td>
<td>20.77</td>
<td>39.07</td>
<td>45.68</td>
<td>4.85</td>
<td>26.57</td>
</tr>
<tr>
<td>SECOND</td>
<td>39.54</td>
<td>74.53</td>
<td>15.16</td>
<td>20.36</td>
<td>43.10</td>
<td>73.58</td>
<td>31.32</td>
<td>72.35</td>
<td>58.91</td>
<td>9.83</td>
<td>57.96</td>
</tr>
<tr>
<td>PartA2_anchor</td>
<td>34.03</td>
<td>58.94</td>
<td>15.50</td>
<td>22.49</td>
<td>44.42</td>
<td>69.52</td>
<td>25.78</td>
<td>54.48</td>
<td>49.61</td>
<td>63.83</td>
<td>346.09</td>
</tr>
<tr>
<td>Centerpoint</td>
<td>41.69</td>
<td>49.09</td>
<td>13.66</td>
<td>20.21</td>
<td>8.61</td>
<td>9.64</td>
<td>20.16</td>
<td>44.42</td>
<td>29.20</td>
<td>8.89</td>
<td>58.23</td>
</tr>
<tr>
<td>Voxel RCNN</td>
<td>35.72</td>
<td>47.39</td>
<td>15.24</td>
<td>18.13</td>
<td>5.77</td>
<td>8.65</td>
<td>18.40</td>
<td>29.41</td>
<td>25.17</td>
<td>16.75</td>
<td>108.41</td>
</tr>
<tr>
<td>PillarNet</td>
<td>12.67</td>
<td>38.28</td>
<td>7.36</td>
<td>15.68</td>
<td>11.27</td>
<td>25.65</td>
<td>23.86</td>
<td>44.51</td>
<td>28.95</td>
<td>11.00</td>
<td>73.20</td>
</tr>
<tr>
<td>PV-RCNN</td>
<td>21.44</td>
<td>54.12</td>
<td>0.06</td>
<td>2.20</td>
<td>9.07</td>
<td>35.65</td>
<td>23.85</td>
<td>39.03</td>
<td>32.20</td>
<td>13.08</td>
<td>76.10</td>
</tr>
<tr>
<td><bold>Rail-PillarNet</bold></td>
<td><bold>45.42</bold></td>
<td><bold>72.52</bold></td>
<td><bold>19.95</bold></td>
<td><bold>24.30</bold></td>
<td><bold>67.60</bold></td>
<td><bold>82.60</bold></td>
<td><bold>26.01</bold></td>
<td><bold>58.97</bold></td>
<td><bold>58.51</bold></td>
<td><bold>5.49</bold></td>
<td><bold>31.77</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In addition, the proposed method is tested using a single TITAN RTX GPU (24 G) with an input point cloud size of 189,069 &#x00D7; 4, where 189,069 is the total number of points in the point cloud and 4 is the feature carried by each point (real world coordinates and reflectivity). In the <xref ref-type="table" rid="table-3">Table 3</xref>, <italic>T</italic><sub><italic>avg</italic></sub> is the average inference time in milliseconds (ms) for 500 repetitions. The inference time of the proposed method on TITAN RTX GPU is 31.77 ms, which is ahead of the SECOND model with the best accuracy and also significantly ahead of other models (PartA2, Centerpoint, PV-RCNN, etc.). In conclusion, the proposed method has satisfactory effect accuracy and real-time performance.</p>

</sec>
<sec id="s5_3">
<label>5.3</label>
<title>Qualitative Experiments</title>
<p>Qualitative experiments are conducted to further evaluate Rail-PillarNet. As shown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>, the detection results of Rail-PillarNet under different scenes [<xref ref-type="bibr" rid="ref-29">29</xref>] in the validation set are shown. The upper left part of the figure shows the corresponding camera image under the scene, the lower left part shows the 3D object detection results, and the right part shows the detection results under the bird&#x2019;s eye view, where the ground-truth box is green and the predicted box is red. As shown in <xref ref-type="fig" rid="fig-6">Fig. 6a</xref>, for the scene with sparse objects and small objects at a long distance, this paper&#x2019;s algorithm is able to detect most of the objects, but for some small objects at a long distance, this paper&#x2019;s method also has some leakage detection. As shown in <xref ref-type="fig" rid="fig-6">Fig. 6b</xref>, for the scene with denser objects and close distance, this paper&#x2019;s method has better detection results. However, due to the near point cloud is too dense, resulting in other objects similar to the target, also produces a certain amount of misdetection.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Visual inspection results of Rail-PillarNet in different scenarios. (a) Long distance scene, (b) Close distance scene, (c) Foggy weather scene, (d) Thick smoke scene
</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_54525-fig-6.tif"/>
</fig>
<p>In addition, we perform robustness tests in different weather lighting conditions. As shown in <xref ref-type="fig" rid="fig-6">Fig. 6c</xref>, under foggy weather conditions, the proposed method is able to better detect pedestrians located on the platform. As shown in <xref ref-type="fig" rid="fig-6">Fig. 6d</xref>, Rail-PillarNet is also able to detect foreign objects under the condition of thick smoke obscuration. In conclusion, Rail-PillarNet achieves satisfactory results under different lighting conditions.</p>

<p>The detection results of Rail-PillarNet are compared with the PointPillars in different scenarios. As shown in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>, where the first column is the corresponding camera image in the scene, the second column is the 3D object detection result of the PointPillars algorithm, and the third column is the 3D object detection effect of Rail-PillarNet, where the ground-truth box is green and the predicted box is red. From the first and second rows of the figure, it can be seen that the PointPillars algorithm has more misdetections in the sparse object and small object far away scenarios, as shown in the dashed red box in the figure. Compared to PointPillars, Rail-PillarNet reduces the false alarm rate in the long range scenarios. From the third and fourth rows of the figure, it can be seen that the PointPillars algorithm has a certain amount of leakage and misdetection in the scenes with denser objects and close distances, as shown in the dashed red box in the figure, Rail-PillarNet reduces the false and missed detections in the close range scenario.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Comparison of visual inspection results between Rail-PillarNet and PointPillars</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_54525-fig-7.tif"/>
</fig>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusions</title>
<p>In this paper, a LiDAR railway object detection method, Rail-PillarNet, is proposed. Firstly, the PAPE is proposed to mitigate the loss of fine-grained information during the PointPillars point cloud encoding, obtain finer pillar features, and effectively improve detection accuracy. Secondly, a fine backbone network is designed by combining the LIDAR point cloud coding characteristics and the residual structure, and effectively improves the feature extraction capability of the network. Finally, the initialisation weight parameters of the model are optimized using the transfer learning training method, which further improves the detection accuracy.</p>
<p>In summary, the experimental results on the OsDar23 dataset demonstrate the effectiveness of the above method. the algorithm compared with PointPillars accuracy increased by 12.83%, and the number of parameters increased by only 0.64 M. In addition, the proposed method also achieved a satisfactory performance compared to other mainstream 3D object detection models.</p>
<p>However, the methodology in this paper focuses on exploring a perceptual approach, which has limitations in terms of decision-making for trains. Future work will consider the detection of track regions based on a priori knowledge to delineate the intrusion limit regions. Next, the detection results of potential foreign objects are combined to determine their location in the real world. Finally, based on the location of the potential foreign object and the delineated intrusion limit regions, it is determined whether an intrusion has occurred. The foreign object intrusion information is sent to the control room for train control commands.</p>
</sec>
</body>
<back>
<ack>
<p>Thanks are extended to the editors and reviewers.</p>
</ack>
<sec><title>Funding Statement</title>
<p>This work was supported by a grant from the National Key Research and Development Project (2023YFB4302100), Key Research and Development Project of Jiangxi Province (No. 20232ACE01011) and Independent Deployment Project of Ganjiang Innovation Research Institute, Chinese Academy of Sciences (E255J001).</p>
</sec>
<sec><title>Author Contributions</title>
<p>Conceptualization: Fan Li and Jie Yang; Data curation: Zhichao Chen and Zhicheng Feng; Investigation, Fan Li and Shuyao Zhang; Methodology, Fan Li and and Shuyao Zhang; Writing original draft, Fan Li and Jie Yang; Writing&#x2014;review: Zhichao Chen and Zhicheng Feng. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability"><title>Availability of Data and Materials</title>
<p>The experiments were performed on OSDaR23 and KITTI datasets. Here are the links to the OSDaR23 and KITTI datasets, respectively. data.fid-move.de/dataset/ osdar23 (accessed on 1 April 2024). <ext-link ext-link-type="uri" xlink:href="http://www.cvlibs.net/datasets/kitti">http://www.cvlibs.net/datasets/kitti</ext-link> (accessed on 1 April 2024).</p>
</sec>
<sec><title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare they have no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Ye</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Zhao</surname></string-name>, and <string-name><given-names>F.</given-names> <surname>Zhou</surname></string-name></person-group>, &#x201C;<article-title>Foreign body detection in rail transit based on a multi-mode feature enhanced convolutional neural network</article-title>,&#x201D; <source>IEEE Trans. Intell. Transp. Syst.</source>, vol. <volume>23</volume>, no. <issue>10</issue>, pp. <fpage>18051</fpage>&#x2013;<lpage>18063</lpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.1109/TITS.2022.3154751</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Feng</surname></string-name>, and <string-name><given-names>L.</given-names> <surname>Jia</surname></string-name></person-group>, &#x201C;<article-title>Efficient railway track region segmentation algorithm based on lightweight neural network and cross-fusion decoder</article-title>,&#x201D; <source>Autom. Constr.</source>, vol. <volume>155</volume>, <year>2023</year>, <comment>Art. no. 105069</comment>. doi: <pub-id pub-id-type="doi">10.1016/j.autcon.2023.105069</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Feng</surname></string-name>, and <string-name><given-names>H.</given-names> <surname>Zhu</surname></string-name></person-group>, &#x201C;<article-title>RailFOD23: A dataset for foreign object detection on railroad transmission lines</article-title>,&#x201D; <source>Sci. Data</source>, vol. <volume>11</volume>, no. <issue>1</issue>, <year>2024</year>, <comment>Art. no. 72</comment>. doi: <pub-id pub-id-type="doi">10.1038/s41597-024-02918-9</pub-id>; <pub-id pub-id-type="pmid">38228610</pub-id></mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Feng</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Chen</surname></string-name>, and <string-name><given-names>Z.</given-names> <surname>Kang</surname></string-name></person-group>, &#x201C;<article-title>LRseg: An efficient railway region extraction method based on lightweight encoder and self-correcting decoder</article-title>,&#x201D; <source>Expert. Syst. Appl.</source>, vol. <volume>238</volume>, <year>2024</year>, <comment>Art. no. 122386</comment>. doi: <pub-id pub-id-type="doi">10.1016/j.eswa.2023.122386</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Du</surname></string-name>, and <string-name><given-names>H.</given-names> <surname>Bao</surname></string-name></person-group>, &#x201C;<article-title>Impact of automation at different cognitive stages on high-speed train driving performance</article-title>,&#x201D; <source>IEEE Trans. Intell. Transp. Syst.</source>, vol. <volume>23</volume>, no. <issue>12</issue>, pp. <fpage>24599</fpage>&#x2013;<lpage>24608</lpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.1109/TITS.2022.3211709</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Vatavu</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>From particles to self-localizing tracklets: A multilayer particle filter-based estimation for dynamic grid maps</article-title>,&#x201D; <source>IEEE Intell. Transp. Syst. Mag.</source>, vol. <volume>12</volume>, no. <issue>4</issue>, pp. <fpage>149</fpage>&#x2013;<lpage>168</lpage>, <year>2020</year>. doi: <pub-id pub-id-type="doi">10.1109/MITS.2020.3014428</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Xie</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Xu</surname></string-name>, and <string-name><given-names>R.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>Obstacle detection and tracking method for autonomous vehicle based on three dimensional lidar</article-title>,&#x201D; <source>Int. J. Adv. Robot. Syst.</source>, vol. <volume>16</volume>, no. <issue>2</issue>, <year>2019</year>. doi: <pub-id pub-id-type="doi">10.1177/1729881419831587</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Yan</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Mao</surname></string-name>, and <string-name><given-names>B.</given-names> <surname>Li</surname></string-name></person-group>, &#x201C;<article-title>SECOND: Sparsely embedded convolutional detection</article-title>,&#x201D; <source>Sensors</source>, vol. <volume>18</volume>, no. <issue>10</issue>, <year>2018</year>, <comment>Art. no. 3337</comment>. doi: <pub-id pub-id-type="doi">10.3390/s18103337</pub-id>; <pub-id pub-id-type="pmid">30301196</pub-id></mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A. H.</given-names> <surname>Lang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Vora</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Caesar</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Yang</surname></string-name> and <string-name><given-names>O.</given-names> <surname>Beijbom</surname></string-name></person-group>, &#x201C;<article-title>PointPillars: Fast encoders for object detection from point clouds</article-title>,&#x201D; in <conf-name>2019 IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR)</conf-name>, <publisher-loc>Long Beach, CA, USA</publisher-loc>, <year>2019</year>, pp. <fpage>12689</fpage>&#x2013;<lpage>12697</lpage>. doi: <pub-id pub-id-type="doi">10.1109/CVPR.2019.01298</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Shi</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Shi</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Wang</surname></string-name>, and <string-name><given-names>H.</given-names> <surname>Li</surname></string-name></person-group>, &#x201C;<article-title>From points to parts: 3D object detection from point cloud with part aware and part-aggregation network</article-title>,&#x201D; <year>2020</year>, <comment><italic>arXiv:1907.03670</italic>.</comment></mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Shi</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>PV-RCNN: Point-voxel feature set abstraction for 3D object detection</article-title>,&#x201D; in <conf-name>2020 IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR)</conf-name>, <publisher-loc>Seattle, WA, USA</publisher-loc>, <year>2020</year>, pp. <fpage>10526</fpage>&#x2013;<lpage>10535</lpage>. doi: <pub-id pub-id-type="doi">10.1109/CVPR42600.2020.01054</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Yuan</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Yu</surname></string-name>, and <string-name><given-names>X.</given-names> <surname>Liu</surname></string-name></person-group>, &#x201C;<article-title>A rail detection method using bird&#x2019;s eye view of LiDAR point clouds</article-title>,&#x201D; in <conf-name>2023 6th Int. Conf. Electron. Technol. (ICET)</conf-name>, <publisher-loc>Chengdu, China</publisher-loc>, <year>2023</year>, pp. <fpage>120</fpage>&#x2013;<lpage>126</lpage>. doi: <pub-id pub-id-type="doi">10.1109/ICET58434.2023.10211715</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Yuwen</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Wan</surname></string-name>, and <string-name><given-names>Y.</given-names> <surname>Luo</surname></string-name></person-group>, &#x201C;<article-title>A lightweight lidar-camera sensing method of obstacles detection and classification for autonomous rail rapid transit</article-title>,&#x201D; <source>IEEE Trans. Intell. Transp. Syst.</source>, vol. <volume>23</volume>, no. <issue>12</issue>, pp. <fpage>23043</fpage>&#x2013;<lpage>23058</lpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.1109/TITS.2022.3194553</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Yu</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Li</surname></string-name>, and <string-name><given-names>D.</given-names> <surname>Li</surname></string-name></person-group>, &#x201C;<article-title>A camera and lidar data fusion method for railway object detection</article-title>,&#x201D; <source>IEEE Sens. J.</source>, vol. <volume>21</volume>, no. <issue>12</issue>, pp. <fpage>13442</fpage>&#x2013;<lpage>13454</lpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.1109/JSEN.2021.3066714</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Wen</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Peng</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Lin</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Gan</surname></string-name>, and <string-name><given-names>R.</given-names> <surname>Tan</surname></string-name></person-group>, &#x201C;<article-title>Multi-modal contrastive learning for lidar point cloud rail-obstacle detection in complex weather</article-title>,&#x201D; <source>Electronics</source>, vol. <volume>13</volume>, no. <issue>1</issue>, <year>2024</year>, <comment>Art. no. 220</comment>. doi: <pub-id pub-id-type="doi">10.3390/electronics13010220</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Neri</surname></string-name> and <string-name><given-names>F.</given-names> <surname>Battisti</surname></string-name></person-group>, &#x201C;<article-title>3D object detection on synthetic point clouds for railway applications</article-title>,&#x201D; in <conf-name>2022 10th Eur. Workshop Vis. Inform. Process. (EUVIP)</conf-name>, <publisher-name>IEEE</publisher-name>, <year>2022</year>, pp. <fpage>1</fpage>&#x2013;<lpage>6</lpage>. doi: <pub-id pub-id-type="doi">10.1109/EUVIP53989.2022.9922901</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Wisultschew</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Mujica</surname></string-name>, <string-name><given-names>J. M.</given-names> <surname>Lanza-Gutierrez</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Portilla</surname></string-name></person-group>, &#x201C;<article-title>3D-lidar based object detection and tracking on the edge of IoT for railway level crossing</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>9</volume>, pp. <fpage>35718</fpage>&#x2013;<lpage>35729</lpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.1109/ACCESS.2021.3062220</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Tagiew</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>OSDaR23: Open sensor data for rail 2023</article-title>,&#x201D; in <conf-name>2023 8th Int. Conf. Robot. Autom. Eng. (ICRAE)</conf-name>, <publisher-loc>Singapore</publisher-loc>, <year>2023</year>, pp. <fpage>270</fpage>&#x2013;<lpage>276</lpage>. doi: <pub-id pub-id-type="doi">10.1109/ICRAE59816.2023.10458449</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Geiger</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Lenz</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Stiller</surname></string-name>, and <string-name><given-names>R.</given-names> <surname>Urtasun</surname></string-name></person-group>, &#x201C;<article-title>Vision meets robotics: The KITTI dataset</article-title>,&#x201D; <source>Int. J. Rob. Res.</source>, vol. <volume>32</volume>, pp. <fpage>1231</fpage>&#x2013;<lpage>1237</lpage>, <year>2013</year>. doi: <pub-id pub-id-type="doi">10.1177/0278364913491297</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>R. Q.</given-names> <surname>Charles</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Su</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Kaichun</surname></string-name>, and <string-name><given-names>L. J.</given-names> <surname>Guibas</surname></string-name></person-group>, &#x201C;<article-title>PointNet: Deep learning on point sets for 3D classification and segmentation</article-title>,&#x201D; in <conf-name>2017 IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR)</conf-name>, <publisher-loc>Honolulu, HI, USA</publisher-loc>, <year>2017</year>, pp. <fpage>77</fpage>&#x2013;<lpage>85</lpage>. doi: <pub-id pub-id-type="doi">10.1109/CVPR.2017.16</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Yang</surname></string-name>, and <string-name><given-names>C.</given-names> <surname>Wang</surname></string-name></person-group>, &#x201C;<article-title>PillarNet&#x002B;&#x002B;: Pillar-based 3-D object detection with multiattention</article-title>,&#x201D; <source>IEEE Sens. J.</source>, vol. <volume>23</volume>, no. <issue>22</issue>, pp. <fpage>27733</fpage>&#x2013;<lpage>27743</lpage>, <year>2023</year>. doi: <pub-id pub-id-type="doi">10.1109/JSEN.2023.3323368</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Simonyan</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Zisserman</surname></string-name></person-group>, &#x201C;<article-title>Very deep convolutional networks for large-scale image recognition</article-title>,&#x201D; <year>2014</year>, <comment><italic>arXiv:1409.1556</italic></comment>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>He</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Ren</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Sun</surname></string-name></person-group>, &#x201C;<article-title>Deep residual learning for image recognition</article-title>,&#x201D; in <conf-name>2016 IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR)</conf-name>, <publisher-loc>Las Vegas, NV, USA</publisher-loc>, <year>2016</year>, pp. <fpage>770</fpage>&#x2013;<lpage>778</lpage>. doi: <pub-id pub-id-type="doi">10.1109/CVPR.2016.90</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Shi</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Li</surname></string-name>, and <string-name><given-names>C.</given-names> <surname>Ma</surname></string-name></person-group>, &#x201C;<article-title>PillarNet: Real-time and high-performance pillar-based 3D object detection</article-title>,&#x201D; in <conf-name>Comput. Vis.&#x2013;ECCV (Eur. Conf. Comput. Vis.) 2022</conf-name>, <year>2022</year>, vol. <volume>13670</volume>, pp. <fpage>35</fpage>&#x2013;<lpage>52</lpage>. doi: <pub-id pub-id-type="doi">10.1007/978-3-031-20080-9_3</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Zhou</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>FastPillars: A deployment-friendly pillar-based 3D detector</article-title>,&#x201D; <year>2023</year>, <comment><italic>arXiv:2302.02367</italic></comment>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Deng</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Shi</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Li</surname></string-name></person-group>, &#x201C;<article-title>Voxel R-CNN: Towards high performance voxel-based 3D object detection</article-title>,&#x201D; in <conf-name> Proc. AAAI (Assoc. Adv. Artif. Intell.) Conf. Artif. Intell.</conf-name>, <year>2021</year>, vol. <volume>35</volume>, pp. <fpage>1201</fpage>&#x2013;<lpage>1209</lpage>. doi: <pub-id pub-id-type="doi">10.1609/aaai.v35i2.16207</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Liu</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>SSD: Single shot multibox detector</article-title>,&#x201D; in <conf-name>Comput. Vis.&#x2013;ECCV (Eur. Conf. Comput. Vis.) 2016</conf-name>, <year>2016</year>, vol. <volume>9905</volume>, pp. <fpage>21</fpage>&#x2013;<lpage>37</lpage>.doi: <pub-id pub-id-type="doi">10.1007/978-3-319-46448-0_2</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Yin</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Zhou</surname></string-name>, and <string-name><given-names>P.</given-names> <surname>Kr&#x00E4;henb&#x00FC;hl</surname></string-name></person-group>, &#x201C;<article-title>Center-based 3D object detection and tracking</article-title>,&#x201D; in <conf-name>2021 IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR)</conf-name>, <publisher-loc>Nashville, TN, USA</publisher-loc>, <year>2021</year>, pp. <fpage>11779</fpage>&#x2013;<lpage>11788</lpage>. doi: <pub-id pub-id-type="doi">10.1109/CVPR46437.2021.01161</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z. C.</given-names> <surname>Chen</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Fast vehicle detection algorithm in traffic scene based on improved SSD</article-title>,&#x201D; <source>Measurement</source>, vol. <volume>201</volume>, <year>2022</year>, <comment>Art. no. 111655</comment>. doi: <pub-id pub-id-type="doi">10.1016/j.measurement.2022.111655</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>