<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMES</journal-id>
<journal-id journal-id-type="nlm-ta">CMES</journal-id>
<journal-id journal-id-type="publisher-id">CMES</journal-id>
<journal-title-group>
<journal-title>Computer Modeling in Engineering &#x0026; Sciences</journal-title>
</journal-title-group>
<issn pub-type="epub">1526-1506</issn>
<issn pub-type="ppub">1526-1492</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">20995</article-id>
<article-id pub-id-type="doi">10.32604/cmes.2022.020995</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>A Multi Moving Target Recognition Algorithm Based on Remote Sensing Video</article-title>
<alt-title alt-title-type="left-running-head">A Multi Moving Target Recognition Algorithm Based on Remote Sensing Video</alt-title>
<alt-title alt-title-type="right-running-head">A Multi Moving Target Recognition Algorithm Based on Remote Sensing Video</alt-title>
</title-group>
<contrib-group content-type="authors">
<contrib id="author-1" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Zheng</surname><given-names>Huanhuan</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><email>zhenghuan@yulinu.edu.cn</email></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Bai</surname><given-names>Yuxiu</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Tian</surname><given-names>Yurun</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<aff id="aff-1"><label>1</label><institution>School of Information Engineering, Yulin University</institution>, <addr-line>Yulin</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>ZTE Communication Co., Ltd.</institution>, <addr-line>Xi&#x0027;an</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Huanhuan Zheng. Email: <email>zhenghuan@yulinu.edu.cn</email></corresp>
</author-notes>
<pub-date pub-type="epub" date-type="pub" iso-8601-date="2022-08-11"><day>11</day>
<month>08</month>
<year>2022</year></pub-date>
<volume>134</volume>
<issue>1</issue>
<fpage>585</fpage>
<lpage>597</lpage>
<history>
<date date-type="received"><day>11</day><month>12</month><year>2021</year></date>
<date date-type="accepted"><day>11</day><month>2</month><year>2022</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2023 Zheng et al.</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Zheng et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMES_20995.pdf"></self-uri>
<abstract>
<p>The Earth observation remote sensing images can display ground activities and status intuitively, which plays an important role in civil and military fields. However, the information obtained from the research only from the perspective of images is limited, so in this paper we conduct research from the perspective of video. At present, the main problems faced when using a computer to identify remote sensing images are: They are difficult to build a fixed regular model of the target due to their weak moving regularity. Additionally, the number of pixels occupied by the target is not enough for accurate detection. However, the number of moving targets is large at the same time. In this case, the main targets cannot be recognized completely. This paper studies from the perspective of Gestalt vision, transforms the problem of moving target detection into the problem of salient region probability, and forms a Saliency map algorithm to extract moving targets. On this basis, a convolutional neural network with global information is constructed to identify and label the target. And the experimental results show that the algorithm can extract moving targets and realize moving target recognition under many complex conditions such as target&#x0027;s long-term stay and small-amplitude movement.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Deep learning</kwd>
<kwd>remote sensing images</kwd>
<kwd>moving target</kwd>
<kwd>recognition</kwd>
<kwd>salient</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1"><label>1</label><title>Introduction</title>
<p>Remote sensing images can intuitively display wide view scene information, and are widely applied in the fields. It can be applied to multi-target recognition and tracking scenes such as battlefield reconnaissance, border patrol, post-disaster rescue, public transportation and more [<xref ref-type="bibr" rid="ref-1">1</xref>]. However, it is difficult to obtain spaceborne remote sensing images, which limits the in-depth research [<xref ref-type="bibr" rid="ref-2">2</xref>]. In recent years, with the development and popularization of UAV technology, UAV technology presents the characteristics of high efficiency, flexibility and low cost, and the imaging equipment carried tends to be mature. It has been widely used in battlefield reconnaissance, border patrol, post disaster rescue, public transportation, etc. [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-4">4</xref>], which makes the aerial photography data present a blowout situation. In the face of so many aerial remote sensing data, how we can obtain useful information from aerial video is an important direction in the field of computer vision [<xref ref-type="bibr" rid="ref-5">5</xref>].</p>
<p>Target detection is the basis of tracking, recognition and image interpretation [<xref ref-type="bibr" rid="ref-6">6</xref>]. There are multiple moving targets in the video, just because the video field angle captured by UAV monitoring system is large. Thus, it is difficult to detect moving targets quickly and accurately. Xiao et al. [<xref ref-type="bibr" rid="ref-7">7</xref>] proposed a low frame rate aerial video vehicle detection and tracking method based on joint probability relation graph according to probability relation. Andriluka et al. [<xref ref-type="bibr" rid="ref-8">8</xref>] applied UAV technology to rescue and search. Cheng et al. [<xref ref-type="bibr" rid="ref-9">9</xref>] established a dynamic Bayesian network from the perspective of pixels to realize vehicle detection. Gaszczak et al. [<xref ref-type="bibr" rid="ref-10">10</xref>] established a model from the perspective of thermal imaging to detect pedestrians and vehicles in aerial images. Lin et al. [<xref ref-type="bibr" rid="ref-11">11</xref>] improved the traditional Hough transform to realize road detection. Rodr&#x00ED;guez-Canosa et al. [<xref ref-type="bibr" rid="ref-12">12</xref>] established a model to solve the problem of aerial image jitter to a certain extent. Zheng et al. [<xref ref-type="bibr" rid="ref-13">13</xref>] established GIS vector map to realize vehicle detection. Liang et al. [<xref ref-type="bibr" rid="ref-14">14</xref>] guided target detection according to background information. Prokaj et al. [<xref ref-type="bibr" rid="ref-15">15</xref>] built a dual tracker to realize the construction of background and prospect. Teutsch et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] established a model under the condition of low contrast to realize vehicle detection in aerial monitoring images. Chen et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] optimized the tracking window to realize multi-target detection. Jiang et al. [<xref ref-type="bibr" rid="ref-18">18</xref>] introduced the prediction module to realize target tracking in order to enhance the stability of tracking. Poostchi et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] established semantic depth map fusion for moving vehicle detection. Tang et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] established single revolutionary neural networks to predict vehicle direction. Aguilar et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] used cascade classifiers with mean shift to realize pedestrian detection. Xu et al. [<xref ref-type="bibr" rid="ref-22">22</xref>] established a prediction and feature point selection model to find moving targets in multi-scale on infrared aerial images. Hamsa et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] established a cascaded support vector machine and Gaussian mixture model to realize vehicle detection (SVM &#x002B; GMM). Ma et al. [<xref ref-type="bibr" rid="ref-24">24</xref>] established the rotation invariant cascaded forest (RICF) to meet the target detection in complex background. Mandal et al. [<xref ref-type="bibr" rid="ref-25">25</xref>] established simple short and shallow network to realize rapid target detection. Song et al. [<xref ref-type="bibr" rid="ref-26">26</xref>] built a model according to the time and space relationship of the target, and then built the regulated AdaBoost recognition model to realize target recognition. Qiu et al. [<xref ref-type="bibr" rid="ref-27">27</xref>] built a deep learning network to track moving targets. Feng et al. [<xref ref-type="bibr" rid="ref-28">28</xref>] used mean shift algorithm to track high-speed targets. Wan et al. [<xref ref-type="bibr" rid="ref-29">29</xref>] used Keystone Transform and Modified Second-Order Keystone Transform to achieve moving target tracking. Lin et al. [<xref ref-type="bibr" rid="ref-30">30</xref>] used multiple drones to achieve multi-target tracking.</p>
<p>The above algorithms have achieved certain results in target detection, but there are still deficiencies: The model considers the limited interference of noise, tree disturbance and other factors, which leads to the inaccurate extraction of moving targets. Insufficient mining target attributes lead to inaccurate recognition. The innovation of our algorithm shown as following aspects: firstly, from the perspective of Gestalt vision, we propose a target motion extraction algorithm based on the saliency graph theory; secondly, in order to achieve fast and accurate target recognition, we constructed the convolutional neural network structure with global information, especially take global information into consideration.</p>
</sec>
<sec id="s2"><label>2</label><title>Algorithm</title>
<p>Gestalt visual school believes that the reason why things are perceived is the result of the public action of eyes and brain. Firstly, the image is obtained through the eyes, and then the objects are combined according to some rules to form an easy to understand unity. If it cannot be combined, it will appear in a disordered state, resulting in incorrect cognition. Based on this principle, a multi-target recognition process in accordance with the Gestalt vision principle is established, as shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. Firstly, the salient graph mechanism is established to extract moving targets, and then the convolution neural network structure based on global information is used to realize target recognition.</p>
<fig id="fig-1"><label>Figure 1</label><caption><title>Algorithm pipeline</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_20995-fig-1.png"/></fig>
<sec id="s2_1"><label>2.1</label><title>Moving Target Extraction Based on Salient Gragh</title>
<p>In the video scene, a large amount of short-term motion information is included between consecutive image frames, and ignores the information that does not move temporarily. A video with long time contains a lot of long-time motion information, so the conditional motion salient graph includes the motion saliency of the target and the motion saliency of the background. In addition, the interference of noise makes it difficult to distinguish the significance.</p>
<p>A time series group containing short-term motion information and long-term motion information is constructed on the time scale. By calculating the motion significance probability, the significance of the moving target is highlighted and the background significance is suppressed, as shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>.</p>
<fig id="fig-2"><label>Figure 2</label><caption><title>Pipeline of moving target extraction algorithm based on salient graph</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_20995-fig-2.png"/></fig>
<p>Temporal Fourier transform (TFT) is a motion saliency detection algorithm based on time information which uses pixels at the same position in consecutive frames to integrate in time to form a time series. The waveform is reconstructed by Fourier transform and inverse Fourier transform, and the maximum point is marked on the salient graph. The significant values indicate the probability that the target belongs to the foreground. Let the time series be
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msubsup><mml:mi>D</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>+</mml:mo><mml:mi>i</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>Its corresponding Fourier transform is obtained as
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>F</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mi>D</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msubsup><mml:mi>p</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>g</mml:mi><mml:mi>l</mml:mi><mml:mi>e</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <italic>F</italic> represents Fourier transform and <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msubsup><mml:mi>p</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represents the phase spectrum of <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msubsup><mml:mi>I</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>g</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mrow><mml:msup><mml:mi>F</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mi>p</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <italic>F<sup>&#x2212;</sup></italic><sup>1</sup> represents inverse Fourier transform, and <italic>g</italic>(<italic>t</italic>) is Gaussian filter. The larger the amplitude of <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msubsup><mml:mi>I</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msubsup></mml:math></inline-formula> change, the larger the time scale change. Thus, the motion significance of the time series is constructed as
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msubsup><mml:mi>&#x03C6;</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msubsup><mml:mo>=</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:mrow><mml:msubsup><mml:mi>I</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>where &#x003C6;<sup><italic>i</italic></sup><sub><italic>x,y</italic></sub> represents the significant value of the pixel (<italic>x</italic>, <italic>y</italic>) in the <italic>i</italic>-th viewing angle sequence,&#x007C; &#x007C; is F2 norm.</p>
<p>Conditional motion saliency probability refers to the probability that pixels belong to the foreground in the time series. To unify the scale, <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msubsup><mml:mi>&#x03C6;</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msubsup></mml:math></inline-formula> is normalized as
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>F</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msubsup><mml:mi>D</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msubsup><mml:mi>&#x03C6;</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msup><mml:mi>&#x03C6;</mml:mi><mml:mi>i</mml:mi></mml:msup></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msup><mml:mi>&#x03C6;</mml:mi><mml:mi>i</mml:mi></mml:msup></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msup><mml:mi>&#x03C6;</mml:mi><mml:mi>i</mml:mi></mml:msup></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>F</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msubsup><mml:mi>D</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represents the normalized value of <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msubsup><mml:mi>&#x03C6;</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msubsup></mml:math></inline-formula>, and <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msubsup><mml:mi>&#x03C6;</mml:mi><mml:mrow></mml:mrow><mml:mi>i</mml:mi></mml:msubsup></mml:math></inline-formula> is the time series significant image composed of the current frame at time t and the <italic>i</italic>-th time slice. The motion saliency value is positively correlated with the probability that the corresponding point belongs to the foreground.</p>
<p>The motion saliency probability represents the probability that the pixel belongs to the foreground by the motion saliency of the pixel. Under the guidance of the full probability formula, it can be calculated as follows:
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>F</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>l</mml:mi></mml:munderover><mml:mrow><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>F</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo fence="false" stretchy="false">|</mml:mo><mml:msubsup><mml:mi>D</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msubsup><mml:mi>D</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula></p>
<p>The motion saliency probability graph uses the long-term and short-term motion information to enhance the saliency of the moving target in the current frame and suppress the motion saliency of the background and historical frame. Due to the background interference, there is false detection in the detection results, which makes the local significance high and difficult to remove by traditional methods. We use the correlation between adjacent pixels to construct a histogram algorithm to segment the saliency graph, then model the spatial information, and use the spatial information modeling to calculate the displacement probability, which has achieved the purpose of eliminating interference.</p>
<p>The histogram based threshold method is used to segment the motion saliency probability graph to obtain the candidate pixels,
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mi>S</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left right" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>F</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2265;</mml:mo><mml:mi>T</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>where <italic>T</italic> is obtained by the traditional Ostu algorithm. When <italic>S&#x2009;</italic>&#x003D;<italic>&#x2009;</italic>1, it indicates foreground candidate pixels. When <italic>S&#x2009;</italic>&#x003D;<italic>&#x2009;</italic>0, it indicates background candidate pixels. When the background, i.e., trees, moves, there is a risk of detection as the target. However, its inter frame motion amplitude is limited. We construct the function:
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mtable columnalign='left'><mml:mtr><mml:mtd><mml:msub><mml:mi>P</mml:mi><mml:mi>N</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munder><mml:mrow><mml:mi>max</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mi>N</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:munder><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>B</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>P</mml:mi><mml:mi>B</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>P</mml:mi><mml:mi>F</mml:mi></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>B</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the probability that the pixel (<italic>x</italic>, <italic>y</italic>) belongs to the background. The above algorithm can suppress the background, but it will also eliminate the real moving target pixels. We limit the detected connected region and measure it by considering the neighborhood position of the whole foreground target. Define the component displacement probability as the probability of a detected connected component:
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="false">&#x220F;</mml:mo><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mi>C</mml:mi></mml:mrow></mml:munder><mml:mrow><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>N</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula></p>
<p>For the connected component of the real target, the probability of the component displacement from the background is very small, and the threshold <italic>t<sub>h</sub></italic> is set to distinguish:
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2264;</mml:mo><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&gt;</mml:mo><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi>h</mml:mi></mml:msub></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>where <italic>S<sub>c</sub></italic>(<italic>x</italic>, <italic>y</italic>) &#x003D; 1 represents the foreground and <italic>S<sub>c</sub></italic>(<italic>x</italic>, <italic>y</italic>) &#x003D; 0 represents the background.</p>
</sec>
<sec id="s2_2"><label>2.2</label><title>Convolution Neural Network Based on Global Information</title>
<p>Convolutional neural network (CNN) is a kind of feedforward neural network, whose neurons carry out corresponding control on the units within the coverage. It has excellent performance in the field of large-scale image processing. Therefore, we process remote sensing images based on traditional CNN.</p>
<p>The basic structure of CNN includes feature extraction layer and feature mapping layer. Feature extraction layer: The input of each neuron is connected to the local acceptance domain of the previous layer, and the local features are extracted. When the local feature is extracted, the position relationship between it and other features is also determined. Feature mapping layer: Each computing layer of the network is composed of multiple feature maps. Each feature mapping can be regarded as a plane, and the weights of all neurons on the plane are equal. Because the feature detection layer of CNN learns from the training data, it avoids explicit feature extraction and implicitly learns from the training data. Because the weights of neurons on the same feature mapping surface are the same, the network can learn in parallel. Subsequent scholars have carried out a lot of research on the basis of CNN: Bayar et al. [<xref ref-type="bibr" rid="ref-31">31</xref>] constrained the convolution layer to meet the image target detection. Li et al. [<xref ref-type="bibr" rid="ref-32">32</xref>] used a double-layer CNN structure to detect targets. The above algorithms have improved CNN from different aspects and achieved certain results.</p>
<p>According to the particularity of remote sensing video, traditional CNN cannot be directly applied to multi-target tracking based on remote sensing video, which is not enough to capture global information. Therefore, a new convolution neural network framework based on global information is proposed by effectively combining global average pool and Atrous convolution, as shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>.</p>
<fig id="fig-3"><label>Figure 3</label><caption><title>Convolutional neural network graph based on global information</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_20995-fig-3.png"/></fig>
<p>The image sequence is convoluted and pooled to reduce the size of the image and increase the receiving domain. The merged image is restored by up sampling to the original size prediction of the image. The information in the original image is lost while zooming out and adjusting. To solve this problem, Atrous Convolution is applied, as shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>. Atrus revolution is a convolution idea proposed to solve the problem of image semantic segmentation in which down sampling will reduce image resolution and lose information. The advantages are: on the condition of loss information without pooling and the same calculation conditions, the receptive field is increased so that each convolution output contains a large range.</p>
<fig id="fig-4"><label>Figure 4</label><caption><title>Atrous convolution</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_20995-fig-4.png"/></fig>
<p>Global information plays a key role in image classification or target detection. In order to obtain more global context information, the global average pool is combined with the atrus revolution. <xref ref-type="fig" rid="fig-5">Fig. 5</xref> shows the main structure of the network.</p>
<fig id="fig-5"><label>Figure 5</label><caption><title>The network structure</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_20995-fig-5.png"/></fig>
<p>When extracting global features for feature fusion, due to different data scales at different levels, it is necessary to regularize the feature data.
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mrow><mml:mrow><mml:mi mathvariant="bold">y</mml:mi></mml:mrow></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mi mathvariant="bold">X</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="bold">X</mml:mi></mml:mrow></mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow></mml:mrow><mml:mn>2</mml:mn></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula>
 We use L2-normalize to normalize the input feature x, where y represents the normalized output vector. &#x007C;&#x007C;&#x2022; &#x007C; is L2 norm.</p>
<p>Because the value distribution of the eigenvector is uneven, the scale parameter is introduced as
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mrow><mml:mrow><mml:mi mathvariant="bold">z</mml:mi></mml:mrow></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="bold">y</mml:mi></mml:mrow></mml:mrow></mml:math></disp-formula>
 In the training process, L2 norm propagation is used to calculate the scale parameters through the chain method:
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>z</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle><mml:mi>&#x03B1;</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mi>&#x03B1;</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mi>j</mml:mi></mml:munder><mml:mrow><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mrow><mml:msub><mml:mi>z</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi mathvariant="normal">&#x2202;</mml:mi><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:msubsup><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow></mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mn>2</mml:mn><mml:mn>2</mml:mn></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>i</mml:mi><mml:mn>2</mml:mn></mml:msubsup></mml:mrow><mml:mrow><mml:msubsup><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow></mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mn>2</mml:mn><mml:mn>3</mml:mn></mml:msubsup></mml:mrow></mml:mfrac></mml:mstyle></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>where <italic>l</italic> is the objective function. L2 norm normalizes the extracted features and the added global features.</p>
<p>The median frequency balance strategy is used, and the cross entropy loss function is
<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:mi>l</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:msup><mml:mi>p</mml:mi><mml:mo>&#x2217;</mml:mo></mml:msup></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>n</mml:mi></mml:mfrac><mml:munder><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mi>i</mml:mi></mml:munder><mml:mrow><mml:mrow><mml:msup><mml:mi>p</mml:mi><mml:mo>&#x2217;</mml:mo></mml:msup></mml:mrow><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mspace width="thinmathspace" /><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mrow><mml:munder><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mi>c</mml:mi></mml:munder><mml:mrow><mml:msubsup><mml:mi>A</mml:mi><mml:mi>i</mml:mi><mml:mi>c</mml:mi></mml:msubsup></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula>as the objective function in the network training process. It is to determine the distance between the actual output P and the desired output p<sup>&#x002A;</sup>, where c is the class of the tag.</p>
<p>Because the number of pixels in each category in the training data varies greatly, different weights are required according to the actual category. In order to obtain better results, the median frequency balance is proposed, and <xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref> is rewritten as follows:
<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mi>l</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>p</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:msup><mml:mi>p</mml:mi><mml:mo>&#x2217;</mml:mo></mml:msup></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mi>n</mml:mi></mml:mfrac></mml:mstyle><mml:munder><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mi>i</mml:mi></mml:munder><mml:mrow><mml:mrow><mml:msup><mml:mi>p</mml:mi><mml:mo>&#x2217;</mml:mo></mml:msup></mml:mrow><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mi>c</mml:mi></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi>M</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>c</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo fence="false" stretchy="false">|</mml:mo><mml:mi>c</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>C</mml:mi></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>c</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mstyle></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula>where <italic>w</italic><sub>c</sub> is the adjustment weight, <italic>f</italic>(<italic>c</italic>) is the proportion of pixel value c to the total number of pixels. Through the different loss weights of real classes in the data to balance the categories, we can achieve better classification results.</p>
<p>In the process of deep neural network training, dropout is used in the full connection layer of convolutional network to prevent network over fitting [<xref ref-type="bibr" rid="ref-33">33</xref>]. However, in FCN, dropout layer cannot improve the network generalization ability. To solve this problem, dropblock is introduced
<disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:mi>&#x03B3;</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>k</mml:mi><mml:mi>p</mml:mi></mml:mrow><mml:mrow><mml:mi>b</mml:mi><mml:mrow><mml:msup><mml:mi>s</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow></mml:mfrac><mml:mfrac><mml:mrow><mml:mi>f</mml:mi><mml:mrow><mml:msup><mml:mi>s</mml:mi><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>f</mml:mi><mml:mi>s</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>b</mml:mi><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula>where <italic>&#x03B3;</italic> is used to control the number of channels removed from each convolution result. <italic>bs</italic> is used to control the block size to 0. <italic>kp</italic> is the dropout parameter. <italic>fs</italic> is the dimension of the characteristic diagram. dropblock makes the training network learn more robust features and greatly improves the generalization ability of the network.</p>
</sec>
</sec>
<sec id="s3"><label>3</label><title>Experiment and Result Analysis</title>
<p>The algorithm proposed in this paper is programmed and applied in the window system, VS2010 platform. Window 7, Intel&#x00AE; Core i5-6500 CPU, 3.20 GHZ, 16.0 GB and uses the deep network to extract features. We normalized the image to 512&#x2009;&#x00D7;&#x2009;512. At present, the average processing time of a single frame image is 1.2 s, which temporarily cannot meet the needs of real-time computing, and further research will be conducted in the future.</p>
<sec id="s3_1"><label>3.1</label><title>Database</title>
<p>3 datasets: The UA-DETRAC [<xref ref-type="bibr" rid="ref-34">34</xref>] dataset is a 10 h video shot by Canon EOS 550D camera at 24 different locations with a frame frequency of 25 fps and a resolution of 960&#x2009;&#x00D7;&#x2009;540 pixels. The UAV-DATA [<xref ref-type="bibr" rid="ref-35">35</xref>] dataset contains scenes such as trees and highways, including the characteristics of large-scale and multiple moving targets. The total video is 4.89 GB, and the minimum and maximum image resolution are 1920&#x2009;&#x00D7;&#x2009;1080 and 3840&#x2009;&#x00D7;&#x2009;2160, respectively. The minimum and maximum frame rate is 4 fps and 25 fps, respectively. The campus environment database is the shooting data over the playground, which contains a large number of small targets, as shown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>.</p>
<fig id="fig-6"><label>Figure 6</label><caption><title>Data display</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_20995-fig-6.png"/></fig>
</sec>
<sec id="s3_2"><label>3.2</label><title>Detection Accuracy</title>
<p>To measure the effectiveness of the algorithm, AOM and ROC curves are introduced
<disp-formula id="eqn-18"><label>(18)</label><mml:math id="mml-eqn-18" display="block"><mml:mrow><mml:mtext>AOM(</mml:mtext></mml:mrow><mml:mi>&#x03B3;</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03BD;</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mtext>S(</mml:mtext></mml:mrow><mml:mi>&#x03B3;</mml:mi><mml:mo>&#x2229;</mml:mo><mml:mi>&#x03BD;</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>S(</mml:mtext></mml:mrow><mml:mi>&#x03B3;</mml:mi><mml:mo>&#x222A;</mml:mo><mml:mi>&#x03BD;</mml:mi><mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mn>100</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></disp-formula>where &#x03B3; represents the result of manual annotation and &#x03BD; represents the detection result of the algorithm.</p>
<p>Based on the dataset described above, we have divided the data into <bold>Data 1:</bold> The first frame is a pure background image, and the target is moving all the time. <bold>Data 2:</bold> Small moving targets. <bold>Data 3:</bold> It remains stationary for a long time after the target moves.</p>
<p>As shown in <xref ref-type="table" rid="table-1">Table 1</xref> and <xref ref-type="fig" rid="fig-7">Fig. 7</xref>, with the complexity of the environment, the performance of the algorithm shows a downward trend. SVM &#x002B; GMM algorithm uses GMM model to extract motion region, and SVM is used to judge whether the region is foreground. RICF detects the target according to the rotation invariance of the target. Time space [<xref ref-type="bibr" rid="ref-26">26</xref>] extracts the moving region according to the relationship between time and space, and establishes AdaBoost model to realize multi-scale target recognition. Under the guidance of gestalt vision, the proposed algorithm establishes a saliency graph mechanism to extract moving targets, and then realizes target recognition based on the convolution neural network structure of global information. Although the operation speed is slightly lower than SVM &#x002B; GMM, AOM is the highest.</p>
<table-wrap id="table-1"><label>Table 1</label><caption><title>AOM and operation time</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Algorithm</th>
<th align="left">Data 1</th>
<th align="left">Data 2</th>
<th align="left">Data 3</th>
<th align="left">The average time (S/frame)</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">SVM&#x002B;GMM</td>
<td align="left">80.1</td>
<td align="left">76.2</td>
<td align="left">72.5</td>
<td align="left">0.03</td>
</tr>
<tr>
<td align="left">RICF</td>
<td align="left">81.8</td>
<td align="left">78.5</td>
<td align="left">74.8</td>
<td align="left">0.08</td>
</tr>
<tr>
<td align="left">Time-Space</td>
<td align="left">82.5</td>
<td align="left">79.2</td>
<td align="left">75.6</td>
<td align="left">0.09</td>
</tr>
<tr>
<td align="left">OURs</td>
<td align="left">83.1</td>
<td align="left">80.3</td>
<td align="left">78.1</td>
<td align="left">0.05</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="fig-7"><label>Figure 7</label><caption><title>Region of interest curves</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_20995-fig-7.png"/></fig>
</sec>
<sec id="s3_3"><label>3.3</label><title>Effect of Target Extraction</title>
<p>In order to intuitively show the effect of our algorithm, some detection results are selected, as shown in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>. The proposed algorithm can effectively detect the target, and has good detection effect in the face of complex background, unstable target motion and small targets.</p>
<fig id="fig-8"><label>Figure 8</label><caption><title>Detection results images</title></caption><graphic mimetype="image" mime-subtype="png" xlink:href="CMES_20995-fig-8.png"/></fig>
</sec>
</sec>
<sec id="s4"><label>4</label><title>Conclusion</title>
<p>Remote sensing images can obtain ground and target&#x0027;s information intuitively so as to provide accurate basis for decision-making. To solve the problem that it is difficult to accurately extract moving objects under complex conditions such as long-term stay and small-amplitude motion of moving objects in remote sensing videos, so we proposed a multi-moving object recognition algorithm based on remote sensing videos. Firstly, the problem of moving target detection is transformed into the problem of salient region probability, and the saliency map is constructed to extract moving targets. Secondly, by analyzing the global and local information of multiple targets, a convolutional neural network with global information is constructed to identify the target. Experiments show that the research results have a better effect on multi-target extraction in complex environments, and provide a new method for multi-target tracking in remote sensing images. So, on this basis, follow-up research on ground feature analysis can be carried out accordingly.</p>
</sec>
</body>
<back>
<fn-group>
<fn fn-type="other"><p><bold>Funding Statement:</bold> This work is supported by Yulin Science and Technology Association Youth Talent Promotion Program (Grant No. 20200212).</p></fn>
<fn fn-type="conflict"><p><bold>Conflicts of Interest:</bold> The authors declare that they have no conflicts of interest to report regarding the present study.</p></fn>
</fn-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>1.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shao</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Deng</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Remote sensing image super-resolution using sparse representation and coupled sparse autoencoder</article-title>. <source>IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing</source><italic>,</italic> <volume>12</volume><issue>(8)</issue><italic>,</italic> <fpage>2663</fpage>&#x2013;<lpage>2674</lpage>. DOI <pub-id pub-id-type="doi">10.1109/JSTARS.2019.2925456</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>2.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ma</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Jiang</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Zhou</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Zhao</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Guo</surname>, <given-names>X.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Guided locality preserving feature matching for remote sensing image registration</article-title>. <source>IEEE Transactions on Geoscience and Remote Sensing</source><italic>,</italic> <volume>56</volume><issue>(8)</issue><italic>,</italic> <fpage>4435</fpage>&#x2013;<lpage>4447</lpage>. DOI <pub-id pub-id-type="doi">10.1109/TGRS.2018.2820040</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>3.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cao</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Jiang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Yan</surname>, <given-names>P.</given-names></string-name></person-group> (<year>2016</year>). <article-title>Correlation-based tracking of multiple targets with hierarchical layered structure</article-title>. <source>IEEE Transactions on Cybernetics</source><italic>,</italic> <volume>48</volume><issue>(1)</issue><italic>,</italic> <fpage>90</fpage>&#x2013;<lpage>102</lpage>. DOI <pub-id pub-id-type="doi">10.1109/TCYB.2016.2625320</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>4.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Farmani</surname>, <given-names>N.</given-names></string-name>, <string-name><surname>Sun</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Pack</surname>, <given-names>D. J.</given-names></string-name></person-group> (<year>2017</year>). <article-title>A scalable multitarget tracking system for cooperative unmanned aerial vehicles</article-title>. <source>IEEE Transactions on Aerospace and Electronic Systems</source><italic>,</italic> <volume>53</volume><issue>(4)</issue><italic>,</italic> <fpage>1947</fpage>&#x2013;<lpage>1961</lpage>. DOI <pub-id pub-id-type="doi">10.1109/TAES.2017.2677746</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>5.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chao</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Cao</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>Y.</given-names></string-name></person-group> (<year>2010</year>). <article-title>Autopilots for small unmanned aerial vehicles: A survey</article-title>. <source>International Journal of Control, Automation and Systems</source><italic>,</italic> <volume>8</volume><issue>(1)</issue><italic>,</italic> <fpage>36</fpage>&#x2013;<lpage>44</lpage>. DOI <pub-id pub-id-type="doi">10.1007/s12555-010-0105-z</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>6.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Gleason</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Nefian</surname>, <given-names>A. V.</given-names></string-name>, <string-name><surname>Bouyssounousse</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Fong</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Bebis</surname>, <given-names>G.</given-names></string-name></person-group> (<year>2011</year>). <article-title>Vehicle detection from aerial imagery</article-title>. <conf-name>2011 IEEE International Conference on Robotics and Automation</conf-name>, pp. 2065&#x2013;2070. Shanghai, China, IEEE. DOI <pub-id pub-id-type="doi">10.1109/ICRA.2011.5979853</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>7.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Xiao</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Cheng</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Sawhney</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Han</surname>, <given-names>F.</given-names></string-name></person-group> (<year>2010</year>). <article-title>Vehicle detection and tracking in wide field-of-view aerial video</article-title>. <conf-name>2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition</conf-name>, pp. 679&#x2013;684. San Francisco, USA, IEEE. DOI <pub-id pub-id-type="doi">10.1109/CVPR.2010.5540151</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>8.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Andriluka</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Schnitzspan</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Meyer</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Kohlbrecher</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Petersen</surname>, <given-names>K.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2010</year>). <article-title>Vision based victim detection from unmanned aerial vehicles</article-title>. <conf-name>2010 IEEE/RSJ International Conference on Intelligent Robots and Systems</conf-name>, pp. 1740&#x2013;1747. Taipei, Taiwan, IEEE. DOI <pub-id pub-id-type="doi">10.1109/IROS.2010.5649223</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>9.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cheng</surname>, <given-names>H. Y.</given-names></string-name>, <string-name><surname>Weng</surname>, <given-names>C. C.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>Y. Y.</given-names></string-name></person-group> (<year>2011</year>). <article-title>Vehicle detection in aerial surveillance using dynamic Bayesian networks</article-title>. <source>IEEE Transactions on Image Processing</source><italic>,</italic> <volume>21</volume><issue>(4)</issue><italic>,</italic> <fpage>2152</fpage>&#x2013;<lpage>2159</lpage>. DOI <pub-id pub-id-type="doi">10.1109/TIP.2011.2172798</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>10.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Gaszczak</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Breckon</surname>, <given-names>T. P.</given-names></string-name>, <string-name><surname>Han</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2011</year>). <chapter-title>Real-time people and vehicle detection from UAV imagery</chapter-title>. <source>Intelligent Robots and Computer Vision XXVIII: Algorithms and Techniques</source><italic>,</italic> vol. 7878. <publisher-name>San Francisco Airport, California, USA</publisher-name>. DOI <pub-id pub-id-type="doi">10.1117/12.876663</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>11.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lin</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Saripalli</surname>, <given-names>S.</given-names></string-name></person-group> (<year>2012</year>). <article-title>Road detection from aerial imagery</article-title>. <conf-name>2012 IEEE International Conference on Robotics and Automation</conf-name>, pp. 3588&#x2013;3593. Saint Paul, MN, USA, IEEE. DOI <pub-id pub-id-type="doi">10.1109/ICRA.2012.6225112</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>12.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rodr&#x00ED;guez-Canosa</surname>, <given-names>G. R.</given-names></string-name>, <string-name><surname>Thomas</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Del Cerro</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Barrientos</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>MacDonald</surname>, <given-names>B.</given-names></string-name></person-group> (<year>2012</year>). <article-title>A Real-time method to detect and track moving objects (DATMO) from unmanned aerial vehicles (UAVs) using a single camera</article-title>. <source>Remote Sensing</source><italic>,</italic> <volume>4</volume><issue>(4)</issue><italic>,</italic> <fpage>1090</fpage>&#x2013;<lpage>1111</lpage>. DOI <pub-id pub-id-type="doi">10.3390/rs4041090</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>13.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zheng</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Zhou</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>X.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2013</year>). <article-title>A novel vehicle detection method with high resolution highway aerial image</article-title>. <source>IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing</source><italic>,</italic> <volume>6</volume><issue>(6)</issue><italic>,</italic> <fpage>2338</fpage>&#x2013;<lpage>2343</lpage>. DOI <pub-id pub-id-type="doi">10.1109/JSTARS.2013.2266131</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>14.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Liang</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Ling</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Blasch</surname>, <given-names>E.</given-names></string-name>, <string-name><surname>Seetharaman</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Shen</surname>, <given-names>D.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2013</year>). <article-title>Vehicle detection in wide area aerial surveillance using temporal context</article-title>. <conf-name>Proceedings of the 16th International Conference on Information Fusion</conf-name>, pp. 181&#x2013;188. Istanbul, Turkey, IEEE.</mixed-citation></ref>
<ref id="ref-15"><label>15.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Prokaj</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Medioni</surname>, <given-names>G.</given-names></string-name></person-group> (<year>2014</year>). <article-title>Persistent tracking for wide area aerial surveillance</article-title>. <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, pp. 1186&#x2013;1193. Columbus, USA.</mixed-citation></ref>
<ref id="ref-16"><label>16.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Teutsch</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Kr&#x00FC;ger</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Beyerer</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2014</year>). <article-title>Evaluation of object segmentation to improve moving vehicle detection in aerial videos</article-title>. <conf-name>2014 11th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS)</conf-name>, pp. 265&#x2013;270. Seoul, Korea, IEEE. DOI <pub-id pub-id-type="doi">10.1109/AVSS.2014.6918679</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>17.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chen</surname>, <given-names>B. J.</given-names></string-name>, <string-name><surname>Medioni</surname>, <given-names>G.</given-names></string-name></person-group> (<year>2015</year>). <article-title>Motion propagation detection association for multi-target tracking in wide area aerial surveillance</article-title>. <conf-name>2015 12th IEEE International Conference on Advanced Video and Signal Based Surveillance (AVSS)</conf-name>, pp. 1&#x2013;6. Karlsruhe, Germany, IEEE. DOI <pub-id pub-id-type="doi">10.1109/AVSS.2015.7301766</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>18.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Jiang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Cao</surname>, <given-names>X.</given-names></string-name></person-group> (<year>2016</year>). <article-title>Surveillance from above: A detection-and-prediction based multiple target tracking method on aerial videos</article-title>. <conf-name>2016 Integrated Communications Navigation and Surveillance (ICNS)</conf-name>, pp. 4D2&#x2013;1. Herndon, USA, IEEE. DOI <pub-id pub-id-type="doi">10.1109/ICNSURV.2016.7486348</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>19.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Poostchi</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Aliakbarpour</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Viguier</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Bunyak</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Palaniappan</surname>, <given-names>K.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2016</year>). <article-title>Semantic depth map fusion for moving vehicle detection in aerial video</article-title>. <conf-name>Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops</conf-name>, pp. 32&#x2013;40. Las Vegas, USA.</mixed-citation></ref>
<ref id="ref-20"><label>20.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tang</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Zhou</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Deng</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Lei</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Zou</surname>, <given-names>H.</given-names></string-name></person-group> (<year>2017</year>). <article-title>Arbitrary-oriented vehicle detection in aerial imagery with single convolutional neural networks</article-title>. <source>Remote Sensing</source><italic>,</italic> <volume>9</volume><issue>(11)</issue><italic>,</italic> <fpage>1170</fpage>. DOI <pub-id pub-id-type="doi">10.3390/rs9111170</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>21.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Aguilar</surname>, <given-names>W. G.</given-names></string-name>, <string-name><surname>Luna</surname>, <given-names>M. A.</given-names></string-name>, <string-name><surname>Moya</surname>, <given-names>J. F.</given-names></string-name>, <string-name><surname>Abad</surname>, <given-names>V.</given-names></string-name>, <string-name><surname>Parra</surname>, <given-names>H.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2017</year>). <article-title>Pedestrian detection for UAVs using cascade classifiers with meanshift</article-title>. <conf-name>2017 IEEE 11th International Conference on Semantic Computing (ICSC)</conf-name>, pp. 509&#x2013;514. San Diego, USA, IEEE. DOI <pub-id pub-id-type="doi">10.1109/ICSC.2017.83</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>22.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xu</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Zhong</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Yan</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Wu</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>W.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Moving object detection in aerial infrared images with registration accuracy prediction and feature points selection</article-title>. <source>Infrared Physics &#x0026; Technology</source><italic>,</italic> <volume>92</volume><italic>,</italic> <fpage>318</fpage>&#x2013;<lpage>326</lpage>. DOI <pub-id pub-id-type="doi">10.1016/j.infrared.2018.06.023</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>23.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Hamsa</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Panthakkan</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Al Mansoori</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Alahamed</surname>, <given-names>H.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Automatic vehicle detection from aerial images using cascaded support vector machine and Gaussian mixture model</article-title>. <conf-name>2018 International Conference on Signal Processing and Information Security ICSPIS</conf-name>, pp. 1&#x2013;4. Dubai, United Arab Emirates, IEEE. DOI <pub-id pub-id-type="doi">10.1109/CSPIS.2018.8642716</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>24.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ma</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Jiang</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Yan</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Yuan</surname>, <given-names>J.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2019</year>). <article-title>Vehicle detection in aerial images using rotation-invariant cascaded forest</article-title>. <source>IEEE Access</source><italic>,</italic> <volume>7</volume><italic>,</italic> <fpage>59613</fpage>&#x2013;<lpage>59623</lpage>. DOI <pub-id pub-id-type="doi">10.1109/ACCESS.2019.2915368</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>25.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Mandal</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Shah</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Meena</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Vipparthi</surname>, <given-names>S. K.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Sssdet: Simple short and shallow network for resource efficient vehicle detection in aerial scenes</article-title>. <conf-name>2019 IEEE International Conference on Image Processing</conf-name>, pp. 3098&#x2013;3102. Taipei, Taiwan, IEEE. DOI <pub-id pub-id-type="doi">10.1109/ICIP.2019.8803262</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>26.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Song</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Si</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Zhou</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Yuan</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>E.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2020</year>). <article-title>Feature extraction and target recognition of moving image sequences</article-title>. <source>IEEE Access</source><italic>,</italic> <volume>8</volume><italic>,</italic> <fpage>147148</fpage>&#x2013;<lpage>147161</lpage>. DOI <pub-id pub-id-type="doi">10.1109/ACCESS.2020.3015261</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>27.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Qiu</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Cheng</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Cui</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Zhou</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Guo</surname>, <given-names>Q.</given-names></string-name></person-group> (<year>2020</year>). <article-title>A moving vehicle tracking algorithm based on deep learning</article-title>. <source>Journal of Ambient Intelligence and Humanized Computing</source><italic>,</italic> <fpage>1</fpage>&#x2013;<lpage>7</lpage>. DOI <pub-id pub-id-type="doi">10.1007/s12652-020-02352-w</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>28.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Feng</surname>, <given-names>Z.</given-names></string-name></person-group> (<year>2021</year>). <article-title>High speed moving target tracking algorithm based on mean shift for video human motion</article-title>. <source>Journal of Physics: Conference Series</source><italic>,</italic> <volume>1744</volume><issue>(4)</issue><italic>,</italic> <fpage>42180</fpage>. IOP Publishing.</mixed-citation></ref>
<ref id="ref-29"><label>29.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wan</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Tan</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>Q.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2021</year>). <article-title>Refocusing of ground moving targets with Doppler ambiguity using keystone transform and modified second-order keystone transform for synthetic aperture radar</article-title>. <source>Remote Sensing</source><italic>,</italic> <volume>13</volume><issue>(2)</issue><italic>,</italic> <fpage>177</fpage>. DOI <pub-id pub-id-type="doi">10.3390/rs13020177</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>30.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Lin</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Shi</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Huang</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>W.</given-names></string-name></person-group> (<year>2022</year>). <chapter-title>Tracking strategy of multiple moving targets using multiple UAVs</chapter-title>. <source>Advances in Guidance, Navigation and Control</source><italic>,</italic> vol. 644, pp. <fpage>1417</fpage>&#x2013;<lpage>1425</lpage>. <publisher-loc>Singapore</publisher-loc>: <publisher-name>Springer</publisher-name>. DOI <pub-id pub-id-type="doi">10.1007/978-981-15-8155-7_118</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>31.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bayar</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Stamm</surname>, <given-names>M. C.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Constrained convolutional neural networks: A new approach towards general purpose image manipulation detection</article-title>. <source>IEEE Transactions on Information Forensics and Security</source><italic>,</italic> <volume>13</volume><issue>(11)</issue><italic>,</italic> <fpage>2691</fpage>&#x2013;<lpage>2706</lpage>. DOI <pub-id pub-id-type="doi">10.1109/TIFS.2018.2825953</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>32.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Dong</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Fu</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Yu</surname>, <given-names>L.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Large-scale oil palm tree detection from high-resolution satellite images using two-stage convolutional neural networks</article-title>. <source>Remote Sensing</source><italic>,</italic> <volume>11</volume><issue>(1)</issue><italic>,</italic> <fpage>11</fpage>. DOI <pub-id pub-id-type="doi">10.3390/rs11010011</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>33.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Achille</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Soatto</surname>, <given-names>S.</given-names></string-name></person-group> (<year>2018</year>). <article-title>Information dropout: Learning optimal representations through noisy computation</article-title>. <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source><italic>,</italic> <volume>40</volume><issue>(12)</issue><italic>,</italic> <fpage>2897</fpage>&#x2013;<lpage>2905</lpage>. DOI <pub-id pub-id-type="doi">10.1109/TPAMI.2017.2784440</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>34.</label><mixed-citation publication-type="web">UA-DETRAC. <uri xlink:href="http://detrac-db.rit.albany.edu/">http://detrac-db.rit.albany.edu/</uri>.</mixed-citation></ref>
<ref id="ref-35"><label>35.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Bai</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>F.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>S.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2019</year>). <article-title>Multiple-object-tracking algorithm based on dense trajectory voting in aerial videos</article-title>. <source>Remote Sensing</source><italic>,</italic> <volume>11</volume><issue>(19)</issue><italic>,</italic> <fpage>2278</fpage>. DOI <pub-id pub-id-type="doi">10.3390/rs11192278</pub-id>.</mixed-citation></ref>
</ref-list>
</back>
</article>