<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CSSE</journal-id>
<journal-id journal-id-type="nlm-ta">CSSE</journal-id>
<journal-id journal-id-type="publisher-id">CSSE</journal-id>
<journal-title-group>
<journal-title>Computer Systems Science &#x0026; Engineering</journal-title>
</journal-title-group>
<issn pub-type="ppub">0267-6192</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">24118</article-id>
<article-id pub-id-type="doi">10.32604/csse.2023.024118</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Image Captioning Using Detectors and Swarm Based Learning Approach for Word Embedding Vectors</article-title><alt-title alt-title-type="left-running-head">Image Captioning Using Detectors and Swarm Based Learning Approach for Word Embedding Vectors</alt-title><alt-title alt-title-type="right-running-head">Image Captioning Using Detectors and Swarm Based Learning Approach for Word Embedding Vectors</alt-title>
</title-group>
<contrib-group content-type="authors">
<contrib id="author-1" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Lalitha</surname><given-names>B.</given-names></name>
<xref ref-type="aff" rid="aff-1">1</xref><email>lalli_j@yahoo.com</email>
</contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Gomathi</surname><given-names>V.</given-names></name>
<xref ref-type="aff" rid="aff-2">2</xref>
</contrib>
<aff id="aff-1"><label>1</label><institution>CSE Department, Sethu Institute of Technology</institution>, <addr-line>Pulloor, Kariapatti, 626115</addr-line>, <country>India</country></aff>
<aff id="aff-2"><label>2</label><institution>CSE Department, National Engineering College</institution>, <addr-line>K.R. Nagar, Kovilpatti, 628503</addr-line>, <country>India</country></aff>
</contrib-group><author-notes><corresp id="cor1"><label>&#x002A;</label>Corresponding Author: B. Lalitha. Email: <email>lalli_j@yahoo.com</email></corresp></author-notes>
<pub-date pub-type="epub" date-type="pub" iso-8601-date="2022-05-24"><day>24</day>
<month>05</month>
<year>2022</year></pub-date>
<volume>44</volume>
<issue>1</issue>
<fpage>173</fpage>
<lpage>189</lpage>
<history>
<date date-type="received"><day>05</day><month>10</month><year>2021</year></date>
<date date-type="accepted"><day>16</day><month>12</month><year>2021</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2023 Lalitha and Gomathi</copyright-statement>
<copyright-year>2023</copyright-year>
<copyright-holder>Lalitha and Gomathi</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CSSE_24118.pdf"></self-uri>
<abstract>
<p><bold>IC</bold> (Image Captioning) is a crucial part of Visual Data Processing and aims at understanding for providing captions that verbalize an image&#x2019;s important elements. However, in existing works, because of the complexity in images, neglecting major relation between the object in an image, poor quality image, labelling it remains a big problem for researchers. Hence, the main objective of this work attempts to overcome these challenges by proposing a novel framework for IC. So in this research work the main contribution deals with the framework consists of three phases that is image understanding, textual understanding and decoding. Initially, the image understanding phase is initiated with image pre-processing to enhance image quality. Thereafter, object has been detected using IYV3MMDs (Improved YoloV3 Multishot Multibox Detectors) in order to relate the interrelation between the image and the object, and then it is followed by MBFOCNNs (Modified Bacterial Foraging Optimization in Convolution Neural Networks), which encodes and provides final feature vectors. Secondly, the textual understanding phase is performed based on an image which is initiated with preprocessing of text where unwanted words, phrases, punctuations are removed in order to provide a healthy text. It is then followed by MGloVEs (Modified Global Vectors for Word Representation), which provides a word embedding of features with the highest priority towards the object present in an image. Finally, the decoding phase has been performed, which decodes the image whether it may be a normal or complex scene image and provides an accurate text by its learning ability using MDAA (Modified Deliberate Adaptive Attention). The experimental outcome of this work shows better accuracy of shows 96.24&#x0025; when compared to existing and similar methods while generating captions for images.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Denoising</kwd>
<kwd>improved YoloV3 multishot multibox detector (IYV3MMD)</kwd>
<kwd>modified bacterial foraging optimization in convolutional neural network (MBFOCNN)</kwd>
<kwd>modified global vectors for word representation (MGloVE)</kwd>
<kwd>modified deliberate adaptive attention (MDAA)</kwd>
<kwd>encoder</kwd>
<kwd>decoder</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>In the current world driven by social media, people easily produce and share rich social multimedia material online including photos and videos as they are allowed by social media players like Twitter, Amazon, Facebook and Google News. This voluminous data is explored by researches in terms of displays, retrievals, and alterations of multimedia content, specifically in IC as it tries to evaluate the visual content of input images and creates captions that verbalize images based on their most essential elements [<xref ref-type="bibr" rid="ref-1">1</xref>]. Few examples of multimodal applications used in data processing include Visual question answering, multimodal event extractions, video captioning, and cross-modal image retrievals [<xref ref-type="bibr" rid="ref-2">2</xref>,<xref ref-type="bibr" rid="ref-3">3</xref>].</p>
<p>Unlike image classifications or object identifications in n computer vision, ICs are multi-modal tasks that demand capture important features of images. This study links NLPs (Natural Language Processing) to ICs for describing them. ICs are processed using encoder-decoder architectures [<xref ref-type="bibr" rid="ref-4">4</xref>,<xref ref-type="bibr" rid="ref-5">5</xref>] where dense pixel-level information is encoded using the CNNs encoder and translated by decoders.</p>
<p>Specific visual data processing models are based on CNN&#x2019;s LSTMs (Long Short Term Memories) framework which can be trained [<xref ref-type="bibr" rid="ref-6">6</xref>,<xref ref-type="bibr" rid="ref-7">7</xref>]. Many collaborative CNN models have also been proposed for ICs: R-CNNs (Regional CNNs) [<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-9">9</xref>], Fast R-CNNs (fast regional CNNs) [<xref ref-type="bibr" rid="ref-10">10</xref>], Mask R-CNNs (mask regional CNNs) [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-12">12</xref>], and FCN-LSTMs (Fully Convolution Network Long Short Term Memory) [<xref ref-type="bibr" rid="ref-13">13</xref>]. Mechanisms using CNN feature maps for Spatial processing have been studied more recently [<xref ref-type="bibr" rid="ref-14">14</xref>,<xref ref-type="bibr" rid="ref-15">15</xref>], where they produce spatial maps emphasizing on important picture regions and matched to words.</p>
<p>In order to execute ICs, these present computer vision tasks must not only collect information contained in images but also extract semantic associations of acquired visual information corresponding to verbal expressions [<xref ref-type="bibr" rid="ref-16">16</xref>,<xref ref-type="bibr" rid="ref-17">17</xref>]. Though these approaches listed above have shown substantial results, they also have resulted in a number of flaws [<xref ref-type="bibr" rid="ref-18">18</xref>,<xref ref-type="bibr" rid="ref-19">19</xref>]. In order to handle these issues, this research work proposes an efficient IC processing technique using MDAAs with the help of the Image feature and word embedding vectors using MBFOCNNs and MGloVEs.</p>
<p>The following is how the rest of the article is organized: Section 2 reviews and analyzes relevant ICs work, Section 3 details the suggested technique, and Section 4 explains the experimental evaluation. Lastly, Section 5 concludes planned study with future prospects.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Literature Survey</title>
<p>Xiao et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] proposed deep hierarchical encoder-decoder networks for ICs, in which the encoder and decoder operations were separated using a deep hierarchical structure. The method was able to leverage deep networks&#x2019; representation capabilities to combine high-level semantics of vision and language to create captions. Visual representations at the highest abstraction level were studied at the same time, and each of these levels was assigned to a single LSTM. The textual inputs were encoded using the bottom-most LSTM. The intermediate layer was used in the encoder-decoder to improve the decoder capabilities of the top-most LSTM. Tests on three benchmark databases showed that study&#x2019;s approach worked effectively beating existing modern techniques: Flickr8K, Flickr30K, and MSCOCO. The issue was in using it for complicated situations with many targets.</p>
<p>For the remote sensing issues in ICs, Shen et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] suggested VRTMM. Initially, it worked with the Varied Auto encoder to fine-tune the CNNs. Secondly, the text description was created by the Transformers using both geographical and semantic data. The quality of the produced phrases was then improved using Reinforcement Learning. For Remote Sensing Image Caption Database, their method outperformed others by significant margin on entire seven values. Findings of experiment indicate that appraoch worked well with remote sensing ICs and produced new outcomes as other captioning methods ignored interrelationships between items in images.</p>
<p>Kinghorn et al. [<xref ref-type="bibr" rid="ref-22">22</xref>] created deep networks that included two critical phases for picture description creations as well as initial regional based developments. Their Region Proposal Networks from Faster R-CNNs generated initial regional proposals. The technique created areas of interest, which were subsequently utilized to annotate and categorize human and object characteristics. System&#x2019;s initial major phase involved creating label descriptions for each location of interest. In the second step, they used encoders-decoders based on RNNs (Recurrent Neural Networks) to convert these regional descriptions into a comprehensive image description. Their empirical findings showed that their strategy was equivalent to many previous studies while outperforming many approaches. Furthermore, when the number of time steps grew, RNNs had gradient vanishing issues.</p>
<p>Su et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] integrated visual and high-level semantic information in their proposal for ICs. The bottom layer and top layers of a hierarchical DNN were created for caption creations where the former collected visual and high-level semantic information from identified areas in images while the latter combined using an adaptive attention method. On the MSCOCO dataset, their experimental findings performed competitively in comparison to other techniques. It was challenging to extract the essential characteristics from images, which required a mix of visual and language data.</p>
<p>Zhao et al. [<xref ref-type="bibr" rid="ref-24">24</xref>] developed a Multimodal fusion technique for generating descriptions that describe image information. The study used CNNs for image feature extractions and an attention model for image attributes extractions and language CNNs to model sentences, and a recurrent networks like LSTM for word predictions. When compared to other approaches, to model long-term interdependences of historical words, the study used image characteristics to enhance image representations and handle all prior words. However, the multimodal fusion was created using a single-layer network, which proved incapable of performing complex tasks.</p>
<p>ICs also include an attention function. Xu et al. [<xref ref-type="bibr" rid="ref-25">25</xref>] employed deterministic soft attention scheme while creating different words and stochastic hard attention to assist decoder&#x2019;s attention on highly important image regions and thus improving sentence creation quality.</p>
<p>Anderson et al. [<xref ref-type="bibr" rid="ref-26">26</xref>] used a bottom-up and top-down approach to enhance attention modules. Attention mechanism efficiently plugged gaps between the visual and language domains. As a result, this technique is frequently used in ICs activities.</p>
<p>Vaswani et al. [<xref ref-type="bibr" rid="ref-27">27</xref>] developed a Transformer model where attention blocks were stacked completely without convolutions or repetitions. The study included an encoder and a decoder. Encoders had a self-attention and a position-wise feed-forward block, whereas decoders had a self-attention and a cross-attention layer.</p>
<p>Zhu et al. [<xref ref-type="bibr" rid="ref-28">28</xref>] proposed a framework by modifying Transformer architecture where CNNs replace the Transformer&#x2019;s encoder. The spatial connections of the R-CNN identified item pairs are incorporated into the attention block in the study by Herdade et al. [<xref ref-type="bibr" rid="ref-29">29</xref>] while Huang et al. [<xref ref-type="bibr" rid="ref-30">30</xref>] presented &#x201C;Attention on Attention&#x201D; module to describe interactions between image objects on encoders and also refined decoder with self-attentions.</p>
<p>Pan et al. [<xref ref-type="bibr" rid="ref-31">31</xref>] offer a unified X-Linear attention block which fully utilizes bilinear pooling to appropriately exploit diverse visual inputs. Cornia et al. [<xref ref-type="bibr" rid="ref-32">32</xref>] proposed a Meshed Transformer with IC operations memory.</p>
<p>Based on the above studies, this work leverages on low/high image characteristics, for employing mesh like connections in decoding images.</p>
</sec>
<sec id="s3">
<label>3</label>
<title>Proposed Encoder-Decoder IC Framework</title>
<p>IC applications have several uses including Image retrieval, assisting visually handicapped, and intelligent human-computer interaction as they automatically create language descriptions for images. It was a difficult cross-disciplinary project that required both computer vision and NLPs in processing. Various DL algorithms has recently developed, however focusing on certain areas or objects of interest in an image is a complex issue specifically during sentence productions while disregarding previously created time steps that constitute sentences. Models may pay more attention to image&#x2019;s same locations in time steps and thus jeopardizing IC performances. This research work proposes an efficient IC processing technique (Illustrated in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>) by utilizing MDAAs with the aid of the image features and word embedding vectors which use MBFOCNNs and MGloves to handle aforesaid issues.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Proposed IC framework</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CSSE_24118-fig-1.png"/>
</fig>
<sec id="s3_1">
<label>3.1</label>
<title>Image Understanding</title>
<p>In Image understanding phase, initially a raw image will be processed for enhancement and size alignment. After standardizing the size as 512&#x002A;512, it will undergo object detection stage and followed by encoding the detector outcomes as the text caption is generated.</p>
<sec id="s3_1_1">
<label>3.1.1</label>
<title>Object Detection</title>
<p>Object detections identify items in images for focusing on closest ones while keeping farther ones out of focus. For accurate IC operations, it is necessary to identify minute things where IYV3MMDs are used in this work. IYV3MMDs use a convolution feed-forward networks to generate bounding boxes of fixed-size array and boxes are scored for the presence of objects, proceeded by non-maximum suppression phase for final detection outcomes. This work&#x2019;s object detections are based on selective training on a default detection boxe group and sizes, along with hard negative mining and data augmentation methods. The IYV3MMD training stages are detailed below:</p>
<p><italic>Step 1: Matching Strategy</italic></p>
<p>Determine default boxes that relate to basic truth detections during training and train networks accordingly. Default boxe groups for each ground truth box was selected in the study based on varying positions, aspect ratios, and scales. Initially, every ground truth box is coordinated to default box with greatest jaccard overlaps followed by matching default boxes to any ground truths with jaccard overlap values greater than threshold and unlike Multi-Boxes (0.5). This strategy simplifies learning as it allows networks to forecast maximum values for many overlying default boxes instead of just ones with greatest overlaps. This is depicted as <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>.<disp-formula id="eqn-1"><label>(1)</label>
<mml:math id="mml-eqn-1" display="block"><mml:mi>M</mml:mi><mml:mi>a</mml:mi><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>S</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mi mathvariant="normal">&#x0393;</mml:mi></mml:msubsup><mml:mo>&#x2265;</mml:mo><mml:mn>0.5</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math>
</disp-formula>where, <inline-formula id="ieqn-1">
<mml:math id="mml-ieqn-1"><mml:mi>M</mml:mi><mml:mi>a</mml:mi><mml:msubsup><mml:mi>t</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>S</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msubsup></mml:math>
</inline-formula> stands for matching strategies and <inline-formula id="ieqn-2">
<mml:math id="mml-ieqn-2"><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mi mathvariant="normal">&#x0393;</mml:mi></mml:msubsup></mml:math>
</inline-formula> implies matching <italic>i</italic><sup><italic>th</italic></sup> default box to <italic>j</italic><sup><italic>th</italic></sup> ground truth box of &#x0393; image category.</p>
<p><italic>Step 2: Training Objective</italic></p>
<p>Training is necessary to detect multiple items. Assuming <inline-formula id="ieqn-3">
<mml:math id="mml-ieqn-3"><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x0393;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo fence="false" stretchy="false">}</mml:mo></mml:mrow></mml:math>
</inline-formula> is matching indicator for i-th default box and j-th ground truth box in category p, then <inline-formula id="ieqn-4">
<mml:math id="mml-ieqn-4"><mml:munder><mml:mrow><mml:mo movablelimits="false">&#x2211;</mml:mo></mml:mrow><mml:mi>i</mml:mi></mml:munder><mml:mo>&#x2061;</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x0393;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2265;</mml:mo><mml:mn>1</mml:mn></mml:math>
</inline-formula>. The total objective loss function is represented by <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref> as the weighted sum of localization loss (loc) as well as confidence loss (conf).<disp-formula id="eqn-2"><label>(2)</label>
<mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mi>L</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mi>C</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mi>P</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mrow><mml:mi mathvariant="normal">T</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mi>C</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi>&#x03B2;</mml:mi><mml:msub><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mrow><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mi>P</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mrow><mml:mi mathvariant="normal">T</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:mstyle></mml:math>
</disp-formula>here, <italic>N</italic> implies coordinated default boxes. When <italic>N</italic>&#x2009;&#x003D;&#x2009;0, loss is 0. Loss due to localizations is a smoothened loss among predicted boxes (<italic>P</italic>) and ground truth boxes (T). Regressing offsets of the centre (<italic>Cx</italic>, <italic>Cy</italic>) in default bounding box (<italic>d</italic>) with width (<italic>w</italic>) and height (<italic>h</italic>) can be depicted using the following Equations.<disp-formula id="eqn-3"><label>(3)</label>
<mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mrow><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mi>P</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mrow><mml:mi mathvariant="normal">T</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mrow><mml:munder><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="normal">M</mml:mi></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mi>C</mml:mi><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>C</mml:mi><mml:mi>y</mml:mi><mml:mo>,</mml:mo><mml:mi>w</mml:mi><mml:mo>,</mml:mo><mml:mi>h</mml:mi></mml:mrow><mml:mo fence="false" stretchy="false">}</mml:mo></mml:mrow></mml:munder><mml:mrow><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mi>k</mml:mi></mml:msubsup><mml:mi>s</mml:mi><mml:mi>m</mml:mi><mml:mi>o</mml:mi><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msubsup><mml:mi>P</mml:mi><mml:mi>i</mml:mi><mml:mi>M</mml:mi></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="normal">T</mml:mi></mml:mrow><mml:mi>j</mml:mi><mml:mi>M</mml:mi></mml:msubsup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:math>
</disp-formula><disp-formula id="eqn-4"><label>(4)</label>
<mml:math id="mml-eqn-4" display="block"><mml:msub><mml:mi mathvariant="normal">&#x03A9;</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>f</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mi>C</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mrow><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mi mathvariant="normal">&#x0393;</mml:mi></mml:msubsup><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>C</mml:mi><mml:mi>i</mml:mi><mml:mi mathvariant="normal">&#x0393;</mml:mi></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2212;</mml:mo></mml:mrow><mml:munder><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>N</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:munder><mml:mrow><mml:mi>log</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>C</mml:mi><mml:mi>i</mml:mi><mml:mn>0</mml:mn></mml:msubsup><mml:mo stretchy="false">)</mml:mo><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mi>w</mml:mi><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:msubsup><mml:mi>C</mml:mi><mml:mi>i</mml:mi><mml:mi mathvariant="normal">&#x0393;</mml:mi></mml:msubsup><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mrow><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>C</mml:mi><mml:mi>i</mml:mi><mml:mn>0</mml:mn></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mi>p</mml:mi></mml:msub><mml:mrow><mml:mi>exp</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>C</mml:mi><mml:mi>i</mml:mi><mml:mn>0</mml:mn></mml:msubsup><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mrow></mml:mstyle></mml:math>
</disp-formula>where <italic>&#x03B2;</italic> stands for a term&#x2019;s weight and is equal to 1 in cross validations.</p>
<p><italic>Step 3: Selecting Scales and Aspect Ratio for Default Boxes</italic></p>
<p>The work employs scale selections in lower/higher feature maps to identify default box size fluctuations or in images to smoothed sizes. The default boxes need not match each layer&#x2019;s real receptive fields. This work aims at tiling default boxes in such a way that particular mappings of functions learn to adapt to different object sizes. Scaling is done for each pixel&#x2019;s feature map, computed using <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref><disp-formula id="eqn-5"><label>(5)</label>
<mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mi mathvariant="normal">&#x0393;</mml:mi><mml:mi>k</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x0393;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x0393;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x0393;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>M</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mi>k</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mi>M</mml:mi><mml:mo stretchy="false">]</mml:mo></mml:mstyle></mml:math>
</disp-formula>here, &#x0393;<sub>max</sub> stands for upper bound values and &#x0393;<sub>min</sub> for lower bound values and layers in between these bounds are evenly spaced.</p>
<p><italic>Step 4: Hard Negative Mining</italic></p>
<p>Most of the default boxes show negativity after the matching stage, particularly whenever there are a lot of default boxes to choose from. The results in an imbalanced training and significant losses. To overcome this issue, this research work uses sorting that selects the largest confidence loss for every default box and thus resulting in a negative-to-positive ratio value greater than 3:1 and faster optimizations/constant training.</p>
<p><italic>Step 5: Augmentation of Data</italic></p>
<p>To make the system highly tolerant to varying input item sizes and shapes, every training image is sampled randomly using one of following parameters:<list list-type="bullet"><list-item>
<p>Using complete original input images as starting points.</p></list-item><list-item>
<p>Sampling patches with objects and a minimum jaccard overlap of 0.1, 0.3, 0.5, 0.7</p></list-item><list-item>
<p>Selecting patches randomly.</p></list-item><list-item>
<p>Sampling patches in the interval [0.1, 1] of actual image size, with an aspect ratio ranging from 1/2 to 2.</p></list-item><list-item>
<p>Preserving centre of overlapping ground truth boxes in sampled patches.</p></list-item><list-item>
<p>Enlarging each sampled patch to a predetermined size and horizontally flipping it with a probability of 0.5 after executing the aforementioned sampling steps.</p></list-item></list></p>
</sec>
<sec id="s3_1_2">
<label>3.1.2</label>
<title>Encoder</title>
<p>Encoding contributes towards a rich representation of image via encoding input image content to a fixed-length vector using an internal representation. The existing work mainly used RNN for processing the encoding but due to inaccurate representation of the image, the work has developed a MBFOCNN as shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>. By integrating the image to a fixed-length vector, the proposed encoder delivers a significant quality improvement of image. MBFOCNN is performed based on the input (<inline-formula id="ieqn-5">
<mml:math id="mml-ieqn-5"><mml:msubsup><mml:mi>&#x03B6;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>.</mml:mo><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>O</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msubsup></mml:math>
</inline-formula>).<list list-type="simple"><list-item>
<p><italic>(a) Convolutional Layer</italic></p></list-item></list></p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>MBFOCNN encoder</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CSSE_24118-fig-2.png"/>
</fig>
<p>CNNs are a deep learning method that takes an input picture and extracts information by convolving it using filters or kernels. The input image is filtered, as well as the convolution technique learns similar feature over whole image. Without pooling, <xref ref-type="disp-formula" rid="eqn-6">Eq. (6)</xref> yields size of the resulting matrix:<disp-formula id="eqn-6"><label>(6)</label>
<mml:math id="mml-eqn-6" display="block"><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x0393;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">]</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="normal">&#x0393;</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03BB;</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:math>
</disp-formula></p>
<p>The window advances with every action, and indeed the feature maps discover the features. To capture the image&#x2019;s local receptive field, the feature maps employ common weights and biases. The convolution process is described by <xref ref-type="disp-formula" rid="eqn-7">Eq. (7)</xref>:<disp-formula id="eqn-7"><label>(7)</label>
<mml:math id="mml-eqn-7" display="block"><mml:msub><mml:mi>O</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mi>O</mml:mi><mml:mi>N</mml:mi><mml:mi>V</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>m</mml:mi><mml:mi>o</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mi>b</mml:mi><mml:mo>+</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mn>2</mml:mn></mml:munderover><mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mn>2</mml:mn></mml:munderover><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi mathvariant="normal">&#x03A8;</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>+</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>+</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula></p>
<p>Now, initialization of the weight is done using Modified Bacteria Foraging Optimization. BFA is an optimization approach inspired on E. coli bacteria&#x2019;s foraging behaviour. Discusses the biological features of bacterial hunting tactics and motile behaviour, and also their decision-making systems. BFA is meant to handle complicated and non-differentiable objective functions and handle non-gradient optimization issues. Chemotaxis, reproduction, and elimination dispersal activities are the three main processes used to search hyperspace. Swimming and tumbling are used in the chemotaxis mechanism. The bacteria spends its whole existence switching between these two motions.</p>
<p>The fundamental BFA&#x2019;s unit step length is fixed, ensuring good searching outcomes for modest optimization tasks. If applied to difficult situations with great dimensionality, nevertheless, it performs poorly. The run length option is crucial for managing the BFA&#x2019;s local and global search capabilities. Modifying the run-length unit may, in this case, be used to balance the exploration and exploitation of search.</p>
<p>The algorithm performs a certain mechanism to find out initial weight. The MBFO algorithm follows three important mechanisms that are chemotaxis, reproduction, and elimination-dispersal. Initialization of weight is evaluated based on the position change of the bacterium, and it is given by <xref ref-type="disp-formula" rid="eqn-8">Eq. (8)</xref>:<disp-formula id="eqn-8"><label>(8)</label>
<mml:math id="mml-eqn-8" display="block"><mml:msup><mml:mi>w</mml:mi><mml:mi>j</mml:mi></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mi>z</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:msup><mml:mi>w</mml:mi><mml:mi>j</mml:mi></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mi>z</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:mi>l</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi mathvariant="normal">&#x0398;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:msqrt><mml:msup><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>T</mml:mi></mml:msup><mml:mo stretchy="false">(</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:msqrt></mml:mrow></mml:mfrac></mml:mrow></mml:mstyle></mml:math>
</disp-formula>Here, &#x0394;(<italic>j</italic>) is <italic>k</italic><sup><italic>th</italic></sup> chemotactic step direction vector, <italic>w</italic><sup><italic>j</italic></sup>(<italic>k</italic>, <italic>z</italic>, <italic>l</italic>) indicates bacterium at the <italic>k</italic><sup><italic>th</italic></sup>chemotactic, <italic>z</italic><sup><italic>th</italic></sup> reproductive, <italic>l</italic><sup><italic>th</italic></sup> elimination dispersal step. &#x0398; (<italic>j</italic>) is chemotactic step size while every run or tumble which is formulated by using central angle formulae given by <xref ref-type="disp-formula" rid="eqn-9">Eq. (9)</xref>:<disp-formula id="eqn-9"><label>(9)</label>
<mml:math id="mml-eqn-9" display="block"><mml:mi mathvariant="normal">&#x0398;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>j</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi mathvariant="normal">A</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mi>r</mml:mi></mml:mfrac></mml:mrow></mml:mstyle></mml:math>
</disp-formula>where <italic>S</italic><sub>A</sub> denotes the arc length and <italic>r</italic> is the radius length</p>
<p>For convolutional layer, the kernals/filters are initialized with MBFO optimizer<list list-type="simple"><list-item>
<p><italic>(b) Pooling Layer</italic></p></list-item></list></p>
<p>Pooling is used in order to preserve input image size. Output image size for &#x2018;SAME&#x2019; Pooling is the same as the input image size and there is no Pooling for &#x2018;True&#x2019; Pooling. The size of the Pooling output matrix is illustrated as <xref ref-type="disp-formula" rid="eqn-10">Eq. (10)</xref>:<disp-formula id="eqn-10"><label>(10)</label>
<mml:math id="mml-eqn-10" display="block"><mml:mo stretchy="false">[</mml:mo><mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x0393;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">]</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x0393;</mml:mi><mml:mo>+</mml:mo><mml:mn>2</mml:mn><mml:mi>p</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03BB;</mml:mi></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03B3;</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math>
</disp-formula>Here, <italic>O</italic><sub><italic>CONV</italic></sub> is the output, p is the Pooling, <italic>&#x03B3;</italic><sub><italic>s</italic></sub> is the stride, <italic>&#x03B4;</italic><sub><italic>sigmoid</italic></sub> is sigmoid activation function, <italic>w</italic><sub><italic>i</italic>,<italic>j</italic></sub> weight matrix of shared weights and &#x03A8;<sub><italic>a</italic>&#x002B;<italic>i</italic>,<italic>b</italic>&#x002B;<italic>j</italic></sub> is input activation at location <italic>i</italic>, <italic>j</italic>.</p>
<p>After Pooling of output matrix, the convolution layer obtains a feature map for the text matrix as well as for image data. The obtained feature map is provided to fully connected layer.<list list-type="simple"><list-item>
<p><italic>(c) Fully Connected Layer</italic></p></list-item></list></p>
<p>Result of previous segmentation mask layer is smoothed and sent into fully connected layer as an input. Flattened vector is used to train fully connected layer, that is similar to an ANN. <xref ref-type="disp-formula" rid="eqn-11">Eq. (11)</xref> used in the training of vectors<disp-formula id="eqn-11"><label>(11)</label>
<mml:math id="mml-eqn-11" display="block"><mml:msubsup><mml:mi mathvariant="normal">&#x0393;</mml:mi><mml:mi>I</mml:mi><mml:mi>T</mml:mi></mml:msubsup><mml:mo>=</mml:mo><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:munderover><mml:mrow><mml:msub><mml:mi>w</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mrow><mml:mi>f</mml:mi><mml:mi>l</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>e</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x2135;</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math>
</disp-formula>where, <inline-formula id="ieqn-6">
<mml:math id="mml-ieqn-6"><mml:msub><mml:mi mathvariant="normal">&#x2135;</mml:mi><mml:mi>b</mml:mi></mml:msub></mml:math>
</inline-formula> implies bias initialized randomly, <italic>w</italic><sub><italic>i</italic></sub> implies weights of respective input nodes, <italic>act</italic> implies activation function. To obtain vector values of input image, fully connected layer uses softmax activation function and reflects activation function. The vector image is then sent to the decoder, which handles the captioning. The of the suggested encoder approach, MBFOCNN, is outlined and depicted as pseudo-code in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Pseudo code for MBFOCNN</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CSSE_24118-fig-3.png"/>
</fig>
</sec>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Caption Generation</title>
<p>Caption generation helps in training the model based on the text so as to obtain the respective captioning. Before giving a text as it is, from a dataset, may cause a high error rate for captioning an image. In order to overcome the flaws, the work was developed with a two major step i.e.,<list list-type="bullet"><list-item>
<p>Preprocessing of textual content</p></list-item><list-item>
<p>Word embedding</p></list-item></list></p>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>Preprocessing of Textual Content</title>
<p>Preprocessing of textual content helps to enhance text quality to minimize the error rate.</p>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>Word Embedding</title>
<p>Word embeddings convert single words into fully valued vectors within specified vector spaces where every word is mapped as a single vector, and vector values are learnt in a form that resemble neural networks or classifying approaches of deep learning. In order to provide the vector values for the pre-processed text, this work uses MGloVes word embedding technique. MGloVe provides word embeddings by combining both global statistics of matrix factorization approaches such as LSA with Word2Vec &#x2018;s local context dependent learning. The MGloVe performs certain steps:</p>
<p><bold>Step 1:</bold> Collect co-occurrence word statistics as the matrix of word co-occurrence matrix &#x03A8;. Every component of the matrix &#x03A8;<sub><italic>ij</italic></sub> represents how much appears in a single word&#x2019;s meaning. It normally looks for background words in a certain area determined by size of the window before and after the word.&#x2019; For more distant words, it normally assigns less weight using the following formula in <xref ref-type="disp-formula" rid="eqn-12">Eq. (12)</xref>:<disp-formula id="eqn-12"><label>(12)</label>
<mml:math id="mml-eqn-12" display="block"><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi><mml:mi>a</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>f</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi></mml:math>
</disp-formula></p>
<p><bold>Step 2:</bold> A soft constraint has been defined for each word pair using <xref ref-type="disp-formula" rid="eqn-13">Eq. (13)</xref>:<disp-formula id="eqn-13"><label>(13)</label>
<mml:math id="mml-eqn-13" display="block"><mml:msubsup><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi><mml:mi>T</mml:mi></mml:msubsup><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03A8;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math>
</disp-formula>where, <italic>&#x03C9;</italic><sub><italic>i</italic></sub> denotes key word, <italic>&#x03C9;</italic><sub><italic>j</italic></sub> indicates vector for context word, <italic>b</italic><sub><italic>i</italic></sub> and <italic>b</italic><sub><italic>j</italic></sub> are scalar biases for main and context words. &#x03A8;<sub><italic>ij</italic></sub> denotes the input embedded vector.</p>
<p><bold>Step 3:</bold> cost function evaluation are done using <xref ref-type="disp-formula" rid="eqn-14">Eq. (14)</xref>:<disp-formula id="eqn-14"><label>(14)</label>
<mml:math id="mml-eqn-14" display="block"><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mrow><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mrow><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo></mml:mrow></mml:mrow><mml:msub><mml:mi>&#x03C2;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msubsup><mml:mi>&#x03C9;</mml:mi><mml:mi>i</mml:mi><mml:mi>T</mml:mi></mml:msubsup><mml:msub><mml:mi>&#x03C9;</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03C2;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mn>2</mml:mn></mml:msup></mml:math>
</disp-formula>Here, <italic>N</italic> is vocabulary size, <italic>f</italic> is weighting feature that aids us to prevent learning from very common word pairs. MGloVe select the following feature is based on <xref ref-type="disp-formula" rid="eqn-15">Eq. (15)</xref>:<disp-formula id="eqn-15"><label>(15)</label>
<mml:math id="mml-eqn-15" display="block"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03C2;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mi>&#x03C2;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03C2;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mrow></mml:mstyle></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mi>&#x2113;</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:msub><mml:mi>&#x03C2;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x003C;</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>h</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:msub><mml:mi>&#x03C2;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub></mml:math>
</disp-formula></p>
<p><bold>Step 4:</bold> Word embedding produced must be compared to obtain semantically resemblance among two vectors. The similarity between the two vectors is found out using <xref ref-type="disp-formula" rid="eqn-16">Eq. (16)</xref><disp-formula id="eqn-16"><label>(16)</label>
<mml:math id="mml-eqn-16" display="block"><mml:mi>O</mml:mi><mml:mi>L</mml:mi><mml:mi>S</mml:mi><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>e</mml:mi><mml:mi>f</mml:mi><mml:mi>N</mml:mi><mml:mi>o</mml:mi><mml:mi>I</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>p</mml:mi><mml:mi>t</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03C2;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:msub><mml:mi>&#x03C2;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mrow><mml:mfrac><mml:mrow><mml:mo fence="false" stretchy="false">&#x27E8;</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03C2;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thickmathspace" /><mml:msub><mml:mi>&#x03C2;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo></mml:mrow></mml:msub></mml:mrow><mml:mo fence="false" stretchy="false">&#x27E9;</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mrow><mml:msub><mml:mi>&#x03C2;</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mrow></mml:mstyle></mml:math>
</disp-formula></p>
</sec>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Decoder</title>
<p>The decoder is done in order to provide a caption for an image. In the decoding process, the text gets trained up by following the textual understanding and thereafter gets the captioning of the image. Several decoding models have been developed up-to-date but accurately captioning an image remains to be a challenge and computationally complex. Existing methods/models are effective in their IC processes, but unable to determine using visual signals or language models correspondingly. In order to overcome the flaws, the work has proposed MDAAs for visual sentinels as shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>MDAA decoder</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CSSE_24118-fig-4.png"/>
</fig>
<p>MDAA provides a decoding process by accepting the image vector as well as embedding text vectors to caption an image. The MDAA uses a Gated Recurrent Unit (GRU) as a decoder, which avoids the gradient vanishing or exploding problems and remains to be less complex as compared to the LSTM.</p>
<sec id="s3_3_1">
<label>3.3.1</label>
<title>Reset and Update Gate</title>
<p>current time step is first supplied as input to both the reset and update gates in GRU, whereas prior time step is given to hidden state. Two fully-connected layers with sigmoid activation function provide outputs of two gates.</p>
<p>If the input is a minibatch <inline-formula id="ieqn-7">
<mml:math id="mml-ieqn-7"><mml:msub><mml:mi mathvariant="normal">&#x2111;</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211C;</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:msup></mml:math>
</inline-formula> (count of instances: <italic>n</italic>, count of inputs: <italic>d</italic>) and the previous time step&#x2019;s <italic>t</italic> hidden state is <inline-formula id="ieqn-8">
<mml:math id="mml-ieqn-8"><mml:msub><mml:mrow><mml:mi mathvariant="normal">H</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211C;</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>h</mml:mi></mml:mrow></mml:msup></mml:math>
</inline-formula> (total hidden units: <italic>h</italic>), then reset and update gates <inline-formula id="ieqn-9">
<mml:math id="mml-ieqn-9"><mml:msub><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211C;</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>h</mml:mi></mml:mrow></mml:msup></mml:math>
</inline-formula> and <inline-formula id="ieqn-10">
<mml:math id="mml-ieqn-10"><mml:msub><mml:mi mathvariant="normal">&#x03A6;</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211C;</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>h</mml:mi></mml:mrow></mml:msup></mml:math>
</inline-formula> are calculated as follows:<disp-formula id="eqn-17"><label>(17)</label>
<mml:math id="mml-eqn-17" display="block"><mml:msub><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x2111;</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>X</mml:mi><mml:mi mathvariant="normal">&#x2200;</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="normal">H</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>h</mml:mi><mml:mi mathvariant="normal">&#x2200;</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi mathvariant="normal">&#x2200;</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math>
</disp-formula><disp-formula id="eqn-18"><label>(18)</label>
<mml:math id="mml-eqn-18" display="block"><mml:msub><mml:mi mathvariant="normal">&#x03A6;</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x2111;</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>X</mml:mi><mml:mi mathvariant="normal">&#x03A6;</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="normal">H</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>h</mml:mi><mml:mi mathvariant="normal">&#x03A6;</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi mathvariant="normal">&#x03A6;</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math>
</disp-formula>where, <inline-formula id="ieqn-11">
<mml:math id="mml-ieqn-11"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>X</mml:mi><mml:mi mathvariant="normal">&#x2200;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>X</mml:mi><mml:mi mathvariant="normal">&#x03A6;</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211C;</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>h</mml:mi></mml:mrow></mml:msup></mml:math>
</inline-formula> and <inline-formula id="ieqn-12">
<mml:math id="mml-ieqn-12"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>h</mml:mi><mml:mi mathvariant="normal">&#x2200;</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>h</mml:mi><mml:mi mathvariant="normal">&#x03A6;</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211C;</mml:mi><mml:mrow><mml:mi>h</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>h</mml:mi></mml:mrow></mml:msup></mml:math>
</inline-formula> is weight parameters <inline-formula id="ieqn-13">
<mml:math id="mml-ieqn-13"><mml:msub><mml:mi>b</mml:mi><mml:mi mathvariant="normal">&#x2200;</mml:mi></mml:msub><mml:msub><mml:mi>b</mml:mi><mml:mi mathvariant="normal">&#x03A6;</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211C;</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>h</mml:mi></mml:mrow></mml:msup></mml:math>
</inline-formula> is bias, <italic>&#x03C3;</italic> is the sigmoid function to transform the input interval from (0, 1).</p>
</sec>
<sec id="s3_3_2">
<label>3.3.2</label>
<title>Candidate Hidden State</title>
<p>In candidate hidden state, reset gate is integrated with regular latent state updating mechanism, which is given by<disp-formula id="eqn-19"><label>(19)</label>
<mml:math id="mml-eqn-19" display="block"><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mi mathvariant="normal">H</mml:mi></mml:mrow><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>tanh</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x2111;</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>&#x03B6;</mml:mi><mml:mi>h</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msub><mml:mi mathvariant="normal">&#x2200;</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2297;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="normal">H</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>h</mml:mi><mml:mi>h</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>h</mml:mi></mml:msub></mml:math>
</disp-formula>where, <inline-formula id="ieqn-14">
<mml:math id="mml-ieqn-14"><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mi mathvariant="normal">H</mml:mi></mml:mrow><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211C;</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>h</mml:mi></mml:mrow></mml:msup></mml:math>
</inline-formula> is the candidate hidden state, <inline-formula id="ieqn-15">
<mml:math id="mml-ieqn-15"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>&#x03B6;</mml:mi><mml:mi>h</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211C;</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>h</mml:mi></mml:mrow></mml:msup></mml:math>
</inline-formula>, <inline-formula id="ieqn-16">
<mml:math id="mml-ieqn-16"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>h</mml:mi><mml:mi>h</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211C;</mml:mi><mml:mrow><mml:mi>h</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>h</mml:mi></mml:mrow></mml:msup></mml:math>
</inline-formula> are weight factors <inline-formula id="ieqn-17">
<mml:math id="mml-ieqn-17"><mml:msub><mml:mi>b</mml:mi><mml:mi>h</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211C;</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>h</mml:mi></mml:mrow></mml:msup></mml:math>
</inline-formula> is bias and &#x2297; is element-wise product operation. The candidate hidden state uses tanh to guarantee that all values remain between interval [&#x2212;1, 1].</p>
</sec>
<sec id="s3_3_3">
<label>3.3.3</label>
<title>Hidden State</title>
<p>Lastly, update gate &#x03A6;<sub><italic>t</italic></sub> has to be included as it tells on new hidden states <inline-formula id="ieqn-18">
<mml:math id="mml-ieqn-18"><mml:msub><mml:mrow><mml:mi mathvariant="normal">H</mml:mi></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi mathvariant="normal">&#x211C;</mml:mi><mml:mrow><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>h</mml:mi></mml:mrow></mml:msup></mml:math>
</inline-formula> and candidate states <inline-formula id="ieqn-19">
<mml:math id="mml-ieqn-19"><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mi mathvariant="normal">H</mml:mi></mml:mrow><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:math>
</inline-formula> which is used by old states H<sub><italic>t</italic>&#x2212;1</sub>. Hence, &#x03A6;<sub><italic>t</italic></sub> is used for element wise convex combinations of H<sub><italic>t</italic>&#x2212;1</sub> and <inline-formula id="ieqn-20">
<mml:math id="mml-ieqn-20"><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mi mathvariant="normal">H</mml:mi></mml:mrow><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:math>
</inline-formula>. T resulting in final GRU update <xref ref-type="disp-formula" rid="eqn-20">Eq. (20)</xref>:<disp-formula id="eqn-20"><label>(20)</label>
<mml:math id="mml-eqn-20" display="block"><mml:msub><mml:mrow><mml:mi mathvariant="normal">H</mml:mi></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03A6;</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2297;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="normal">H</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi mathvariant="normal">&#x03A6;</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2297;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mi mathvariant="normal">H</mml:mi></mml:mrow><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:math>
</disp-formula></p>
<p>Old state is simply retained when the update gate &#x03A6;<sub><italic>t</italic></sub> is near to 1, In this case, data from <italic>&#x03B6;</italic><sub><italic>t</italic></sub> in the dependency chain is basically ignored and effectively skip the time step<italic>t</italic>. In comparison, the current latent state H<sub><italic>t</italic></sub> reaches the latent candidate state <inline-formula id="ieqn-21">
<mml:math id="mml-ieqn-21"><mml:msub><mml:mrow><mml:mover><mml:mrow><mml:mi mathvariant="normal">H</mml:mi></mml:mrow><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub></mml:math>
</inline-formula> once &#x03A6;<sub><italic>t</italic></sub> is close to 0. These designs address RNN&#x2019;s vanishing gradient issue while resulting in better capturing dependencies for long time-step distance sequences. If the update gate has been close to 1 for all of a subsequence&#x2019;s time steps, for instance, the old hidden state will be disclosed. Will be easily kept and transferred to its end at the time stage of its beginning, regardless of the subsequence&#x2019;s duration. Finally, the model decodes the image and provides the captioning. Thus, the overall outline for suggested decoder approach that is MDAAs is illustrated in pseudo-code form as stated in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Pseudo code for MDAA</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CSSE_24118-fig-5.png"/>
</fig>
</sec>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Results and Discussion</title>
<p>This work&#x2019;s suggested framework&#x2019;s performance in ICs is evaluated on publicly available datasets with current methodologies and using various performance indicators to determine its efficacy. The framework is implemented in PYTHON. The performance of the proposed MBFOCNN encoder and MDAA decoder is analyzed with the existing methods: BRNNs (Bidirectional Recurrent Neural Networks) [<xref ref-type="bibr" rid="ref-33">33</xref>]; CNN-LSTM [<xref ref-type="bibr" rid="ref-34">34</xref>]; Fast RCNNs [<xref ref-type="bibr" rid="ref-35">35</xref>] and CNN-AAs (CNNs with Adaptive Attentions) [<xref ref-type="bibr" rid="ref-36">36</xref>]. The analysis has been made based on the metrics: Accuracy, Specificity, Sensitivity, F-Measure, precision; NPVs (Negative predictive values); MCCs (Matthew&#x2019;s Correlation Coefficients);FPRs (False Predictive Rates), and FNRs (False Negative Rates). Based on the metrics, such as Accuracy, Specificity, Sensitivity, and Precision, the evaluation has been done for the various techniques. The evaluation of the techniques has been given in <xref ref-type="table" rid="table-1">Tab. 1</xref>.</p>
<table-wrap id="table-1"><label>Table 1</label>
<caption>
<title>Inference of existing methods for breast cancer detection</title></caption>
<table><colgroup><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">S. no</th>
<th align="left">Author details</th>
<th align="left">Methodology</th>
<th align="left">Merits</th>
<th align="left">Demerits</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">1.</td>
<td align="left">Kinghorn et al. (2017)</td>
<td align="left">Region proposal network from faster R-CNN</td>
<td align="left">Selective search, for example, is less reliable.</td>
<td align="left">However it has issue with time complexity</td>
</tr>
<tr>
<td align="left">2.</td>
<td align="left">Su et al. (2020)</td>
<td align="left">Hierarchical deep neural network for IC</td>
<td align="left">Solves the computational problems from different &#xFB01;elds</td>
<td align="left">Fast prediction is not possible</td>
</tr>
<tr>
<td align="left">3</td>
<td align="left">Anderson et al. (2018)</td>
<td align="left">Faster convolutional neural net<break/>(CNN) R-CNN</td>
<td align="left">High detection rate</td>
<td align="left">Time consumption is high</td>
</tr>
<tr>
<td align="left">4</td>
<td align="left">Zhu et al. (2018)</td>
<td align="left">CNN classifier</td>
<td align="left">High sensitivity and accuracy</td>
<td align="left">Requires high number of mammogram images for training process</td>
</tr>
<tr>
<td align="left">5</td>
<td align="left">Xu et al. (2015)</td>
<td align="left">Deep learning algorithm</td>
<td align="left">It achieves high level of classification accuracy<break/>Overall efficiency is greater</td>
<td align="left">This technique has issue with longer training time</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-2">Tab. 2</xref> shows how the proposed MBFOCNN-MDAA was evaluated using performance criteria and current methods specified above. Fast RCNN-RNN, and CNN-AA. For obtaining an efficient IC, the encoder and decoder should work together better in order to minimize the error and to obtain an accurate captioning of an image. The proposed model is also evaluated with the metrics by achieving an Accuracy value of 96.24&#x0025;, Specificity value of 96.49&#x0025;, Sensitivity value of 97.63&#x0025;, and Precision value of 97.63&#x0025;, which ranges between 96.24&#x0025;&#x2013;97.63&#x0025;. But the existing BRNN, CNN-LSTM, Fast RCNN-RNN, and CNN-AA methods often achieve a value for the metrics ranging from 81.39 percent to 92.52 percent, which is lower than the suggested technique. By limiting the incidence of mistakes, the suggested MBFOCNN-MDAA technique provides superior encoding and decoding for creating an image-based caption, and it is proven to be efficient when compared to existing methods. <xref ref-type="fig" rid="fig-6">Fig. 6</xref> depicts a graphical depiction of the suggested technique alongside several current methods.</p>
<table-wrap id="table-2"><label>Table 2</label>
<caption>
<title>Analysis of suggested MBFOCNN-MDAA depending on accuracy, specificity, sensitivity, and precision</title></caption>
<table><colgroup><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Performance metrics/techniques</th>
<th align="left">BRNN</th>
<th align="left">CNN-LSTM</th>
<th align="left">Fast RCNN&#x2013;RNN</th>
<th align="left">CNN-AA</th>
<th align="left">Proposed MBFOCNN-MDAA</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Accuracy</td>
<td align="left">90.86</td>
<td align="left">88.19</td>
<td align="left">81.39</td>
<td align="left">83.81</td>
<td align="left">96.24</td>
</tr>
<tr>
<td align="left">Specificity</td>
<td align="left">91.48</td>
<td align="left">88.92</td>
<td align="left">81.82</td>
<td align="left">84.14</td>
<td align="left">96.49</td>
</tr>
<tr>
<td align="left">Sensitivity</td>
<td align="left">92.22</td>
<td align="left">91.71</td>
<td align="left">88.71</td>
<td align="left">84.88</td>
<td align="left">97.63</td>
</tr>
<tr>
<td align="left">Precision</td>
<td align="left">92.52</td>
<td align="left">91.71</td>
<td align="left">88.23</td>
<td align="left">84.86</td>
<td align="left">97.63</td>
</tr>
</tbody>
</table>
</table-wrap>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Graphical analysis of proposed MBFOCNN-MDAA based on accuracy, specificity, sensitivity, and precision</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CSSE_24118-fig-6.png"/>
</fig>
<p><xref ref-type="fig" rid="fig-7">Figs. 7a</xref> and <xref ref-type="fig" rid="fig-7">7b</xref> gives a graphical representation of the proposed MBFOCNN-MDAA method evaluated with the existing BRNN, CNN-LSTM, Fast RCNN-RNN, and CNN-AA methods. Graphically, it can be stated that proposed methods tend to achieve a high metrics value as compared to existing methods. Thereafter, the BRNN techniques perform better after the proposed technique, and the remaining CNN-LSTM, Fast RCNN-RNN, and CNN-AA differ by a wide range while compared with the suggested approach. MBFOCNN-MDAA is analyzed depending on metrics, such as F-Measure, NPV, MCC, FPR, and FNR, and evaluated with the existing methods, such as BRNN, CNN-LSTM, Fast RCNN-RNN, and CNN-AA. The evaluation of the techniques has been tabulated in <xref ref-type="table" rid="table-3">Tab. 3</xref>.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Graphical analysis of suggested MBFOCNN-MDAA formed on (a) F&#x2013;Measure, NPV, MCC (b) FPR, FNR</title></caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CSSE_24118-fig-7.png"/>
</fig>
<table-wrap id="table-3"><label>Table 3</label>
<caption>
<title>Analysis of the proposed MBFOCNN-MDAA based on F-Measure, NPV, MCC, FPR, and KNN</title></caption>
<table><colgroup><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/><col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Performance metrics/techniques</th>
<th align="left">BRNN</th>
<th align="left">CNN-LSTM</th>
<th align="left">Fast RCNN&#x2013;RNN</th>
<th align="left">CNN-AA</th>
<th align="left">Proposed MBFOCNN-MDAA</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">F-measure</td>
<td align="left">92.22</td>
<td align="left">91.71</td>
<td align="left">88.71</td>
<td align="left">84.88</td>
<td align="left">97.63</td>
</tr>
<tr>
<td align="left">NPV</td>
<td align="left">96.48</td>
<td align="left">95.92</td>
<td align="left">95.82</td>
<td align="left">96.24</td>
<td align="left">97.49</td>
</tr>
<tr>
<td align="left">MCC</td>
<td align="left">92.22</td>
<td align="left">91.71</td>
<td align="left">89.71</td>
<td align="left">84.58</td>
<td align="left">97.63</td>
</tr>
<tr>
<td align="left">FPR</td>
<td align="left">14.59</td>
<td align="left">35.92</td>
<td align="left">34.24</td>
<td align="left">21.59</td>
<td align="left">2.65</td>
</tr>
<tr>
<td align="left">FNR</td>
<td align="left">21.11</td>
<td align="left">51.41</td>
<td align="left">35.43</td>
<td align="left">25.14</td>
<td align="left">13.04</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Based on the metrics, such as F-Measure, NPV, MCC, FPR, and FNR, the evaluation of the proposed MBFOCNN-MDAA has been done with various existing techniques, such as BRNN, CNN-LSTM, Fast RCNN-RNN, and CNN-AA as shown in <xref ref-type="table" rid="table-3">Tab. 3</xref>. In order to depict a better model, not every measure acquired must be of greater value; nevertheless, achieving lower values for specific measures, like FPR and FNR, can help to achieve an effective system. As per this, suggested system attains an FPR of 2.65 percent and a FNR of 13.04 percent, while current approaches try to obtain FPR and FNR values between 14.59 percent to 51.41 percent, indicating a high level of false detection of the text to the respective image when contrasted with suggested technique. Suggested approach, conversely, achieves high F-Measure, NPV, and MCC value of 97.63 percent, 97.49 percent, and 97.63 percent, respectively, whereas the existing method achieves F-Measure, NPV, and MCC values between 84.88 percent to 96.48 percent, indicating lower model efficiency when compared to the proposed MBFOCNN-MDAA method. The suggested MBFOCNN-MDAA technique improves encoding quality by eliminating erroneous detection, and decoding for creating an image-based caption, it was found to be more efficient than previous techniques which is due to faster prediction and good convergence. <xref ref-type="fig" rid="fig-7">Fig. 7</xref> depicts a graphical depiction of the suggested technique alongside several current methods.</p>
<p><xref ref-type="fig" rid="fig-7">Fig. 7</xref> displays the suggested MBFOCNN-MDAA method&#x2019;s graphical analysis depending on performance metrics. <xref ref-type="fig" rid="fig-7">Fig. 7a</xref> indicates graphical analysis of techniques depending on F&#x2013;Measure, NPV, MCC, which states that the proposed technique achieves a better metrics value as compared to the existing techniques. The CNN-LSTM technique tends to give better results after the proposed method and the remaining existing techniques vary within a huge difference. <xref ref-type="fig" rid="fig-7">Fig. 7b</xref> states the FPR and FNR metrics analysis for the proposed technique with the existing methodologies. The proposed method tends to achieve a lower value of FPR and FNR as compared to existing methods, and it is considered to be efficient due to avoiding the false prediction for IC. Thus, the proposed MBFOCNN-MDAA method outperforms the existing methods and remains to perform accurate IC.</p>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>The proposed work has developed a framework for ICs consisting of two important components that are image encoding using MBFOCNN and decoding of the image using MDAA. The two major components provide an accurate IC and also extract semantic correlations of acquired visual information with corresponding language expressions. In addition to that, the proposed work provides accurate IC for complex scenes and concentrates on every feature by keeping in mind the quality of the image. Contrasted with modern method, suggested system attends to interrelationships between objects in an image and provides Automatic caption or description generation from images. The experimental outcome showed that suggested approach achieved 96.24&#x0025; accuracy, 96.49&#x0025; Specificity, 97.63&#x0025; Precision, and obtains a minimized false detection by achieving an FPR and FNR value of 2.65&#x0025; and 13.04&#x0025; respectively. In the future, the work will be directed in improving the encoder and decoder mechanism for captioning an image.</p>
</sec>
</body>
<back>
<ack>
<p>We thank anonymous referees for their helpful suggestions.</p>
</ack><fn-group>
<fn fn-type="other">
<p><bold>Funding Statement:</bold> The authors received no specific funding for this study.</p>
</fn>
<fn fn-type="conflict">
<p><bold>Conflicts of Interest:</bold> The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</fn>
</fn-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>P.</given-names> <surname>He</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Skitmore</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Manifesting construction activity scenes via image captioning</article-title>,&#x201D; <source>Automation in Construction</source>, vol. <volume>119</volume>, no. <issue>6</issue>, pp. <fpage>01</fpage>&#x2013;<lpage>19</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Li</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Jiang</surname></string-name></person-group>, &#x201C;<article-title>Know more say less: Image captioning based on scene graphs</article-title>,&#x201D; <source>IEEE Transactions on Multimedia</source>, vol. <volume>21</volume>, no. <issue>8</issue>, pp. <fpage>2117</fpage>&#x2013;<lpage>2130</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Q.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Shen</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Dick</surname></string-name> and <string-name><given-names>A. V. D.</given-names> <surname>Hengel</surname></string-name></person-group>, &#x201C;<article-title>IC and visual question answering based on attributes and external knowledge</article-title>,&#x201D; <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>, vol. <volume>40</volume>, no. <issue>6</issue>, pp. <fpage>1367</fpage>&#x2013;<lpage>1381</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>He</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Lu</surname></string-name></person-group>, &#x201C;<article-title>A modularized architecture of multi-branch convolutional neural network for image captioning</article-title>,&#x201D; <source>Electronics</source>, vol. <volume>8</volume>, no. <issue>12</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>15</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Hoxha</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Melgani</surname></string-name> and <string-name><given-names>B.</given-names> <surname>Demir</surname></string-name></person-group>, &#x201C;<article-title>Toward remote sensing image retrieval under a deep IC perspective</article-title>,&#x201D; <source>IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing</source>, vol. <volume>13</volume>, no. <issue>3</issue>, pp. <fpage>4462</fpage>&#x2013;<lpage>4475</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Ye</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Han</surname></string-name> and <string-name><given-names>N.</given-names> <surname>Liu</surname></string-name></person-group>, &#x201C;<article-title>Attentive linear transformation for IC</article-title>,&#x201D; <source>IEEE Transactions on Image Processing</source>, vol. <volume>27</volume>, no. <issue>11</issue>, pp. <fpage>5514</fpage>&#x2013;<lpage>5524</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Yu</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Yu</surname></string-name> and <string-name><given-names>Q.</given-names> <surname>Huang</surname></string-name></person-group>, &#x201C;<article-title>Multimodal transformer with multi-view visual representation for image captioning</article-title>,&#x201D; <source>IEEE Transactions on Circuits and Systems for Video Technology</source>, vol. <volume>30</volume>, no. <issue>12</issue>, pp. <fpage>4467</fpage>&#x2013;<lpage>4480</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>D. D.</given-names> <surname>Feng</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Learning visual relationship and context-aware attention for IC</article-title>,&#x201D; <source>Pattern Recognition</source>, vol. <volume>98</volume>, no. <issue>10</issue>, pp. <fpage>01</fpage>&#x2013;<lpage>11</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>M. Z.</given-names> <surname>Hossain</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Sohel</surname></string-name>, <string-name><given-names>M. F.</given-names> <surname>Shiratuddin</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Laga</surname></string-name> and <string-name><surname>Bennamoun</surname> <given-names>M.</given-names></string-name></person-group>, &#x201C;<chapter-title>Bi-SAN-CAP:Bi-directional self-attention for image captioning</chapter-title>,&#x201D; in <source>IEEE 2019 Digital Image Computing: Techniques and Applications (DICTA)</source>, pp. <fpage>1</fpage>&#x2013;<lpage>7</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Cao</surname></string-name>, <string-name><given-names>G.</given-names> <surname>An</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Zheng</surname></string-name> and <string-name><given-names>Q.</given-names> <surname>Ruan</surname></string-name></person-group>, &#x201C;<article-title>Interactions guided generative adversarial network for unsupervised IC</article-title>,&#x201D; <source>Neurocomputing</source>, vol. <volume>417</volume>, no. <issue>12</issue>, pp. <fpage>419</fpage>&#x2013;<lpage>431</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Xia</surname></string-name>, <string-name><given-names>J.</given-names> <surname>He</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Yin</surname></string-name></person-group>, &#x201C;<article-title>Boosting image caption generation with feature fusion module</article-title>,&#x201D; <source>Multimedia Tools and Applications</source>, vol. <volume>79</volume>, no. <issue>33</issue>, pp. <fpage>24225</fpage>&#x2013;<lpage>24239</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Yang</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Meinel</surname></string-name></person-group>, &#x201C;<article-title>IC with deep bidirectional LSTMs and multi-task learning</article-title>,&#x201D; <source>ACM Transactions on Multimedia Computing, Communications, and Applications (TOMM)</source>, vol. <volume>14</volume>, no. <issue>2s</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>20</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J. H.</given-names> <surname>Tan</surname></string-name>, <string-name><given-names>C. S.</given-names> <surname>Chan</surname></string-name> and <string-name><given-names>J. H.</given-names> <surname>Chuah</surname></string-name></person-group>, &#x201C;<article-title>COMIC: Toward a compact IC model with attention</article-title>,&#x201D; <source>IEEE Transactions on Multimedia</source>, vol. <volume>21</volume>, no. <issue>10</issue>, pp. <fpage>2686</fpage>&#x2013;<lpage>2696</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Zhao</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Feng</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Zhao</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Multitask learning for cross-domain IC</article-title>,&#x201D; <source>IEEE Transactions on Multimedia</source>, vol. <volume>21</volume>, no. <issue>4</issue>, pp. <fpage>1047</fpage>&#x2013;<lpage>1061</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Wang</surname></string-name> and <string-name><given-names>F.</given-names> <surname>Chen</surname></string-name></person-group>, &#x201C;<article-title>High-quality IC with fine-grained and semantic-guided visual attention</article-title>,&#x201D; <source>IEEE Transactions on Multimedia</source>, vol. <volume>21</volume>, no. <issue>7</issue>, pp. <fpage>1681</fpage>&#x2013;<lpage>1693</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Xiao</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Ding</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Xiang</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Pan</surname></string-name></person-group>, &#x201C;<article-title>Dense semantic embedding network for IC</article-title>,&#x201D; <source>Pattern Recognition</source>, vol. <volume>90</volume>, pp. <fpage>285</fpage>&#x2013;<lpage>296</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Ding</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Lin</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Shan</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Image captioning with memorized knowledge</article-title>,&#x201D; <source>Cognitive Computation</source>, vol. <volume>13</volume>, no. <issue>4</issue>, pp. <fpage>807</fpage>&#x2013;<lpage>820</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>He</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Song</surname></string-name>, <string-name><given-names>R. W. H.</given-names> <surname>Lau</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Jiao</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>IC via semantic element embedding</article-title>,&#x201D; <source>Neurocomputing</source>, vol. <volume>395</volume>, no. <issue>6</issue>, pp. <fpage>212</fpage>&#x2013;<lpage>221</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Xiao</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Gong</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Shen</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Li</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>DAA: Dual LSTMs with adaptive attention for IC</article-title>,&#x201D; <source>Neurocomputing</source>, vol. <volume>364</volume>, no. <issue>10</issue>, pp. <fpage>322</fpage>&#x2013;<lpage>329</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Xiao</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Ding</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Xiang</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Pan</surname></string-name></person-group>, &#x201C;<article-title>Deep hierarchical encoder&#x2013;decoder network for IC</article-title>,&#x201D; <source>IEEE Transactions on Multimedia</source>, vol. <volume>21</volume>, no. <issue>11</issue>, pp. <fpage>2942</fpage>&#x2013;<lpage>2956</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Shen</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Zhao</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Liu</surname></string-name></person-group>, &#x201C;<article-title>Remote sensing IC via variational autoencoder and reinforcement learning</article-title>,&#x201D; <source>Knowledge-Based Systems</source>, vol. <volume>203</volume>, no. <issue>4</issue>, pp. <fpage>01</fpage>&#x2013;<lpage>11</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Kinghorn</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Zhang</surname></string-name> and <string-name><given-names>L.</given-names> <surname>Shao</surname></string-name></person-group>, &#x201C;<article-title>A hierarchical and regional deep learning architecture for image description generation</article-title>,&#x201D; <source>Pattern Recognition Letters</source>, vol. <volume>119</volume>, no. <issue>9</issue>, pp. <fpage>77</fpage>&#x2013;<lpage>85</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Su</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Xu</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Liu</surname></string-name></person-group>, &#x201C;<article-title>Hierarchical deep neural network for IC</article-title>,&#x201D; <source>Neural Processing Letters</source>, vol. <volume>52</volume>, no. <issue>2</issue>, pp. <fpage>1057</fpage>&#x2013;<lpage>1067</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Zhao</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Chang</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Guo</surname></string-name></person-group>, &#x201C;<article-title>A multimodal fusion approach for IC</article-title>,&#x201D; <source>Neurocomputing</source>, vol. <volume>329</volume>, pp. <fpage>476</fpage>&#x2013;<lpage>485</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Ba</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Kiros</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Cho</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Courville</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Show, attend and tell: Neural image caption generation with visual attention</article-title>,&#x201D; in <conf-name>Int. Conf. on Machine Learning</conf-name>, <conf-loc>Lille, France</conf-loc>, pp. <fpage>2048</fpage>&#x2013;<lpage>2057</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Anderson</surname></string-name>, <string-name><given-names>X.</given-names> <surname>He</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Buehler</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Teney</surname></string-name> <string-name><given-names>M.</given-names> <surname>Johnson</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>Bottom-up and top-down attention for image captioning and visual question answering</article-title>,&#x201D; in <conf-name>Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition</conf-name>, <comment>Salt Lake City, UT, USA</comment>, pp. <fpage>6077</fpage>&#x2013;<lpage>6086</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Vaswani</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Shazeer</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Parmar</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Uszkoreit</surname></string-name> <string-name><given-names>L.</given-names> <surname>Jones</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<chapter-title>Attention is all you need. MIT Press</chapter-title>,&#x201D; in <source>Advances in Neural Information Processing Systems</source>, pp. <fpage>5998</fpage>&#x2013;<lpage>6008</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Zhu</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Peng</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Niu</surname></string-name></person-group>, &#x201C;<article-title>Captioning transformer with stacked attention modules</article-title>,&#x201D; <source>Applied Sciences</source>, vol. <volume>8</volume>, no. <issue>5</issue>, pp. <fpage>01</fpage>&#x2013;<lpage>11</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Herdade</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Kappeler</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Boakye</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Soares</surname></string-name></person-group>, &#x201C;<chapter-title>IC: Transforming objects into words</chapter-title>,&#x201D; <source>MIT press in Advances in Neural Information Processing Systems</source>, <publisher-loc>Cambridge</publisher-loc>, <publisher-loc>MA</publisher-loc>, <publisher-loc>USA</publisher-loc>, pp. <fpage>11135</fpage>&#x2013;<lpage>11145</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Chen</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Wei</surname></string-name></person-group>, &#x201C;<article-title>Attention on attention for IC</article-title>,&#x201D; in <conf-name>Proc. of the IEEE Int. Conf. on Computer Vision</conf-name>, <conf-loc>Seoul, Korea</conf-loc>, pp. <fpage>4634</fpage>&#x2013;<lpage>4643</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Pan</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Yao</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Li</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Tao</surname></string-name></person-group>, &#x201C;<article-title>X-linear attention networks for IC</article-title>,&#x201D; in <conf-name>Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>Nashville, TN, USA</conf-loc>, pp. <fpage>10971</fpage>&#x2013;<lpage>10980</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Cornia</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Stefanini</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Baraldi</surname></string-name> and <string-name><given-names>R.</given-names> <surname>Cucchiara</surname></string-name></person-group>, &#x201C;<article-title>Meshed-memory transformer for IC</article-title>,&#x201D; in <conf-name>Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>Nashville, TN, USA</conf-loc>, pp. <fpage>10578</fpage>&#x2013;<lpage>10587</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Fan</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Qian</surname></string-name>, <string-name><given-names>F. L.</given-names> <surname>Xie</surname></string-name> and <string-name><given-names>F. K.</given-names> <surname>Soong</surname></string-name></person-group>, &#x201C;<article-title>TTS synthesis with bidirectional LSTM based recurrent neural networks</article-title>,&#x201D; in <conf-name>Fifteenth Annual Conf. of the Int. Speech Communication Association</conf-name>, <conf-loc>Singapore</conf-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>5</lpage>, <year>2014</year>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Soh</surname></string-name></person-group>, &#x201C;<source>Learning CNN-LSTM Architectures for Image Caption Generation</source>,&#x201D; <publisher-name>Department of Computer Science Stanford University</publisher-name>, <publisher-loc>CA</publisher-loc>, <publisher-loc>USA</publisher-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>9</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z. C.</given-names> <surname>Fei</surname></string-name></person-group>, &#x201C;<article-title>Fast image caption generation with position alignment</article-title>,&#x201D; <source>Computer Vision and Pattern Recognition</source>, pp. <fpage>1</fpage>&#x2013;<lpage>8</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Hani</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Tagougui</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Kherallah</surname></string-name></person-group>, &#x201C;<article-title>Image caption generation using a deep architecture</article-title>,&#x201D; in <conf-name>Int. Arab Conf. on Information Technology (ACIT)</conf-name>, <conf-loc>Al Ain, United Arab Emirates</conf-loc>, pp. <fpage>246</fpage>&#x2013;<lpage>251</lpage>, <year>2019</year>.</mixed-citation></ref>
</ref-list>
</back>
</article>