<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">42963</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2023.042963</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Efficient Object Segmentation and Recognition Using Multi-Layer Perceptron Networks</article-title>
<alt-title alt-title-type="left-running-head">Efficient Object Segmentation and Recognition Using Multi-Layer Perceptron Networks</alt-title>
<alt-title alt-title-type="right-running-head">Efficient Object Segmentation and Recognition Using Multi-Layer Perceptron Networks</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Naseer</surname><given-names>Aysha</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Almujally</surname><given-names>Nouf Abdullah</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Alotaibi</surname><given-names>Saud S.</given-names></name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Alazeb</surname><given-names>Abdulwahab</given-names></name><xref ref-type="aff" rid="aff-4">4</xref></contrib>
<contrib id="author-5" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Park</surname><given-names>Jeongmin</given-names></name><xref ref-type="aff" rid="aff-5">5</xref><email>jmpark@tukorea.ac.kr</email></contrib>
<aff id="aff-1"><label>1</label><institution>Department of Computer Science, Air University</institution>, <addr-line>Islamabad, 44000</addr-line>, <country>Pakistan</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of Information Systems, College of Computer and Information Sciences, Princess Nourah bint Abdulrahman University</institution>, <addr-line>P.O. Box 84428, Riyadh, 11671</addr-line>, <country>Saudi Arabia</country></aff>
<aff id="aff-3"><label>3</label><institution>Information System Department, Umm Al-Qura University</institution>, <addr-line>Makkah</addr-line>, <country>Saudi Arabia</country></aff>
<aff id="aff-4"><label>4</label><institution>Department of Computer Science, College of Computer Science and Information System, Najran University</institution>, <addr-line>Najran, 55461</addr-line>, <country>Saudi Arabia</country></aff>
<aff id="aff-5"><label>5</label><institution>Department of Computer Engineering, Tech University of Korea, Gyeonggi-do</institution>, <addr-line>15073</addr-line>, <country>South Korea</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Jeongmin Park. Email: <email>jmpark@tukorea.ac.kr</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic"><year>2024</year></pub-date>
<pub-date date-type="pub" publication-format="electronic"><day>30</day><month>1</month><year>2024</year></pub-date>
<volume>78</volume>
<issue>1</issue>
<fpage>1381</fpage>
<lpage>1398</lpage>
<history>
<date date-type="received"><day>17</day><month>6</month><year>2023</year></date>
<date date-type="accepted"><day>13</day><month>11</month><year>2023</year></date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2024 Naseer et al.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Naseer et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_42963.pdf"></self-uri>
<abstract>
<p>Object segmentation and recognition is an imperative area of computer vision and machine learning that identifies and separates individual objects within an image or video and determines classes or categories based on their features. The proposed system presents a distinctive approach to object segmentation and recognition using Artificial Neural Networks (ANNs). The system takes RGB images as input and uses a k-means clustering-based segmentation technique to fragment the intended parts of the images into different regions and label them based on their characteristics. Then, two distinct kinds of features are obtained from the segmented images to help identify the objects of interest. An Artificial Neural Network (ANN) is then used to recognize the objects based on their features. Experiments were carried out with three standard datasets, MSRC, MS COCO, and Caltech 101 which are extensively used in object recognition research, to measure the productivity of the suggested approach. The findings from the experiment support the suggested system&#x2019;s validity, as it achieved class recognition accuracies of 89&#x0025;, 83&#x0025;, and 90.30&#x0025; on the MSRC, MS COCO, and Caltech 101 datasets, respectively.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>K-region fusion</kwd>
<kwd>segmentation</kwd>
<kwd>recognition</kwd>
<kwd>feature extraction</kwd>
<kwd>artificial neural network</kwd>
<kwd>computer vision</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>MSIT (Ministry of Science and ICT), Korea, under the ITRC (Information Technology Research Center) Support Program</funding-source>
<award-id>IITP-2023-2018-0-01426</award-id>
</award-group>
<award-group id="awg2">
<funding-source>Princess Nourah bint Abdulrahman University</funding-source>
<award-id>PNURSP2023R410</award-id>
</award-group>
<award-group id="awg3">
<funding-source>Deanship of Scientific Research at Najran University</funding-source>
<award-id>NU/RG/SERC/12/6</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1"><label>1</label><title>Introduction</title>
<p>Segmenting and recognizing objects of interest in images and videos can be vital for various applications, including video and security surveillance [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-2">2</xref>] hyperspectral imaging [<xref ref-type="bibr" rid="ref-3">3</xref>], human detection [<xref ref-type="bibr" rid="ref-4">4</xref>], video streaming [<xref ref-type="bibr" rid="ref-5">5</xref>], emotion recognition [<xref ref-type="bibr" rid="ref-6">6</xref>], and traffic flow prediction [<xref ref-type="bibr" rid="ref-7">7</xref>], medical field [<xref ref-type="bibr" rid="ref-8">8</xref>], Self-driving cars [<xref ref-type="bibr" rid="ref-9">9</xref>], etc. This makes posture recognition [<xref ref-type="bibr" rid="ref-10">10</xref>] and scene understanding [<xref ref-type="bibr" rid="ref-11">11</xref>] sizzling issues in artificial intelligence (AI) and computer vision (CV). The purpose of the field is in the direction of teaching machines to understand (recognize) the content of images in the same way as humans. The extent of this research is constrained to object segmentation and recognition. While there has been significant progress in object detection and segmentation techniques, more can yet be done. Researchers are driven to create algorithms that are more precise, reliable, and scalable than those used today to overcome the drawbacks and difficulties of existing methods. To produce innovative findings and receive respect from the academic community, researchers in the fields of object detection and segmentation compete to surpass one another in benchmarks and challenges. The outcome of this research could influence a variety of areas and lead to better monitoring and security systems, safer autonomous vehicles, and increased industrial automation. To comprehend visual scenes, acquire data, and come to conclusions, segmentation sets object borders while object detection identifies specific objects in images or videos.</p>
<p>We outline a thorough approach to accurate object recognition in tropical settings. Through a series of carefully thought-out steps, our technique ensures accurate and successful object identification. To decrease noise while keeping important edge features, we first preprocess all RGB images using spatial domain filtering. Then, using our special &#x201C;k-region fusion&#x201D; method, which combines region-based segmentation and k-means clustering, we execute image segmentation to extract necessary objects from the background. By creating more meaningful and coherent object portions, this fusion technique raises the quality of segmentation. Then, we extract features utilizing two distinct descriptors: SIFT (Scale Invariant Feature Transform) and ORB (Oriented FAST and Rotated BRIEF). Finally, object classification is accomplished using an artificial neural network. Large, publically accessible datasets were used in the experiment, including MSRC-v2, MSCOCO 2017, and Caltech 101. This work aims to address the issue of the need for improved object identification under difficult tropical environments. Current solutions sometimes lack the precision and durability required for critical applications like privacy and vehicle autonomy. The motivations behind this include the industry&#x2019;s competitiveness, its potential for widespread influence, and the need to enhance critical performance metrics. Over time, these elements will contribute to the development of safer and more efficient systems in several fields.</p>
<p>Our research&#x2019;s primary insights can be summed up as follows:
<list list-type="bullet">
<list-item><p>The incorporation of a spatial domain filter during pre-processing successfully reduces noise in the images while maintaining crucial edge information, producing improved segmentation results.</p></list-item>
<list-item><p>The development of &#x201C;<bold>k-region fusion</bold>&#x201D; illustrates the potency of merging region-based segmentation with k-means clustering as a strategy for segmenting images.</p></list-item>
<list-item><p>Our robust object recognition approach is demonstrated by our high-performance classification system, which is driven by an Artificial Neural Network and features that originate from ORB and SIFT descriptors.</p></list-item>
<list-item><p>We have substantially elevated the precision, sensitivity, F1 score, and mean accuracy performance measures for object recognition when compared to prior approaches.</p></list-item>
<list-item><p>In the experimental results, the suggested model&#x2019;s impact has been confirmed across three publically accessible datasets, displaying exceptional performance.</p></list-item>
</list></p>
<p>The remaining part of this article is structured into several units that provide a comprehensive overview of the proposed system, its approach, and experimental results. In <xref ref-type="sec" rid="s2">Section 2</xref>, we discuss and analyze relevant investigated work related to the presented system, providing a comprehensive review of the existing literature. <xref ref-type="sec" rid="s2_1">Section 3</xref> defends the whole methodology of our system, which includes a general pre-classification procedure. <xref ref-type="sec" rid="s3">Section 4</xref> examines the datasets used in our recommended approach and demonstrates the structure&#x2019;s strength over various tests. Finally, in <xref ref-type="sec" rid="s4">Section 5</xref>, we summarise our significant results and contributions to the research. Overall, this work presents a thorough description of our proposed approach and its potential consequences for the research field.</p>
</sec>
<sec id="s2"><label>2</label><title>Related Work</title>
<p>Object detection and recognition have been progressively developed by researchers for several years [<xref ref-type="bibr" rid="ref-12">12</xref>]. They have investigated the possibility of complex images for anomaly detection [<xref ref-type="bibr" rid="ref-13">13</xref>], videos in depth and RGB&#x2009;&#x002B;&#x2009;D [<xref ref-type="bibr" rid="ref-14">14</xref>] (Red, Green, Blue, Depth) films to improve the effectiveness of their processes in addition to ordinary RGB images. The complete image is typically utilized as feedback, and characteristics are dug out from it, which is a modest and effective method of identifying multiple objects in a single image. Segmentation is an essential first step in many techniques, including mine. This involves dividing the image into distinctive, meaningful regions with identical characteristics based on image component incoherence or similarities. The performance of succeeding processes is strongly dependent on the accuracy of segmentation findings. Furthermore, tools for segmentation and classification for feature detection method [<xref ref-type="bibr" rid="ref-15">15</xref>]. In the proposed approach, objects are generated through an image segmentation process, wherein pixels with similar spectral characteristics are grouped to form a segment. Neural network implementation for object recognition then leads to more precise results. As a result, associated work can be classified into object recognition using Red, Green, Blue (RGB), and depth images.</p>
<sec id="s2_1"><label>2.1</label><title>Object Detection and Recognition over RGB Images</title>
<p>In the past, image-based approaches were frequently exploited. Chaturvedi et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] used a combination of different classifiers with the Viola-Jones algorithm and the You Look Only Once version 3 (YOLOv3) algorithm for object recognition. The focus of their research was on selecting an algorithm that provides us with a good balance of accuracy and efficiency. Li et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] introduced an object identification system centered around the best Bag of Words model and Area of Interest (AOI). They estimated the Region of Interest (ROI) using a saliency map and a Shi-Thomasi corner. To recognize and classify objects, they used Scale Invariant Feature Transform (SIFT) feature descriptors, a visual codebook, a Gaussian Mixture Model (GMM), and a Support Vector Machine (SVM). Deshmukh et al. [<xref ref-type="bibr" rid="ref-18">18</xref>] employed a novel processing strategy to find objects in original images by merging the object detection API, a mixture of identified edges, and an edge detection algorithm.</p>
</sec>
<sec id="s2_2"><label>2.2</label><title>Object Detection and Recognition over Depth Images</title>
<p>Many researchers have engaged in identifying the objects of interest in an image over the last couple of years. As depth images are insensitive to lighting variations and intrinsically integrate 3D data, numerous intensity-based detection methods have been suggested. Lin et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] presented a strong and reliable system for object identification in their research article. Their approach includes applying probabilistic image segmentation to remove the background from images. Cupec et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] developed a method for recognizing fruits using depth image analysis. This strategy builds a group of triangles from depth images using Delaunay triangulation. Delaunay triangulation is a geometry technique aimed at generating a standard triangular grid from a given point collection. Then, using a region growth mechanism, convex surfaces were formed by joining triangles, each of which represented a possible fruit. Ahmed et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] recommended a different approach by a using Histogram of Oriented Gradients (HOG) to excerpt features and detect an object by applying Nearest Neighbor search (NNS). Finally, they used the Hough voting algorithm to recognize the objects.</p>
</sec>
</sec>
<sec id="s3"><label>3</label><title>Suggested Method</title>
<p>This section outlines the proposed object recognition context. The schematic architecture is portrayed in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. Each RGB image was pre-processed. To smooth the images, the spatial domain filter is utilized. The images were then fragmented into foreground and background using the k-region fusion technique. The desired objects are present in the resulting foreground. For feature extraction, two kinds of features Oriented Fast and Robust Brief (ORB) and Scale Invariant Feature Transform (SIFT) were fused after extraction. Finally, the requested objects are classified using an artificial neural network. The subsections that follow describe each phase of the framework .</p>
<fig id="fig-1"><label>Figure 1</label><caption><title>An overall description of the suggested system</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_42963-fig-1.tif"/></fig>
<sec id="s3_1"><label>3.1</label><title>Image Pre-Processing General</title>
<p>As part of the pre-processing, all RGB images in both datasets have undergone a filtering approach. <xref ref-type="fig" rid="fig-2">Fig. 2</xref> illustrates the outcome of the filtered images, demonstrating that pre-processing at this stage increases the entirety of the system&#x2019;s efficiency. Furthermore, the image normalization and median filter employed in pre-processing are explored in further detail in the subsection that follows.</p>
<fig id="fig-2"><label>Figure 2</label><caption><title>Original images (a) MS COCO (b) MSRC-v2</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_42963-fig-2.tif"/></fig>
<sec id="s3_1_1"><label>3.1.1</label><title>Image Normalization</title>
<p>Normalization of an image is the process of altering the intensity values of pixels within an image to increase the image&#x2019;s contrast. Datasets with the initial images are gathered underneath various conditions during pre-processing, such as radiance variations and dispersion of contrast [<xref ref-type="bibr" rid="ref-22">22</xref>], yielding more objects, greater intensity values, and different object scales in the images. To eliminate this unwanted information, we initially minimized the resolution to 213&#x2009;&#x00D7;&#x2009;213 by using fixed window resizing. <xref ref-type="fig" rid="fig-2">Fig. 2</xref> represents normalized images.</p>
</sec>
<sec id="s3_1_2"><label>3.1.2</label><title>Noise Removal</title>
<p>A median filter has been employed to enhance image quality as well as minimize noise. A median filter smoothed the images while keeping all of the objects&#x2019; edges [<xref ref-type="bibr" rid="ref-23">23</xref>]. The median filter is a nonlinear digital filter that is applied to eradicate distortion from an image or signal. This form of noise reduction is a popular pre-processing approach for improving the results of subsequent processing (such as edge identification on an image) [<xref ref-type="bibr" rid="ref-24">24</xref>]. This approach works by replacing every pixel value with the median value calculated from the adjacent pixels. <xref ref-type="disp-formula" rid="eqn-1">Eqs. (1)</xref> and <xref ref-type="disp-formula" rid="eqn-2">(2)</xref> can be used to define the smooth image obtained after applying the median filter. <xref ref-type="fig" rid="fig-3">Fig. 3</xref> is showing the pre-processed images of some classes from the mentioned datasets.
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mi>J</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>k</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext mathvariant="italic">sorted</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext mathvariant="italic">flatten</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>J</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>k</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mi>J</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext mathvariant="italic">filtered</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>M</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mrow><mml:mo>(</mml:mo><mml:mi>J</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>k</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represent the submatrix centered at <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>J</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>k</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext mathvariant="italic">sorted</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mtext>&#x00A0;</mml:mtext></mml:math></inline-formula> represent the sorted vector of pixel values achieved by flattening the submatrix <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mrow><mml:mo>(</mml:mo><mml:mi>J</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>k</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>J</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext mathvariant="italic">filtered</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the output filtered image and <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>M</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the median value of the sorted vector.</p>
<fig id="fig-3"><label>Figure 3</label><caption><title>Outcomes after pre-processing (a) MS COCO (b) MSRC-v2</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_42963-fig-3.tif"/></fig>
</sec>
</sec>
<sec id="s3_2"><label>3.2</label><title>K-Mean Clustering</title>
<p>K-means clustering is a standard unsupervised machine learning technique for gathering information based on similarities. It can be used on a variety of data types, including images. K-means clustering can be used in image processing to group similar pixels together to simplify the image or extract useful information from it. K-means clustering can be used in image segmentation to group similar pixels together based on their color or texture, and thus separate regions of the image that have different color or texture characteristics [<xref ref-type="bibr" rid="ref-24">24</xref>]. To cluster homogeneous color regions, the k-mean algorithm is used, and it only requires the number of clusters k at the start, with no other prior knowledge required [<xref ref-type="bibr" rid="ref-25">25</xref>,<xref ref-type="bibr" rid="ref-26">26</xref>]. K-means clustering uses Euclidean distance as in <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref> to find similarities between objects.
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mi>D</mml:mi><mml:mo>=</mml:mo><mml:msqrt><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:msqrt></mml:math></disp-formula></p>
<p>Assigning random centroids to clusters and updating them based on the mean of the objects in the cluster until convergence are the steps involved in K-means clustering [<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-28">28</xref>]. <xref ref-type="fig" rid="fig-4">Fig. 4</xref> is showing a combined resultant flowchart after preprocessing and clustering on incorporated datasets.</p>
<fig id="fig-4"><label>Figure 4</label><caption><title>Combined flowchart for preprocessing and clustering</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_42963-fig-4.tif"/></fig>
</sec>
<sec id="s3_3"><label>3.3</label><title>Object Segmentation</title>
<p>Image segmentation was carried out after pre-processing the images. The goal of region-based segmentation is to create a set of homogeneous regions based on these criteria [<xref ref-type="bibr" rid="ref-16">16</xref>]. In this article, in order to obtain better segmentation results, we initially apply k-region fusion (clustering using K-means to an image and then performing region-based segmentation on the generated clusters). Huda et al. suggested a method for region-merging segmentation [<xref ref-type="bibr" rid="ref-29">29</xref>]. K-means clustering can be used to cluster similar pixels together based on their color or texture reducing image complexity and improving the effectiveness of region-based segmentation [<xref ref-type="bibr" rid="ref-30">30</xref>]. We can identify regions with similar color or texture characteristics by grouping similar pixels into clusters, which can then be used as inputs for region-based segmentation (see <xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref>).
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mrow><mml:mo>|</mml:mo><mml:mi>I</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mi>a</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>I</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mi>a</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mo>&#x003C;</mml:mo><mml:mrow><mml:mtext mathvariant="italic">Thresh</mml:mtext></mml:mrow></mml:math></disp-formula></p>
<p>The similarities between neighboring pixels (i, j) are ascertained using region-based segmentation. Pixels that have similar properties will form a unique region [<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-20">20</xref>]. Ahmed et al. also used regions to detect objects rather than the traditional sliding window method [<xref ref-type="bibr" rid="ref-21">21</xref>]. Adjacent pixels in an image are compared to the region&#x2019;s reference intensity values at each pixel [<xref ref-type="bibr" rid="ref-22">22</xref>]. The adjacent pixel is chosen if the difference is less than or equal to the difference threshold. <xref ref-type="fig" rid="fig-5">Fig. 5</xref> displays the resultant clustered segmented images.</p>
<fig id="fig-5"><label>Figure 5</label><caption><title>Resultant segmentation of k-region fusion</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_42963-fig-5.tif"/></fig>
</sec>
<sec id="s3_4"><label>3.4</label><title>Feature Extraction</title>
<p>In this article, we incorporate the use of both Oriented Fast and Robust Brief (ORB) features and Scale Invariant Feature Transform (SIFT) features. The following subsections define the specifics and outcomes of the mentioned features. The process of identifying and extracting useful information, or features, from an image for further analysis or processing is known as feature extraction on an image. Feature extraction aims to reduce the number of features in a dataset by creating new features from the existing ones (and subsequently removing the original features). This new, more condensed collection of features should then serve as a representation of the vast majority of the details in the original set of features. The extracted features should capture important aspects of the image, such as texture, color, shape, or edges [<xref ref-type="bibr" rid="ref-31">31</xref>].</p>
<sec id="s3_4_1"><label>3.4.1</label><title>Oriented Fast and Robust Brief</title>
<p>The Oriented Fast and Robust Brief (ORB) [<xref ref-type="bibr" rid="ref-32">32</xref>] is a high-performance feature detector that combines the FAST (Features from Accelerated Segment Test) keypoint detector&#x2019;s orientation and rotational resilience with the description of visual appearance BRIEF stands for &#x201C;Binary Robust Independent Elementary Features&#x201D; [<xref ref-type="bibr" rid="ref-33">33</xref>]. It detects key characteristics efficiently and provides a quick and reliable solution for feature extraction in computer vision applications. Location of key points determined by <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref>.
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>u</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mi>u</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup><mml:msup><mml:mi>v</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msup><mml:mi>I</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>u</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where i and j are the pixel intensity analyzed and determined at a and b, respectively. &#x201C;u&#x201D; and &#x201C;v&#x201D; is near the FAST feature point, a circle with a radius &#x201C;r&#x201D; of the vicinity u, v &#x025B; [&#x2212;r, r]. Then figure out the center of mass, as presented in <xref ref-type="disp-formula" rid="eqn-6">Eq. (6)</xref> which is also referred to as its &#x201C;center of mass&#x201D;.
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mrow><mml:mtext mathvariant="italic">Centroid</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mn>10</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mn>00</mml:mn></mml:mrow></mml:msub></mml:mfrac><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mfrac><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mn>01</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mn>00</mml:mn></mml:mrow></mml:msub></mml:mfrac></mml:math></disp-formula></p>
<p>We will calculate binary descriptors from BRIEF using <xref ref-type="disp-formula" rid="eqn-7">Eqs. (7)</xref> and <xref ref-type="disp-formula" rid="eqn-8">(8)</xref>.
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mi>&#x03C4;</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mi>I</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mo>:</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>i</mml:mi><mml:mi>f</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>I</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x003C;</mml:mo><mml:mi>I</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn><mml:mtext>&#x00A0;</mml:mtext><mml:mo>,</mml:mo><mml:mrow><mml:mtext mathvariant="italic">otherwise</mml:mtext></mml:mrow></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /><mml:mspace width="thinmathspace" /></mml:math></disp-formula>
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>I</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>i</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup><mml:mtext>&#x00A0;</mml:mtext><mml:mi>&#x03C4;</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mi>I</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mo>:</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <italic>I</italic> is an image and (<italic>a, b</italic>) are the pixel values. We can figure out the patch&#x2019;s orientation and generate a vector from the center (<xref ref-type="disp-formula" rid="eqn-9">Eq. (9)</xref>) of the corner to the centroid.
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mi>&#x03B8;</mml:mi><mml:mo>=</mml:mo><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mfrac><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mn>01</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>z</mml:mi><mml:mrow><mml:mn>10</mml:mn></mml:mrow></mml:msub></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mtext>atan</mml:mtext></mml:mrow><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>u</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi></mml:mrow></mml:munder><mml:mi>v</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>I</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>u</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>u</mml:mi><mml:mo>,</mml:mo><mml:mi>v</mml:mi></mml:mrow></mml:munder><mml:mi>u</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mi>I</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>u</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mo>,</mml:mo><mml:mi>v</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p><xref ref-type="fig" rid="fig-6">Fig. 6</xref> depicts the retrieved characteristics using ORB.</p>
<fig id="fig-6"><label>Figure 6</label><caption><title>Features extracted using ORB on images from both datasets</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_42963-fig-6.tif"/></fig>
</sec>
<sec id="s3_4_2"><label>3.4.2</label><title>Scale Invariant Feature Transform</title>
<p>To construct the set of image features, SIFT (see Algorithm 1) computed the following points, where H (a, b) is an input image, i and j are distances from points a and b [<xref ref-type="bibr" rid="ref-34">34</xref>], respectively, and is the Gaussian scale (see <xref ref-type="disp-formula" rid="eqn-10">Eq. (10)</xref>). After fitting a model to determine scale and location (<xref ref-type="disp-formula" rid="eqn-11">Eq. (11)</xref>), key points are chosen based on stability [<xref ref-type="bibr" rid="ref-35">35</xref>].
</p>
<fig id="fig-8">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_42963-fig-8.tif"/>
</fig>
<p>SIFT helps to regulate a direction for each key point to define a feature vector for each key [<xref ref-type="bibr" rid="ref-36">36</xref>].
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mi>G</mml:mi><mml:mi>u</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mn>2</mml:mn><mml:mi>&#x03C0;</mml:mi><mml:msup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mtext>&#x00A0;</mml:mtext><mml:mi>e</mml:mi><mml:mfrac><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>i</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mi>j</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mn>2</mml:mn><mml:msup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mi>Z</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>G</mml:mi><mml:mi>u</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi><mml:mi>&#x03C3;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>G</mml:mi><mml:mi>u</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi><mml:mo>,</mml:mo><mml:mi>&#x03C3;</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mi>H</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula>where H (a, b) is an input image, i and j are the distances from point a, b and is the scale of Gaussian, respectively. The SIFT technique (Algorithm 1) is useful for 3D reconstruction and object detection. It can withstand variations in illumination, rotation, and image scale. Each key point&#x2019;s direction is normalized by SIFT, creating a feature vector [<xref ref-type="bibr" rid="ref-37">37</xref>]. A key point has an orientation to maintain robustness against rotation variations and calculate gradient magnitude (<xref ref-type="disp-formula" rid="eqn-12">Eq. (12)</xref>) G.M (a, b) and gradient rotation (<xref ref-type="disp-formula" rid="eqn-13">Eq. (13)</xref>) G.R (a, b) [<xref ref-type="bibr" rid="ref-28">28</xref>] around collected key points.
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mi>G</mml:mi><mml:mo>.</mml:mo><mml:msub><mml:mi>M</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msqrt><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:msqrt></mml:math></disp-formula>
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mi>G</mml:mi><mml:mo>.</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mn>2</mml:mn><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>H</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p><xref ref-type="fig" rid="fig-7">Fig. 7</xref> represents the extracted features by SIFT.</p>
<fig id="fig-7"><label>Figure 7</label><caption><title>Features extracted using SIFT on images from both datasets</title></caption><graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_42963-fig-7.tif"/></fig>
</sec>
<sec id="s3_4_3"><label>3.4.3</label><title>Feature Fusion</title>
<p>While SIFT is dependable and invariant, handling scale, rotation, and lighting changes using histograms of gradient magnitudes and orientations, ORB uses binary strings for effective and condensed keypoint encoding. The binary character of ORB and the stability of SIFT are combined in feature fusion to provide a potent keypoint representation that captures crucial visual information including texture, edges, and shape. Using these combined attributes, operations like object detection and image matching are greatly enhanced. The concatenated feature vector F_combined is given by <xref ref-type="disp-formula" rid="eqn-14">Eq. (14)</xref>.
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mi>F</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext mathvariant="italic">combined</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi>F</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>S</mml:mi><mml:mi>I</mml:mi><mml:mi>F</mml:mi><mml:mi>T</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi>F</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>O</mml:mi><mml:mi>R</mml:mi><mml:mi>B</mml:mi><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>F</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>S</mml:mi><mml:mi>I</mml:mi><mml:mi>F</mml:mi><mml:mi>T</mml:mi></mml:math></inline-formula> is the feature vector obtained from the SIFT descriptor, with dimensionality D_SIFT and <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>F</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>O</mml:mi><mml:mi>R</mml:mi><mml:mi>B</mml:mi></mml:math></inline-formula> is the feature vector obtained from the ORB descriptor, with dimensionality D_ORB.</p>
<p>Here, [F_SIFT, F_ORB] denotes the association of the SIFT feature vector F_SIFT with the ORB feature vector F_ORB to create a single feature vector with a combined dimensionality of D_combined in <xref ref-type="disp-formula" rid="eqn-15">Eq. (15)</xref>.
<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:mi>D</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mrow><mml:mtext mathvariant="italic">combined</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi>D</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>S</mml:mi><mml:mi>I</mml:mi><mml:mi>F</mml:mi><mml:mi>T</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi>D</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>O</mml:mi><mml:mi>R</mml:mi><mml:mi>B</mml:mi><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula></p>
</sec>
</sec>
<sec id="s3_5"><label>3.5</label><title>Artificial Neural Network</title>
<p>Neurons in the input (X), hidden, and output (Y) layers make up Artificial Neural Networks (ANNs). They can be divided into three groups: feedback networks, multi-layer feedforward networks [<xref ref-type="bibr" rid="ref-38">38</xref>], and single-layer feedforward networks. In this study, artificial neural networks are used to process data using artificial neurons. A neuron&#x2019;s activation function, which is comparable to decision-making in the brain, is the result of inputs moving from one neuron to the next. Interconnected layers of nodes (neurons) make up ANNs, which receive inputs, process them, and then output the results [<xref ref-type="bibr" rid="ref-39">39</xref>]. Weights, which reflect neural connections and are modified throughout learning to improve task performance. In Algorithm 2, the functionality of ANN is described.
</p>
<fig id="fig-9">
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_42963-fig-9.tif"/>
</fig>
<p>The principle of neurons can be presented as <xref ref-type="disp-formula" rid="eqn-16">Eqs. (16)</xref> and <xref ref-type="disp-formula" rid="eqn-17">(17)</xref>.
<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:msubsup><mml:mrow><mml:mo>&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mtext>&#x00A0;</mml:mtext></mml:mrow></mml:msub><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi><mml:mo>=</mml:mo></mml:mrow></mml:msub><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mtext>&#x00A0;</mml:mtext><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula>
<disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mo>&#x2211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula>where &#x201C;<inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>&#x201D; is the weighted sum for the kth neuron, &#x201C;n&#x201D; is the input features. &#x201C;<inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>&#x201D;: weight connecting the ith input to the kth neuron, &#x201C;<inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mrow><mml:mtext>a</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>&#x201D; value of ith input feature, &#x201C;<inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mi>b</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>&#x201D; is the bias term. Output &#x201C;<inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>&#x201D; of a neuron is defined in <xref ref-type="disp-formula" rid="eqn-18">Eq. (18)</xref>.
<disp-formula id="eqn-18"><label>(18)</label><mml:math id="mml-eqn-18" display="block"><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>z</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:math></disp-formula>where &#x201C;<inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mi>o</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>&#x201D;: the output of the kth neuron, and &#x201C;e&#x201D; is Euler&#x2019;s number (approx. 2.71).</p>
</sec>
</sec>
<sec id="s4"><label>4</label><title>Evaluation Metrics</title>
<p>The study assesses the performance of the suggested system using three datasets for object recognition. It contrasts the proposed system with current object recognition technologies.</p>
<sec id="s4_1"><label>4.1</label><title>Dataset Overview</title>
<sec id="s4_1_1"><label>4.1.1</label><title>Microsoft Common Objects in Context (MS COCO)</title>
<p>To segment and identify objects in images, Microsoft developed the MS COCO dataset [<xref ref-type="bibr" rid="ref-40">40</xref>]. In 330,000 images, there are 2.5 million object instances organized into 80 classes, including zebras, bears, and other typical objects. With significant occurrences of each of the categories, the dataset is utilized to evaluate the recognition tasks.</p>
</sec>
<sec id="s4_1_2"><label>4.1.2</label><title>Microsoft Research in Cambridge (MSRC-v2)</title>
<p>The MSRC-v2 dataset entails 591 high-resolution images [<xref ref-type="bibr" rid="ref-41">41</xref>], 21 different object categories (cow, sheep, grass, tree, horse, car, bicycle, plane, face) as well as one backdrop category [<xref ref-type="bibr" rid="ref-42">42</xref>] and ten distinct classes. Each 213&#x2009;&#x00D7;&#x2009;320-pixel image has a different color scheme and context. Due to the complex backgrounds and illumination, the dataset is demanding.</p>
</sec>
<sec id="s4_1_3"><label>4.1.3</label><title>Caltech 101</title>
<p>The images included in the Caltech 101 dataset have several categories and are divided into object and background categories. The resolution of individual image is roughly 300&#x2009;&#x00D7;&#x2009;200 pixels. Numerous object classes, such as the camera, barrel, cup, bike, panda, chair, rhino, airplane, tree, and water are included in the Caltech 101 dataset.</p>
</sec>
</sec>
<sec id="s4_2"><label>4.2</label><title>Experimentations and Results</title>
<p>Python (3.7) has been employed for training and evaluating the system on an Intel Core i7 PC running 64-bit Windows 10. The machine has 16 GB of RAM (random access memory) and a 5 (GHz) CPU.</p>
<sec id="s4_2_1"><label>4.2.1</label><title>Experimental I: Class Recognition Accuracy</title>
<p>The classification accuracies of the employed datasets by ten arbitrarily selected classes are shown in <xref ref-type="table" rid="table-1">Table 1</xref> as a confusion matrix. for MSRC-v2, <xref ref-type="table" rid="table-2">Table 2</xref> as a confusion matrix for MS COCO, and <xref ref-type="table" rid="table-3">Table 3</xref> as a confusion matrix for Caltech 101.</p>
<table-wrap id="table-1"><label>Table 1</label><caption><title>Recognition accuracy confusion matrix over MSRC-v2 using ANN</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Object classes</th>
<th align="left">House</th>
<th align="left">Cow</th>
<th align="left">Horse</th>
<th align="left">Sheep</th>
<th align="left">Tree</th>
<th align="left">Car</th>
<th align="left">Plane</th>
<th align="left">Face</th>
<th align="left">Duck</th>
<th align="left">Bird</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">House</td>
<td align="left"><bold>0.89</bold></td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Cow</td>
<td align="left">0.0</td>
<td align="left"><bold>0.91</bold></td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Horse</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left"><bold>0.92</bold></td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Sheep</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left"><bold>0.85</bold></td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Tree</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left"><bold>0.90</bold></td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.02</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Car</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.01</td>
<td align="left"><bold>0.90</bold></td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Plane</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left"><bold>0.89</bold></td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Face</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.03</td>
<td align="left">0.0</td>
<td align="left"><bold>0.92</bold></td>
<td align="left">0.0</td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Duck</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.01</td>
<td align="left">0.02</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.01</td>
<td align="left"><bold>0.88</bold></td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Bird</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.01</td>
<td align="left"><bold>0.92</bold></td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-2"><label>Table 2</label><caption><title>Recognition accuracy confusion matrix over MS COCO using ANN</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Object classes</th>
<th align="left">Person</th>
<th align="left">Bicycle</th>
<th align="left">Car</th>
<th align="left">Motor cycle</th>
<th align="left">Air plane</th>
<th align="left">Bus</th>
<th align="left">Truck</th>
<th align="left">Train</th>
<th align="left">Boat</th>
<th align="left">Traffic light</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Person</td>
<td align="left"><bold>0.91</bold></td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Bicycle</td>
<td align="left">0.0</td>
<td align="left"><bold>0.83</bold></td>
<td align="left">0.01</td>
<td align="left">0.03</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Car</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left"><bold>0.82</bold></td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.01</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Motorcycle</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.30</td>
<td align="left"><bold>0.86</bold></td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Airplane</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left"><bold>0.81</bold></td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.5</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Bus</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.01</td>
<td align="left"><bold>0.85</bold></td>
<td align="left">0.1</td>
<td align="left">0.0</td>
<td align="left">0.01</td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Truck</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.02</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.01</td>
<td align="left"><bold>0.88</bold></td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Train</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.03</td>
<td align="left">0.0</td>
<td align="left"><bold>0.83</bold></td>
<td align="left">0.0</td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Boat</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.01</td>
<td align="left"><bold>0.88</bold></td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Traffic light</td>
<td align="left">0.05</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left"><bold>0.81</bold></td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-3"><label>Table 3</label><caption><title>Recognition accuracy confusion matrix over Caltech 101 using ANN</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Object classes</th>
<th align="left">Camera</th>
<th align="left">Barrel</th>
<th align="left">Cup</th>
<th align="left">Bike</th>
<th align="left">Panda</th>
<th align="left">Chair</th>
<th align="left">Rhino</th>
<th align="left">Airplane</th>
<th align="left">Tree</th>
<th align="left">Water</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Camera</td>
<td align="left"><bold>0.89</bold></td>
<td align="left">0.07</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Barrel</td>
<td align="left">0.0</td>
<td align="left"><bold>0.93</bold></td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.01</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Cup</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left"><bold>0.92</bold></td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Bike</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left"><bold>0.85</bold></td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Panda</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left"><bold>0.90</bold></td>
<td align="left">0.0</td>
<td align="left">0.05</td>
<td align="left">0.02</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Chair</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.01</td>
<td align="left"><bold>0.93</bold></td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Rhino</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.05</td>
<td align="left">0.01</td>
<td align="left"><bold>0.89</bold></td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Airplane</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.03</td>
<td align="left">0.0</td>
<td align="left"><bold>0.92</bold></td>
<td align="left">0.0</td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Tree</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.01</td>
<td align="left">0.02</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.01</td>
<td align="left"><bold>0.88</bold></td>
<td align="left">0.0</td>
</tr>
<tr>
<td align="left">Water</td>
<td align="left">0.0</td>
<td align="left">0.01</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.0</td>
<td align="left">0.00</td>
<td align="left"><bold>0.92</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2_2"><label>4.2.2</label><title>Experimental II: Precision, Sensitivity, and F1 Measure</title>
<p>We report the precision, recall, and F measures for ten randomly selected classes from the datasets in this section. The results demonstrate that the presented recognition system is highly precise at identifying a variety of complicated objects. <xref ref-type="disp-formula" rid="eqn-19">Eqs. (19)</xref>&#x2013;<xref ref-type="disp-formula" rid="eqn-21">(21)</xref> were used to figure out the precision, recall, and F1 scores for each object class in the datasets.
<disp-formula id="eqn-19"><label>(19)</label><mml:math id="mml-eqn-19" display="block"><mml:mrow><mml:mtext mathvariant="italic">Precision</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-20"><label>(20)</label><mml:math id="mml-eqn-20" display="block"><mml:mrow><mml:mtext mathvariant="italic">Recall</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="italic">Sensitivity</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-21"><label>(21)</label><mml:math id="mml-eqn-21" display="block"><mml:mi>F</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:mtext mathvariant="italic">measure</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mtext mathvariant="italic">Precision</mml:mtext></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mrow><mml:mtext mathvariant="italic">Recall</mml:mtext></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext mathvariant="italic">Precision</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mtext mathvariant="italic">Recall</mml:mtext></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p><xref ref-type="table" rid="table-4 table-5 table-6">Tables 4&#x2013;6</xref> indicate the precision, sensitivity, and F measure using ANN for all datasets, i.e., MSRC-v2, MS COCO, and Caltech 101, respectively, as TP stands for True positive, FP stands for False positive, and FN is for False negative.</p>
<table-wrap id="table-4"><label>Table 4</label><caption><title>Precision, recall, F1 measure and computation time for MSRC-v2 dataset</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="center" colspan="5">MSRC-v2</th>
</tr>
<tr>
<th align="left">Classes</th>
<th align="left">Precision</th>
<th align="left">Recall</th>
<th align="left">F1 measure</th>
<th align="left">Computing time</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">House</td>
<td align="left">0.54</td>
<td align="left">0.40</td>
<td align="left">0.93</td>
<td align="left">90.7</td>
</tr>
<tr>
<td align="left">Cow</td>
<td align="left">0.50</td>
<td align="left">0.54</td>
<td align="left">0.87</td>
<td align="left">75.5</td>
</tr>
<tr>
<td align="left">Horse</td>
<td align="left">0.35</td>
<td align="left">0.66</td>
<td align="left">0.75</td>
<td align="left">96.1</td>
</tr>
<tr>
<td align="left">Sheep</td>
<td align="left">0.65</td>
<td align="left">0.80</td>
<td align="left">0.45</td>
<td align="left">71.2</td>
</tr>
<tr>
<td align="left">Tree</td>
<td align="left">0.85</td>
<td align="left">0.70</td>
<td align="left">0.98</td>
<td align="left">73.5</td>
</tr>
<tr>
<td align="left">Car</td>
<td align="left">0.50</td>
<td align="left">0.79</td>
<td align="left">0.35</td>
<td align="left">92.2</td>
</tr>
<tr>
<td align="left">Airplane</td>
<td align="left">0.35</td>
<td align="left">0.39</td>
<td align="left">0.93</td>
<td align="left">81.8</td>
</tr>
<tr>
<td align="left">Face</td>
<td align="left">0.45</td>
<td align="left">0.22</td>
<td align="left">0.87</td>
<td align="left">81.5</td>
</tr>
<tr>
<td align="left">Duck</td>
<td align="left">0.17</td>
<td align="left">0.44</td>
<td align="left">0.75</td>
<td align="left">98.3</td>
</tr>
<tr>
<td align="left">Bird</td>
<td align="left">1.00</td>
<td align="left">0.00</td>
<td align="left">0.45</td>
<td align="left">78.2</td>
</tr>
<tr>
<td align="left">MEAN</td>
<td align="left">0.86</td>
<td align="left">0.83</td>
<td align="left">0.89</td>
<td align="left">83.90&#x2005;s</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-5"><label>Table 5</label><caption><title>Precision, recall, F1 measure and computation time for MS COCO dataset</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="center" colspan="5">MS COCO</th>
</tr>
<tr>
<th align="left">Classes</th>
<th align="left">Precision</th>
<th align="left">Recall</th>
<th align="left">F1 measure</th>
<th align="left">Computing time</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Person</td>
<td align="left">0.93</td>
<td align="left">0.54</td>
<td align="left">0.54</td>
<td align="left">131.2</td>
</tr>
<tr>
<td align="left">Bicycle</td>
<td align="left">0.87</td>
<td align="left">0.50</td>
<td align="left">0.50</td>
<td align="left">114.2</td>
</tr>
<tr>
<td align="left">Car</td>
<td align="left">0.75</td>
<td align="left">0.35</td>
<td align="left">0.35</td>
<td align="left">188.9</td>
</tr>
<tr>
<td align="left">Motorcycle</td>
<td align="left">0.45</td>
<td align="left">0.54</td>
<td align="left">0.77</td>
<td align="left">170.3</td>
</tr>
<tr>
<td align="left">Airplane</td>
<td align="left">0.98</td>
<td align="left">0.01</td>
<td align="left">0.79</td>
<td align="left">105.9</td>
</tr>
<tr>
<td align="left">Bus</td>
<td align="left">0.35</td>
<td align="left">0.50</td>
<td align="left">0.80</td>
<td align="left">156.3</td>
</tr>
<tr>
<td align="left">Train</td>
<td align="left">0.93</td>
<td align="left">0.35</td>
<td align="left">0.35</td>
<td align="left">199.2</td>
</tr>
<tr>
<td align="left">Truck</td>
<td align="left">0.87</td>
<td align="left">0.45</td>
<td align="left">0.45</td>
<td align="left">157.0</td>
</tr>
<tr>
<td align="left">Boat</td>
<td align="left">0.75</td>
<td align="left">0.54</td>
<td align="left">0.17</td>
<td align="left">162.7</td>
</tr>
<tr>
<td align="left">Traffic light</td>
<td align="left">0.45</td>
<td align="left">0.50</td>
<td align="left">0.97</td>
<td align="left">113.2</td>
</tr>
<tr>
<td align="left">Mean</td>
<td align="left">0.80</td>
<td align="left">0.82</td>
<td align="left">0.87</td>
<td align="left">149.89&#x2005;s</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-6"><label>Table 6</label><caption><title>Precision, recall, F1 measure and computation time for Caltech 101 dataset</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="center" colspan="5">Caltech 101</th>
</tr>
<tr>
<th align="left">Classes</th>
<th align="left">Precision</th>
<th align="left">Recall</th>
<th align="left">F1 measure</th>
<th align="left">Computing time</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Camera</td>
<td align="left">0.78</td>
<td align="left">0.73</td>
<td align="left"><bold>0.75</bold></td>
<td align="left">112.0</td>
</tr>
<tr>
<td align="left">Barrel</td>
<td align="left">0.83</td>
<td align="left">0.75</td>
<td align="left"><bold>0.80</bold></td>
<td align="left">96.5</td>
</tr>
<tr>
<td align="left">Cup</td>
<td align="left">0.81</td>
<td align="left">0.79</td>
<td align="left"><bold>0.78</bold></td>
<td align="left">171.0</td>
</tr>
<tr>
<td align="left">Bike</td>
<td align="left">0.83</td>
<td align="left">0.76</td>
<td align="left"><bold>0.77</bold></td>
<td align="left">150.2</td>
</tr>
<tr>
<td align="left">Panda</td>
<td align="left">0.85</td>
<td align="left">0.70</td>
<td align="left"><bold>0.73</bold></td>
<td align="left">94.1</td>
</tr>
<tr>
<td align="left">Chair</td>
<td align="left">0.73</td>
<td align="left">0.72</td>
<td align="left"><bold>0.70</bold></td>
<td align="left">133.2</td>
</tr>
<tr>
<td align="left">Rhino</td>
<td align="left">0.79</td>
<td align="left">0.77</td>
<td align="left"><bold>0.72</bold></td>
<td align="left">170.9</td>
</tr>
<tr>
<td align="left">Airplane</td>
<td align="left">0.85</td>
<td align="left">0.70</td>
<td align="left"><bold>0.75</bold></td>
<td align="left">131.2</td>
</tr>
<tr>
<td align="left">Tree</td>
<td align="left">0.82</td>
<td align="left">0.76</td>
<td align="left"><bold>0.79</bold></td>
<td align="left">135.0</td>
</tr>
<tr>
<td align="left">Water</td>
<td align="left">0.76</td>
<td align="left">0.71</td>
<td align="left"><bold>0.75</bold></td>
<td align="left">97.5</td>
</tr>
<tr>
<td align="left">Mean</td>
<td align="left">0.80</td>
<td align="left">0.82</td>
<td align="left"><bold>0.87</bold></td>
<td align="left">129.16&#x2005;s</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In this paper, the comparison has been made among different classifiers on three datasets, i.e., MS COCO, MSRC-v2, and Caltech 101. Artificial Neural Network gives us better results among all three. Results produced from all three classifiers have been shown below in <xref ref-type="table" rid="table-7">Table 7</xref>.</p>
<table-wrap id="table-7"><label>Table 7</label><caption><title>Comparison of ANN, RF, and Adaboost</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Classifiers</th>
<th align="left">MSRC-v2</th>
<th align="left">MS COCO</th>
<th align="left">Caltech101</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Adaboost</td>
<td align="left">81&#x0025;</td>
<td align="left">80&#x0025;</td>
<td align="left">80.5&#x0025;</td>
</tr>
<tr>
<td align="left">RF</td>
<td align="left">80&#x0025;</td>
<td align="left">75&#x0025;</td>
<td align="left">81.9&#x0025;</td>
</tr>
<tr>
<td align="left">ANN</td>
<td align="left">89&#x0025;</td>
<td align="left">83&#x0025;</td>
<td align="left">90.3&#x0025;</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Finally, <xref ref-type="table" rid="table-8 table-9 table-10">Tables 8&#x2013;10</xref> contrast the proposed system&#x2019;s functionality for object recognition with other state-of-the-art methodologies over the mentioned RGB object datasets.</p>
<table-wrap id="table-8"><label>Table 8</label><caption><title>A comparative analysis against contemporary methods over the MSRC dataset</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Authors</th>
<th align="left">Object recognition accuracy</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Ahmed et al. [<xref ref-type="bibr" rid="ref-42">42</xref>]</td>
<td align="left">86.1&#x0025;</td>
</tr>
<tr>
<td align="left">Ahmed et al. [<xref ref-type="bibr" rid="ref-43">43</xref>]</td>
<td align="left">86.0&#x0025;</td>
</tr>
<tr>
<td align="left">Bansal et al. [<xref ref-type="bibr" rid="ref-44">44</xref>]</td>
<td align="left">82.7&#x0025;</td>
</tr>
<tr>
<td align="left">Ours</td>
<td align="left">89&#x0025;</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-9"><label>Table 9</label><caption><title>A comparative analysis against contemporary methods over the MS COCO dataset</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Authors</th>
<th align="left">Object recognition accuracy</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Kong et al. [<xref ref-type="bibr" rid="ref-45">45</xref>]</td>
<td align="left">80.7&#x0025;</td>
</tr>
<tr>
<td align="left">Kim et al. [<xref ref-type="bibr" rid="ref-46">46</xref>]</td>
<td align="left">58.3&#x0025;</td>
</tr>
<tr>
<td align="left">Tan et al. [<xref ref-type="bibr" rid="ref-47">47</xref>]</td>
<td align="left">61.6&#x0025;</td>
</tr>
<tr>
<td align="left">Ours</td>
<td align="left">83.0&#x0025;</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-10"><label>Table 10</label><caption><title>A comparative analysis against contemporary methods over the Caltech 101 dataset</title></caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th align="left">Authors</th>
<th align="left">Object recognition accuracy</th>
</tr>
</thead>
<tbody>
<tr>
<td align="left">Rahmawati et al. [<xref ref-type="bibr" rid="ref-48">48</xref>]</td>
<td align="left">74.28&#x0025;</td>
</tr>
<tr>
<td align="left">Srivasti et al. [<xref ref-type="bibr" rid="ref-49">49</xref>]</td>
<td align="left">79.0&#x0025;</td>
</tr>
<tr>
<td align="left">Jalal et al. [<xref ref-type="bibr" rid="ref-50">50</xref>]</td>
<td align="left">89.26&#x0025;</td>
</tr>
<tr>
<td align="left">Ours</td>
<td align="left">90.30&#x0025;</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
</sec>
<sec id="s5"><label>5</label><title>Conclusion</title>
<p>This study presents a useful technique for identifying intricate real-world objects. RGB images are first normalized and median filtered, and then the targeted objects are segmented using K-means clustering and segmentation jointly called k-region fusion. Then, to extract important details from the segmented objects, ORB and SIFT are fused to extract the key points. Finally, object labeling and recognition are accomplished using an Artificial Neural Network (ANN). Comparative analyses against cutting-edge systems illustrate how better our suggested approach is, highlighting its outstanding performance on object recognition tasks. The proposed solution is intended to work with a variety of real-world applications such as security systems, the medical field, self-driving cars, assisted living, and online learning. Including depth, information improves object segmentation and identification. Depth adds a new dimension, which enhances spatial comprehension. It aids in separating objects at various distances, managing occlusion situations, and lessening the effect on recognition. Segmentation is streamlined by localization in 3D space. To provide an accurate representation, depth-based features are added to RGB data. It improves scene comprehension and resistance to changes in lighting. Effective for dealing with objects without textures. In general, depth integration enhances accuracy in challenging situations.</p>
</sec>
</body>
<back>
<ack>
<p>The authors are thankful to Princess Nourah bint Abdulrahman University Researchers Supporting Project, Princess Nourah bint Abdulrahman University, Riyadh, Saudi Arabia.</p>
</ack>
<sec><title>Funding Statement</title>
<p>This research was supported by the MSIT (Ministry of Science and ICT), Korea, under the ITRC (Information Technology Research Center) Support Program (IITP-2023-2018-0-01426) supervised by the IITP (Institute for Information &#x0026; Communications Technology Planning &#x0026; Evaluation). The funding for this work was provided by Princess Nourah bint Abdulrahman University Researchers Supporting Project Number (PNURSP2023R410), Princess Nourah bint Abdulrahman University, Riyadh, Saudi Arabia. The authors are thankful to the Deanship of Scientific Research at Najran University for funding this work under the Research Group Funding Program Grant Code (NU/RG/SERC/12/6).</p></sec>
<sec><title>Author Contributions</title>
<p>Study conception and design: Aysha Naseer, Jeongmin Park, data collection: Nouf Abdullah Almujally; analysis and interpretation of results: Aysha Naseer, Saud S. Alotaidi and Abdulwahab Alazeb; draft manuscript preparation: Aysha Naseer. All authors reviewed the results and approved the final version of the manuscript.</p></sec>
<sec sec-type="data-availability"><title>Availability of Data and Materials</title>
<p>All publicly available datasets are used in the study.</p></sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare that they have no conflicts of interest to report regarding the present study.</p></sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Song</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Tang</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Tan</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Ren</surname></string-name></person-group>, &#x201C;<article-title>ISHD: Intelligent standing human detection of video surveillance for the smart examination environment</article-title>,&#x201D; <source>Computer Modeling in Engineering &#x0026; Sciences</source>, vol. <volume>137</volume>, no. <issue>1</issue>, pp. <fpage>509</fpage>&#x2013;<lpage>526</lpage>, <year>2023</year>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Varma</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Sreeraj</surname></string-name></person-group>, &#x201C;<article-title>Object detection and classification in surveillance system</article-title>,&#x201D; <source>2013 IEEE Recent Advances in Intelligent Computational Systems (RAICS)</source>, <publisher-loc>Trivandrum, India</publisher-loc>, vol. <volume>2013</volume>, pp. <fpage>299</fpage>&#x2013;<lpage>303</lpage>, <year>2013</year>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Laghari, A.</given-names> <surname>Ali</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Yin</surname></string-name></person-group>, &#x201C;<article-title>How to collect and interpret medical pictures captured in highly challenging environments that range from nanoscale to hyperspectral imaging</article-title>,&#x201D; <source>Current Medical Imaging</source>, vol. <volume>54</volume>, pp. <fpage>36582065</fpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Zhou</surname></string-name> and <string-name><given-names>L.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>SA-FPN: An effective feature pyramid network for crowded human detection</article-title>,&#x201D; <source>Applied Intelligence</source>, vol. <volume>52</volume>, no. <issue>11</issue>, pp. <fpage>12556</fpage>&#x2013;<lpage>12568</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Laghari, A.</given-names> <surname>Ali</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Shahid</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Yadav</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Karim</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>The state of art and review on video streaming</article-title>,&#x201D; <source>Journal of High Speed Networks</source>, vol. <volume>22</volume>, no. <issue>3</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>26</lpage>, <year>2023</year>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Xia</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Self-training maximum classifier discrepancy for EEG emotion recognition</article-title>,&#x201D; <source>CAAI Transactions on Intelligence Technology</source>, vol. <volume>38</volume>, pp. <fpage>12174</fpage>, <year>2023</year>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Wen</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Yan</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Feng</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Xia</surname></string-name></person-group>, &#x201C;<article-title>A hybrid-convolution spatial&#x2013;temporal recurrent network for traffic flow prediction</article-title>,&#x201D; <source>The Computer Journal</source>, vol. <volume>76</volume>, pp. <fpage>171</fpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. L.</given-names> <surname>Giger</surname></string-name></person-group>, &#x201C;<article-title>Machine learning in medical imaging</article-title>,&#x201D; <source>Journal of the American College of Radiology</source>, vol. <volume>15</volume>, no. <issue>15</issue>, pp. <fpage>512</fpage>&#x2013;<lpage>520</lpage>, <year>2018</year>; <pub-id pub-id-type="pmid">29398494</pub-id></mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Park</surname></string-name> and <string-name><given-names>J. Y.</given-names> <surname>Choi</surname></string-name></person-group>, &#x201C;<article-title>Malware detection in self-driving vehicles using machine learning algorithms</article-title>,&#x201D; <source>Journal of Advanced Transportation</source>, vol. <volume>2020</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>9</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Jiang</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Hu</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Wang</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>A survey on artificial intelligence in posture recognition</article-title>,&#x201D; <source>Computer Modeling in Engineering &#x0026; Sciences</source>, vol. <volume>137</volume>, no. <issue>1</issue>, pp. <fpage>35</fpage>&#x2013;<lpage>82</lpage>, <year>2023</year>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R. P.</given-names> <surname>de Lima</surname></string-name> and <string-name><given-names>K.</given-names> <surname>Marfurt</surname></string-name></person-group>, &#x201C;<article-title>Convolutional neural network for remote-sensing scene classification: Transfer learning analysis</article-title>,&#x201D; <source>Remote Sensing</source>, vol. <volume>12</volume>, no. <issue>1</issue>, pp. <fpage>86</fpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Elhoseny</surname></string-name></person-group>, &#x201C;<article-title>Multi-object detection and tracking (MODT) machine learning model for real-time video surveillance systems</article-title>,&#x201D; <source>Circuits, Systems and Signal Processing</source>, vol. <volume>39</volume>, pp. <fpage>611</fpage>&#x2013;<lpage>630</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Hu</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Sun</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Liu</surname></string-name></person-group>, &#x201C;<article-title>Hyperspectral anomaly detection using ensemble and robust collaborative representation</article-title>,&#x201D; <source>Information Sciences</source>, vol. <volume>624</volume>, pp. <fpage>748</fpage>&#x2013;<lpage>760</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Wen</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Xu</surname></string-name> and <string-name><given-names>L. V.</given-names> <surname>Gool</surname></string-name></person-group>, &#x201C;<article-title>Visual recognition in RGB images and videos by learning from RGB-D data</article-title>,&#x201D; <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>, vol. <volume>40</volume>, no. <issue>8</issue>, pp. <fpage>2030</fpage>&#x2013;<lpage>2036</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Ross</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Donahue</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Darrell</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Malik</surname></string-name></person-group>, &#x201C;<article-title>Region-based convolutional networks for accurate object detection and segmentation</article-title>,&#x201D; <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>, vol. <volume>38</volume>, pp. <fpage>142</fpage>&#x2013;<lpage>158</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Chaturvedi</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Kaur</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Rakesh</surname></string-name> and <string-name><given-names>P.</given-names> <surname>Nand</surname></string-name></person-group>, &#x201C;<article-title>Object recognition using image segmentation</article-title>,&#x201D; in <conf-name>Proc of 6th Int. Conf. of Parallel, Distributed and Grid Computing</conf-name>, <conf-loc>Solan, India</conf-loc>, pp. <fpage>550</fpage>&#x2013;<lpage>556</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Li</surname></string-name> and <string-name><given-names>P.</given-names> <surname>Dong</surname></string-name></person-group>, &#x201C;<article-title>Object recognition based on the region of interest and optimal bag of words model</article-title>,&#x201D; <source>Neurocomputing</source>, vol. <volume>172</volume>, pp. <fpage>271</fpage>&#x2013;<lpage>280</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Deshmukh</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Moh</surname></string-name></person-group>, &#x201C;<article-title>Fine object detection in automated solar panel layout generation</article-title>,&#x201D; in <conf-name>Proc. of 17th IEEE Int. Conf. of Machine Learning and Applications</conf-name>, <conf-loc>Florida, USA</conf-loc>, pp. <fpage>1402</fpage>&#x2013;<lpage>1407</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Lin</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Tang</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Zou</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Xiong</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Fang</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Color-, depth-, and shape-based 3D fruit detection</article-title>,&#x201D; <source>Remote Sensing</source>, vol. <volume>21</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>17</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Cupec</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Filko</surname></string-name>, <string-name><given-names>I.</given-names> <surname>Vidovi&#x0107;</surname></string-name>, <string-name><given-names>E. K.</given-names> <surname>Nyarko</surname></string-name>, <string-name><given-names>&#x017D;.</given-names> <surname>Hocenski</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Point cloud segmentation to approximately convex surfaces for fruit recognition</article-title>,&#x201D; in <conf-name>Proc. of the Croatian Computer Vision Workshop</conf-name>, <conf-loc>Zargeb, Croatia</conf-loc>, pp. <fpage>56</fpage>&#x2013;<lpage>61</lpage>, <year>2014</year>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Ahmed</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Jalal</surname></string-name> and <string-name><given-names>K.</given-names> <surname>Kim</surname></string-name></person-group>, &#x201C;<article-title>RGB-D images for object segmentation, localization and recognition in indoor scenes using feature descriptor and hough voting</article-title>,&#x201D; in <conf-name>Proc. of IEEE17th Int. Bhurban Conf. on Applied Sciences and Technology</conf-name>, <conf-loc>Bhurban</conf-loc>, <conf-loc>Pakistan</conf-loc>, pp. <fpage>290</fpage>&#x2013;<lpage>295</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Maheswari</surname></string-name> and <string-name><given-names>V.</given-names> <surname>Radha</surname></string-name></person-group>, &#x201C;<article-title>Noise removal in compound image using median filter</article-title>,&#x201D; <source>Computer Science and Engineering</source>, vol. <volume>2</volume>, no. <issue>4</issue>, pp. <fpage>1359</fpage>&#x2013;<lpage>1362</lpage>, <year>2010</year>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y. K.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>S. H.</given-names> <surname>Zu</surname></string-name>, <string-name><given-names>Y. F.</given-names> <surname>Wang</surname></string-name> and <string-name><given-names>X. H.</given-names> <surname>Chen</surname></string-name></person-group>, &#x201C;<article-title>Deblending of simultaneous source data using a structure-oriented space-varying median filter</article-title>,&#x201D; <source>Geophysical Journal International</source>, vol. <volume>222</volume>, no. <issue>3</issue>, pp. <fpage>1805</fpage>&#x2013;<lpage>1823</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Castro, Ery</surname></string-name> and <string-name><given-names>D. L.</given-names> <surname>Donoho</surname></string-name></person-group>, &#x201C;<article-title>Does median filtering truly preserve edges better than linear filtering?</article-title>&#x201D; vol. <issue>37</issue>, no. <issue>3</issue>, pp. <fpage>1172</fpage>&#x2013;<lpage>1206</lpage>, <year>2009</year>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K. P.</given-names> <surname>Sinaga</surname></string-name> and <string-name><given-names>M. S.</given-names> <surname>Yang</surname></string-name></person-group>, &#x201C;<article-title>Unsupervised K-means clustering algorithm</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>8</volume>, pp. <fpage>80716</fpage>&#x2013;<lpage>80727</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Yang</surname></string-name> and <string-name><given-names>G.</given-names> <surname>Sheng</surname></string-name></person-group>, &#x201C;<article-title>Research on Chinese synergies based on K-means clustering algorithm and correspondence analysis</article-title>,&#x201D; in <conf-name>Proc.of IEEE Conf. on Telecommunications, Optics and Computer Science (TOCS)</conf-name>, <publisher-loc>Henan, China</publisher-loc>, pp. <fpage>387</fpage>&#x2013;<lpage>390</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Niazi</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Rahbar</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Sheikhan</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Khademi</surname></string-name></person-group>, &#x201C;<article-title>Entropy-based kernel graph cut with weighted K-means for textural image region segmentation</article-title>,&#x201D; <source>Signal Processing and Renewable Energy</source>, vol. <volume>7</volume>, no. <issue>3</issue>, pp. <fpage>13</fpage>&#x2013;<lpage>29</lpage>, <year>2023</year>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Patel</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Patel</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Patel</surname></string-name></person-group>, &#x201C;<article-title>Data analysis in shopping mall data using K-means clusterin</article-title>,&#x201D; in <conf-name>2022 4th Int. Conf. on Advances in Computing, Communication Control and Networking (ICAC3N)</conf-name>, <publisher-loc>Greater Noida, India</publisher-loc>, pp. <fpage>349</fpage>&#x2013;<lpage>352</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Huda</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Peng</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Yang</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Algburi</surname></string-name></person-group>, &#x201C;<article-title>Object scale selection of hierarchical image segmentation with deep seeds</article-title>,&#x201D; <source>Image Processing</source>, vol. <volume>15</volume>, no. <issue>1</issue>, pp. <fpage>191</fpage>&#x2013;<lpage>205</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>B. R.</given-names> <surname>Chughtai</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Jalal</surname></string-name></person-group>, &#x201C;<article-title>Object detection and segmentation for scene understanding via random forest</article-title>,&#x201D; in <conf-name>Proc. of 4th Int. Conf. on Advancements in Computational Sciences (ICACS)</conf-name>, <conf-loc>Lahore, Pakistan</conf-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>6</lpage>, <year>2023</year>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M. R.</given-names> <surname>Satpute</surname></string-name> and <string-name><given-names>S. M.</given-names> <surname>Jagdale</surname></string-name></person-group>, &#x201C;<article-title>Color, size, volume, shape and texture feature extraction techniques for fruits: A review</article-title>,&#x201D; <source>International Research Journal of Engineering and Technology</source>, vol. <volume>3</volume>, pp. <fpage>703</fpage>&#x2013;<lpage>708</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J. C.</given-names> <surname>Kavitha</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Suruliandi</surname></string-name></person-group>, &#x201C;<article-title>Texture and color feature extraction for classification of melanoma using SVM</article-title>,&#x201D; in <conf-name>Int. Conf. on Computing Technologies and Intelligent Data Engineering</conf-name>, <conf-loc>Kovilpatti, India</conf-loc>, pp. <fpage>1</fpage>&#x2013;<lpage>6</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Siddiqui</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Zafar</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Khan</surname></string-name> and <string-name><given-names>N.</given-names> <surname>Iftekhar</surname></string-name></person-group>, &#x201C;<article-title>Computer vision analysis of BRIEF and ORB feature detection algorithms</article-title>,&#x201D;in <conf-name>Proc.of Int. Conf. on Computing in Engineering &#x0026; Technology</conf-name>, <conf-loc>Singapore, Springer Nature Singapore</conf-loc>, pp. <fpage>425</fpage>&#x2013;<lpage>433</lpage>, <year>2023</year>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Bansal</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Kumar</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Kumar</surname></string-name></person-group>, &#x201C;<article-title>2D object recognition: A comparative analysis of SIFT, SURF and ORB feature descriptors</article-title>,&#x201D; <source>Multimedia Tools and Applications</source>, vol. <volume>80</volume>, pp. <fpage>18839</fpage>&#x2013;<lpage>18857</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Gupta</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Kumar</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Garg</surname></string-name></person-group>, &#x201C;<article-title>Improved object recognition results using SIFT and ORB feature detector</article-title>,&#x201D; <source>Multimedia Tools and Applications</source>, vol. <volume>78</volume>, pp. <fpage>34157</fpage>&#x2013;<lpage>34171</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Chhabra</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Garg</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Kumar</surname></string-name></person-group>, &#x201C;<article-title>Content-based image retrieval system using ORB and SIFT features</article-title>,&#x201D; <source>Neural Computing and Application</source>, vol. <volume>32</volume>, pp. <fpage>2725</fpage>&#x2013;<lpage>2733</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Sharma</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Singh</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Singh</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Gupta</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Goyal</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>A novel approach of object detection using point feature matching technique for colored images</article-title>,&#x201D; in <conf-name>Proc. of ICRIC Recent Innovations in Computing</conf-name>, <publisher-name>Springer International Publishing</publisher-name>, <publisher-loc>Jammu and Kashmir, India</publisher-loc>, pp. <fpage>561</fpage>&#x2013;<lpage>576</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>N. M.</given-names> <surname>Trieu</surname></string-name> and <string-name><given-names>N. T.</given-names> <surname>Thinh</surname></string-name></person-group>, &#x201C;<article-title>A study of combining knn and ann for classifying dragon fruits automatically</article-title>,&#x201D; <source>Journal of Image and Graphics</source>, vol. <volume>10</volume>, no. <issue>1</issue>, pp. <fpage>28</fpage>&#x2013;<lpage>35</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y. Y.</given-names> <surname>Ghadi</surname></string-name>, <string-name><given-names>A. A.</given-names> <surname>Rafique</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Al Shloul</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Alsuhibany</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Jalal</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Robust object categorization and scene classification over remote sensing images via features fusion and fully convolutional network</article-title>,&#x201D; <source>Remote Sensing</source>, vol. <volume>14</volume>, no. <issue>7</issue>, pp. <fpage>1550</fpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Lin</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Maire</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Belongie</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Bourdev</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Girshick</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Microsoft common object in context</article-title>,&#x201D; <source>Computer Vision</source>, vol. <volume>13</volume>, no. <issue>5</issue>, pp. <fpage>740</fpage>&#x2013;<lpage>755</lpage>, <year>2014</year>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Dong</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Liu</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Deep object co-segmentation and co-saliency detection via high-order spatial-semantic network modulation</article-title>,&#x201D; <source>IEEE Transactions on Multimedia</source>, vol. <volume>25</volume>, pp. <fpage>5773</fpage>&#x2013;<lpage>5746</lpage>, <year>2022</year>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Ahmed</surname></string-name> and <string-name><given-names>A.</given-names> <surname>Jalal</surname></string-name></person-group>, &#x201C;<article-title>Salient segmentation based object detection and recognition using hybrid genetic transform</article-title>,&#x201D; in <conf-name>Proc of IEEE Int. Conf. on Applied and Engineering Mathematics</conf-name>, <conf-loc>Taxila, Pakistan</conf-loc>, pp. <fpage>203</fpage>&#x2013;<lpage>208</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Ahmed</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Jalal</surname></string-name> and <string-name><given-names>K.</given-names> <surname>Kim</surname></string-name></person-group>, &#x201C;<article-title>RGB-D images for object segmentation, localization and recognition in indoor scenes using feature descriptor and hough voting</article-title>,&#x201D; in <conf-name>Proc of 17th Int. Bhurban Conf. Application of Science and Technologies</conf-name>, <conf-loc>Bhurban, Pakistan</conf-loc>, pp. <fpage>290</fpage>&#x2013;<lpage>295</lpage>, <year>2020</year>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Bansal</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Kumar</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Kumar</surname></string-name> and <string-name><given-names>K.</given-names> <surname>Kumar</surname></string-name></person-group>, &#x201C;<article-title>An efficient technique for object recognition using Shi-Tomasi corner detection algorithm</article-title>,&#x201D; <source>Soft Computing</source>, vol. <volume>25</volume>, pp. <fpage>4423</fpage>&#x2013;<lpage>4432</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Kong</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Sun</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Yao</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Lu</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>RON: Reverse connection with objectness prior networks for object detection</article-title>,&#x201D; in <conf-name>Proc. of the IEEE Conf. on Computer Vision and Pattern Recognition</conf-name>, <conf-loc>Honolulu, HI, USA</conf-loc>, pp. <fpage>5936</fpage>&#x2013;<lpage>5944</lpage>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J. U.</given-names> <surname>Kim</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Man Ro</surname></string-name></person-group>, &#x201C;<article-title>Attentive layer separation for object classification and object localization in object detection</article-title>,&#x201D; in <conf-name>IEEE Int. Conf. on Image Processing</conf-name>, <conf-loc>Taipei, Taiwan</conf-loc>, pp. <fpage>3995</fpage>&#x2013;<lpage>3999</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Tan</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Nie</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Qian</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Li</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Li</surname></string-name></person-group>, &#x201C;<article-title>Learning to rank proposals for object detection</article-title>,&#x201D; in <conf-name>Proc. of the IEEE/CVF Int. Conf. on Computer Vision</conf-name>, <conf-loc>Seoul, South Korea</conf-loc>, pp. <fpage>8273</fpage>&#x2013;<lpage>8281</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Rahmawati</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Devita</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Zain</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Rianti</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Lubis</surname></string-name> <etal>et al.,</etal></person-group> &#x201C;<article-title>Prewitt and canny methods on inversion image edge detection: An evaluation</article-title>,&#x201D; <source>Journal of Physics: Conference Series</source>, vol. <volume>1933</volume>, pp. <fpage>012039</fpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Srivastava</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Bakthula</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Agarwal</surname></string-name></person-group>, &#x201C;<article-title>Image classification using SURF and a bag of LBP features constructed by clustering with fixed centers</article-title>,&#x201D; <source>Multimedia Tools and Applications</source>, vol. <volume>78</volume>, no. <issue>11</issue>, pp. <fpage>14129</fpage>&#x2013;<lpage>14153</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Jalal</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Ahmed</surname></string-name>, <string-name><given-names>A. A.</given-names> <surname>Rafique</surname></string-name> and <string-name><given-names>K.</given-names> <surname>Kim</surname></string-name></person-group>, &#x201C;<article-title>Scene semantic recognition based on modified fuzzy C-mean and maximum entropy using object-to-object relations</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>9</volume>, pp. <fpage>27758</fpage>&#x2013;<lpage>27772</lpage>, <year>2021</year>.</mixed-citation></ref>
</ref-list>
</back></article>