<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">57118</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2024.057118</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Enhancing Building Facade Image Segmentation via Object-Wise Processing and Cascade U-Net</article-title>
<alt-title alt-title-type="left-running-head">Enhancing Building Facade Image Segmentation via Object-Wise Processing and Cascade U-Net</alt-title>
<alt-title alt-title-type="right-running-head">Enhancing Building Facade Image Segmentation via Object-Wise Processing and Cascade U-Net</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Jung</surname><given-names>Haemin</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Park</surname><given-names>Heesung</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Jung</surname><given-names>Hae Sun</given-names></name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<contrib id="author-4" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Lee</surname><given-names>Kwangyon</given-names></name><xref ref-type="aff" rid="aff-4">4</xref><email>kylee@ssu.ac.kr</email></contrib>
<aff id="aff-1"><label>1</label><institution>Department of Industrial &#x0026; Management Engineering, Korea National University of Transportation</institution>, <addr-line>Chungju, 27469</addr-line>, <country>Republic of Korea</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of Industrial Engineering, Yonsei University</institution>, <addr-line>Seoul, 03722</addr-line>, <country>Republic of Korea</country></aff>
<aff id="aff-3"><label>3</label><institution>Department of Applied Artificial Intelligence, Sungkyunkwan University</institution>, <addr-line>Seoul, 03063</addr-line>, <country>Republic of Korea</country></aff>
<aff id="aff-4"><label>4</label><institution>School of Electronic Engineering, Soongsil University</institution>, <addr-line>Seoul, 06978</addr-line>, <country>Republic of Korea</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Kwangyon Lee. Email: <email>kylee@ssu.ac.kr</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2024</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>18</day><month>11</month><year>2024</year>
</pub-date>
<volume>81</volume>
<issue>2</issue>
<fpage>2261</fpage>
<lpage>2279</lpage>
<history>
<date date-type="received">
<day>08</day>
<month>8</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>07</day>
<month>10</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2024 The Authors.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_57118.pdf"></self-uri>
<abstract>
<p>The growing demand for energy-efficient solutions has led to increased interest in analyzing building facades, as buildings contribute significantly to energy consumption in urban environments. However, conventional image segmentation methods often struggle to capture fine details such as edges and contours, limiting their effectiveness in identifying areas prone to energy loss. To address this challenge, we propose a novel segmentation methodology that combines object-wise processing with a two-stage deep learning model, Cascade U-Net. Object-wise processing isolates components of the facade, such as walls and windows, for independent analysis, while Cascade U-Net incorporates contour information to enhance segmentation accuracy. The methodology involves four steps: object isolation, which crops and adjusts the image based on bounding boxes; contour extraction, which derives contours; image segmentation, which modifies and reuses contours as guide data in Cascade U-Net to segment areas; and segmentation synthesis, which integrates the results obtained for each object to produce the final segmentation map. Applied to a dataset of Korean building images, the proposed method significantly outperformed traditional models, demonstrating improved accuracy and the ability to preserve critical structural details. Furthermore, we applied this approach to classify window thermal loss in real-world scenarios using infrared images, showing its potential to identify windows vulnerable to energy loss. Notably, our Cascade U-Net, which builds upon the relatively lightweight U-Net architecture, also exhibited strong performance, reinforcing the practical value of this method. Our approach offers a practical solution for enhancing energy efficiency in buildings by providing more precise segmentation results.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Building facade image</kwd>
<kwd>image segmentation</kwd>
<kwd>edge detection</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Korea Institute for Advancement of Technology (KIAT)</funding-source>
<award-id>P0017123</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>As global efforts to improve energy efficiency intensify, interest in buildings, the primary consumers of energy in urban environments, has surged. In the urban areas of advanced economies, buildings account for 38% of total energy consumption and 76% of electricity usage, underscoring the significance of effective building energy management [<xref ref-type="bibr" rid="ref-1">1</xref>]. Thermal energy loss plays a critical role in building energy consumption, significantly affecting heating and cooling costs while contributing to overall carbon emissions. In 2022, buildings were responsible for 34% of global energy demand and 37% of energy and process-related carbon dioxide emissions [<xref ref-type="bibr" rid="ref-2">2</xref>]. This has led to a focus on research aimed at identifying areas of thermal loss in building envelopes. By identifying areas vulnerable to thermal loss in buildings, heat leakage can be reduced through solutions such as replacing materials or adding insulation in targeted areas.</p>
<p>However, directly measuring thermal loss in building envelopes not only requires experts but is also highly challenging and inefficient. Traditional methods such as <italic>in-situ</italic> thermal measurements, blower door tests, and manual infrared thermography are time-consuming and labor-intensive. These methods require specialized equipment and trained personnel, making them costly and impractical for large-scale assessments. Additionally, accessing certain parts of a building, especially in high-rise structures, can pose safety risks. To address this issue, recent studies have proposed a process for detecting thermal loss through the analysis of building facade images. This process is also referred to as building facades parsing [<xref ref-type="bibr" rid="ref-3">3</xref>]. Notably, Park et al. [<xref ref-type="bibr" rid="ref-4">4</xref>] segmented building facade images, applied infrared images to the results, and analyzed the temperature distribution to detect abnormal thermal losses in building envelopes.</p>
<p>A fundamental element of this study is the segmentation of building facade images, which effectively differentiates between the building area and the non-building background, while also pinpointing objects within the building area. This segmentation is critical, as it precisely isolates areas of interest that are susceptible to thermal loss within the building. Using a segmentation algorithm not only streamlines the process of identifying these key objects but also significantly improves the efficiency of detecting thermal loss.</p>
<p>In the aforementioned research, the DeepLab V3&#x002B; [<xref ref-type="bibr" rid="ref-5">5</xref>], a prominent image segmentation method based on convolutional neural network (CNN), was utilized. While this methodology demonstrates commendable performance, achieving more accurate segmentation necessitates the development of an improved algorithm that takes into account the unique characteristics of building facade images.</p>
<p>This study proposes a methodology aimed at achieving higher building image segmentation performance. The input consists of building facade images taken from the front, while the output is a segmentation map containing the background, building, and windows. We focus on windows as key objects of interest, as they contribute significantly to thermal loss. This method features two distinctive characteristics. The first is object-wise processing, which separately addresses the two primary components of a building facade image: the building (or exterior wall of the building) and the windows. We use two neural network models, rather than a single model, to focus on characteristics of each object type. The second characteristic is the Cascade U-Net, which is a variation on the existing U-Net [<xref ref-type="bibr" rid="ref-6">6</xref>] model. The Cascade U-Net comprises two consecutive U-Net structures and improves the final segmentation performance by using contours of the objects as guide data for the model.</p>
<p>The methodology unfolds over four stages. In the first stage, we isolate the building and windows in bounding boxes from the input images by object detection. In the second stage, edge detection is employed to extract the edge map within the bounding boxes and to identify the contour, which is a closed curve that encompasses the largest area. Subsequently, the edge map serves as input data while the contour serves as guide data for the Cascade U-Net, our proposed image segmentation network. Finally, the results for the building and windows are synthesized to generate the ultimate segmentation map.</p>
<p>To evaluate the effectiveness of the method, we conducted experiments on a dataset of 586 Korean building images. The results demonstrated that our methodology surpasses the segmentation performance of end-to-end methods, including DeepLab V3&#x002B;. The results indicated that processing target objects individually proved to be more effective than attempting to segment areas across the entire image in a single step. Furthermore, incorporating contours as additional information within the Cascade U-Net demonstrated clear advantages: by using contours as direct guidance, it was evident that the network was able to produce segmentation outcomes more aligned with our objectives. We anticipate that improving the segmentation of building facade images through our methodology will significantly aid in thermal loss detection and, consequently, in building energy management.</p>
<p>In summary, the contributions of this research are summarized as follows:
<list list-type="bullet">
<list-item>
<p>We propose an object-wise processing methodology that isolates building facades and windows using bounding boxes, enhancing segmentation focus and accuracy.</p></list-item>
<list-item>
<p>We introduce the Cascade U-Net with contour guidance, which leverages contour data to improve segmentation performance through a two-stage U-Net network.</p></list-item>
<list-item>
<p>We demonstrate superior performance compared to end-to-end methods on a dataset of Korean building images. Additionally, we provide detailed evaluations of the contributions of each component to validate our methodology and its practical applications.</p></list-item>
</list></p>
<p>This section precedes <xref ref-type="sec" rid="s2">Section 2</xref>, which provides a review of existing image segmentation models. <xref ref-type="sec" rid="s3">Section 3</xref> delves into our methodology in greater detail, followed by a discussion of the experimental results in <xref ref-type="sec" rid="s4">Section 4</xref>. The paper concludes with <xref ref-type="sec" rid="s5">Section 5</xref>, where we discuss the study&#x2019;s implications and outline directions for future research.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<sec id="s2_1">
<label>2.1</label>
<title>Image Segmentation Methods</title>
<p>Image segmentation is the process of dividing a given image <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>I</mml:mi></mml:math></inline-formula> into <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>N</mml:mi></mml:math></inline-formula> segments <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, where each segment represents different objects or regions within the image. The primary objective is to determine which segment <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> each pixel <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> in the image <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>I</mml:mi></mml:math></inline-formula> belongs. This can be mathematically expressed as follows:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mi>I</mml:mi><mml:mo>=</mml:mo><mml:msubsup><mml:mrow><mml:mo>&#x222A;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2061;</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula>where, <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2229;</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mi>&#x2205;</mml:mi></mml:mrow></mml:math></inline-formula> for all <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>i</mml:mi><mml:mo>&#x2260;</mml:mo><mml:mi>j</mml:mi></mml:math></inline-formula> (i.e., no two segments overlap), and each <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is a set of contiguous pixels.</p>
<p>In a pixel-wise classification approach to segmentation, the goal is to learn a function <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>f</mml:mi></mml:math></inline-formula> that assigns a label <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> to each pixel:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>p</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<p>Here, <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>l</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is the label of the pixel <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>p</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, and <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mi>f</mml:mi></mml:math></inline-formula> is a function that takes a pixel as input and outputs the label of the segment it belongs to.</p>
<p>Typical image segmentation models utilize an end-to-end architecture that takes an image as input and outputs a segmentation map through a deep learning model, specifically using a CNN. While these models perform well, each has its own distinct advantages and limitations. U-Net introduced a novel approach by adding intermediate feature maps from its contracting path (encoder) to its expanding path (decoder) using skip connections. This model uses a patch-based recognition approach instead of the conventional sliding window method to improve speed and accuracy. U-Net has been utilized in studies such as automated crack segmentation, demonstrating strong performance [<xref ref-type="bibr" rid="ref-7">7</xref>]. U-Net has the advantage of a simple architecture and is highly effective when working with limited data. However, it tends to lose fine boundary details, particularly in complex or noisy images. DeepLab V3&#x002B;, which is an improved version of DeepLab V3 [<xref ref-type="bibr" rid="ref-8">8</xref>], minimizes the loss of original image information and enriches pixel representation through the implementation of depth-wise separable convolution and the Atrous Spatial Pyramid Pooling (ASPP) module in its decoder. DeepLab V3&#x002B; offers superior performance by capturing multi-scale context, but its increased complexity and higher computational demands make it less suitable for resource-constrained environments. High-Resolution Network (HR-Net) [<xref ref-type="bibr" rid="ref-9">9</xref>] features a unique architecture that avoids downscaling, thereby preserving information fidelity. It maintains high-resolution feature maps while concurrently applying lower-resolution maps, facilitating continuous inter-scale information exchange within subnetworks. This model is optimally utilized for addressing heat map-based problems, such as pose estimation. However, a downside of HR-Net is its increased memory and computational demand, as high-resolution features must be preserved throughout the network. This can pose challenges for scalability and efficiency.</p>
<p>Such models, while powerful, may not fully address the challenges posed by building facade images, which often contain repetitive patterns and subtle features. Performance limitations arise when models rely solely on raw images as input, as they may struggle to accurately segment fine details and boundaries. This observation motivates the exploration of methods that incorporate supplementary information to enhance segmentation accuracy. Advancements like Bounding Box UNet (BB-UNet) [<xref ref-type="bibr" rid="ref-10">10</xref>] and Holistically-nested Edge Detection UNet (HED-UNet) [<xref ref-type="bibr" rid="ref-11">11</xref>] build upon U-Net&#x2019;s foundational strengths, introducing additional information for performance. BB-UNet, based on U-Net&#x2019;s structure, uses bounding boxes as masks to enhance the segmentation of specific areas. It concatenates these masks to the skip connections for decoding. While this approach improves segmentation for pre-identified regions, its performance may be limited when objects lack clear bounding boxes or when precise localization of the region of interest is challenging. HED-UNet also uses U-Net, emphasizing the importance of contour information in segmentation results. It predicts contours and segmentation results together, applying contour prediction results to the segmentation outcome in the form of attention [<xref ref-type="bibr" rid="ref-12">12</xref>] to achieve the final area segmentation. However, like BB-UNet, the reliance on additional contour data increases the complexity of model training and may not be applicable to datasets where contour annotations are scarce or difficult to obtain.</p>
<p>Motivated by the recent insight that leveraging supplementary information can significantly improve segmentation outcomes, our research aims to harness contour information of objects within images for enhanced segmentation accuracy. In addition, by isolating objects within the images, we focus the segmentation process on each object individually, allowing for more concentrated and accurate segmentation results.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Edge Detection Methods</title>
<p>The methodology proposed in this study involves extracting edge maps and obtaining contours from images. For edge detection, studies have used algorithm-based methods like Canny edge detection [<xref ref-type="bibr" rid="ref-13">13</xref>], as well as deep learning-based models such as Holistically-nested Edge Detection (HED) [<xref ref-type="bibr" rid="ref-14">14</xref>], and Dense Extreme Inception Network for Edge Detection (DexiNed) [<xref ref-type="bibr" rid="ref-15">15</xref>].</p>
<p>Canny edge detection remains the most prevalent algorithm for contour detection due to its efficiency and independence from training. It operates through multiple stages, initially focusing on noise reduction&#x2014;a crucial step as noise significantly affects contour detection. It uses a Gaussian filter to minimize image noise, followed by the application of Sobel kernels in both horizontal and vertical directions to compute the gradient magnitudes. By establishing a threshold, it discerns the significant gradient variations among neighboring pixels, thereby isolating the precise edges. While effective at detecting clear object contours in low-noise images, its performance diminishes with increased background noise.</p>
<p>Xie et al. [<xref ref-type="bibr" rid="ref-14">14</xref>] sought to address comprehensive image learning and prediction challenges, including multi-level feature learning, through their contour detection model. The HED model they proposed learns hierarchical representations to mitigate the ambiguity often encountered in contour and object boundary detection. Demonstrating commendable performance on the Berkeley Segmentation Data Set 500 (BSDS 500) dataset [<xref ref-type="bibr" rid="ref-16">16</xref>]&#x2014;a benchmark for contour detection&#x2014;it stands out for its speed (0.4 s per image), surpassing many contemporary CNN-based algorithms in efficiency.</p>
<p>DexiNed is a state-of-the-art deep learning model for contour detection, notable for producing finer contours compared to predecessors like HED. Structurally, DexiNed comprises an encoder with six blocks, each linked through convolutional blocks, and includes an auxiliary network enabling information exchange among these blocks.</p>
<p>Due to the need for clean edge maps and precise contours in this research, DexiNed was chosen for its superior noise reduction and contour detection capabilities, making it a crucial part of our edge detection process.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Building Facade Image Segmentation Methodology</title>
<p>This section provides a detailed explanation of our algorithm, which is distinguished by two principal innovations. The first innovation is object-wise processing; instead of using the original images directly, the approach employs object-wise bounding boxes as inputs. Unlike end-to-end models, the facade and window are processed separately and combined together in the final step for a comprehensive segmentation output. The second key innovation is the use of a neural network structure called Cascade U-Net, which improves on the conventional U-Net model by integrating two U-Net architectures. This configuration utilizes contours extracted from the images as guiding data, enhancing the segmentation performance of the model. <xref ref-type="fig" rid="fig-1">Fig. 1</xref> presents an overview of the proposed methodology, illustrating the process from the initial building facade image to the final segmentation map.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Overview of proposed method</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57118-fig-1.tif"/>
</fig>
<p>The methodology is composed of four stages, with each step playing a specific role:</p>
<p>Object Isolation: In this initial stage, an object detection model is used to isolate the building&#x2019;s facades and windows using bounding boxes, followed by resizing to standardize the input data for further processing.</p>
<p>Edge and Contour Extraction: Subsequently, an edge map is generated, facilitating the identification of precise contours within the delineated object boxes, thus laying the groundwork for detailed segmentation.</p>
<p>Object-Wise Segmentation: Central to our methodology, the Cascade U-Net structure incorporates two sequentially connected U-Net networks. The initial network is tasked with refining object contours for accuracy, while the subsequent network leverages these refined contours to execute the segmentation.</p>
<p>Output Synthesis: The final stage integrates the segmented outputs for facades and windows processed through the respective Cascade U-Net, culminating in a unified segmentation map that accurately reflects the architectural elements of the building facade.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Object Isolation</title>
<p>In the object isolation step, the facade and window objects are separated from the original image and ready to be processed in parallel for the next step. This step begins with object detection for the target objects. The bounding box obtained as a result of object detection represents the minimal, rectangular-shaped area enclosing the object, consisting of a position defined by x and y coordinates, as well as width and height. Here we assume that each input image contains only one building; therefore, because of object detection, we obtain one bounding box for the facade and multiple bounding boxes for the windows.</p>
<p>After identifying the bounding boxes for the objects, these areas are cropped from the original image, and their coordinates in the image are recorded. This record-keeping is crucial for the final step, where the final segmentation map is synthesized.</p>
<p>Given the variability in the sizes of the cropped object images, a standardization process is undertaken to resize them to a uniform input shape suitable for deep learning analysis. We chose the shape to approximate the average size of the cropped images, resizing facade images to 512 &#x00D7; 512 pixels and window images to 64 &#x00D7; 64 pixels. The rationale behind these specific resizing decisions is further elaborated upon in <xref ref-type="sec" rid="s4">Section 4</xref>.</p>
<p>This strategy to use bounding boxes as inputs for the model, rather than the original images, was inspired by the previous works including BB-UNet [<xref ref-type="bibr" rid="ref-10">10</xref>], Kervadec et al. [<xref ref-type="bibr" rid="ref-17">17</xref>], and Lee et al. [<xref ref-type="bibr" rid="ref-18">18</xref>].</p>
<p>To maintain clarity and focus on our explanation, illustrations and examples provided henceforth will primarily address the processing of facade. It should be noted, however, that the same procedural framework is applied to windows as well.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Contour Extraction</title>
<p>In this step, the contours of objects are extracted. A contour is the out-most closed curve that represents the target object, which we consider as a key latent factor for segmentation. The contour serves as guiding data in the subsequent Cascade U-Net stage to enhance the accuracy of building image segmentation. <xref ref-type="fig" rid="fig-2">Fig. 2</xref> displays the intermediate outputs of this process.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Contour extraction</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57118-fig-2.tif"/>
</fig>
<p>We extract the edge map from the object image using DexiNed [<xref ref-type="bibr" rid="ref-15">15</xref>], the edge detection model that demonstrates the best performance. We utilized weights trained on the Berkeley Segmentation Dataset, specifically BSDS500 for this model.</p>
<p>We select DexiNed not only because it represents the state-of-the-art but also because it proved to be the most suitable for the building facade image dataset in experimental results. The Canny method [<xref ref-type="bibr" rid="ref-13">13</xref>] has the advantage of not requiring training, but it performed poorly due to its sensitivity to various noises, such as power lines or clouds around buildings. Deep learning-based contour detection models like HED [<xref ref-type="bibr" rid="ref-14">14</xref>] and HED-UNet [<xref ref-type="bibr" rid="ref-11">11</xref>] were also tested; though they were less sensitive, they often produced outputs with gaps instead of closed curves in experiments. Therefore, DexiNed was selected as the edge detection model for this study. Among the detected edges, we choose the closed curve with most extensive area as the object&#x2019;s contour using the OpenCV library [<xref ref-type="bibr" rid="ref-19">19</xref>]. This approach of contour extraction is based on the assumption that the closed curve with the largest area within the bounding box of the target object is most likely to represent the object.</p>
<p>The second and the third images in <xref ref-type="fig" rid="fig-2">Fig. 2</xref> illustrate the edge map and the resulting contour of the target object, respectively. As can be seen in the area marked with a red circle (<xref ref-type="fig" rid="fig-2">Fig. 2c</xref>), the contour of the object is not accurate due to the background noises present in the building image. In the subsequent step, the front network of the Cascade U-Net will address this issue.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Object-Wise Segmentation</title>
<p>The Cascade U-Net builds on the traditional U-Net architecture but introduces two distinct differences. First, instead of using a single U-Net, it connects two networks in sequence to refine the process and generate intermediate results, aiming for improved performance. Another difference is the utilization of contours as guide data to steer the results in the desired direction for each network. Cascade U-Net is illustrated in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Cascade U-Net</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57118-fig-3.tif"/>
</fig>
<p>The Cascade U-Net progresses through two sub-tasks. First sub-task is contour modification generates a refined contour from edge map, using the contour derived from the previous step. The next sub-task is image segmentation, which produces a segmentation map from a grayscale image, using the modified contour as guiding data.</p>
<p>The first network of the Cascade U-Net, called the contour refinement network, aims to refine the contours representing the object. The contour found in the previous step is used as guiding data to derive the modified contour from the edge map.</p>
<p>This network fundamentally follows the U-Net model, taking an edge map as input and outputting a modified contour. The distinctive aspect proposed in this study is the use of the contour as guide data, adding it twice: once at the input part of the encoding path and again at the predict layer, as indicated by orange arrows in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>.</p>
<p>The second network, the image segmentation network, aims to produce an accurate segmentation map for the input object. Its overall structure including the application of guide data mirrors the Contour Refinement Network. However, the input for the network is changed to the grayscale image from the edge map. Additionally, the guide data here is the refined contours from the former network to ensure results are aligned with the segmentation objective.</p>
<p>In both networks, the guide data is concatenated with the primary input along the channel dimension at the encoder&#x2019;s input layer. This concatenation increases the number of input channels from one to two, allowing the network to process both the image intensity and contour information simultaneously.</p>
<p>In summary, the edge map and initial contour flow into the first U-Net, which outputs a refined contour. This refined contour then flows into the second U-Net along with the grayscale image to produce the segmentation map.</p>
<p>The process of incorporating guide data occurs twice, utilizing a layer named the Merging and Prediction Layer (MPL) before the prediction layer (<xref ref-type="fig" rid="fig-4">Fig. 4</xref>). The traditional U-Net&#x2019;s prediction layer, indicated in blue (output of U-Net), carries a feature map of width &#x00D7; height &#x00D7; 64, which is processed through a 3 &#x00D7; 3 convolution layer and a 1 &#x00D7; 1 convolution layer (i.e., fully connected layer) to produce the final segmentation output. In this study, to effectively reflect the guide data, the feature map of width &#x00D7; height &#x00D7; 64 is reduced to two channels (width &#x00D7; height &#x00D7; 2) through a convolution layer, concatenated with guide data of width &#x00D7; height &#x00D7; 1, and then passed through two 3 &#x00D7; 3 convolution layers and a fully convolution layer in the MPL.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Merging and prediction layer</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57118-fig-4.tif"/>
</fig>
<p>This MPL layer ensures that the guide data directly influences the final predictions, improving the network&#x2019;s ability to accurately segment object boundaries.</p>
<p>As the final step, output synthesis process is conducted to compile the final results. This stage reverses the process of object isolation. After passing through the Cascade U-Net, the results for each object, namely the segmentation map of the facade and the windows, are obtained. The segmented results for each object are restored to their original sizes, and then pasted back into their original positions, respectively. The sequence involves first attaching the exterior wall and then overlaying the windows onto them. The last segmentation map in <xref ref-type="fig" rid="fig-1">Fig. 1</xref> displays this integrated final output.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiment</title>
<sec id="s4_1">
<label>4.1</label>
<title>Experimental Design</title>
<p>To demonstrate the effectiveness of the proposed methodology, we compare its performance against other methodologies using a dataset of Korean building images captured from the front.</p>
<sec id="s4_1_1">
<label>4.1.1</label>
<title>Dataset</title>
<p>The CMP FACADE dataset [<xref ref-type="bibr" rid="ref-20">20</xref>] is a benchmark dataset frequently used for building image analysis. However, FACADE presented challenges for use in this study for three main reasons. Firstly, it is uncommon for the images within FACADE to capture the entire building, which is crucial for our study as it requires contour extraction and images showing the complete outline of the building facade. Secondly, the dataset predominantly features European-style buildings, which often include architectural elements not typically found in Korean buildings, such as column-like structures. Also, the dataset lacks features common in Korean buildings, such as external air conditioning units. Therefore, it was not suitable for out method which targets Korean building images. Lastly, inaccuracies in the segmentation ground truths within the FACADE dataset, particularly due to the ambiguous facades of older buildings could lead to incorrect identification of object types or areas on the segmentation maps. These factors made the dataset impractical for our purposes.</p>
<p>As a result, we collected our dataset, consisting of 586 images of Korean buildings, carefully selected to ensure each image included the complete outline of the buildings. Additionally, we included a diverse range of building types commonly seen, such as commercial shops, villas, and office buildings, to ensure a comprehensive dataset. <xref ref-type="fig" rid="fig-5">Fig. 5</xref> showcases examples from the dataset we utilized.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Data examples</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57118-fig-5.tif"/>
</fig>
<p>For the Cascade U-Net, we cropped images of facades or windows as inputs, using 586 cropped facade images and 1420 cropped window images. We employed the open-source tool labelIMG (<ext-link ext-link-type="uri" xlink:href="https://github.com/tzutalin/labelImg">https://github.com/tzutalin/labelImg</ext-link> (accessed on 08 August 2024)) for manually marking the locations of objects in a box shape on the training data. We applied a margin of 5 pixels around the bounding boxes during cropping to avoid accidentally cutting parts of the objects.</p>
<p>For data augmentation, we used only flip and rotate techniques. While blur [<xref ref-type="bibr" rid="ref-21">21</xref>], mixup [<xref ref-type="bibr" rid="ref-22">22</xref>], and cutmix [<xref ref-type="bibr" rid="ref-23">23</xref>] are popular in recent studies, they generate images in ways that could conflict with our study&#x2019;s requirement for detecting a single object within a box.</p>
<p>We split the dataset into training, validation, and test sets with an 8:1:1 ratio. For segmentation labeling, domain experts manually annotated the images to provide ground truth. For contours, we used the edges of the segmentation labels.</p>
</sec>
<sec id="s4_1_2">
<label>4.1.2</label>
<title>Evaluation Metric</title>
<p>To assess segmentation performance, we utilized widely recognized metrics: pixel-wise accuracy (PA) and intersection over union (IoU). Pixel-wise accuracy measures how correct each pixel is by comparing the predicted image with the ground truth, calculated as follows:
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mi>P</mml:mi><mml:mi>A</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:math></disp-formula>where, <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the total count of pixels predicted as class <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mi>j</mml:mi></mml:math></inline-formula> when they actually belong to class <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mi>i</mml:mi></mml:math></inline-formula>.</p>
<p>IoU evaluates the proportion of overlap between the predicted and actual areas, calculated as follows:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo>|</mml:mo><mml:mi>P</mml:mi><mml:mo>&#x2229;</mml:mo><mml:mi>G</mml:mi><mml:mo>|</mml:mo></mml:mrow><mml:mrow><mml:mo>|</mml:mo><mml:mi>P</mml:mi><mml:mo>&#x222A;</mml:mo><mml:mi>G</mml:mi><mml:mo>|</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>Here, <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mi>P</mml:mi></mml:math></inline-formula> represents the predicted area, and <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mi>G</mml:mi></mml:math></inline-formula> is the ground truth area. Both metrics yield values between 0 and 1, with higher scores indicating either more accurately predicted pixels or greater overlap between the predicted and actual areas, thus denoting superior performance. Specifically, we differentiated our IoU analysis into mean IoU and wall IoU, where mean IoU averages the individual IoU scores for exterior walls and windows, and wall IoU exclusively focuses on the exterior wall&#x2019;s IoU score.</p>
</sec>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Results</title>
<sec id="s4_2_1">
<label>4.2.1</label>
<title>Quantitative Analysis</title>
<p>In the quantitative analysis, we compared the segmentation map generated by each algorithm against the ground truth and evaluated performance. For baseline models, we selected end-to-end architectures: U-Net, DeepLab V3&#x002B;, and HED-UNet. These models were given a grayscale version of the original images without object isolation (i.e., cropping into bounding boxes) as an input. For the training, U-Net and our proposed method were set to run for 100 epochs. In contrast, DeepLab V3&#x002B; and HED-UNet followed the training parameters recommended in their original publications.</p>
<p>The outcomes of these evaluations are summarized in <xref ref-type="table" rid="table-1">Table 1</xref>. In the table, our method (the Cascade U-Net with object isolation) demonstrated the highest performance among the tested models. Compared to the U-Net, there was a significant improvement in performance. Specifically, it achieved a 2.90% increase over DeepLab V3&#x002B; and a 1.88% increase over HED-UNet based on the Wall IoU metric. These results clearly illustrate the superiority of the Cascade U-Net over traditional models.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Performance comparison with baseline models</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>Segmentation model</th>
<th>PA (%)</th>
<th>Wall IoU (%)</th>
<th>Mean IoU (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="3">End-to-end segmentation</td>
<td>U-Net</td>
<td>63.4</td>
<td>65.0</td>
<td>65.6</td>
</tr>
<tr>
<td>DeepLab V3&#x002B;</td>
<td>68.7</td>
<td>66.1</td>
<td>70.7</td>
</tr>
<tr>
<td>HED-UNet</td>
<td>70.8</td>
<td>67.1</td>
<td>70.8</td>
</tr>
<tr>
<td>Ours</td>
<td>Cascade U-Net</td>
<td>72.3</td>
<td>69.0</td>
<td>73.8</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Here, DeepLab V3&#x002B; also outperformed the standard U-Net when applied in an end-to-end manner. However, for our proposed methodology, we selected U-Net to construct the Cascade U-Net architecture instead of cascading DeepLab V3&#x002B;. This was because U-Net has a simpler and more straightforward architecture with fewer parameters compared to DeepLab V3&#x002B;. When cascading networks for each object, computational complexity and resource demands increase significantly. Utilizing U-Net allows us to cascade two networks without rendering the model impractically large or slow.</p>
<p>To gain a deeper understanding of our methodology, we conducted experiments to evaluate the impact of each component of the methodology on performance. The results are summarized in <xref ref-type="table" rid="table-2">Table 2</xref>, with descriptions for each setting as follows:</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Performance impact of methodology components (Overall) (OD: Object Detection, OI: Object Isolation, CE: Contour Extraction, OS: Output Synthesis, ISN: Image Segmentation Network)</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Method</th>
<th>PA (%)</th>
<th>Wall IoU (%)</th>
<th>Mean IoU (%)</th>
</tr>
</thead>
<tbody>
<tr>
<td>(1) OD</td>
<td>66.2</td>
<td>58.8</td>
<td>55.0</td>
</tr>
<tr>
<td>(2) OI, U-Net (grayscale)</td>
<td>69.2</td>
<td>63.6</td>
<td>67.6</td>
</tr>
<tr>
<td>(3) OI, CE, OS</td>
<td>66.3</td>
<td>53.8</td>
<td>60.0</td>
</tr>
<tr>
<td>(4) OI, CE, U-Net (edge map)</td>
<td>71.9</td>
<td>67.8</td>
<td>71.2</td>
</tr>
<tr>
<td>(5) OI, CE, ISN (contour, grayscale), OS</td>
<td>71.8</td>
<td>67.0</td>
<td>71.1</td>
</tr>
<tr>
<td>(6) OI, CE, Cascade U-Net without MPL, OS</td>
<td>72.0</td>
<td>67.7</td>
<td>71.4</td>
</tr>
<tr>
<td>(7) OI, CE, Cascade U-Net, OS</td>
<td>72.3</td>
<td>69.0</td>
<td>73.8</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>(1) Bounding boxes obtained through object detection in the Object Isolation stage were considered as the object areas themselves. (2) The part of the building&#x2019;s bounding box obtained through Object Isolation was converted to grayscale and fed into U-Net to obtain a segmentation map. (3) The internal areas of contours obtained for buildings and windows were considered as the areas of each object, and synthesizing these yielded a segmentation map. (4) An edge map of the building&#x2019;s bounding box obtained through Object Isolation was fed into U-Net to obtain a segmentation map. (5) Only the Image Segmentation Network part of Cascade U-Net was used to obtain and synthesize the segmentation map. (6) Cascade U-Net was applied, but without the additional application of MPL to guide contours. (7) Our methodology was fully applied.</p>
<p>Bounding boxes obtained through object detection showed poor performance (1), and contours alone were insufficient for segmentation (3). U-Net performs better when receiving an edge map rather than a grayscale image (2, 4), indicating the value of edge maps as additional information aiding segmentation. However, directly feeding contours extracted from the edge map into U-Net introduced noise, reducing performance (4, 5). Providing more accurate contours through contour refinement led to higher performance (6) and reflecting guidance in the last layer through MPL further improved performance (7).</p>
<p>The experimental results demonstrate that each component incrementally contributes to the overall performance improvement.
<list list-type="bullet">
<list-item>
<p>Object Isolation (OI): By cropping images to focus on individual objects, we reduce background noise and allow the model to concentrate on relevant features, leading to an initial performance boost.</p></list-item>
<list-item>
<p>Contour Extraction (CE): Extracting contours provides structural information about the objects, which is essential for accurate boundary delineation.</p></list-item>
<list-item>
<p>Cascade U-Net Architecture: Introducing a two-stage U-Net allows for the refinement of contours before they are used as guide data in the segmentation process. This sequential processing ensures that the guidance provided to the segmentation network is accurate and reliable.</p></list-item>
<list-item>
<p>Merging and Prediction Layer (MPL): The MPL enhances the model&#x2019;s ability to integrate guide data effectively. By merging the refined contours with the network&#x2019;s feature maps, it ensures that the final segmentation output benefits from precise boundary information.</p></list-item>
</list></p>
<p>By analyzing each component&#x2019;s contribution, we confirm that the combination of object-wise processing, contour refinement, and guided segmentation in the Cascade U-Net architecture leads to significant improvements.</p>
</sec>
<sec id="s4_2_2">
<label>4.2.2</label>
<title>Qualitative Analysis</title>
<p>For a qualitative comparison, we juxtaposed the segmentation ground truths with the results from DeepLab V3&#x002B; and our proposed method, as depicted in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Qualitative analysis results comparing the ground truth with outcomes from DeepLab V3&#x002B; and our method</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57118-fig-6.tif"/>
</fig>
<p>A notable observations from the qualitative analysis is the continuity within the predicted areas. Observing the second row of DeepLab V3&#x002B;&#x2019;s results in <xref ref-type="fig" rid="fig-6">Fig. 6</xref> reveals white spaces indicating the background surrounded by the building and windows. Since buildings typically lack openings in the middle, such discontinuities detract from the overall performance. In addition, when there are reflections on the windows, the performance gap between DeepLab V3&#x002B; and our methodology becomes more pronounced. Specifically, in cases where other buildings are reflected on the windows, DeepLab V3&#x002B; tends to misinterpret those areas as part of the building&#x2019;s exterior wall. Similarly, when the sky is reflected, it often mistakes those regions for the background outside the building. Additionally, using bounding boxes reduces unnecessary window detection, enhancing prediction precision. Thus, we conclude that our method achieves results that accurately reflect the common characteristics of building facade images.</p>

</sec>
<sec id="s4_2_3">
<label>4.2.3</label>
<title>Hyperparameter Optimization</title>
<p>Experiments were conducted on two key optimizable parameters: the input size of windows and the number of convolution layers within the MPL.</p>
<p>Due to the variability in bounding box sizes of objects, a standardized square input size is needed for the Cascade U-Net model. We experimented with window size w set to 32, 64, and 128 pixels, maintaining two layers in the MPL throughout. The experiments utilized these window sizes as inputs specifically for the window component of our proposed model. As shown in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>, the optimal performance was observed with <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mi>w</mml:mi><mml:mo>=</mml:mo><mml:mn>64</mml:mn></mml:math></inline-formula>, yielding the highest wall IoU across the model. Performance was significantly lower with <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi>w</mml:mi><mml:mo>=</mml:mo><mml:mn>32</mml:mn></mml:math></inline-formula>, at 60.7%. At <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mi>w</mml:mi><mml:mo>=</mml:mo><mml:mn>128</mml:mn></mml:math></inline-formula>, there was a slight performance decrease to 68.6%.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Performance relative to window image size</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57118-fig-7.tif"/>
</fig>
<p>To ascertain why 64 &#x00D7; 64 offered the best performance, we checked the pixel count of window images cropped from the original images. The window images&#x2019; average pixel count was 3339 pixels, closely aligning with the 64 &#x00D7; 64 size. This suggests that downsizing the window images to 32 &#x00D7; 32 compressed the information too much, leading to worse performance. This underscores the importance of tailoring the input bounding box size to accurately reflect the specific traits of the target dataset.</p>
<p>For the number of convolution layers, we explored the optimal number needed to merge the U-Net&#x2019;s standard output of <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mi>w</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>&#x00D7;</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>h</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>&#x00D7;</mml:mo><mml:mspace width="thinmathspace" /><mml:mn>64</mml:mn></mml:math></inline-formula> with the guide data of <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mi>w</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>&#x00D7;</mml:mo><mml:mspace width="thinmathspace" /><mml:mi>h</mml:mi><mml:mspace width="thinmathspace" /><mml:mo>&#x00D7;</mml:mo><mml:mspace width="thinmathspace" /><mml:mn>1</mml:mn></mml:math></inline-formula> for accurate segmentation. To avoid potential bias from directly concatenating the original U-Net&#x2019;s 64-channel output with the contour, we reduced it to 2 channels.</p>
<p>The experimental setup was designed to assess performance across one to four 3 &#x00D7; 3 convolution layers. The outcomes are depicted in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>. The scenario with a single layer, although closest to the original U-Net configuration, was inadequately trained with the addition of guide data, hence not shown in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>. The findings revealed that two convolution layers delivered the best performance, achieving a wall IoU of 69.01%. Introducing three or four layers led to performance reductions of approximately 2% and 7%, respectively. As the number of layers increased, performance declined, likely due to the unnecessary complexity causing information loss, consistent with insights from He et al. [<xref ref-type="bibr" rid="ref-24">24</xref>].</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Performance with respect to number of convolution layers</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57118-fig-8.tif"/>
</fig>
</sec>
<sec id="s4_2_4">
<label>4.2.4</label>
<title>Practical Scenario: Window Thermal Loss Classification</title>
<p>We explored the practical application of the proposed methodology by classifying windows based on their vulnerability to thermal loss. <xref ref-type="fig" rid="fig-9">Fig. 9</xref> illustrates an example of how we analyzed thermal loss in windows using the segmentation map generated from our research and infrared images.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Window thermal loss classification based on infrared imaging and cascade U-Net segmentation</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57118-fig-9.tif"/>
</fig>
<p>In this scenario, the infrared image was divided into 279 distinct classes representing relative temperature levels, and the pixel frequencies for each class were visualized in a histogram. Since there is no predefined threshold for thermal loss, we selected the midpoint of the observed bimodal distribution as the threshold. Using this threshold, we classified windows based on the proportion of pixels in the segmented window regions that exceeded the threshold. Specifically, if more than 20% of a window&#x2019;s pixels were classified above the threshold, the window was considered highly vulnerable to thermal loss. Windows with 5%&#x2013;20% of pixels above the threshold were classified as moderately vulnerable, while those with fewer than 5% were classified as having minimal thermal loss. We applied this method to several buildings for which we had infrared images, confirming the potential of this approach for practical use. This demonstrates that our methodology can be used in real-world applications, such as identifying areas of concern in building energy management.</p>
</sec>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Discussion</title>
<p>As shown in <xref ref-type="table" rid="table-1">Table 1</xref>, the proposed methodology demonstrated superior performance over other methodologies, including DeepLab V3&#x002B;, on our constructed dataset of 586 images. Notably, the segmentation of building area exhibited high performance, with a mean IoU of over 73%.</p>

<p>The superior performance of our methodology over end-to-end models suggests that object-wise processing, which separates target objects into smaller bounding boxes rather than processing the original input image, is more effective. Furthermore, we explored the benefits of utilizing contours with the cascade U-Net. Contours, a key latent information embedded within the input image, act as a direct guide for the segmentation model, steering the segmentation process towards the desired outcomes.</p>
<p>Furthermore, our qualitative analysis revealed that the segmentation map produced by our model exhibited greater continuity compared to those from other models. Once again, this success can be attributed to the use of object-wise processing, generating cropped images by bounding boxes in the object isolation step.</p>
<p>It is important to note that the dataset used in this study is relatively small compared to standard benchmark datasets. Future efforts should focus on expanding the dataset for further testing. It is deemed valuable to not only increase the count of building facade images but also to examine the performance on non-frontal images, including side view images and partially obscured building images.</p>
<p>Despite achieving high performance, our methodology carries certain drawbacks. The use of two interconnected U-Net structures for the segmentation of walls and windows introduces increased complexity in the model architecture, a greater number of parameters to train, and extended training durations. While these may be minor issues when prioritizing performance, future studies should explore solutions to alleviate these challenges.</p>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusion</title>
<p>In this study, we proposed a methodology enhancing the segmentation of the building and the windows within building facade images designed to facilitate the analysis of thermal loss in buildings. Our method has two key characteristics.</p>
<p>Firstly, instead of directly extracting the final segmentation map from the input image, we utilized a method that separately segments the objects comprising the facade image before synthesizing them. This object-wise processing confines the input image to bounding boxes, allowing the image segmentation model to focus on the target object&#x2019;s features. It also aids in maintaining the continuity of areas in the segmentation output.</p>
<p>Secondly, we proposed the Cascade U-Net, composed of two consecutive U-Nets. It treats contours as a key latent factor in image segmentation and uses them as guide data. This approach steers the model towards producing the desired outcomes.</p>
<p>Our experiments on the custom-built dataset of Korean building images demonstrated that these features contributed to achieving higher segmentation performance. If this approach enables better segmentation of building facade images, it is anticipated to greatly assist in building energy management when linked with thermal loss detection.</p>
<p>The main advantage of our methodology is its improved segmentation accuracy, outperforming other end-to-end methodologies including DeepLab V3&#x002B;. By object-wise processing, each model was able to concentrate on specific objects, leading to further improvements in segmentation performance. Additionally, the incorporation of contours as guide data significantly enhanced the model&#x2019;s ability to accurately delineate object boundaries, a critical aspect of image segmentation.</p>
<p>However, our approach also presents some disadvantages. Although U-Net is considered a relatively lightweight model compared to more recent architectures, the use of multiple Cascade U-Net structures for different objects increases the model&#x2019;s complexity, resulting in a larger number of parameters to train and higher computational demands. This added complexity also results in longer training and inference times compared to simpler models, which may limit its use in scenarios requiring fast processing. Moreover, the success of our method relies on accurate contour extraction, meaning that it may require larger and more diverse datasets to generalize effectively across different building types and conditions.</p>
<p>For future research directions, we are considering the diversification of input images and the variety of objects to be detected. The input images in this study were building facade images that fully encompass the front of the buildings, but there&#x2019;s no guarantee that buildings will always be photographed in this manner. Moreover, since there are various objects constituting the building envelope beyond just windows, we plan to research methods capable of segmenting into multiple classes based on a wider variety of inputs. Finally, to ensure the practicality of this approach in real-world scenarios, we plan to evaluate the model&#x2019;s performance on edge devices. While we select U-Net due to its lightweight nature, we did not conduct specific experiments on such devices. This evaluation will be crucial in determining its efficiency and feasibility in real-time applications in resource-constrained environments.</p>
</sec>
</body>
<back>
<ack>
<p>The authors would like to express sincere gratitude to Smart Systems Lab for their invaluable support and guidance throughout this research.</p>
</ack>
<sec><title>Funding Statement</title>
<p>This work was supported by Korea Institute for Advancement of Technology (KIAT): P0017123, the Competency Development Program for Industry Specialist.</p>
</sec>
<sec><title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Conceptualization: Haemin Jung, Heesung Park and Kwangyon Lee; Analysis and interpretation of results: Haemin Jung and Kwangyon Lee; Draft manuscript preparation: Haemin Jung, Heesung Park, Hae Sun Jung and Kwangyon Lee. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability"><title>Availability of Data and Materials</title>
<p>The datasets used and/or analysed during the current study available from the corresponding author on reasonable request.</p>
</sec>
<sec><title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>U.S. Department of Energy</collab></person-group>, &#x201C;<article-title>Chapter 5&#x2014;Increasing efficiency of building systems and technologies</article-title>,&#x201D; <year>Sep. 2015</year>. <comment>Accessed: Aug. 8, 2024</comment>. [Online]. Available: <ext-link ext-link-type="uri" xlink:href="https://www.energy.gov/articles/chapter-5-increasing-efficiency-buildings-systems-and-technologies">https://www.energy.gov/articles/chapter-5-increasing-efficiency-buildings-systems-and-technologies</ext-link></mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><collab>United Nations Environment Programme</collab></person-group>, &#x201C;<article-title>Global status report for buildings and construction&#x2014;Beyond foundations: Mainstreaming sustainable solutions to cut emissions from the buildings sector</article-title>,&#x201D; in <source>Technical Reports</source>. <publisher-loc>Nairobi</publisher-loc>: <publisher-name>United Nations Environment Programme</publisher-name>, <year>2024</year>. doi: <pub-id pub-id-type="doi">10.59117/20.500.11822/45095</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Zhu</surname></string-name>, and <string-name><given-names>S. C.</given-names> <surname>Hoi</surname></string-name></person-group>, &#x201C;<chapter-title>DeepFacade: A deep learning approach to facade parsing</chapter-title>,&#x201D; <article-title>presented at the 26th Int. Joint Conf. Artif. Intell. (IJCAI-17)</article-title>, <publisher-loc>Melbourne, Australia</publisher-loc>, <year>Aug. 19&#x2013;25, 2017</year>, pp. <fpage>2301</fpage>&#x2013;<lpage>2307</lpage>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Park</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Lee</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Jang</surname></string-name>, and <string-name><given-names>C.</given-names> <surname>Kim</surname></string-name></person-group>, &#x201C;<article-title>Thermal anomaly detection in walls via CNN-based segmentation</article-title>,&#x201D; <source>Autom. Constr.</source>, vol. <volume>125</volume>, <year>2021, Art. no. 103627</year>. doi: <pub-id pub-id-type="doi">10.1016/j.autcon.2021.103627</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>L. C.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zhu</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Papandreou</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Schroff</surname></string-name>, and <string-name><given-names>H.</given-names> <surname>Adam</surname></string-name></person-group>, &#x201C;<chapter-title>Encoder-decoder with atrous separable convolution for semantic image segmentation</chapter-title>,&#x201D; <article-title>presented at the Eur. Conf. Comput. Vis. (ECCV)</article-title>, <publisher-loc>Munich, Germany</publisher-loc>, <year>Sep. 8&#x2013;14, 2018</year>, pp. <fpage>801</fpage>&#x2013;<lpage>818</lpage>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>O.</given-names> <surname>Ronneberger</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Fischer</surname></string-name>, and <string-name><given-names>T.</given-names> <surname>Brox</surname></string-name></person-group>, &#x201C;<chapter-title>U-Net: Convolutional networks for biomedical image segmentation</chapter-title>,&#x201D; <article-title>presented at the Med. Image Comput. Comput.-Assist. Interv. (MICCAI 2015)</article-title>, <publisher-loc>Munich, Germany</publisher-loc>, <year>Oct. 5&#x2013;9, 2015</year>, vol. <volume>9351</volume>, pp. <fpage>234</fpage>&#x2013;<lpage>241</lpage>. doi: <pub-id pub-id-type="doi">10.1007/978-3-319-24574-4_28</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Reichard</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Xu</surname></string-name>, and <string-name><given-names>A.</given-names> <surname>Akanmu</surname></string-name></person-group>, &#x201C;<article-title>Automated crack segmentation in close-range building fa&#x00E7;ade inspection images using deep learning techniques</article-title>,&#x201D; <source>J. Build. Eng.</source>, vol. <volume>43</volume>, <year>2021, Art. no. 102913</year>. doi: <pub-id pub-id-type="doi">10.1016/j.jobe.2021.102913</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>L. C.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Papandreou</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Schroff</surname></string-name>, and <string-name><given-names>H.</given-names> <surname>Adam</surname></string-name></person-group>, &#x201C;<article-title>Rethinking atrous convolution for semantic image segmentation</article-title>,&#x201D; <comment>2017, <italic>arXiv:1706.05587</italic></comment>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Sun</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>High-resolution representations for labeling pixels and regions</article-title>,&#x201D; <comment>2019, <italic>arXiv:1904.04514</italic></comment>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>El Jurdi</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Petitjean</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Honeine</surname></string-name>, and <string-name><given-names>F.</given-names> <surname>Abdallah</surname></string-name></person-group>, &#x201C;<article-title>BB-UNet: U-Net with bounding box prior</article-title>,&#x201D; <source>IEEE J. Sel. Top. Signal Process.</source>, vol. <volume>14</volume>, no. <issue>6</issue>, pp. <fpage>1189</fpage>&#x2013;<lpage>1198</lpage>, <year>2020</year>. doi: <pub-id pub-id-type="doi">10.1109/JSTSP.2020.3001502</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Heidler</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Mou</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Baumhoer</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Dietz</surname></string-name>, and <string-name><given-names>X. X.</given-names> <surname>Zhu</surname></string-name></person-group>, &#x201C;<article-title>HED-UNet: Combined segmentation and edge detection for monitoring the Antarctic coastline</article-title>,&#x201D; <source>IEEE Trans. Geosci. Remote Sens.</source>, vol. <volume>60</volume>, pp. <fpage>1</fpage>&#x2013;<lpage>14</lpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.1109/TGRS.2021.3064606</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Vaswani</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<chapter-title>Attention is all you need</chapter-title>,&#x201D; <article-title>presented at the 31st Conf. Neural Inform. Process. Syst. (NeurIPS 2017)</article-title>, <publisher-loc>Long Beach, CA, USA</publisher-loc>, <year>Dec. 4&#x2013;9, 2017</year>, vol. <volume>30</volume>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Canny</surname></string-name></person-group>, &#x201C;<article-title>A computational approach to edge detection</article-title>,&#x201D; <source>IEEE Trans. Pattern Anal. Mach. Intell.</source>, vol. <volume>PAMI-8</volume>, no. <issue>6</issue>, pp. <fpage>679</fpage>&#x2013;<lpage>698</lpage>, <year>1986</year>. doi: <pub-id pub-id-type="doi">10.1109/TPAMI.1986.4767851</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Xie</surname></string-name> and <string-name><given-names>Z.</given-names> <surname>Tu</surname></string-name></person-group>, &#x201C;<chapter-title>Holistically-nested edge detection</chapter-title>,&#x201D; <article-title>presented at the IEEE Int. Conf. Comput. Vis. (ICCV)</article-title>, <publisher-loc>Santiago, Chile</publisher-loc>, <year>Dec. 11&#x2013;18, 2015</year>, pp. <fpage>1395</fpage>&#x2013;<lpage>1403</lpage>. doi: <pub-id pub-id-type="doi">10.1109/ICCV.2015.164</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>X. S.</given-names> <surname>Poma</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Sappa</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Humanante</surname></string-name>, and <string-name><given-names>A.</given-names> <surname>Arbarinia</surname></string-name></person-group>, &#x201C;<article-title>Dense extreme inception network for edge detection</article-title>,&#x201D; <comment>2021, <italic>arXiv:2112.02250</italic></comment>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Arbelaez</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Maire</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Fowlkes</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Malik</surname></string-name></person-group>, &#x201C;<article-title>Contour detection and hierarchical image segmentation</article-title>,&#x201D; <source>IEEE Trans. Pattern Anal. Mach. Intell.</source>, vol. <volume>33</volume>, no. <issue>5</issue>, pp. <fpage>898</fpage>&#x2013;<lpage>916</lpage>, <year>2010</year>. doi: <pub-id pub-id-type="doi">10.1109/TPAMI.2010.161</pub-id>; <pub-id pub-id-type="pmid">20733228</pub-id></mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Kervadec</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Dolz</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>E.</given-names> <surname>Granger</surname></string-name>, and <string-name><given-names>I. B.</given-names> <surname>Ayed</surname></string-name></person-group>, &#x201C;<chapter-title>Bounding boxes for weakly supervised segmentation: Global constraints get close to full supervision</chapter-title>,&#x201D; <article-title>presented at the Med. Imaging Deep Learn. (MIDL)</article-title>, <publisher-loc>Montreal, QC, Canada</publisher-loc>, <year>Jul. 6&#x2013;8, 2020</year>, pp. <fpage>365</fpage>&#x2013;<lpage>381</lpage>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Lee</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Yi</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Shin</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Yoon</surname></string-name></person-group>, &#x201C;<chapter-title>BBAM: Bounding box attribution map for weakly supervised semantic and instance segmentation</chapter-title>,&#x201D; <article-title>presented at the IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR)</article-title>, <publisher-loc>Nashville, TN, USA</publisher-loc>, <year>Jun. 19&#x2013;25, 2021</year>, pp. <fpage>2643</fpage>&#x2013;<lpage>2652</lpage>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Bradski</surname></string-name></person-group>, &#x201C;<article-title>The OpenCV library</article-title>,&#x201D; <source>Dr. Dobb's J. Softw. Tools</source>, vol. <volume>25</volume>, no. <issue>11</issue>, pp. <fpage>120</fpage>&#x2013;<lpage>123</lpage>, <year>2000</year>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Tyle&#x010D;ek</surname></string-name> and <string-name><given-names>R.</given-names> <surname>&#x0160;&#x00E1;ra</surname></string-name></person-group>, &#x201C;<chapter-title>Spatial pattern templates for recognition of objects with regular structure</chapter-title>,&#x201D; <article-title>presented at the 35th German Conf. Pattern Recognit. (GCPR)</article-title>, <publisher-loc>Saarbr&#x00FC;cken, Germany</publisher-loc>, <year>Sep. 3&#x2013;6, 2013</year>, pp. <fpage>364</fpage>&#x2013;<lpage>374</lpage>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Li</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Jia</surname></string-name></person-group>, &#x201C;<chapter-title>Image partial blur detection and classification</chapter-title>,&#x201D; <article-title>presented at the IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR)</article-title>, <publisher-loc>Anchorage, AK, USA</publisher-loc>, <year>Jun. 23&#x2013;28, 2008</year>, pp. <fpage>1</fpage>&#x2013;<lpage>8</lpage>. doi: <pub-id pub-id-type="doi">10.1109/CVPR.2008.4587465</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Inoue</surname></string-name></person-group>, &#x201C;<article-title>Data augmentation by pairing samples for images classification</article-title>,&#x201D; <comment>2018, <italic>arXiv:1801.02929</italic></comment>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Yun</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Han</surname></string-name>, <string-name><given-names>S. J.</given-names> <surname>Oh</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Chun</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Choe</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Yoo</surname></string-name></person-group>, &#x201C;<chapter-title>CutMix: Regularization strategy to train strong classifiers with localizable features</chapter-title>,&#x201D; <article-title>presented at the IEEE/CVF Int. Conf. Comput. Vis. (ICCV)</article-title>, <publisher-loc>Seoul, Republic of Korea</publisher-loc>, <year>Oct. 27&#x2013;Nov. 2, 2019</year>, pp. <fpage>6023</fpage>&#x2013;<lpage>6032</lpage>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>He</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Ren</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Sun</surname></string-name></person-group>, &#x201C;<chapter-title>Deep residual learning for image recognition</chapter-title>,&#x201D; <article-title>presented at the IEEE Conf. Comput. Vis. Pattern Recognit. (CVPR)</article-title>, <publisher-loc>Las Vegas, NV, USA</publisher-loc>, <year>Jun. 26&#x2013;30, 2016</year>, pp. <fpage>770</fpage>&#x2013;<lpage>778</lpage>.</mixed-citation></ref>
</ref-list>
</back></article>