<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMES</journal-id>
<journal-id journal-id-type="nlm-ta">CMES</journal-id>
<journal-id journal-id-type="publisher-id">CMES</journal-id>
<journal-title-group>
<journal-title>Computer Modeling in Engineering &#x0026; Sciences</journal-title>
</journal-title-group>
<issn pub-type="epub">1526-1506</issn>
<issn pub-type="ppub">1526-1492</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">80624</article-id>
<article-id pub-id-type="doi">10.32604/cmes.2026.080624</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Robust Analog Gauge Reading via Virtual Point-Based Geometric Rectification and P2-YOLO-Pose</article-title>
<alt-title alt-title-type="left-running-head">Robust Analog Gauge Reading via Virtual Point-Based Geometric Rectification and P2-YOLO-Pose</alt-title>
<alt-title alt-title-type="right-running-head">Robust Analog Gauge Reading via Virtual Point-Based Geometric Rectification and P2-YOLO-Pose</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Lee</surname><given-names>Jaekyung</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Kim</surname><given-names>Youngjun</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Ko</surname><given-names>Byungsung</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Kim</surname><given-names>Taewon</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Park</surname><given-names>Jaeheon</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-6" contrib-type="author">
<name name-style="western"><surname>Lee</surname><given-names>Jiwon</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-7" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Kim</surname><given-names>Wonhee</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><email>whkim79@cau.ac.kr</email></contrib>
<aff id="aff-1"><label>1</label><institution>School of Energy Systems Engineering, Chung-Ang University</institution>, <addr-line>Seoul</addr-line>, <country>Republic of Korea</country></aff>
<aff id="aff-2"><label>2</label><institution>KEPCO Research Institute</institution>, <addr-line>Daejeon</addr-line>, <country>Republic of Korea</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Wonhee Kim. Email: <email>whkim79@cau.ac.kr</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>27</day><month>4</month><year>2026</year>
</pub-date>
<volume>147</volume>
<issue>1</issue>
<elocation-id>35</elocation-id>
<history>
<date date-type="received">
<day>13</day>
<month>02</month>
<year>2026</year>
</date>
<date date-type="accepted">
<day>30</day>
<month>03</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMES_80624.pdf"></self-uri>
<abstract>
<p>Automated reading of analog gauges in industrial environments is essential for predictive maintenance and safety monitoring. However, conventional computer vision approaches encounter two fundamental bottlenecks: polar unwrapping techniques induce severe nonlinear scaling distortions under oblique viewing angles and axis-aligned bounding boxes (AABBs) are geometrically inefficient for encapsulating high-aspect-ratio rotating needles. To overcome these limitations, this paper proposes a novel end-to-end framework that innovatively redefines gauge reading as a structural pose estimation task. We model each gauge as a topological five-keypoint skeleton (<inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>), and localize these landmarks using a customized P2-YOLO-Pose architecture. By integrating a high-resolution P2 feature layer (stride 4) while excising the macro-scale P5 layer, the network yields a 40% enhancement in small-gauge detection recall with a negligible (&#x003C;1%) frame-rate degradation. Furthermore, to address the intrinsic lack of salient vertices in circular dials, we introduce a Virtual Point (VP) generation algorithm. This algorithm exploits the point symmetry of the detected keypoints to autonomously synthesize four spatial correspondences, thereby enabling markerless, homography-based perspective rectification for corner-free objects. An adaptive control mechanism based on aspect ratio analysis (<inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>A</mml:mi><mml:mi>R</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mn>1.5</mml:mn></mml:math></inline-formula>) dynamically regulates the geometric warping to prevent algorithmic over-correction. Extensive evaluations on an 11,000-image field dataset acquired from an operational power data center demonstrate a Pose mAP50 of 99.45% and an mAP50-95 of 99.37%. Under severe vertical tilt conditions, the VP-based rectification curtails the absolute reading error from 3.5% to 0.6% compared to the uncorrected baseline, attaining measurement precision commensurate with physical ArUco marker-based ground truths. Operating in real-time at 25.9 FPS, the proposed system is currently deployed within an integrated inspection platform coupled with an autonomous quadruped robot (Boston Dynamics SPOT), facilitating reliable, perspective-invariant visual inspections across 10 distinct classes of analog gauges in an active industrial facility.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Analog gauge</kwd>
<kwd>deep learning</kwd>
<kwd>keypoint detection</kwd>
<kwd>geometric rectification</kwd>
<kwd>Industrial Internet of Things (IIoT)</kwd>
<kwd>pose estimation</kwd>
</kwd-group><funding-group>
<award-group id="awg1">
<funding-source>Korea Electric Power Corporation</funding-source>
<award-id>R25IA04</award-id>
</award-group>
</funding-group></article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>The uninterrupted monitoring of pivotal physical parameters, including pressure, temperature, and flow rate in both analog and digital formats, underpins the operational safety and systemic efficiency of modern industrial infrastructures, most notably power plants and power data management centers [<xref ref-type="bibr" rid="ref-1">1</xref>]. Notwithstanding the rapid proliferation of embedded digital sensors and the emergence of the Industrial Internet of Things (IIoT), external analog gauges remain ubiquitously deployed across diverse industrial sectors. This enduring presence is primarily attributed to their exceptional durability, autonomy from external power sources, and robust reliability in hazardous environments where digital sensors may succumb to electromagnetic interference or extreme temperatures [<xref ref-type="bibr" rid="ref-2">2</xref>]. Typical examples of widely used industrial gauges are shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Typical examples of analog gauges used in industrial facilities, illustrating diverse scale configurations and environmental conditions.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80624-fig-1.tif"/>
</fig>
<p>The inherent analog characteristics of these instruments induce a profound systemic gap; specifically, the absence of native data transmission capabilities restricts the acquisition of indicated physical values to visual inspection.</p>
<p>Consequently, industrial digitalization presents a systemic paradox: whereas the shift toward smart infrastructure is an industry-wide mandate, the persistent reliability of analog instruments in harsh environments creates a substantial bottleneck. Absent native digital interfaces, these gauges necessitate a continued reliance on manual inspection, which is a process characterized by heavy labor requirements and susceptibility to human error. This human-centric data acquisition cycle introduces significant temporal gaps between measurement and analysis, ultimately hindering the implementation of real-time monitoring and robust predictive maintenance systems.</p>
<sec id="s1_1">
<label>1.1</label>
<title>Challenges in Analog Gauge Reading</title>
<p>The persistent reliance on analog instruments creates a critical systemic bottleneck in the digitalization of industrial maintenance. Manual inspection routines, characterized by periodic physical patrols, are inherently labor-intensive and susceptible to human-induced inaccuracies. This dependency not only incurs high operational costs but also introduces significant temporal gaps between data acquisition and analysis, ultimately hindering the realization of real-time monitoring and advanced predictive maintenance systems [<xref ref-type="bibr" rid="ref-3">3</xref>].</p>
<p>To overcome these limitations, automatic gauge reading systems leveraging autonomous agents&#x2014;such as quadruped robots (Boston Dynamics SPOT [<xref ref-type="bibr" rid="ref-4">4</xref>]), Unmanned Aerial Vehicles(UAVs), and fixed CCTV cameras&#x2014;have been actively investigated [<xref ref-type="bibr" rid="ref-2">2</xref>]. In this context, computer vision (CV)-based automatic analog gauge reading has evolved rapidly from classical image processing to advanced deep learning techniques [<xref ref-type="bibr" rid="ref-5">5</xref>].</p>
<p>However, realizing field gauge reading through autonomous agents requires addressing several key challenges:<list list-type="bullet">
<list-item>
<p>Environmental variability: industrial sites present diverse conditions including non-uniform illumination, protective glass reflections, dust and fog interference, and complex backgrounds (piping, cables), all of which severely impair the stable operation of classical image processing techniques.</p></list-item>
<list-item>
<p>Viewing angle diversity: image capture via robots or CCTV cameras does not always occur from the frontal position, making acute oblique viewing angles inevitable. Under such conditions, circular dials appear as ellipses, and conventional polar unwrapping methods introduce nonlinear scaling errors.</p></list-item>
<list-item>
<p>Object representation limitations: gauge needles are elongated objects with an extremely small width-to-length ratio. Conventional Axis-Aligned Bounding Boxes (AABBs) cause the background noise ratio to surge when the needle rotates, impeding precise orientation learning.</p></list-item>
<list-item>
<p>Value conversion accuracy: obtaining the final reading requires precise conversion of the needle&#x2019;s visual position to a physical value (pressure, temperature), which in turn presupposes accurate distortion correction and faithful scale structure reconstruction.</p></list-item>
</list></p>
<p>To systematically address these challenges, this study focuses on two fundamental geometric limitations.</p>
</sec>
<sec id="s1_2">
<label>1.2</label>
<title>The Limitation of Axis-Aligned Bounding Boxes</title>
<p>The second challenge lies in a fundamental limitation of standard object detection frameworks. Previous studies often treat gauge reading as a vanilla object detection task [<xref ref-type="bibr" rid="ref-6">6</xref>], employing Axis-Aligned Bounding Boxes (AABBs) to localize the Region of Interest (ROI). While AABBs are effective for prominent block-shaped objects, they are structurally ill-suited for representing thin, rotating objects such as gauge needles.</p>
<p>As illustrated in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>, when a needle of length <italic>L</italic> and width <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>w</mml:mi></mml:math></inline-formula> (<inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>L</mml:mi><mml:mo>&#x226B;</mml:mo><mml:mi>w</mml:mi></mml:math></inline-formula>) rotates by angle <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula>, the enclosing AABB area expands substantially while the actual object area remains constant. This geometric mismatch causes the Intersection over Union (IoU) between the predicted and ground truth boxes to drop sharply, as the regression targets are inherently coupled with the pointer&#x2019;s spatial tilt. Mathematically, the AABB area and the resulting IoU can be approximated as:<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>A</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>a</mml:mi><mml:mi>b</mml:mi><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2248;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>L</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>cos</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x00D7;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>L</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>sin</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi><mml:mo>&#x2248;</mml:mo><mml:mfrac><mml:mrow><mml:mi>L</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mi>L</mml:mi><mml:mn>2</mml:mn></mml:msup><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>sin</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mi>cos</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>0</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Comparison of object representation methods for analog gauges. (<bold>a</bold>) Typical structure of an analog gauge. (<bold>b</bold>) AABB is efficient for axis-aligned needles. (<bold>c</bold>) As the needle rotates, the AABB expands to include substantial background noise, reducing IoU. (<bold>d</bold>) The proposed keypoint-based approach captures explicit geometric structure regardless of rotation.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80624-fig-2.tif"/>
</fig>
<p>At a <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msup><mml:mi>45</mml:mi><mml:mo>&#x2218;</mml:mo></mml:msup></mml:math></inline-formula> inclination, the ratio of object-relevant pixels to irrelevant environmental noise within the ROI reaches its global minimum, compelling the convolutional filters to extract features from a high-entropy region. This dilution of structural information introduces significant &#x201C;feature pollution&#x201D;, where the stochastic textures of the background (e.g., dial markings, shadows, or dust) dominate the feature maps. Consequently, the gradient signals during backpropagation become increasingly noisy as the spatial overlap decreases, hindering the convergence of orientation-sensitive layers. This structural inefficiency often results in a &#x201C;systemic decoupling&#x201D;, where the model achieves accurate gauge localization (high recall) but fails to maintain the precision required for fine-grained needle angle estimation.</p>
</sec>
<sec id="s1_3">
<label>1.3</label>
<title>The Challenge of Perspective Distortion</title>
<p>The second critical challenge concerns the frontal viewing angle assumption. Most existing systems rely on polar-to-Cartesian unwrapping or vector direction detection for gauge reading. The fundamental limitation of polar unwrapping is that it is valid only when the gauge appears as a perfect circle. In practical scenarios involving mobile robots or fixed CCTV cameras, gauges are frequently captured at oblique angles, causing the circular dial to appear elliptical. Applying standard polar coordinate transformation to elliptical images introduces severe nonlinear scaling errors, as non-uniform arc lengths on an ellipse are treated linearly.</p>
<p>Meanwhile, purely vector-based methods that detect only pointer direction lack the mechanisms to reconstruct the scale structure and convert directions into physical values. They also lack projective distortion correction, causing the detected needle direction to be severely distorted under oblique viewing angles.</p>
<p>In summary, the common limitations are:<list list-type="bullet">
<list-item>
<p>Reliance on circularity causing geometric errors under oblique capture.</p></list-item>
<list-item>
<p>Lack of physical value conversion mechanisms.</p></list-item>
<list-item>
<p>Inability to geometrically rectify circular objects lacking distinct corners.</p></list-item>
</list></p>
</sec>
<sec id="s1_4">
<label>1.4</label>
<title>Proposed Approach and Contributions</title>
<p>This study focuses on achieving high accuracy and robustness under extreme viewing angles and challenging illumination conditions. To simultaneously address the limitations identified above&#x2014;the nonlinearity of polar coordinate conversion, the absence of value conversion in vector methods, and the lack of corner-free rectification&#x2014;we redefine gauge reading not as a simple object detection task but as a <italic>structural pose estimation</italic> problem. Inspired by the Human Pose Estimation (HPE) paradigm, we propose a novel framework that treats each gauge as a skeleton composed of key structural points: five keypoints corresponding to the center, needle tip, and scale start/mid/end positions.</p>
<p>The main contributions of this paper are summarized as follows:<list list-type="bullet">
<list-item>
<p>A robust structural keypoint approach is introduced that defines gauges through structural relationships among specific points (center, needle tip, scale start, scale end). Unlike AABBs, this approach ensures robustness to thin needle geometries and partial occlusion while effectively excluding background noise for precise reading.</p></list-item>
<list-item>
<p>A high-resolution P2 architecture is proposed through a modified YOLOv11-Pose model that integrates a high-resolution P2 layer (stride 4). This architectural enhancement preserves fine spatial information, substantially improving the detection recall for small needles and fine scale markings that are often lost during the downsampling process of standard models.</p></list-item>
<list-item>
<p>A Virtual Point (VP)-based geometric rectification method is proposed, specifically designed for circular objects lacking distinct corners. By exploiting the point symmetry of detected keypoints, the algorithm automatically constructs four correspondence pairs and restores elliptically distorted gauges to a mathematically frontal circle via homography transformation. Unlike the polar coordinate transformations used by GAUREAD or Under Pressure, this approach enables precise angle-to-value conversion without nonlinear distortion even under oblique capture.</p></list-item>
</list></p>
<p>The remainder of this paper is organized as follows. <xref ref-type="sec" rid="s2">Section 2</xref> provides a systematic review of prior work on analog gauge reading, covering traditional image processing, deep learning approaches, pose estimation models, and the limitations of existing methods. <xref ref-type="sec" rid="s3">Section 3</xref> details the proposed methodology, including problem formulation, the virtual point-based geometric rectification algorithm, and the P2-YOLO-Pose architecture. <xref ref-type="sec" rid="s4">Section 4</xref> describes the data collection process and experimental setup in a real-world industrial setting. <xref ref-type="sec" rid="s5">Section 5</xref> presents evaluation metrics, quantitative and qualitative analyses, comparisons with state-of-the-art methods, and ablation study results. <xref ref-type="sec" rid="s6">Section 6</xref> provides an in-depth discussion of the strengths, limitations, failure cases, and practical deployment considerations. <xref ref-type="sec" rid="s7">Section 7</xref> presents use cases in industrial monitoring, smart manufacturing and IIoT, and the energy and utility sector. Finally, <xref ref-type="sec" rid="s8">Section 8</xref> summarizes the contributions and proposes future research directions.</p>
</sec>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>Automatic analog gauge reading is a core challenge in industrial automation. A broad spectrum of methodologies has been proposed, ranging from traditional image processing techniques to deep learning-based approaches, pose estimation model applications, and recent end-to-end frameworks. This section systematically categorizes prior work, analyzes the contributions and limitations of each approach, and clarifies the motivation and contributions of the present study.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Traditional Image Processing Methods for Gauge Reading</title>
<p>Early research on analog gauge reading relied primarily on classical computer vision techniques, employing handcrafted features to detect the structural elements of gauges.</p>
<p>The Hough Transform proposed by Duda and Hart [<xref ref-type="bibr" rid="ref-7">7</xref>] is a classical method for detecting lines and circles in images, and has been widely applied to detect the circular contours of gauge dials [<xref ref-type="bibr" rid="ref-8">8</xref>]. Alegria and Serra [<xref ref-type="bibr" rid="ref-9">9</xref>] extended this approach by extracting the center and radius of the dial using the Circle Hough Transform (CHT), computing the needle angle based on these parameters, and proposing the first automated gauge reading system. Their work is regarded as establishing the foundation of the automatic analog gauge reading field.</p>
<p>Canny&#x2019;s [<xref ref-type="bibr" rid="ref-10">10</xref>] edge detection algorithm has been used to extract the contours of needles and scales by detecting abrupt brightness changes, while Otsu&#x2019;s [<xref ref-type="bibr" rid="ref-11">11</xref>] thresholding technique has played a central role in separating the foreground (needle) from the background. Chi et al. [<xref ref-type="bibr" rid="ref-12">12</xref>] combined these preprocessing techniques into a complete pipeline: edge detection followed by Hough Transform for circular dial detection, binarization for needle segmentation, and angle computation for reading. Ma and Jiang [<xref ref-type="bibr" rid="ref-13">13</xref>] subsequently experimented with various preprocessing combinations based on similar principles to improve reading accuracy.</p>
<p>The camera calibration and multiple view geometry frameworks systematized by Zhang [<xref ref-type="bibr" rid="ref-14">14</xref>] and Hartley and Zisserman [<xref ref-type="bibr" rid="ref-15">15</xref>] provide the mathematical foundation for correcting projective distortion in gauge images. In particular, homography-based projective transformation, which estimates a view transformation matrix from planar correspondences, serves as the key tool for restoring obliquely captured gauges to frontal views.</p>
<p>While classical techniques offer computational efficiency and deterministic behavior, they possess inherent limitations. First, they are extremely sensitive to environmental variables such as non-uniform illumination, protective glass reflections, dust, and complex backgrounds (piping, cables), requiring manual parameter tuning for the Hough Transform on a per-environment basis. Second, binarization-based needle detection produces significant errors when contrast between the background and needle colors is insufficient&#x2014;for example, a black needle on a dark dial face. Third, most of these methods assume frontal capture and lack automatic correction mechanisms for the elliptical distortion caused by oblique viewing angles. These limitations substantially restrict practical deployment in uncontrolled industrial settings.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Deep Learning Approaches for Analog Instrument Recognition</title>
<p>Advances in deep learning have significantly alleviated the environmental sensitivity problems of traditional methods. Through large-scale data and learnable feature extraction, deep learning-based systems demonstrate more robust gauge detection and reading performance across diverse conditions. This subsection provides a detailed analysis of recently proposed methods.</p>
<p>Milana et al. [<xref ref-type="bibr" rid="ref-2">2</xref>] proposed GAUREAD, an end-to-end gauge reading system comprising YOLOv5-based gauge detection, Circle Hough Transform for circular dial detection, ellipse fitting for shape estimation, and polar-to-Cartesian unwrapping for scale/needle detection. GAUREAD achieved a processing time of 800 ms on an NVIDIA Jetson Nano, demonstrating the feasibility of edge-device deployment. However, the system exhibits a reading error of 3% within a <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msup><mml:mi>20</mml:mi><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> viewing angle from the frontal position, which escalates to 9% at <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msup><mml:mi>50</mml:mi><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula>. This degradation stems from the structural limitation that polar unwrapping treats non-uniform arc lengths on an ellipse linearly. Additionally, the Circle Hough Transform can fail on gauges with unclear circular contours, such as semi-circular or fan-shaped gauges.</p>
<p>Dong et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] proposed the Vector Detection Network (VDN), which models gauge pointers as two-dimensional vectors. In VDN, the initial point of the vector corresponds to the needle tip, and the direction follows tail-to-tip. The network estimates a confidence map to determine the initial point (peak pixel) and extracts direction components from a two-layer scalar map at each peak. Evaluated on the self-constructed Pointer-10K dataset, VDN demonstrated strong generalization performance and real-time processing speed across various gauge forms, including circular, semi-circular, and multi-pointer types. However, VDN detects only pointer direction without reconstructing scale structure (start point, end point, range) or providing a mechanism to convert direction information into physical values (pressure, temperature). Furthermore, the absence of projective distortion correction means that the needle direction itself becomes distorted under oblique viewing angles.</p>
<p>Most recently, Reitsma et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] proposed the Under Pressure framework at ETH Zurich ASL. The system follows a step-by-step pipeline of gauge detection, notch detection with ellipse fitting, needle segmentation, scale marker recognition, and unit extraction. A notable advantage is that each stage&#x2019;s potential failures can be diagnosed in an interpretable manner. The system operates without prior knowledge of gauge type or scale range and provides automatic unit extraction. Experimental results achieved relative error below 2%. However, this performance was primarily measured under near-frontal viewing angles, and robustness under obscure notch conditions or severe projective distortion remains unvalidated.</p>
<p>Leon-Alcazar et al. [<xref ref-type="bibr" rid="ref-3">3</xref>] proposed training robust reading models using large-scale synthetic data for diverse gauge forms and conditions. While synthetic data presents a promising approach to reducing data collection costs, domain gaps between synthetic and real-world field data persist, particularly in reproducing subtle geometric distortions and site-specific interference (glass reflections, condensation). Tian et al. [<xref ref-type="bibr" rid="ref-5">5</xref>] proposed GaugeTracker, a hybrid system combining template matching with deep learning, achieving improved reading precision but lacking flexibility for gauge types without predefined templates or severely distorted images. The Programmable Gradient Information (PGI) concept introduced by Wang et al. [<xref ref-type="bibr" rid="ref-18">18</xref>] in YOLOv9 enhances feature learning efficiency for small object detection and represents an architectural advancement applicable to gauge reading technologies.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Pose Estimation Models in Visual Measurement</title>
<p>Human Pose Estimation (HPE) is one of the most actively researched topics in computer vision, aiming to estimate joint positions and reconstruct skeletal structures from images. This study innovatively applies the HPE paradigm to the industrial measurement domain.</p>
<p>Pose estimation is broadly classified into two paradigms. Top-down approaches first detect each object and then estimate keypoints within each detected instance, while bottom-up approaches first detect all keypoints and subsequently group them into individual objects. YOLO-Pose [<xref ref-type="bibr" rid="ref-19">19</xref>] is a representative model that integrates the top-down approach into a single network, simultaneously performing object detection and keypoint regression at real-time speed. This model introduces the Object Keypoint Similarity (OKS) loss function to incorporate structural relationships among keypoints into the learning process. He et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] demonstrated with Mask Regions with Convolutional Neural Networks features (R-CNN) that P2-level feature maps from the Feature Pyramid Network (FPN) are essential for precise keypoint localization, experimentally establishing the importance of high-resolution feature maps.</p>
<p>The Feature Pyramid Network (FPN) proposed by Lin et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] is a key architecture that hierarchically fuses multi-scale feature maps to effectively detect objects of various sizes. YOLO-family models typically use a three-level pyramid comprising P3 (stride 8), P4 (stride 16), and P5 (stride 32). However, for small objects such as gauge needles and fine scale markings, P3 may not preserve sufficient spatial resolution. Adding a P2 (stride 4) layer provides 4<inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> higher spatial resolution, though the trade-off between computational overhead and memory consumption must be considered. This study achieves a balance between high-resolution keypoint detection performance and computational efficiency through a structural optimization that integrates the P2 layer while simultaneously removing the P5 layer.</p>
<p>The keypoint-based structural recognition paradigm established in HPE extends naturally to the industrial measurement domain. Just as human joints define the physical structure of arms, legs, and torso, gauge keypoints (center, needle tip, scale start/mid/end) define the geometric structure of the circular dial-needle system. Based on this analogy, the proposed method encodes the structural relationships among five key gauge keypoints into the loss function and maximizes relative positional accuracy through OKS-based training. This enables precise needle direction estimation and complete scale structure reconstruction that were impossible with AABB-based methods. Furthermore, industrial defect detection benchmarks such as MVTec AD provided by Bergmann et al. [<xref ref-type="bibr" rid="ref-22">22</xref>] underscore the importance of rigorous evaluation methodologies in industrial visual inspection, a principle that this study applies to its experimental design.</p>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Limitations of Existing Methods</title>
<p>Synthesizing the common limitations of the prior work reviewed above, current automatic analog gauge reading technology faces three critical open challenges.</p>
<p>Most methods, including GAUREAD [<xref ref-type="bibr" rid="ref-2">2</xref>] and Under Pressure [<xref ref-type="bibr" rid="ref-17">17</xref>], perform correction based on polar coordinate transformation (polar unwrapping) or ellipse fitting. However, these approaches assume circular or quasi-circular dials, and nonlinear errors increase rapidly as elliptical distortion from oblique viewing intensifies. GAUREAD reports errors of 3% within <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msup><mml:mi>20</mml:mi><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> and 9% at <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msup><mml:mi>50</mml:mi><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula>; such angle-dependent error growth constitutes a critical constraint in real-world robotic patrol scenarios, where it is infeasible for the robot to always stop directly in front of each gauge.</p>
<p>While VDN [<xref ref-type="bibr" rid="ref-16">16</xref>] provides a flexible and generalizable method for pointer direction detection, the pipeline does not include mechanisms to convert detected directions into actual physical values (pressure, temperature). Since the ultimate objective of gauge reading in industrial settings is to obtain quantitative measurements, directional information alone has limited practical utility. Value conversion requires knowledge of scale start points, end points, and the range between them&#x2014;structural information that VDN does not estimate.</p>
<p>The optimal mathematical solution to ensure measurement invariance regardless of the camera viewpoint is homography rectification; however, this requires securing at least four discrete point-to-point correspondence pairs. Unlike rectangular objects with identifiable salient vertices, circular gauges are intrinsically &#x201C;corner-free&#x201D; objects lacking prominent corners. Previous studies attempted to bypass this limitation through curve extraction via CHT or shallow notch matching, but the detection reliability of such alternative features drops precipitously in real-world environments characterized by partial occlusion from piping or irregular notch patterns across manufacturers. This chronic inability to autonomously establish reliable correspondences without external fiducial markers remains the most significant algorithmic barrier to achieving flawless planar rectification.</p>
<p>In conclusion, realizing autonomous inspection in uncontrolled, unstructured industrial environments requires completely departing from the fragmented approaches that treat analog gauges merely as simple bounding boxes or isolated line segments (vectors). To concurrently resolve the three major open challenges that existing methodologies have failed to overcome, (1) the nonlinearity of oblique distortion, (2) the absence of quantitative value conversion, and (3) the inability to rectify corner-free objects, a novel paradigm organically integrating structural topology estimation and geometric rectification is imperative.</p>
<p>Accordingly, in the subsequent <xref ref-type="sec" rid="s3">Section 3</xref> (Proposed Methodology), this study details our uniquely integrated end-to-end framework. This framework seamlessly connects the extraction of keypoint skeletons based on high-resolution P2-YOLO-Pose, the autonomous generation of Virtual Points (VPs) for corner-free objects by mathematically leveraging point symmetry to perform homography rectification, and error-free quantitative value conversion within a linear coordinate system completely devoid of projective distortion.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Proposed Methodology</title>
<sec id="s3_1">
<label>3.1</label>
<title>Problem Definition and System Overview</title>
<sec id="s3_1_1">
<label>3.1.1</label>
<title>Problem Formulation</title>
<p>The analog gauge reading problem is formally defined as estimating the true physical value <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>l</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> indicated by a gauge from a raw 2D input image <italic>I</italic>. Mathematically, this corresponds to optimizing a composite mapping function <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>f</mml:mi></mml:math></inline-formula> parameterized by learnable weights <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mi mathvariant="normal">&#x0398;</mml:mi></mml:math></inline-formula>:<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>l</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>I</mml:mi><mml:mo>;</mml:mo><mml:mi mathvariant="normal">&#x0398;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow><mml:mo>&#x2218;</mml:mo><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow><mml:mo>&#x2218;</mml:mo><mml:mrow><mml:mi>&#x1D49F;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>I</mml:mi><mml:mo>;</mml:mo><mml:mi mathvariant="normal">&#x0398;</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mrow><mml:mi>&#x1D49F;</mml:mi></mml:mrow></mml:math></inline-formula> denotes the structural pose detection function that extracts the topological keypoint skeleton of the gauge, <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mrow><mml:mi>&#x211B;</mml:mi></mml:mrow></mml:math></inline-formula> represents the geometric rectification function that neutralizes perspective distortion via homography, and <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mrow><mml:mi>&#x1D4B1;</mml:mi></mml:mrow></mml:math></inline-formula> is the deterministic vector-based value computation function. This tripartite formulation is explicitly designed to resolve the geometric bottlenecks of existing models: the spatial constraints of bounding boxes, the nonlinear scaling errors of polar unwrapping, and the lack of salient vertices in circular dials.</p>
</sec>
<sec id="s3_1_2">
<label>3.1.2</label>
<title>Overall System Architecture and Workflow</title>
<p>The overall architecture of the proposed system is illustrated in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>. The system consists of three major stages:<list list-type="bullet">
<list-item>
<p>Stage 1&#x2014;High-Resolution Keypoint Extraction via P2-YOLO-Pose: the P2-enhanced YOLOv11-Pose model simultaneously extracts five keypoints (<inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) and bounding boxes from the input image.</p>
</list-item>
<list-item>
<p>Stage 2&#x2014;Virtual Point-Based Adaptive Geometric Rectification: after multi-stage validation of detected keypoints, virtual points are generated using the point symmetry principle, and the distorted elliptical gauge is restored to a frontal circle via homography transformation.</p></list-item>
<list-item>
<p>Stage 3&#x2014;Vector-Based Value Computation in Canonical Metric Space: vector-based angle calculation and a circular distance function are applied in the rectified coordinate system to convert the needle position into a physical value.</p></list-item>
</list></p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>System architecture of the proposed analog gauge reading framework, illustrating the complete pipeline from keypoint detection through geometric rectification to final value computation.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80624-fig-3.tif"/>
</fig>
<p>The operational workflow proceeds as follows. An autonomous robot (Boston Dynamics SPOT) approaches inspection targets along predefined waypoints and acquires high-resolution images through its mounted optical payload. The captured images are processed by P2-YOLO-Pose for keypoint extraction, and only data passing multi-stage validity verification (confidence, structural completeness, physical constraints) proceeds to subsequent processing. Virtual points are generated from verified keypoints, and the aspect ratio (AR) analysis determines whether rectification is performed. Homography rectification is applied when AR <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mo>&#x2264;</mml:mo></mml:math></inline-formula> 1.5; otherwise, values are computed directly from the original image when AR &#x003E; 1.5.</p>
</sec>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Stage 1: High-Resolution Keypoint Extraction via P2-YOLO-Pose</title>
<p>To establish a robust geometric foundation for subsequent perspective correction, precise pixel-level localization of the gauge components is paramount. Therefore, we shift the detection paradigm from regional bounding boxes to topological keypoints.</p>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>Keypoint Detection for Gauge Needle and Scale</title>
<p>We mathematically model the analog gauge as a rigid topological skeleton comprising five semantically distinct keypoints: <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mi>K</mml:mi><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>. The semantic role of each keypoint is explicitly assigned to bypass the noise-susceptible nature of conventional AABBs:<list list-type="bullet">
<list-item>
<p>Scale points: <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>&#x2014;corresponding to the physical minimum, midpoint, and maximum positions on the gauge dial. These three points establish the reference for the angular span of the reading range and serve as the foundational data for virtual point generation.</p></list-item>
<list-item>
<p>Rotation center: <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>&#x2014;represents the mechanical pivot of the needle. It functions as the absolute geometric origin for all trigonometric computations and the invariant anchor for point-symmetric reflection.</p></list-item>
<list-item>
<p>Needle tip: <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>&#x2014;denotes the terminal extremity of the indicator needle that dictates the temporal measured value, demanding the highest sub-pixel localization precision within the network.</p></list-item>
</list></p>
<p>This keypoint-based approach differs from AABB-based detection in three fundamental ways. First, <italic>structural context</italic>: the model infers the needle tip position in relation to the center and scale points, enabling robust localization even for thin needles or under partial occlusion. Second, <italic>background noise suppression</italic>: by focusing on specific coordinates rather than the entire bounding box, background elements such as piping and cables are effectively ignored. Third, <italic>geometric robustness</italic>: the skeleton formed by five keypoints provides complete geometric information for homography rectification.</p>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>YOLO-Pose Architecture with P2 Feature Layer Enhancement</title>
<p>YOLO-Pose [<xref ref-type="bibr" rid="ref-19">19</xref>] is a real-time pose estimation framework that simultaneously performs object detection and keypoint regression in a single network. Unlike conventional top-down approaches, it does not require a separate human detector and introduces the Object Keypoint Similarity (OKS) loss function to incorporate structural relationships among keypoints into the learning process.</p>
<p>This study adopts Ultralytics&#x2019; YOLOv11 [<xref ref-type="bibr" rid="ref-23">23</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>] as the base architecture. YOLOv11 consists of Backbone (feature extraction), Neck (multi-scale feature fusion), and Head (detection and keypoint regression), with the Neck combining Feature Pyramid Network (FPN) [<xref ref-type="bibr" rid="ref-21">21</xref>] and Path Aggregation Network (PAN) structures for efficient multi-scale information integration.</p>
<p>By default, YOLOv11-Pose outputs a three-level pyramid comprising P3 (stride 8), P4 (stride 16), and P5 (stride 32) [<xref ref-type="bibr" rid="ref-21">21</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>]. However, for the few-pixel-wide needles and fine scale markings required in gauge reading, spatial information may not be sufficiently preserved even at P3 (stride 8). He et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] demonstrated with Mask R-CNN that high-resolution feature maps are essential for precise keypoint localization; this study applies this principle to the industrial measurement domain by introducing the P2 layer.</p>
<p>The precision of analog gauge reading depends directly on the accurate localization of the needle tip and scale markings. Since the minimum stride of standard YOLOv11 is 8, the P3 feature map resolution is only 160 <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 160 when the input image is 1280 <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 1280 pixels. At this resolution, a needle only a few pixels wide spans 1&#x2013;2 grid cells, making precise keypoint localization difficult.</p>
<p>This study adds a P2 layer with stride 4, doubling the feature map resolution to 320 <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 320. This enables needle tip regression on a 4<inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> finer grid, significantly reducing keypoint localization error.</p>
<p>Simultaneously, since the detection of large objects (buildings, background elements) is unnecessary for gauge reading, the P5 (stride 32) layer is removed. This structural optimization (P2 addition &#x002B; P5 removal) provides the following benefits:<list list-type="bullet">
<list-item>
<p>High-resolution keypoint detection: the P2 layer preserves spatial information for fine structures (needles, scale markings)</p></list-item>
<list-item>
<p>Computational efficiency: removing P5 offsets the computational overhead introduced by adding P2</p></list-item>
<list-item>
<p>Domain optimization: elimination of large-object detection capability focuses the model on small-to-medium objects relevant to gauge reading</p></list-item>
</list></p>
</sec>
<sec id="s3_2_3">
<label>3.2.3</label>
<title>Loss Functions and Optimization Strategy</title>
<p>The P2-YOLO-Pose model is trained using the following multi-task loss function:<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mi>o</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mi>C</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>f</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>D</mml:mi><mml:mi>F</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>O</mml:mi><mml:mi>K</mml:mi><mml:mi>S</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p>The role of each component is as follows:<list list-type="bullet">
<list-item>
<p><inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>: Complete Intersection over Union loss, which simultaneously regresses the position, size, and aspect ratio of bounding boxes. Compared to standard IoU, it additionally considers center distance and aspect ratio consistency, improving convergence speed and accuracy.</p></list-item>
<list-item>
<p><inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mi>C</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>: Binary Cross-Entropy classification loss for learning gauge class probabilities.</p></list-item>
<list-item>
<p><inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>D</mml:mi><mml:mi>F</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>: Distribution Focal Loss, which represents bounding box coordinates as discrete distributions, enhancing flexibility and precision compared to single-point regression.</p></list-item>
<list-item>
<p><inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>O</mml:mi><mml:mi>K</mml:mi><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>: Object Keypoint Similarity-based keypoint loss that maximizes the similarity between predicted and ground truth keypoints. OKS considers per-keypoint scale, evaluating precision relative to object size rather than absolute pixel error.</p></list-item>
<list-item>
<p><inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>: Keypoint objectness loss that learns the visibility of each keypoint, providing the basis for judgment under partial occlusion conditions.</p></list-item>
</list></p>
<p>A pivotal architectural decision in our optimization strategy is the imposition of an asymmetric weight distribution, where the Object Keypoint Similarity (OKS) loss weight (<inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) is prioritized substantially over the bounding box loss weight (<inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mi>o</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>). In visual measurement tasks, a multi-pixel deviation in a bounding box is functionally negligible; conversely, a sub-pixel coordinate perturbation of the needle tip (<inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) directly induces fatal angular translation errors. Thus, this formulation forces the network&#x2019;s gradient updates to overwhelmingly prioritize structural keypoint fidelity.</p>
</sec>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Stage 2: Virtual Point-Based Geometric Rectification</title>
<sec id="s3_3_1">
<label>3.3.1</label>
<title>Perspective Distortion Correction</title>
<p>Deep learning-based pose estimation is susceptible to hallucinating coordinates under extreme specular reflections or severe partial occlusions. Prior to executing geometric transformations, the detected skeleton undergoes rigorous deterministic filtering to ensure physical plausibility:<list list-type="bullet">
<list-item>
<p>Confidence Threshold: Detections with output confidence below a preset threshold are eliminated to exclude background noise and false positives.</p></list-item>
<list-item>
<p>Topological Completeness: The inference is rejected if any of the five cardinal keypoints are undetected, as incomplete skeletons preclude exact geometric reconstruction.</p></list-item>
<list-item>
<p>Radius Ratio Consistency: Let <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msub><mml:mi>d</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>e</mml:mi></mml:msub></mml:math></inline-formula> denote the Euclidean radial distances from <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> to the respective scale boundaries <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. Under valid projection assumptions, these radii must satisfy a geometric concentricity tolerance:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mrow><mml:mtext>if</mml:mtext></mml:mrow><mml:mspace width="1em" /><mml:mfrac><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>e</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>s</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>m</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>e</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac><mml:mo>&#x003C;</mml:mo><mml:mn>0.4</mml:mn><mml:mspace width="1em" /><mml:mo stretchy="false">&#x21D2;</mml:mo><mml:mspace width="1em" /><mml:mrow><mml:mtext>reject</mml:mtext></mml:mrow></mml:math></disp-formula>Since the three radii should theoretically be equal in a circular gauge, a ratio below 0.4 indicates physically implausible keypoint locations (severe deformation or hallucination). This threshold was established from field data analysis showing that legitimate gauges exhibit minimum ratios above 0.45, while erroneous detections mostly fall below 0.3.</p></list-item>
<list-item>
<p>Boundary Constraints: Detections where <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is located within a marginal threshold (<inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:mi>&#x03B4;</mml:mi><mml:mo>=</mml:mo><mml:mn>10</mml:mn></mml:math></inline-formula> pixels) of the image boundaries are discarded:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mrow><mml:mtext>if</mml:mtext></mml:mrow><mml:mspace width="1em" /><mml:msub><mml:mi>x</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x003C;</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:mtext>or</mml:mtext></mml:mrow><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>x</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:mtext>or</mml:mtext></mml:mrow><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>y</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x003C;</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:mtext>or</mml:mtext></mml:mrow><mml:mtext>&#x00A0;</mml:mtext><mml:msub><mml:mi>y</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mspace width="1em" /><mml:mo stretchy="false">&#x21D2;</mml:mo><mml:mspace width="1em" /><mml:mrow><mml:mtext>reject</mml:mtext></mml:mrow><mml:mspace width="1em" /><mml:mo stretchy="false">(</mml:mo><mml:mi>&#x03B4;</mml:mi><mml:mo>=</mml:mo><mml:mn>10</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula>This is because symmetric point generation from a truncated gauge projects outside the image boundary, causing numerical instability in the homography transformation.</p></list-item>
</list></p>
</sec>
<sec id="s3_3_2">
<label>3.3.2</label>
<title>Autonomous Virtual Point (VP) Generation via Point Symmetry</title>
<p>Computing a planar homography matrix <italic>M</italic> mathematically requires a minimum of four distinct correspondence pairs. However, circular gauges are intrinsically &#x201C;corner-free.&#x201D; To establish reliable correspondences autonomously without relying on ambiguous curve features or unpredictable notches, we exploit the rotational invariance of the gauge center. By mapping the verified scale boundaries <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> through point-symmetric reflection across <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, two mathematically sound Virtual Points (VPs) are synthesized:<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:msubsup><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msubsup><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p>The resulting discrete spatial set <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>r</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> forms an intrinsic &#x201C;rhombus grid&#x201D; parallel to the physical dial plane. This synthesis circumvents the absence of salient vertices by algorithmically inferring exact geometric correspondences.</p>
</sec>
<sec id="s3_3_3">
<label>3.3.3</label>
<title>Adaptive Rectification Control via Aspect Ratio Analysis</title>
<p>Indiscriminate application of forced perspective warping to extreme oblique angles induces aggressive pixel interpolation (stretching), which corrupts fine needle morphology. To mitigate algorithm-induced over-correction, an adaptive geometric control logic evaluates the Aspect Ratio (<italic>AR</italic>) of the synthesized rhombus grid:<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mi>A</mml:mi><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mo>,</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mo>,</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>A high AR indicates that the angle between the camera and the gauge plane is very acute, causing the four correspondence points to converge toward collinearity.</p>
<p>An AR threshold of <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mrow><mml:mo>&#x003E;</mml:mo></mml:mrow><mml:mn>1.5</mml:mn></mml:math></inline-formula> is established as the criterion for bypassing rectification. This value is justified as follows:<list list-type="bullet">
<list-item>
<p>Geometric rationale: AR <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mo>=</mml:mo><mml:mn>1.5</mml:mn></mml:math></inline-formula> corresponds to a 3:2 major-to-minor axis ratio, equivalent to a camera capturing the gauge plane from approximately <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msup><mml:mn>48.2</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula>. Beyond this angle, the condition number of the homography matrix increases sharply, and numerical stability of the resulting matrix cannot be guaranteed.</p></list-item>
<list-item>
<p>Experimental rationale: forced warping at AR <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mo>&#x003E;</mml:mo><mml:mn>1.5</mml:mn></mml:math></inline-formula> induces pixel stretching&#x2014;expanding a limited number of pixels into a larger space&#x2014;which was observed to degrade the fine angular information of the needle.</p></list-item>
</list></p>
<p>Consequently, the algorithm adopts an adaptive strategy: homography rectification is performed when AR <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:mo>&#x2264;</mml:mo><mml:mn>1.5</mml:mn></mml:math></inline-formula>, and values are computed directly from the original image when AR <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mo>&#x003E;</mml:mo><mml:mn>1.5</mml:mn></mml:math></inline-formula>.</p>
</sec>
<sec id="s3_3_4">
<label>3.3.4</label>
<title>Homography Transformation</title>
<p>From the validated keypoints and generated virtual points, the homography matrix between the source and target coordinate systems is computed.</p>
<p>The source coordinates <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>r</mml:mi><mml:mi>c</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are the rhombus grid <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msubsup><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:msubsup><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> in the actual image, while the target coordinates <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:msub><mml:mi>P</mml:mi><mml:mrow><mml:mi>d</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are designed so that the gauge forms a perfect frontal circle in the restored image. Specifically, the target coordinates are symmetrically placed on a circle of radius <italic>R</italic> while preserving the span angle between scale points:<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>p</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>arccos</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>v</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mi>s</mml:mi></mml:msub><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>v</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mi>e</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>v</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mi>s</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>v</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mi>e</mml:mi></mml:msub><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msub><mml:mrow><mml:mover><mml:mi>v</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mi>s</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msub><mml:mrow><mml:mover><mml:mi>v</mml:mi><mml:mo stretchy="false">&#x2192;</mml:mo></mml:mover></mml:mrow><mml:mi>e</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula>where <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>p</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the angular extent of the arc between the scale start and end points. The target <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are placed symmetrically at <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>p</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:math></inline-formula> from a reference angle, which is set to <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:math></inline-formula> (upper) or <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:mo>+</mml:mo><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:math></inline-formula> (lower) depending on the gauge orientation (arc-up or arc-down).</p>
<p>The homography matrix <italic>M</italic> between the two correspondence point sets is defined in homogeneous coordinates as:<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mi>i</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mi>M</mml:mi><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mi>i</mml:mi></mml:msub><mml:mspace width="1em" /><mml:mo stretchy="false">&#x21D2;</mml:mo><mml:mspace width="1em" /><mml:mrow><mml:mo>[</mml:mo><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:msubsup><mml:mi>x</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup><mml:msubsup><mml:mi>y</mml:mi><mml:mi>i</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mtable columnalign="center center center" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mn>11</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mn>12</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mn>13</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mn>21</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mn>22</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mn>23</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mn>31</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mn>32</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mn>33</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable><mml:mo>]</mml:mo></mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>1</mml:mn></mml:mtd></mml:mtr></mml:mtable><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p><italic>M</italic> has 8 independent degrees of freedom and is uniquely determined by the 8 linear equations provided by 4 correspondence pairs. The Direct Linear Transformation (DLT) algorithm is used for matrix computation.</p>
<p>Unlike simple cropping, homography-based restoration mathematically compensates for the &#x201C;asymmetric displacement of scale markings&#x201D; caused by the viewing angle. Under lateral capture, scale markings closer to the lens appear compressed while those farther away appear stretched; the matrix <italic>M</italic> reverses this nonlinear pixel density variation, realigning all scale intervals to be proportional to their actual angles. This process is equivalent to restoring the distorted rhombus grid to a rectangular structure, recovering the orthogonality of coordinate axes and ensuring the geometric linearity required for subsequent vector-based angle computation.</p>
</sec>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Stage 3: Coordinate Transformation and Value Calculation</title>
<p>Once the geometric integrity of the dial is mathematically restored (or conditionally bypassed when <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mi>A</mml:mi><mml:mi>R</mml:mi><mml:mo>&#x003E;</mml:mo><mml:mn>1.5</mml:mn></mml:math></inline-formula>), the ultimate conversion from visual data to physical quantities is executed within this distortion-free linear coordinate system.</p>
<sec id="s3_4_1">
<label>3.4.1</label>
<title>Coordinate Translation and Circular Distance Metric</title>
<p>The top-left origin of the standard digital image coordinate system is translated to the gauge rotation center (<inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>), and the <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:mi>y</mml:mi></mml:math></inline-formula>-axis is inverted to conform to standard Cartesian and trigonometric conventions. The absolute orientation angle <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula> for any valid target point <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is extracted via:<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mi>&#x03B8;</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext>atan2</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>y</mml:mi><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mi>x</mml:mi><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi>c</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x22C5;</mml:mo><mml:mfrac><mml:mn>180</mml:mn><mml:mi>&#x03C0;</mml:mi></mml:mfrac></mml:math></disp-formula></p>
<p>The computed angles are normalized to the range <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:mo stretchy="false">[</mml:mo><mml:msup><mml:mn>0</mml:mn><mml:mo>&#x2218;</mml:mo></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mn>360</mml:mn><mml:mo>&#x2218;</mml:mo></mml:msup><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> using modular arithmetic. Because analog dials frequently traverse the <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:msup><mml:mn>360</mml:mn><mml:mo>&#x2218;</mml:mo></mml:msup></mml:math></inline-formula> boundary (e.g., an abrupt transition from <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:msup><mml:mn>350</mml:mn><mml:mo>&#x2218;</mml:mo></mml:msup></mml:math></inline-formula> to <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:msup><mml:mn>10</mml:mn><mml:mo>&#x2218;</mml:mo></mml:msup></mml:math></inline-formula>), simple arithmetic subtraction yields unphysical discontinuities. We circumvent this by defining a continuous, clockwise relative circular distance function:<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>a</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>b</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mspace width="1em" /><mml:mo stretchy="false">(</mml:mo><mml:mi>mod</mml:mi><mml:mspace width="0.333em" /><mml:mn>360</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
</sec>
<sec id="s3_4_2">
<label>3.4.2</label>
<title>Final Physical Value Interpolation</title>
<p>Given the prerequisite that homography transformation has strictly restored geometric linearity&#x2014;where equivalent angular segments correctly correspond to uniform physical scale ticks&#x2014;the definitive continuous physical reading <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>l</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is deterministically interpolated using the pre-registered physical operational bounds <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> of the specific gauge:<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>u</mml:mi><mml:mi>l</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mfrac><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mi>&#x03B8;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<p>Because this interpolation occurs entirely within the mathematically rectified planar space, the resulting measurement is strictly immunized against perspective-induced nonlinear errors, seamlessly translating unstructured visual capture into an exact quantitative diagnostic.</p>
</sec>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiments</title>
<p>This section describes the experiments conducted to validate the performance and field applicability of the proposed system. The data collection process reflecting diverse environmental variables in a real industrial setting is presented first, followed by the characteristics of the constructed dataset. Subsequently, the training configuration for the P2-YOLO-Pose model and the rationale for each hyperparameter selection are provided. Finally, the ArUco marker-based control group experimental design for objectively verifying the accuracy of the geometric rectification algorithm is described.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Data Collection and Dataset Construction</title>
<sec id="s4_1_1">
<label>4.1.1</label>
<title>Experimental Environment</title>
<p>Field experiments were conducted in the underground infrastructure facilities of an operational power data center for deep learning model training and validation. The experimental sites were selected from facilities subject to regular manual inspections to ensure stable data center operation. The data center&#x2019;s underground infrastructure comprises 17 patrol inspection compartments, of which 9 are accessible to robots. The total number of inspection points is approximately 1200, of which roughly 670 were deemed accessible via robotic inspection. The machine rooms, where analog gauges are most extensively distributed, fire extinguisher facilities, emergency diesel generators, reservoirs, geothermal systems, and drainage pumps. The dataset was constructed from 10 types of analog gauges (7 pressure gauges, 1 ammeter, 1 thermometer, and 1 thermo-hygrometer) found in these facilities (<xref ref-type="fig" rid="fig-4">Fig. 4</xref>).</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Data collection environment and target analog gauge types. (<bold>a</bold>) Fire extinguisher pressure gauge; (<bold>b</bold>) drainage pump room pressure gauge; (<bold>c</bold>) geothermal recovery thermometer; (<bold>d</bold>) pressure gauge Type A; (<bold>e</bold>) pressure gauge Type B; (<bold>f</bold>) pressure gauge Type C.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80624-fig-4.tif"/>
</fig>
<p>Fire extinguisher pressure gauges, which constitute the largest proportion of data center infrastructure inspections, are located near fire extinguisher signage. Fire extinguishers are individually distributed, and the pressure gauges attached to agent storage and actuation vessels are extremely small, demanding precise small-object detection capability. Additionally, numerous pumps related to geothermal heat pump systems and water supply pressurization are present, requiring periodic pressure inspection of pumps and piping under conditions where visual occlusion frequently occurs due to complex piping and cylinder configurations.</p>
</sec>
<sec id="s4_1_2">
<label>4.1.2</label>
<title>Mobile Platform and Image Acquisition</title>
<p>A Boston Dynamics SPOT quadruped robot [<xref ref-type="bibr" rid="ref-4">4</xref>] was employed as the mobile platform for data collection, capable of stable navigation within complex and confined underground facilities. A camera module supporting high-resolution optical zoom and PTZ (Pan-Tilt-Zoom) control was mounted on the robot&#x2019;s back, enabling gauge image capture from various heights (low/high angle) and orientations during autonomous or remote navigation.</p>
<p>A total of 500 high-resolution raw images were acquired through field exploration using the robot.</p>
</sec>
<sec id="s4_1_3">
<label>4.1.3</label>
<title>Data Augmentation and Dataset Split</title>
<p>A two-stage augmentation strategy was adopted to construct a dataset of sufficient scale from the 500 raw images. Stage 1 expands the dataset itself through offline augmentation, while Stage 2 introduces additional diversity during training through online augmentation within the YOLO training pipeline.</p>
<p>An offline augmentation pipeline based on the Albumentations [<xref ref-type="bibr" rid="ref-25">25</xref>] library was constructed using a custom-developed data management tool. A key feature of this pipeline is the simultaneous transformation of image and five-keypoint coordinates, ensuring geometric consistency between augmented images and labels. The following augmentation techniques were applied:<list list-type="bullet">
<list-item>
<p>Random perspective transform: perspective transformation was applied by randomly displacing image corners to simulate the robot&#x2019;s diverse approach angles. This provides a more realistic simulation of actual oblique capture distortion compared to simple 2D rotation.</p></list-item>
<list-item>
<p>Color jittering: brightness, contrast, saturation, and hue were simultaneously randomized to simulate color variations arising from time-of-day and lighting conditions.</p></list-item>
<list-item>
<p>Gaussian blur: blur effects were applied to simulate out-of-focus conditions of the PTZ camera.</p></list-item>
<list-item>
<p>Gaussian noise: signal noise arising from low-light environments or sensor characteristics was simulated.</p></list-item>
<list-item>
<p>Occlusion: to simulate partial gauge occlusion by piping, cables, and other elements common in industrial settings, random rectangular regions were generated within the keypoint bounding area and filled with black masking or blur. The occlusion area was set to approximately 5% of the object area to maintain geometric relationships among keypoints while learning robustness to partial occlusion.</p></list-item>
</list></p>
<p>Through offline augmentation, a final dataset of 11,000 images was constructed. For reliable model training and evaluation, the dataset was split into Training (70%), Validation (15%), and Test (15%) subsets.</p>
</sec>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Model Training Configuration</title>
<p>This subsection describes the specific hyperparameters used for P2-YOLO-Pose model training and the rationale for each setting.</p>
<sec id="s4_2_1">
<label>4.2.1</label>
<title>Model Architecture and Input Settings</title>
<p>The P2 layer-integrated and P5 layer-removed YOLOv11-Pose was used as the base architecture, initialized with pre-trained weights from the COCO-Pose dataset [<xref ref-type="bibr" rid="ref-26">26</xref>].</p>
<p>The input image resolution was set to 1280 <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 1280 pixels to maximize detection capability for small objects such as fire extinguisher gauges, providing 4<inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> the pixel space compared to the 640 <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 640 typically used in standard YOLO training. When combined with the P2 layer (stride 4), the feature map resolution reaches 320 <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 320, enabling stable regression of needle tips only a few pixels wide. Furthermore, to evaluate the maximum bounds of small-object detection capability, inference was performed not only at the native training resolution (1280 px) but also at an upscaled resolution (1920 px).</p>
</sec>
<sec id="s4_2_2">
<label>4.2.2</label>
<title>Loss Function Weight Configuration</title>
<p>The weights for each component of the multi-task loss function are configured as shown in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Loss function weight configuration and rationale.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Loss Function</th>
<th>Weight (<inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:mi>&#x03BB;</mml:mi></mml:math></inline-formula>)</th>
<th>Rationale</th>
</tr>
</thead>
<tbody>
<tr>
<td><inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>I</mml:mi><mml:mi>o</mml:mi><mml:mi>U</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> (Box)</td>
<td>7.5</td>
<td>Stable gauge region localization</td>
</tr>
<tr>
<td><inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mi>C</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> (Cls)</td>
<td>0.5</td>
<td>Reduced classification burden (10 classes)</td>
</tr>
<tr>
<td><inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>D</mml:mi><mml:mi>F</mml:mi><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>1.5</td>
<td>Default weight for coordinate distribution learning</td>
</tr>
<tr>
<td><inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>O</mml:mi><mml:mi>K</mml:mi><mml:mi>S</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> (Pose)</td>
<td>31.0</td>
<td>Keypoint precision prioritized</td>
</tr>
<tr>
<td><inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi><mml:mi>o</mml:mi><mml:mi>b</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td>2.0</td>
<td>Auxiliary keypoint visibility judgment</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Notably, <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>31.0</mml:mn></mml:math></inline-formula>, approximately 4.1<inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> that of <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>b</mml:mi><mml:mi>o</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>7.5</mml:mn></mml:math></inline-formula>. Although the dataset covers 10 types of analog gauges, all classes share an identical five-keypoint skeleton (<inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) and are read through the same geometric pipeline regardless of gauge type; consequently, accurate class identification contributes minimally to the final reading value, justifying a relatively low classification loss weight (<inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0.5</mml:mn></mml:math></inline-formula>). Conversely, since the final reading accuracy is directly determined by the pixel coordinate precision of the needle tip (<inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>), the highest weight is assigned to the keypoint loss. This ratio was determined through a grid search over <inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mi>p</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mn>15</mml:mn><mml:mo>,</mml:mo><mml:mn>17</mml:mn><mml:mo>,</mml:mo><mml:mn>19</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mn>43</mml:mn><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> (step of 2) in preliminary experiments, selecting the value yielding the highest Pose mAP50-95.</p>
</sec>
<sec id="s4_2_3">
<label>4.2.3</label>
<title>Optimizer and Learning Rate Schedule</title>
<p>The optimizer and schedule settings used for training are presented in <xref ref-type="table" rid="table-2">Table 2</xref>.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Summary of model training configuration.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Parameter</th>
<th>Value</th>
<th>Rationale</th>
</tr>
</thead>
<tbody>
<tr>
<td align="center" colspan="3"><italic>Model Configuration</italic></td>
</tr>
<tr>
<td>Model</td>
<td>YOLOv11-Pose (P2&#x002B;, P5&#x2212;)</td>
<td>Small keypoint detection optimization</td>
</tr>
<tr>
<td>Pre-training</td>
<td>COCO-Pose Pretrained</td>
<td>Accelerated feature extraction</td>
</tr>
<tr>
<td>Input resolution</td>
<td>1280 <inline-formula id="ieqn-96"><mml:math id="mml-ieqn-96"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 1280</td>
<td>Small gauge accommodation</td>
</tr>
<tr>
<td align="center" colspan="3"><italic>Optimizer Settings</italic></td>
</tr>
<tr>
<td>Optimizer</td>
<td>AdamW (auto)</td>
<td>Adaptive LR &#x002B; decoupled weight decay</td>
</tr>
<tr>
<td>Initial LR (<inline-formula id="ieqn-97"><mml:math id="mml-ieqn-97"><mml:msub><mml:mi>&#x03B7;</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:math></inline-formula>)</td>
<td><inline-formula id="ieqn-98"><mml:math id="mml-ieqn-98"><mml:mn>4.305</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td>Ultralytics auto-tune</td>
</tr>
<tr>
<td>Final LR ratio (<inline-formula id="ieqn-99"><mml:math id="mml-ieqn-99"><mml:msub><mml:mi>&#x03B7;</mml:mi><mml:mi>f</mml:mi></mml:msub><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:msub><mml:mi>&#x03B7;</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:math></inline-formula>)</td>
<td>0.01</td>
<td>Late-stage fine adjustment</td>
</tr>
<tr>
<td>LR Schedule</td>
<td>Cosine Annealing</td>
<td>Convergence stability</td>
</tr>
<tr>
<td>Momentum</td>
<td>0.937</td>
<td>Gradient inertia retention</td>
</tr>
<tr>
<td>Weight Decay</td>
<td><inline-formula id="ieqn-100"><mml:math id="mml-ieqn-100"><mml:mn>5</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td>Overfitting prevention</td>
</tr>
<tr>
<td align="center" colspan="3"><italic>Training Strategy</italic></td>
</tr>
<tr>
<td>Warm-up</td>
<td>3 epochs, momentum <inline-formula id="ieqn-101"><mml:math id="mml-ieqn-101"><mml:mn>0.8</mml:mn><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>0.937</mml:mn></mml:math></inline-formula></td>
<td>Initial instability prevention</td>
</tr>
<tr>
<td>Total epochs</td>
<td>50</td>
<td>Convergence verified (plateau after epoch 30)</td>
</tr>
<tr>
<td>Batch size</td>
<td>12</td>
<td>GPU VRAM limit (at 1280 px)</td>
</tr>
<tr>
<td>Early Stopping</td>
<td>Patience &#x003D; 15</td>
<td>Early overfitting prevention</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>AdamW was adopted because it decouples weight decay from the learning rate, providing more stable regularization for precise regression tasks such as keypoint coordinate prediction. The initial learning rate <inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:msub><mml:mi>&#x03B7;</mml:mi><mml:mn>0</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn>4.305</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> was determined through Ultralytics&#x2019; auto-tuning mechanism, and cosine annealing gradually reduces it to 1% of <inline-formula id="ieqn-95"><mml:math id="mml-ieqn-95"><mml:msub><mml:mi>&#x03B7;</mml:mi><mml:mn>0</mml:mn></mml:msub></mml:math></inline-formula> for fine-grained adjustments in later training stages.</p>
<p>A batch size of 12 is the maximum feasible setting considering the 1280 <inline-formula id="ieqn-102"><mml:math id="mml-ieqn-102"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 1280 high-resolution input and the additional computational burden of the P2 layer. At 1280 px input with the P2-included model, GPU memory usage reaches approximately 20 GB (on a 24 GB VRAM GPU), and increasing the batch size beyond 12 resulted in OOM (Out-of-Memory) errors. The total training epochs were set to 50, based on learning curve analysis confirming validation loss saturation after approximately epoch 30. Early stopping was employed to prevent overfitting.</p>
</sec>
<sec id="s4_2_4">
<label>4.2.4</label>
<title>Training-Time Data Augmentation Strategy</title>
<p>In addition to offline augmentation (<xref ref-type="sec" rid="s4_1">Section 4.1</xref>), the specific parameters for online augmentation applied in real-time during training are presented in <xref ref-type="table" rid="table-3">Table 3</xref>.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Training-time data augmentation parameter configuration.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Augmentation</th>
<th>Parameter</th>
<th>Simulated Environmental Variable</th>
</tr>
</thead>
<tbody>
<tr>
<td>Mosaic</td>
<td>1.0 (disabled last 10 epochs)</td>
<td>Complex backgrounds (piping, cables)</td>
</tr>
<tr>
<td>MixUp</td>
<td><inline-formula id="ieqn-103"><mml:math id="mml-ieqn-103"><mml:mi>&#x03B1;</mml:mi><mml:mo>=</mml:mo><mml:mn>0.5</mml:mn></mml:math></inline-formula></td>
<td>Decision boundary regularization</td>
</tr>
<tr>
<td>Copy-Paste</td>
<td>1.0 (flip mode)</td>
<td>Gauge appearance contextual diversity</td>
</tr>
<tr>
<td>HSV-Hue</td>
<td><inline-formula id="ieqn-104"><mml:math id="mml-ieqn-104"><mml:mo>&#x00B1;</mml:mo><mml:mn>0.015</mml:mn></mml:math></inline-formula></td>
<td>Illumination color temperature</td>
</tr>
<tr>
<td>HSV-Saturation</td>
<td><inline-formula id="ieqn-105"><mml:math id="mml-ieqn-105"><mml:mo>&#x00D7;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0.3</mml:mn><mml:mo>,</mml:mo><mml:mn>1.7</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula></td>
<td>Saturation (glass reflections, dust)</td>
</tr>
<tr>
<td>HSV-Value</td>
<td><inline-formula id="ieqn-106"><mml:math id="mml-ieqn-106"><mml:mo>&#x00D7;</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:mn>0.6</mml:mn><mml:mo>,</mml:mo><mml:mn>1.4</mml:mn><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula></td>
<td>Brightness (shadows, direct sunlight)</td>
</tr>
<tr>
<td>Scale</td>
<td><inline-formula id="ieqn-107"><mml:math id="mml-ieqn-107"><mml:mi>s</mml:mi><mml:mo>&#x223C;</mml:mo><mml:mi>U</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mn>0.5</mml:mn><mml:mo>,</mml:mo><mml:mn>1.5</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td>Capture distance variation</td>
</tr>
<tr>
<td>Horizontal Flip</td>
<td><inline-formula id="ieqn-108"><mml:math id="mml-ieqn-108"><mml:mi>P</mml:mi><mml:mo>=</mml:mo><mml:mn>0.5</mml:mn></mml:math></inline-formula></td>
<td>Left-right placement differences</td>
</tr>
<tr>
<td>RandAugment</td>
<td>Automatic policy</td>
<td>General invariance</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Mosaic augmentation is effective for inducing rapid convergence by quadrupling per-batch contextual diversity during early training. However, artificially composed images may hinder fine-grained keypoint regression in later stages [<xref ref-type="bibr" rid="ref-27">27</xref>]. Therefore, Mosaic is deactivated during the last 10 epochs, allowing the model to focus on precise keypoint localization in individual images.</p>
<p>The Hue transformation range is conservatively set because gauge scale color differentiation (normal &#x003D; green, danger &#x003D; red) may be utilized for initial reading verification. Conversely, Saturation and Value ranges are set relatively wide to ensure robustness against field saturation and brightness variations caused by glass reflections, dust, and shadows.</p>
<p>Copy-Paste augmentation [<xref ref-type="bibr" rid="ref-28">28</xref>], which copies and pastes gauge instances with flipping, is applied at 100% probability. This simulates the characteristic of industrial sites where identical gauge types are installed against various backgrounds, effectively learning background invariance for keypoint detection.</p>
</sec>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Geometric Rectification Accuracy Validation</title>
<p>To objectively validate the accuracy of the proposed keypoint-based geometric rectification method, a precision comparison experiment was designed using ArUco markers as a control group.</p>
<sec id="s4_3_1">
<label>4.3.1</label>
<title>Control Group Setup</title>
<p>During data collection, ArUco markers (dictionary: <monospace>DICT_ARUCO_ORIGINAL</monospace>, size: 30 mm) were attached adjacent to selected target gauges (within 5 cm on the floor or side). When multiple markers are present in a single image, the Euclidean distance between the detected gauge bounding box center and each marker center is computed, and the nearest marker is automatically selected as the reference marker for that gauge.</p>
<p>The 6-Degrees of Freedom (6-DoF) pose is estimated using the four corner points of the marker, and the <inline-formula id="ieqn-109"><mml:math id="mml-ieqn-109"><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>3</mml:mn></mml:math></inline-formula> rotation matrix <italic>R</italic> is computed:<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mtable columnalign="center center center" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>11</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>12</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>13</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>21</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>22</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>23</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>31</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>32</mml:mn></mml:mrow></mml:msub></mml:mtd><mml:mtd><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>33</mml:mn></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>Roll (<inline-formula id="ieqn-110"><mml:math id="mml-ieqn-110"><mml:mi>&#x03D5;</mml:mi></mml:math></inline-formula>), Pitch (<inline-formula id="ieqn-111"><mml:math id="mml-ieqn-111"><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula>), and Yaw (<inline-formula id="ieqn-112"><mml:math id="mml-ieqn-112"><mml:mi>&#x03C8;</mml:mi></mml:math></inline-formula>) are decomposed from this rotation matrix following the Tait-Bryan angles in Z(Yaw)-Y(Pitch)-X(Roll) sequence:<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:mi>&#x03D5;</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext>atan2</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>21</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>11</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:mi>&#x03B8;</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext>atan2</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>31</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msqrt><mml:msubsup><mml:mi>R</mml:mi><mml:mrow><mml:mn>32</mml:mn></mml:mrow><mml:mn>2</mml:mn></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mi>R</mml:mi><mml:mrow><mml:mn>33</mml:mn></mml:mrow><mml:mn>2</mml:mn></mml:msubsup></mml:msqrt><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:mi>&#x03C8;</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext>atan2</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>32</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mrow><mml:mn>33</mml:mn></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
<p>These RPY values, based on the physical size of the marker, provide reliable ground truth for the tilt of the gauge plane.</p>
</sec>
<sec id="s4_3_2">
<label>4.3.2</label>
<title>Comparison Methodology</title>
<p>The experiment simultaneously computes and compares RPY values and normalized reading values under the following three conditions for the same input image:<list list-type="bullet">
<list-item>
<p>Raw image reading (RAW): RPY and normalized reading values (<inline-formula id="ieqn-113"><mml:math id="mml-ieqn-113"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) are computed directly using keypoints detected by P2-YOLO-Pose in the original distorted image, without any rectification. This serves as the baseline including projective distortion.</p></list-item>
<list-item>
<p>Virtual point-based geometric rectification (RECT): after keypoint-based virtual point generation and homography rectification, RPY and normalized reading values (<inline-formula id="ieqn-114"><mml:math id="mml-ieqn-114"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) are computed in the rectified coordinates. The rectification effect is measured by the RPY change (Delta) relative to RAW.</p></list-item>
<list-item>
<p>ArUco-based reading (ArUco): after rectification using the homography matrix computed from the ArUco marker&#x2019;s four-point correspondences, the same keypoints are transformed and RPY and normalized reading values (<inline-formula id="ieqn-115"><mml:math id="mml-ieqn-115"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) are computed.</p></list-item>
</list></p>
</sec>
<sec id="s4_3_3">
<label>4.3.3</label>
<title>Quantitative Evaluation Criteria</title>
<p>Comparison among the three conditions is performed using two metrics:</p>
<p>The RAW RPY is set as the reference (origin), and the RPY change after rectification by RECT and ArUco is computed as the delta (<inline-formula id="ieqn-116"><mml:math id="mml-ieqn-116"><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:math></inline-formula>):<disp-formula id="eqn-16"><label>(16)</label><mml:math id="mml-eqn-16" display="block"><mml:msub><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>RPY</mml:mtext></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mtext>RPY</mml:mtext></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mspace width="1em" /><mml:msub><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mtext>RPY</mml:mtext></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mtext>RPY</mml:mtext></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p>The smaller the difference between <inline-formula id="ieqn-117"><mml:math id="mml-ieqn-117"><mml:msub><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-118"><mml:math id="mml-ieqn-118"><mml:msub><mml:mi mathvariant="normal">&#x0394;</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, the closer the proposed method achieves to the accuracy of physical marker-based rectification.</p>
<p>The practical effect of rectification is evaluated by the difference in normalized reading values computed under each condition:<disp-formula id="eqn-17"><label>(17)</label><mml:math id="mml-eqn-17" display="block"><mml:msub><mml:mi>&#x03F5;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>u</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>As <inline-formula id="ieqn-119"><mml:math id="mml-ieqn-119"><mml:msub><mml:mi>&#x03F5;</mml:mi><mml:mrow><mml:mi>v</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>u</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> approaches zero, the proposed method&#x2019;s rectification achieves ArUco-level precision.</p>
<p>All comparison results are recorded and can be simultaneously verified through a real-time visual dashboard displaying three-panel views of original/keypoint-rectified/ArUco-rectified images (<xref ref-type="fig" rid="fig-5">Fig. 5</xref>).</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Quantitative validation experimental setup using ArUco markers. (<bold>a</bold>) Capture under severe oblique angle, exhibiting substantial projective distortion. (<bold>b</bold>) Target gauge and attached ArUco marker captured at near-frontal angle. (<bold>c</bold>) Visualization of the ArUco marker detection process used for reference pose and homography matrix computation.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80624-fig-5.tif"/>
</fig>
</sec>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Results</title>
<p>This section presents a quantitative and qualitative analysis of the proposed system&#x2019;s performance. First, the basic performance of keypoint detection is validated, followed by an ablation study on architectural variants. The effect of geometric rectification is analyzed based on RPY and reading values, and the adaptive strategy under extreme conditions is examined. Finally, a comparison with existing state-of-the-art methods is presented.</p>
<sec id="s5_1">
<label>5.1</label>
<title>Feasibility Analysis of Keypoint-Based Approach</title>
<p>The accuracy of detecting the proposed five keypoints on analog gauges was first validated.</p>
<p>Despite the complex background and metallic reflections of the experimental environment, five keypoints were accurately aligned and detected within the very small fire extinguisher pressure gauges (<xref ref-type="fig" rid="fig-6">Fig. 6a</xref>). Furthermore, <inline-formula id="ieqn-120"><mml:math id="mml-ieqn-120"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>(C) and <inline-formula id="ieqn-121"><mml:math id="mml-ieqn-121"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>(N) were accurately connected even under oblique capture angles exhibiting projective distortion (<xref ref-type="fig" rid="fig-6">Fig. 6b</xref>). This demonstrates that the extracted <inline-formula id="ieqn-122"><mml:math id="mml-ieqn-122"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>(S), <inline-formula id="ieqn-123"><mml:math id="mml-ieqn-123"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>(M), and <inline-formula id="ieqn-124"><mml:math id="mml-ieqn-124"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>(E) points provide sufficient foundational data for virtually reconstructing the circular arc trajectory.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Five-keypoint detection results in a power data center environment. (<bold>a</bold>) Detection on a fire extinguisher gauge; (<bold>b</bold>) detection on a piping gauge.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80624-fig-6.tif"/>
</fig>
<p>Learning convergence was assessed by analyzing training curves over 50 epochs. <xref ref-type="table" rid="table-4">Table 4</xref> presents the performance metrics at key epochs.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Training convergence analysis by epoch (training/validation loss and performance metrics).</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Metric</th>
<th>Epoch 1</th>
<th>Epoch 10</th>
<th>Epoch 25</th>
<th>Epoch 50</th>
</tr>
</thead>
<tbody>
<tr>
<td>Pose mAP50</td>
<td>53.15%</td>
<td>99.23%</td>
<td>99.21%</td>
<td>99.45%</td>
</tr>
<tr>
<td>Pose mAP50-95</td>
<td>47.61%</td>
<td>98.60%</td>
<td>99.21%</td>
<td>99.37%</td>
</tr>
<tr>
<td>Box mAP50-95</td>
<td>37.89%</td>
<td>98.60%</td>
<td>93.95%</td>
<td>96.62%</td>
</tr>
<tr>
<td>Precision</td>
<td>52.11%</td>
<td>&#x2013;</td>
<td>98.91%</td>
<td>99.18%</td>
</tr>
<tr>
<td>Recall</td>
<td>53.41%</td>
<td>&#x2013;</td>
<td>99.40%</td>
<td>99.74%</td>
</tr>
<tr>
<td>Train/pose_loss</td>
<td>4.596</td>
<td>0.439</td>
<td>0.238</td>
<td>0.081</td>
</tr>
<tr>
<td>Val/pose_loss</td>
<td>1.091</td>
<td>0.446</td>
<td>0.319</td>
<td>0.256</td>
</tr>
<tr>
<td>Train/box_loss</td>
<td>2.138</td>
<td>0.617</td>
<td>0.468</td>
<td>0.276</td>
</tr>
<tr>
<td>Val/box_loss</td>
<td>0.476</td>
<td>0.986</td>
<td>0.993</td>
<td>0.994</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The final model achieved Pose mAP50 of 99.45% and Pose mAP50-95 of 99.37%. The possibility of overfitting at these high mAP values is analyzed from three perspectives.</p>
<p>As shown in <xref ref-type="table" rid="table-4">Table 4</xref>, val/pose_loss decreased monotonically from 1.091 (Epoch 1) to 0.446 (Epoch 10), 0.319 (Epoch 25), and 0.256 (Epoch 50) throughout the entire 50-epoch training period <italic>without any rebound</italic>. The hallmark of overfitting&#x2014;training loss decreasing while validation loss rebounds&#x2014;was not observed, and Early Stopping (patience &#x003D; 15) was not triggered.</p>

<p>All 10 types of analog gauges in this dataset share a distinctive visual structure of &#x201C;circular frame &#x002B; needle &#x002B; scale&#x201D;, with clear visual differentiation from the background (piping, walls). Furthermore, the data was collected from a controlled environment in a single facility, limiting the range of illumination and background variation compared to natural-scene image datasets (e.g., COCO). This ceiling effect attributable to domain characteristics is the primary cause of the high mAP, and this point is explicitly discussed as a dataset bias limitation.</p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Ablation Study: P2 Feature Enhancement</title>
<p>To validate the effectiveness of the P2 high-resolution feature layer, the performance of the Baseline model (P3&#x2013;P5) and the Proposed model (P2&#x2013;P4) was compared.</p>
<sec id="s5_2_1">
<label>5.2.1</label>
<title>Detection Performance in Standard Conditions</title>
<p>As shown in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>, both models demonstrated excellent detection performance in standard high-resolution conditions, indicating no difference in baseline detection capability.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Detection results in standard high-resolution conditions. Both the baseline and proposed models reliably detect all targets.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80624-fig-7.tif"/>
</fig>
</sec>
<sec id="s5_2_2">
<label>5.2.2</label>
<title>Small Object Detection Sensitivity</title>
<p><xref ref-type="fig" rid="fig-8">Fig. 8</xref> visually demonstrates that the Proposed P2 model possesses superior detection sensitivity (Discovery Index) compared to the Baseline. The Baseline model (top) passively detected only the 3 major gauges specified in the training data ground truth, whereas the Proposed model (bottom) additionally identified 2 small gauges in the background that were not included during the labeling process, comprehensively detecting a total of 5 objects.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Detection sensitivity comparison in a high-resolution environment (1920 px). The Baseline model (top) detected only 3 objects matching the ground truth, while the Proposed model (bottom) detected 5 gauges including missed small gauges.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80624-fig-8.tif"/>
</fig>
</sec>
<sec id="s5_2_3">
<label>5.2.3</label>
<title>Quantitative Performance Comparison</title>
<p><xref ref-type="table" rid="table-5">Table 5</xref> presents a comprehensive performance comparison between the Baseline and Proposed models across different inference resolutions (the native 1280 px and an upscaled 1920 px) to verify the impact of input scaling.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Comprehensive performance comparison of baseline and proposed model (P2 integration) by resolution.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Inference Resolution</th>
<th>Model</th>
<th>S_PCK@10 (%)</th>
<th>S_DI</th>
<th>S_Extra (count)</th>
<th>FPS</th>
</tr>
</thead>
<tbody>
<tr>
<td>1280 px</td>
<td>Baseline</td>
<td>100.00</td>
<td>1.04</td>
<td>34</td>
<td>27.8</td>
</tr>
<tr>
<td>1280 px</td>
<td>Proposed</td>
<td>100.00</td>
<td>1.07</td>
<td>58</td>
<td>28.2</td>
</tr>
<tr>
<td>1920 px</td>
<td>Baseline</td>
<td>100.00</td>
<td>1.06</td>
<td>52</td>
<td>26.1</td>
</tr>
<tr>
<td>1920 px</td>
<td>Proposed</td>
<td>99.88</td>
<td>1.09</td>
<td>73</td>
<td>25.9</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The accuracy of the final gauge reading is evaluated using the following metrics:<list list-type="bullet">
<list-item>
<p><inline-formula id="ieqn-125"><mml:math id="mml-ieqn-125"><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>P</mml:mi><mml:mi>C</mml:mi><mml:mi>K</mml:mi><mml:mrow><mml:mo>@</mml:mo></mml:mrow><mml:mn>10</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>: Percentage of Correct Keypoints at 10% threshold&#x2014;the proportion of predicted keypoints whose Euclidean distance from the ground truth keypoint is within 10% of the bounding box diagonal length.</p></list-item>
<list-item>
<p><inline-formula id="ieqn-126"><mml:math id="mml-ieqn-126"><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>D</mml:mi><mml:mi>I</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> (Discovery Index): the proportion of all gauges successfully detected and read.</p></list-item>
<list-item>
<p><inline-formula id="ieqn-127"><mml:math id="mml-ieqn-127"><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>E</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>: The number of additional objects successfully detected by the proposed model compared to the baseline. This metric quantifies the enhanced detection capability.</p></list-item>
</list></p>
<p><inline-formula id="ieqn-128"><mml:math id="mml-ieqn-128"><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>D</mml:mi><mml:mi>I</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> (Discovery Index) exceeding 1.0 indicates that the model independently identified objects that human annotators missed during the labeling process. This suggests that the model possesses the ability to generalize by learning the intrinsic geometric characteristics of analog gauges, rather than simply overfitting to the training data.</p>
<p>In the <inline-formula id="ieqn-129"><mml:math id="mml-ieqn-129"><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>E</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> metric, the Proposed model detected 58 and 73 additional objects at 1280 and 1920 px, respectively&#x2014;1.7<inline-formula id="ieqn-130"><mml:math id="mml-ieqn-130"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> and 1.4<inline-formula id="ieqn-131"><mml:math id="mml-ieqn-131"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> higher detection capability compared to the Baseline (34/52). This demonstrates that the P2 layer effectively compensates for the resolution loss inherent in the existing P3&#x2013;P5 structure.</p>
<p>Adding high-resolution feature maps typically causes a sharp increase in computational cost. However, the Proposed model&#x2019;s FPS decrease was less than 1% (26.1 <inline-formula id="ieqn-132"><mml:math id="mml-ieqn-132"><mml:mo stretchy="false">&#x2192;</mml:mo></mml:math></inline-formula> 25.9), confirming the effectiveness of the computational offset strategy through P5 removal.</p>
</sec>
</sec>
<sec id="s5_3">
<label>5.3</label>
<title>Rectification and Reading Accuracy Analysis</title>
<sec id="s5_3_1">
<label>5.3.1</label>
<title>RPY-Based Rectification Effect Analysis</title>
<p>The reading accuracy of the proposed keypoint-based rectification (RECT), raw baseline (Raw), and ArUco control group (ArUco) was compared across three major distortion types (vertical tilt, lateral viewpoint, in-plane rotation) that can arise from the robot&#x2019;s travel path and camera mounting position. <xref ref-type="table" rid="table-6">Table 6</xref> shows the comparison results.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Reading accuracy analysis.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center" rowspan="2">Condition</th>
<th align="center" colspan="2">Correction Delta <inline-formula id="ieqn-136"><mml:math id="mml-ieqn-136"><mml:mo stretchy="false">(</mml:mo><mml:mi>R</mml:mi><mml:mo>,</mml:mo><mml:mi>P</mml:mi><mml:mo>,</mml:mo><mml:mi>Y</mml:mi><mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula></th>
<th align="center" colspan="4">Reading Value (Normalized)</th>
</tr>
<tr>
<th>RECT</th>
<th>ArUco</th>
<th>GT</th>
<th>Raw</th>
<th>RECT</th>
<th>ArUco</th>
</tr>
</thead>
<tbody>
<tr>
<td align="center" colspan="7"><italic>1. Roll (Vertical Tilt)</italic></td>
</tr>
<tr>
<td>Frontal</td>
<td><inline-formula id="ieqn-137"><mml:math id="mml-ieqn-137"><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mn mathvariant="bold">3</mml:mn></mml:mrow><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mn>5</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-138"><mml:math id="mml-ieqn-138"><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mn>1</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mn>8</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mn>3</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td>0.520</td>
<td>0.519</td>
<td>0.513</td>
<td>0.520</td>
</tr>
<tr>
<td>Severe</td>
<td><inline-formula id="ieqn-139"><mml:math id="mml-ieqn-139"><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mn mathvariant="bold">30</mml:mn></mml:mrow><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mn>1</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msup><mml:mn>1</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-140"><mml:math id="mml-ieqn-140"><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msup><mml:mn>1</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msup><mml:mn>5</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td>0.370</td>
<td>0.335</td>
<td>0.376</td>
<td>0.380</td>
</tr>
<tr>
<td align="center" colspan="7"><italic>2. Pitch (Lateral Viewpoint)</italic></td>
</tr>
<tr>
<td>Mild</td>
<td><inline-formula id="ieqn-141"><mml:math id="mml-ieqn-141"><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mrow><mml:mo mathvariant="bold">&#x2212;</mml:mo><mml:msup><mml:mn mathvariant="bold">5</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-142"><mml:math id="mml-ieqn-142"><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mn>15</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msup><mml:mn>12</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td>0.161</td>
<td>0.162</td>
<td>0.164</td>
<td>0.159</td>
</tr>
<tr>
<td>Severe</td>
<td><inline-formula id="ieqn-143"><mml:math id="mml-ieqn-143"><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mn>19</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mrow><mml:msup><mml:mn mathvariant="bold">25</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msup><mml:mn>0</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-144"><mml:math id="mml-ieqn-144"><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mn>3</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mn>17</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msup><mml:mn>5</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td>0.139</td>
<td>0.185</td>
<td>0.170</td>
<td>0.135</td>
</tr>
<tr>
<td align="center" colspan="7"><italic>3. Yaw (In-plane Rotation)</italic></td>
</tr>
<tr>
<td>&#x2013;</td>
<td><inline-formula id="ieqn-145"><mml:math id="mml-ieqn-145"><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mn>11</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mrow><mml:mo mathvariant="bold">&#x2212;</mml:mo><mml:msup><mml:mn mathvariant="bold">6</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-146"><mml:math id="mml-ieqn-146"><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mn>9</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mn>1</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td>0.470</td>
<td>0.452</td>
<td>0.446</td>
<td>0.454</td>
</tr>
<tr>
<td>&#x2013;</td>
<td><inline-formula id="ieqn-147"><mml:math id="mml-ieqn-147"><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:msup><mml:mn>9</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msup><mml:mn>0</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mrow><mml:mo mathvariant="bold">&#x2212;</mml:mo><mml:msup><mml:mn mathvariant="bold">8</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-148"><mml:math id="mml-ieqn-148"><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mn>3</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mn>4</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula></td>
<td>0.390</td>
<td>0.388</td>
<td>0.398</td>
<td>0.431</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><list list-type="bullet">
<list-item>
<p>Vertical tilt (Roll) analysis: under frontal capture, Raw, RECT, and ArUco all exhibited similar accuracy. However, under severe tilt conditions where the robot looks upward from below, the Raw reading showed approximately 3.5% error, while the proposed geometric rectification reduced the error to the 0.6% level. This demonstrates that vector-based angle computation in the rectified linear coordinate system is effectively invariant to projective distortion.</p></list-item>
<list-item>
<p>Lateral viewpoint (Pitch) analysis: under mild lateral deviation, all methods produced satisfactory results. However, as lateral deviation increases, the circular gauge is projected as an ellipse, introducing nonlinear scale distortion. The Raw reading exhibited approximately 4.5% error, which was reduced to approximately 3.1% after geometric rectification by restoring the ellipse to a circle. This demonstrates that the proposed method can mitigate nonlinear scale distortion from lateral capture and improve reading accuracy.</p></list-item>
<list-item>
<p>In-plane rotation (Yaw) analysis: when in-plane rotation occurs, the gauge start point (<inline-formula id="ieqn-133"><mml:math id="mml-ieqn-133"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>), end point (<inline-formula id="ieqn-134"><mml:math id="mml-ieqn-134"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>), and needle position (<inline-formula id="ieqn-135"><mml:math id="mml-ieqn-135"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) undergo the same rigid transformation, preserving the spatial relationships among keypoints. Consequently, the performance difference before and after rectification was relatively small compared to vertical or lateral distortion cases.</p></list-item>
</list></p>
</sec>
<sec id="s5_3_2">
<label>5.3.2</label>
<title>Adaptive Strategy Analysis under Extreme Geometric Conditions (AR <inline-formula id="ieqn-149"><mml:math id="mml-ieqn-149"><mml:mo>&#x003E;</mml:mo><mml:mn>1.5</mml:mn></mml:math></inline-formula>)</title>
<p>This section compares the reading accuracy of (b) forced warping and (c) original image retention under extreme conditions where the aspect ratio (AR) of the keypoint-defined Region of Interest (ROI) exceeds 1.5.</p>
<p>Two primary causes lead to AR <inline-formula id="ieqn-150"><mml:math id="mml-ieqn-150"><mml:mo>&#x003E;</mml:mo><mml:mn>1.5</mml:mn></mml:math></inline-formula>. First, extreme height differences between the robot and gauge result in capture at steep angles (severe projective distortion). Second, gauges such as fire extinguisher pressure gauges where <inline-formula id="ieqn-151"><mml:math id="mml-ieqn-151"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-152"><mml:math id="mml-ieqn-152"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are positioned near the horizontal axis of <inline-formula id="ieqn-153"><mml:math id="mml-ieqn-153"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> structurally reduce the rhombus grid height (inherent gauge geometry).</p>
<p>In this experimental environment, AR increases due to projective distortion were excluded through robot path optimization, and AR <inline-formula id="ieqn-154"><mml:math id="mml-ieqn-154"><mml:mo>&#x003E;</mml:mo><mml:mn>1.5</mml:mn></mml:math></inline-formula> was primarily observed in the inherent design of fire extinguisher pressure gauges (narrow operating range, horizontal placement of <inline-formula id="ieqn-155"><mml:math id="mml-ieqn-155"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-156"><mml:math id="mml-ieqn-156"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>).</p>
<p>As shown in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>, applying forced warping to this gauge causes excessive stretching that distorts the needle angle, resulting in a misreading of 0.37 against a GT of 0.49. In contrast, retaining the original image yields a reading of 0.49, exactly matching the GT. This result experimentally validates the AR-based adaptive rectification strategy.</p>
<fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Analysis of high AR cases due to inherent geometric structure (fire extinguisher gauge). (<bold>a</bold>) Original gauge with structurally high AR. (<bold>b</bold>) Reading error due to vertical over-stretching when forced warping is applied (Val: 0.37). (<bold>c</bold>) Accurate reading maintaining geometric structure without warping (Val: 0.49, matches GT).</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80624-fig-9.tif"/>
</fig>
</sec>
</sec>
<sec id="s5_4">
<label>5.4</label>
<title>System Robustness and Field Applicability</title>
<sec id="s5_4_1">
<label>5.4.1</label>
<title>Robustness under Environmental Variations</title>
<p>While illumination conditions in indoor industrial environments are relatively stable, localized specular reflection from the glass covers of analog gauges remains an unavoidable challenge. Experimental results show that the proposed system consistently maintained keypoint extraction stability even under adverse conditions such as high-contrast illumination and severe light reflections (<xref ref-type="fig" rid="fig-10">Fig. 10</xref>).</p>
<fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>Qualitative evaluation of environmental robustness. (<bold>a</bold>) Stable detection under high-contrast conditions; (<bold>b</bold>) Successful inference of structural features despite information loss due to glass surface reflection and occlusion.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80624-fig-10.tif"/>
</fig>
<p>This environmental robustness is attributed to the Position-Sensitive Attention (PSA) mechanism [<xref ref-type="bibr" rid="ref-24">24</xref>] introduced in the YOLOv11 architecture. PSA does not rely solely on local information when extracting features at specific positions of the input feature map; instead, it analyzes correlations (long-range dependencies) with other regions in the image. Thus, even when pixel-level reliability of gauge scale markings or needle segments is degraded by light reflections, the model compensates for occluded features by cross-referencing the curvature information of unoccluded arc segments and the center point position.</p>
</sec>
<sec id="s5_4_2">
<label>5.4.2</label>
<title>Generalization to Field Scenarios</title>
<p>The model demonstrated excellent adaptability to the complex environmental variables of actual industrial sites (<xref ref-type="fig" rid="fig-11">Fig. 11</xref>).</p>
<fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>Generalization performance evaluation across industrial field scenarios. (<bold>a</bold>) Precise detection of a small gauge; (<bold>b</bold>) robust object recognition in complex densely-piped environments.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMES_80624-fig-11.tif"/>
</fig>
<p><xref ref-type="fig" rid="fig-11">Fig. 11a</xref> shows a small gauge mounted on piping, demonstrating that the Proposed model accurately detects it despite occupying a very small ROI relative to the entire image, confirming detection sensitivity for small objects. <xref ref-type="fig" rid="fig-11">Fig. 11b</xref> shows analysis results in a complex background with numerous pipes and gauges densely arranged. Despite the presence of visual patterns similar to gauges (pipe tape, metallic reflectors), the model clearly distinguished background clutter from actual gauges through structural context.</p>
</sec>
<sec id="s5_4_3">
<label>5.4.3</label>
<title>Inference Efficiency Analysis</title>
<p>Inference speed and latency were measured to validate the applicability of the proposed system for real-time monitoring. According to the results in <xref ref-type="table" rid="table-5">Table 5</xref>, the Proposed model recorded approximately 28.2 FPS (average latency approximately 35.4 ms) at the 1280 <inline-formula id="ieqn-157"><mml:math id="mml-ieqn-157"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 1280 high-resolution input condition. This is equivalent to the Baseline&#x2019;s 27.8 FPS, and even at 1920 px, the Proposed model achieved 25.9 FPS&#x2014;only 0.2 FPS below the Baseline (26.1 FPS). The structural optimization of adding P2 and removing P5 substantially improved feature extraction accuracy while satisfying real-time processing requirements.</p>

</sec>
</sec>
<sec id="s5_5">
<label>5.5</label>
<title>Comparison with Existing Methods</title>
<p>A methodological comparison with major recent approaches in the analog gauge reading field is presented. Direct numerical comparison is difficult for the following reasons: (1) the Pointer-10K dataset of VDN [<xref ref-type="bibr" rid="ref-16">16</xref>] is not publicly available; (2) the experimental datasets of GAUREAD [<xref ref-type="bibr" rid="ref-2">2</xref>] and Under Pressure [<xref ref-type="bibr" rid="ref-17">17</xref>] are also not released; and (3) each method addresses different gauge types and evaluation conditions. Therefore, <xref ref-type="table" rid="table-7">Table 7</xref> presents a systematic comparison of methodological characteristics.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Systematic comparison with existing methods.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Characteristic</th>
<th>GAUREAD</th>
<th>VDN</th>
<th>Under Pressure</th>
<th>Ours</th>
</tr>
</thead>
<tbody>
<tr>
<td>Detection method</td>
<td>YOLO &#x002B; CHT</td>
<td>Confidence Map</td>
<td>YOLO &#x002B; Notch</td>
<td>P2-YOLO-Pose</td>
</tr>
<tr>
<td>Gauge representation</td>
<td>Ellipse Fit</td>
<td>2D Vector</td>
<td>Ellipse &#x002B; Notch</td>
<td>5-KP Skeleton</td>
</tr>
<tr>
<td>Projection correction</td>
<td>Polar Unwrap</td>
<td>&#x2013;</td>
<td>Polar Unwrap</td>
<td>VP-Homography</td>
</tr>
<tr>
<td>Corner-free</td>
<td><inline-formula id="ieqn-158"><mml:math id="mml-ieqn-158"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula>(circle req.)</td>
<td>N/A</td>
<td><inline-formula id="ieqn-159"><mml:math id="mml-ieqn-159"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> (notch)</td>
<td>&#x2713;</td>
</tr>
<tr>
<td>Nonlinear scale</td>
<td><inline-formula id="ieqn-160"><mml:math id="mml-ieqn-160"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-161"><mml:math id="mml-ieqn-161"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula></td>
<td>Partial</td>
<td>&#x2713;</td>
</tr>
<tr>
<td>Effective range</td>
<td><inline-formula id="ieqn-162"><mml:math id="mml-ieqn-162"><mml:mrow><mml:mo>&#x2264;</mml:mo></mml:mrow><mml:msup><mml:mn>20</mml:mn><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> (3%)</td>
<td>&#x2013;</td>
<td>Frontal (&#x003C;2%)</td>
<td>AR <inline-formula id="ieqn-163"><mml:math id="mml-ieqn-163"><mml:mo>&#x2264;</mml:mo></mml:math></inline-formula> 1.5 full range</td>
</tr>
<tr>
<td>Real-time FPS</td>
<td><inline-formula id="ieqn-164"><mml:math id="mml-ieqn-164"><mml:mo>&#x223C;</mml:mo></mml:math></inline-formula>1.25</td>
<td>Real-time</td>
<td>Not reported</td>
<td>25.9</td>
</tr>
<tr>
<td>Field validation</td>
<td>DLR facility</td>
<td>Lab (10 K)</td>
<td>ETH Lab</td>
<td>KEPCO 11K</td>
</tr>
<tr>
<td>Robustness thresholds</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>AR, Radius, Boundary</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The three key differentiating factors of this study are as follows:</p>
<p>The polar unwrap approach adopted by GAUREAD and Under Pressure is valid only when the circular gauge appears as a perfect circle from the frontal view. Under oblique capture, nonlinear distortion arises during the ellipse-to-rectangle transformation; GAUREAD reported errors of 3% at <inline-formula id="ieqn-165"><mml:math id="mml-ieqn-165"><mml:msup><mml:mi>20</mml:mi><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> and 9% at <inline-formula id="ieqn-166"><mml:math id="mml-ieqn-166"><mml:msup><mml:mi>50</mml:mi><mml:mrow><mml:mo>&#x2218;</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> viewing angles. The proposed VP-Homography method restores the ellipse to a frontal circle before computing angles, fundamentally eliminating nonlinear distortion.</p>
<p>VDN detects only the pointer direction (vector) without inferring scale range or tick intervals, making conversion to actual physical readings impossible. The proposed method fully reconstructs the scale structure through <inline-formula id="ieqn-167"><mml:math id="mml-ieqn-167"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-168"><mml:math id="mml-ieqn-168"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, and <inline-formula id="ieqn-169"><mml:math id="mml-ieqn-169"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> keypoints and performs physical value conversion via the circular distance function.</p>
<p>GAUREAD relies on the Circle Hough Transform and Under Pressure depends on notch detection, both of which can fail on gauges lacking the corresponding geometric features. The proposed virtual point generation method, based on the point symmetry principle, enables stable homography rectification on any circular gauge without external markers or predefined feature points.</p>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Discussion</title>
<p>This section synthesizes the strengths of the proposed system based on the experimental results, analyzes the limitations of the dataset and methodology, and presents practical considerations for industrial deployment.</p>
<sec id="s6_1">
<label>6.1</label>
<title>Strengths of the Proposed Approach</title>
<p>The contributions of this work can be summarized from three perspectives.</p>
<p>Previous analog gauge reading methods adopt a two-stage pipeline that detects bounding boxes and then applies post-processing (Circle Hough Transform, ellipse fitting, polar unwrapping) to compute values. This approach suffers from a structural vulnerability in which the entire pipeline fails when circularity assumptions break or characteristic features are absent. The proposed method directly regresses five keypoints (<inline-formula id="ieqn-170"><mml:math id="mml-ieqn-170"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) from a single network, unifying detection and geometric inference into an end-to-end framework. As a result, the physical structure of the gauge&#x2014;the scale arc, needle, and rotation center&#x2014;is explicitly encoded as a keypoint skeleton, ensuring that subsequent rectification and value computation proceed on a mathematically stable foundation.</p>
<p>The virtual point generation strategy relies solely on point symmetry, enabling homography rectification without requiring external markers (ArUco, checkerboard) or prior detection of gauge-specific features (notches, tick marks). Experimental results show that the normalized reading difference between this approach and the ArUco-based ground truth averages within 0.02, achieving correction precision comparable to physical markers. This addresses the practical constraint that individual markers cannot be attached to every gauge in large-scale industrial facilities.</p>
<p>Rather than relying on a single fixed pipeline, the proposed system adaptively switches its processing strategy according to input data conditions. Multi-stage validity checks&#x2014;including radius ratio verification (<inline-formula id="ieqn-171"><mml:math id="mml-ieqn-171"><mml:mrow><mml:mtext>min/max</mml:mtext></mml:mrow><mml:mo>&#x003C;</mml:mo><mml:mn>0.4</mml:mn></mml:math></inline-formula>), boundary constraints (<inline-formula id="ieqn-172"><mml:math id="mml-ieqn-172"><mml:mi>&#x03B4;</mml:mi><mml:mo>=</mml:mo><mml:mn>10</mml:mn></mml:math></inline-formula> px), and aspect ratio assessment (AR <inline-formula id="ieqn-173"><mml:math id="mml-ieqn-173"><mml:mo>&#x2264;</mml:mo><mml:mn>1.5</mml:mn></mml:math></inline-formula>)&#x2014;ensure that homography is applied only when rectification is valid, and direct reading from the original image is performed otherwise. This adaptive strategy fundamentally prevents the over-correction problem that can arise in methods that unconditionally apply rectification.</p>
</sec>
<sec id="s6_2">
<label>6.2</label>
<title>Limitations and Failure Cases</title>
<sec id="s6_2_1">
<label>6.2.1</label>
<title>Dataset Bias</title>
<p>The dataset comprises 10 types of analog gauges collected from a single power data center facility. This introduces the following generalization limitations:<list list-type="bullet">
<list-item>
<p>Limited facility diversity: the lighting conditions, background characteristics, and gauge placement patterns of a single facility dominate the training data. Performance degradation is expected when directly transferring to gauge environments in other industries (petrochemical, manufacturing, or power generation).</p></list-item>
<list-item>
<p>Restricted gauge types: The 10 gauge types included are exclusively circular analog gauges, as these reflect the specific hardware inventory of the target deployment environment. Because non-circular gauges (semi-circular, linear, fan-shaped), multi-needle gauges, and digital-analog hybrid gauges were absent from the facility, they were not included in the current validation.</p></list-item>
<list-item>
<p>Absence of extreme conditions: the dataset does not include outdoor environments, weather variables (rain, fog, dust), or severe occlusion (occlusion <inline-formula id="ieqn-174"><mml:math id="mml-ieqn-174"><mml:mo>&#x003E;</mml:mo><mml:mn>50</mml:mn><mml:mi mathvariant="normal">&#x0025;</mml:mi></mml:math></inline-formula>).</p></list-item>
</list></p>
<p>Despite these limitations, the <inline-formula id="ieqn-175"><mml:math id="mml-ieqn-175"><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>D</mml:mi><mml:mi>I</mml:mi></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mn>1.0</mml:mn></mml:math></inline-formula> result in <xref ref-type="table" rid="table-5">Table 5</xref> suggests that the model has learned an upper-level geometric pattern of &#x201C;circular gauges&#x201D; rather than being confined to the specific gauge types in the training data. Future work on multi-facility dataset construction and domain adaptation techniques is expected to mitigate this limitation.</p>

</sec>
<sec id="s6_2_2">
<label>6.2.2</label>
<title>Structural Limitations for Non-Circular Gauges</title>
<p>The virtual point generation strategy is based on point symmetry (<inline-formula id="ieqn-176"><mml:math id="mml-ieqn-176"><mml:msubsup><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="normal">&#x2032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>), which is geometrically valid only when the scale arc forms a circular arc. The following issues arise for non-circular gauges:<list list-type="bullet">
<list-item>
<p>Semi-circular gauges: the symmetric counterparts of <inline-formula id="ieqn-177"><mml:math id="mml-ieqn-177"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-178"><mml:math id="mml-ieqn-178"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> project outside the gauge body, increasing the condition number of the homography matrix and destabilizing the warp result.</p></list-item>
<list-item>
<p>Linear gauges: when scales are arranged linearly, the definition of <inline-formula id="ieqn-179"><mml:math id="mml-ieqn-179"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> itself becomes ambiguous, rendering point-symmetric virtual point generation meaningless.</p></list-item>
<list-item>
<p>Multi-needle gauges: the current five-keypoint skeleton is designed for a single needle and cannot simultaneously track multiple needles.</p></list-item>
</list></p>
<p>To overcome these structural limitations, a more flexible approach is required. For linear gauges, the model could bypass <inline-formula id="ieqn-180"><mml:math id="mml-ieqn-180"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-181"><mml:math id="mml-ieqn-181"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> to adopt a distance-based decoding logic instead of polar transformation. For multi-needle gauges, the current global keypoint assignment should be evolved into an instance-specific keypoint grouping strategy. Ultimately, a hierarchical architecture that first classifies the gauge geometry and then adaptively assigns the optimal keypoint configuration and rectification strategy will be the focus of future work.</p>
</sec>
<sec id="s6_2_3">
<label>6.2.3</label>
<title>Threshold Dependency</title>
<p>The validity verification thresholds&#x2014;radius ratio 0.4, AR 1.5, and boundary margin 10 px&#x2014;were established based on experimental and geometric rationale but are inherently domain-specific parameters. Readjustment of these thresholds may be necessary when applying the system to different industrial environments or gauge types.</p>
<p>However, each threshold can be interpreted as a continuous quality measure rather than a binary pass/fail judgment. Future research should explore soft-thresholding strategies that compute a continuous quality score combining keypoint confidence and geometric consistency, thereby adjusting the rectification intensity continuously rather than relying on fixed thresholds.</p>
</sec>
</sec>
<sec id="s6_3">
<label>6.3</label>
<title>Practical Deployment Considerations</title>
<sec id="s6_3_1">
<label>6.3.1</label>
<title>Computing Resources and Real-Time Constraints</title>
<p>The inference speeds reported in <xref ref-type="table" rid="table-5">Table 5</xref> were measured in a GPU environment (NVIDIA RTX series). Depending on the deployment scenario in industrial settings, the following considerations apply:<list list-type="bullet">
<list-item>
<p>Edge device deployment: inference FPS may decrease on the SPOT robot&#x2019;s onboard computer (NVIDIA Jetson series). Dynamic adjustment of input resolution (1280 px <inline-formula id="ieqn-182"><mml:math id="mml-ieqn-182"><mml:mo stretchy="false">&#x2194;</mml:mo></mml:math></inline-formula> 640 px) or model optimization through TensorRT may be required to improve throughput.</p>
</list-item>
<list-item>
<p>Server-based processing: in architectures where the robot transmits images for server-side inference, network latency is added. Given the 35.4 ms inference time plus network delay, periodic capture-and-analyze mode is more practical than real-time video stream processing.</p></list-item>
</list></p>
</sec>
<sec id="s6_3_2">
<label>6.3.2</label>
<title>Operation Mode Strategies</title>
<p>The proposed system supports two operational modes:<list list-type="bullet">
<list-item>
<p>Patrol mode: the SPOT robot autonomously navigates predefined routes, capturing and reading gauges at each inspection point. In this mode, reading accuracy takes priority over real-time FPS, and high-resolution input of 1280 px or above is appropriate.</p></list-item>
<list-item>
<p>Live monitoring mode: reading values are displayed in real-time via fixed cameras or during remote control. This mode requires 20 FPS or higher and demands a trade-off between resolution and accuracy.</p></list-item>
</list></p>
</sec>
<sec id="s6_3_3">
<label>6.3.3</label>
<title>Scale Range Pre-Registration</title>
<p>In the value computation process, <inline-formula id="ieqn-183"><mml:math id="mml-ieqn-183"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-184"><mml:math id="mml-ieqn-184"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> (physical minimum and maximum values) must be registered in advance because the scale range differs for each gauge. The current implementation relies on manual input by operators during inspection point registration. Future research should explore automatic recognition of gauge face numerals using OCR (Optical Character Recognition) or Vision-Language Models (VLMs) to automate scale range configuration. Under Pressure [<xref ref-type="bibr" rid="ref-17">17</xref>] presents an initial attempt in this direction, and combining it with the proposed keypoint-based framework could enable a fully autonomous reading system.</p>
</sec>
<sec id="s6_3_4">
<label>6.3.4</label>
<title>Confidence-Based Decision Support</title>
<p>In safety-related readings at industrial sites, explicitly reporting &#x201C;reading unavailable&#x201D; is more important than providing an inaccurate value. The multi-stage validity verification of the proposed system&#x2014;confidence, keypoint completeness, radius ratio, boundary constraints, and AR&#x2014;implements this philosophy by returning &#x201C;reading unavailable&#x201D; for data that fails any verification stage, prompting manual review by operators. This <italic>fail-safe</italic> design is essential for ensuring the reliability of unmanned inspection systems for industrial certification.</p>
</sec>
</sec>
</sec>
<sec id="s7">
<label>7</label>
<title>Applications and Use Cases</title>
<p>This section discusses the practical application of the proposed system in an operational power data center and its potential for expansion across broader industrial fields.</p>
<p>In the currently operating architecture, data collection and data processing are decoupled to ensure system stability and scalability. For data collection, an autonomous quadruped robot (Boston Dynamics SPOT [<xref ref-type="bibr" rid="ref-4">4</xref>]) is utilized to navigate predefined underground infrastructure patrol routes and acquire visual inspection images of various instruments.</p>
<p>The captured images are transmitted to a centralized integrated inspection platform server for processing. This central platform reads analog gauges utilizing the P2-YOLO-Pose based algorithm proposed in this paper. Furthermore, it operates comprehensively, encompassing modules for digital gauge reading and switch/LED status determination. This integrated system enables the simultaneous evaluation of all types of instrument states during a single robot patrol cycle, facilitating comprehensive monitoring of equipment health, real-time dispatch of anomaly alarms, and long-term degradation trend analysis for predictive maintenance.</p>
<p>The automation of this entire pipeline resolves the fundamental limitations of existing manual inspections, which demanded significant manpower, posed risks of incorrect entry due to handwritten records, and placed a heavy burden during night shifts. Particularly for critical safety equipment like fire suppression pressure gauges, where regular inspections are mandated by regulations, the automated, systematic recording of inspection results in a database guarantees high reliability.</p>
<p>Furthermore, the scale-independent geometric rectification capability of the proposed system provides excellent extensibility to the general energy and utility industry sectors. Homography-based distortion correction enables accurate angle interpolation even for industrial gauges with non-uniform, non-linear scales, significantly enhancing the reliability of automated meter reading data. Future research is underway to integrate Large Language Models (LLMs) to fully automate the recognition of gauge face units and maximum/minimum ranges without human intervention, which will serve as the foundation for more universal unmanned inspection automation.</p>
</sec>
<sec id="s8">
<label>8</label>
<title>Conclusion</title>
<p>This study presented a novel framework for automatic analog gauge reading, comprising two core modules: P2-YOLO-Pose-based five-keypoint detection and virtual point-based geometric rectification. The proposed system achieves robust automatic reading in real-world industrial environments where projective distortion is prevalent.</p>
<p>The primary contributions of this work are threefold. First, the geometric structure of analog gauges is modeled as a five-keypoint skeleton (<inline-formula id="ieqn-185"><mml:math id="mml-ieqn-185"><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>m</mml:mi><mml:mi>i</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>c</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>), fundamentally overcoming the limitations of axis-aligned bounding box representations. Second, point-symmetric virtual point generation combined with adaptive homography rectification achieves geometric correction precision comparable to ArUco control markers, without requiring any external fiducial markers. Third, the integration of the high-resolution P2 feature layer and the removal of P5 improve the detection sensitivity for small gauges (<inline-formula id="ieqn-186"><mml:math id="mml-ieqn-186"><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>E</mml:mi><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>) by approximately 40%, with less than 1% FPS degradation.</p>
<p>Experiments on an 11,000-image field dataset collected from a power data center demonstrate Pose mAP50 of 99.45% and Pose mAP50-95 of 99.37%. The consistent decrease in validation loss and <inline-formula id="ieqn-187"><mml:math id="mml-ieqn-187"><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>D</mml:mi><mml:mi>I</mml:mi></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mn>1.0</mml:mn></mml:math></inline-formula> confirm robust generalization without overfitting. The RPY-based rectification analysis shows that the proposed method significantly reduces reading error from 3.5% to 0.6% under severe vertical tilt conditions compared to the raw baseline. Furthermore, the multi-stage validity verification establishes a fail-safe design that outputs a &#x201C;reading unavailable&#x201D; status in highly uncertain scenarios instead of producing erroneous readings, ensuring industrial reliability. Systematic comparisons with GAUREAD, VDN, and Under Pressure validate the distinctive advantages of the corner-free design, nonlinear scale handling, and adaptive rectification strategy.</p>
<p>The main limitations of this study include the dataset bias resulting from a single facility and the restricted applicability to circular analog gauges. To overcome these limitations and achieve more universal industrial applicability, several future research directions are proposed. First, it is necessary to construct a large-scale industrial benchmark dataset encompassing multiple facilities and diverse instrument types. Second, research on more adaptive keypoint topological models is required to extend the diagnostic scope to non-circular surfaces and multi-needle gauges. Third, a soft-thresholding rectification strategy based on continuous quality scores should be introduced to complement the existing fixed-threshold decision logic, thereby controlling the rectification process more precisely. Fourth, an architecture that integrates Vision-Language Models (VLMs) is needed to fully automate the recognition of gauge units and scale ranges without human intervention. Ultimately, the objective is to implement these comprehensive capabilities via real-time distributed inference optimization (e.g., TensorRT) between edge devices and the central control server.</p>
<p>The proposed framework is currently operating within an architecture that combines mobile data acquisition via an autonomous quadruped robot (Boston Dynamics SPOT) with a centralized integrated inspection platform in an operational power data center, providing a highly scalable and practical solution for the universal realization of unmanned inspection automation.</p>
</sec>
</body>
<back>
<ack>
<p>This research was supported by the Korea Electric Power Corporation (KEPCO).</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This research was funded by Korea Electric Power Corporation, grant number R25IA04.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Conceptualization, Jaekyung Lee and Wonhee Kim; methodology, Jaekyung Lee and Youngjun Kim; software, Jaekyung Lee and Byungsung Ko; validation, Jaekyung Lee, Taewon Kim, Jaeheon Park, and Jiwon Lee; formal analysis, Jaekyung Lee and Youngjun Kim; investigation, Jaekyung Lee, Byungsung Ko, and Taewon Kim; data curation, Jaekyung Lee and Jaeheon Park; writing&#x2014;original draft preparation, Jaekyung Lee; writing&#x2014;review and editing, Jaekyung Lee, Taewon Kim and Wonhee Kim; supervision, Wonhee Kim. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The datasets used and/or analyzed during the current study are not publicly available due to confidentiality agreements with the Korea Electric Power Corporation (KEPCO), but are available from the corresponding author upon reasonable request and with permission from KEPCO.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Compare</surname> <given-names>M</given-names></string-name>, <string-name><surname>Baraldi</surname> <given-names>P</given-names></string-name>, <string-name><surname>Zio</surname> <given-names>E</given-names></string-name></person-group>. <article-title>Challenges to IoT-enabled predictive maintenance for industry 4.0</article-title>. <source>IEEE Internet Things J</source>. <year>2020</year>;<volume>7</volume>(<issue>5</issue>):<fpage>4585</fpage>&#x2013;<lpage>97</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jiot.2019.2957029</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Milana</surname> <given-names>E</given-names></string-name>, <string-name><surname>Ram&#x00ED;rez-Agudelo</surname> <given-names>OH</given-names></string-name>, <string-name><surname>Estevam Schmiedt</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Autonomous reading of gauges in unstructured environments</article-title>. <source>Sensors</source>. <year>2022</year>;<volume>22</volume>(<issue>17</issue>):<fpage>6681</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s22176681</pub-id>; <pub-id pub-id-type="pmid">36081139</pub-id></mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Leon-Alcazar</surname> <given-names>J</given-names></string-name>, <string-name><surname>Alnumay</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>C</given-names></string-name>, <string-name><surname>Trigui</surname> <given-names>H</given-names></string-name>, <string-name><surname>Patel</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ghanem</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Learning to read analog gauges from synthetic data</article-title>. <comment>arXiv:2308.14583. 2023</comment>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Boston Dynamics</collab></person-group>. <article-title>Spot</article-title>. <comment>2021 [cited 2021 Jul 2]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.bostondynamics.com/spot">https://www.bostondynamics.com/spot</ext-link>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Tian</surname> <given-names>B</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>H</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>B</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>GaugeTracker: AI-powered cost-effective analog gauge monitoring system</article-title>. In: <conf-name>Proceedings of the 2024 IEEE 7th International Conference on Multimedia Information Processing and Retrieval (MIPR); 2024 Aug 7&#x2013;9</conf-name>; <publisher-loc>San Jose, CA, USA</publisher-loc>. p. <fpage>477</fpage>&#x2013;<lpage>83</lpage>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ren</surname> <given-names>S</given-names></string-name>, <string-name><surname>He</surname> <given-names>K</given-names></string-name>, <string-name><surname>Girshick</surname> <given-names>R</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Faster R-CNN: towards real-time object detection with region proposal networks</article-title>. <comment>arXiv:1506.01497. 2016</comment>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Duda</surname> <given-names>RO</given-names></string-name>, <string-name><surname>Hart</surname> <given-names>PE</given-names></string-name></person-group>. <article-title>Use of the Hough transformation to detect lines and curves in pictures</article-title>. <source>Commun ACM</source>. <year>1972</year>;<volume>15</volume>(<issue>1</issue>):<fpage>11</fpage>&#x2013;<lpage>5</lpage>. doi:<pub-id pub-id-type="doi">10.1145/361237.361242</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zou</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>R</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Automatic recognition reading method of pointer meter based on YOLOv5-MR model</article-title>. <source>Sensors</source>. <year>2023</year>;<volume>23</volume>(<issue>14</issue>):<fpage>6644</fpage>. doi:<pub-id pub-id-type="doi">10.1117/12.2637498</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Alegria</surname> <given-names>FC</given-names></string-name>, <string-name><surname>Serra</surname> <given-names>AC</given-names></string-name></person-group>. <article-title>Automatic calibration of analog and digital measuring instruments using computer vision</article-title>. <source>IEEE Trans Instru Meas</source>. <year>2000</year>;<volume>49</volume>(<issue>1</issue>):<fpage>94</fpage>&#x2013;<lpage>9</lpage>. doi:<pub-id pub-id-type="doi">10.1109/19.836317</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Canny</surname> <given-names>J</given-names></string-name></person-group>. <article-title>A computational approach to edge detection</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>1986</year>;<volume>8</volume>(<issue>6</issue>):<fpage>679</fpage>&#x2013;<lpage>98</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tpami.1986.4767851</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Otsu</surname> <given-names>N</given-names></string-name></person-group>. <article-title>A threshold selection method from gray-level histograms</article-title>. <source>IEEE Trans Syst Man Cybern</source>. <year>1979</year>;<volume>9</volume>(<issue>1</issue>):<fpage>62</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tsmc.1979.4310076</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chi</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Machine vision based automatic detection method of indicating values of a pointer gauge</article-title>. <source>Math Probl Eng</source>. <year>2015</year>;<volume>2015</volume>(<issue>1</issue>):<fpage>283629</fpage>. doi:<pub-id pub-id-type="doi">10.1155/2015/283629</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ma</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>A robust and high-precision automatic reading algorithm of pointer meters based on machine vision</article-title>. <source>Meas Sci Technol</source>. <year>2019</year>;<volume>30</volume>(<issue>1</issue>):<fpage>015401</fpage>. doi:<pub-id pub-id-type="doi">10.1088/1361-6501/ab7487</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>A flexible new technique for camera calibration</article-title>. <source>IEEE Trans Pattern Anal Mach Intell</source>. <year>2000</year>;<volume>22</volume>(<issue>11</issue>):<fpage>1330</fpage>&#x2013;<lpage>4</lpage>. doi:<pub-id pub-id-type="doi">10.1109/34.888718</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Hartley</surname> <given-names>R</given-names></string-name>, <string-name><surname>Zisserman</surname> <given-names>A</given-names></string-name></person-group>. <source>Multiple view geometry in computer vision</source>. <publisher-loc>Cambridge, UK</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>; <year>2003</year>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dong</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Vector detection network: an application study on robots reading analog meters in the wild</article-title>. <source>IEEE Trans Artif Intell</source>. <year>2021</year>;<volume>2</volume>(<issue>5</issue>):<fpage>394</fpage>&#x2013;<lpage>403</lpage>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Reitsma</surname> <given-names>M</given-names></string-name>, <string-name><surname>Keller</surname> <given-names>J</given-names></string-name>, <string-name><surname>Blomqvist</surname> <given-names>K</given-names></string-name>, <string-name><surname>Siegwart</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Under pressure: learning-based analog gauge reading in the wild</article-title>. <comment>arXiv:2404.08785. 2024</comment>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>CY</given-names></string-name>, <string-name><surname>Yeh</surname> <given-names>IH</given-names></string-name>, <string-name><surname>Liao</surname> <given-names>HYM</given-names></string-name></person-group>. <article-title>YOLOv9: learning what you want to learn using programmable gradient information</article-title>. <comment>arXiv:2402.13616. 2024</comment>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Maji</surname> <given-names>D</given-names></string-name>, <string-name><surname>Nagori</surname> <given-names>S</given-names></string-name>, <string-name><surname>Mathew</surname> <given-names>M</given-names></string-name>, <string-name><surname>Poddar</surname> <given-names>D</given-names></string-name></person-group>. <article-title>YOLO-Pose: enhancing YOLO for multi person pose estimation using object keypoint similarity loss</article-title>. <comment>arXiv:2204.06806. 2022</comment>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>K</given-names></string-name>, <string-name><surname>Gkioxari</surname> <given-names>G</given-names></string-name>, <string-name><surname>Doll&#x00E1;r</surname> <given-names>P</given-names></string-name>, <string-name><surname>Girshick</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Mask R-CNN</article-title>. <comment>arXiv:1703.06870. 2018</comment>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Lin</surname> <given-names>TY</given-names></string-name>, <string-name><surname>Doll&#x00E1;r</surname> <given-names>P</given-names></string-name>, <string-name><surname>Girshick</surname> <given-names>R</given-names></string-name>, <string-name><surname>He</surname> <given-names>K</given-names></string-name>, <string-name><surname>Hariharan</surname> <given-names>B</given-names></string-name>, <string-name><surname>Belongie</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Feature pyramid networks for object detection</article-title>. <comment>arXiv:1612.03144. 2017</comment>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Bergmann</surname> <given-names>P</given-names></string-name>, <string-name><surname>Fauser</surname> <given-names>M</given-names></string-name>, <string-name><surname>Sattlegger</surname> <given-names>D</given-names></string-name>, <string-name><surname>Steger</surname> <given-names>C</given-names></string-name></person-group>. <article-title>MVTec AD&#x2014;a comprehensive real-world dataset for unsupervised anomaly detection</article-title>. In: <conf-name>Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; 2019 Jun 15&#x2013;20</conf-name>; <publisher-loc>Long Beach, CA, USA</publisher-loc>. p. <fpage>9584</fpage>&#x2013;<lpage>92</lpage>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Jocher</surname> <given-names>G</given-names></string-name>, <string-name><surname>Qiu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chaurasia</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Ultralytics YOLO. Ultralytics</article-title>. <comment>2023 [cited 2026 Mar 29]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://github.com/ultralytics/ultralytics">https://github.com/ultralytics/ultralytics</ext-link>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Khanam</surname> <given-names>R</given-names></string-name>, <string-name><surname>Hussain</surname> <given-names>M</given-names></string-name></person-group>. <article-title>YOLOv11: an overview of the key architectural enhancements</article-title>. <comment>arXiv:2410.17725. 2024</comment>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Buslaev</surname> <given-names>A</given-names></string-name>, <string-name><surname>Iglovikov</surname> <given-names>VI</given-names></string-name>, <string-name><surname>Khvedchenya</surname> <given-names>E</given-names></string-name>, <string-name><surname>Parinov</surname> <given-names>A</given-names></string-name>, <string-name><surname>Druzhinin</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kalinin</surname> <given-names>AA</given-names></string-name></person-group>. <article-title>Albumentations: fast and flexible image augmentations</article-title>. <source>Information</source>. <year>2020</year>;<volume>11</volume>(<issue>2</issue>):<fpage>125</fpage>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lin</surname> <given-names>TY</given-names></string-name>, <string-name><surname>Maire</surname> <given-names>M</given-names></string-name>, <string-name><surname>Belongie</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hays</surname> <given-names>J</given-names></string-name>, <string-name><surname>Perona</surname> <given-names>P</given-names></string-name>, <string-name><surname>Ramanan</surname> <given-names>D</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Microsoft coco: common objects in context</article-title>. In: <conf-name>Proceedings of the Computer Vision&#x2013;ECCV 2014: 13th European Conference; 2014 Sep 6&#x2013;12</conf-name>. <publisher-loc>Zurich, Switzerland</publisher-loc>. p. <fpage>740</fpage>&#x2013;<lpage>55</lpage>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Bochkovskiy</surname> <given-names>A</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>CY</given-names></string-name>, <string-name><surname>Liao</surname> <given-names>HYM</given-names></string-name></person-group>. <article-title>YOLOv4: optimal speed and accuracy of object detection</article-title>. <comment>arXiv:2004.10934. 2020</comment>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ghiasi</surname> <given-names>G</given-names></string-name>, <string-name><surname>Cui</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Srinivas</surname> <given-names>A</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>R</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>TY</given-names></string-name>, <string-name><surname>Cubuk</surname> <given-names>ED</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Simple copy-paste is a strong data augmentation method for instance segmentation</article-title>. <comment>arXiv:2012.07177. 2021</comment>.</mixed-citation></ref>
</ref-list>
</back></article>