<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">SDHM</journal-id>
<journal-id journal-id-type="nlm-ta">SDHM</journal-id>
<journal-id journal-id-type="publisher-id">SDHM</journal-id>
<journal-title-group>
<journal-title>Structural Durability &#x0026; Health Monitoring</journal-title>
</journal-title-group>
<issn pub-type="epub">1930-2991</issn>
<issn pub-type="ppub">1930-2983</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">66098</article-id>
<article-id pub-id-type="doi">10.32604/sdhm.2025.066098</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Diff-Fastener: A Few-Shot Rail Fastener Anomaly Detection Framework Based on Diffusion Model</article-title>
<alt-title alt-title-type="left-running-head">Diff-Fastener: A Few-Shot Rail Fastener Anomaly Detection Framework Based on Diffusion Model</alt-title>
<alt-title alt-title-type="right-running-head">Diff-Fastener: A Few-Shot Rail Fastener Anomaly Detection Framework Based on Diffusion Model</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Sun</surname><given-names>Peng</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Yao</surname><given-names>Dechen</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref><email>yaodechen@bucea.edu.cn</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Yang</surname><given-names>Jianwei</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Long</surname><given-names>Quanyu</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<aff id="aff-1"><label>1</label><institution>School of Mechanical-Electrical and Vehicle Engineering, Beijing University of Civil Engineering and Architecture</institution>, <addr-line>Beijing, 100044</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>Key Laboratory of Performance Guarantee on Urban Rail Transit Vehicles, Beijing University of Civil Engineering and Architecture</institution>, <addr-line>Beijing, 100044</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Dechen Yao. Email: <email>yaodechen@bucea.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2025</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>05</day><month>09</month><year>2025</year>
</pub-date>
<volume>19</volume>
<issue>5</issue>
<fpage>1221</fpage>
<lpage>1239</lpage>
<history>
<date date-type="received">
<day>29</day>
<month>3</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>26</day>
<month>5</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 The Authors.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="_SDHM_66098.pdf"></self-uri>
<abstract>
<p>Supervised learning-based rail fastener anomaly detection models are limited by the scarcity of anomaly samples and perform poorly under data imbalance conditions. However, unsupervised anomaly detection methods based on diffusion models reduce the dependence on the number of anomalous samples but suffer from too many iterations and excessive smoothing of reconstructed images. In this work, we have established a rail fastener anomaly detection framework called Diff-Fastener, the diffusion model is introduced into the fastener detection task, half of the normal samples are converted into anomaly samples online in the model training stage, and One-Step denoising and canonical guided denoising paradigms are used instead of iterative denoising to improve the reconstruction efficiency of the model while solving the problem of excessive smoothing. DACM (Dilated Attention Convolution Module) is proposed in the middle layer of the reconstruction network to increase the detail information of the reconstructed image; meanwhile, Sparse-Skip connections are used instead of dense connections to reduce the computational load of the model and enhance its scalability. Through exhaustive experiments on MVTec, VisA, and railroad fastener datasets, the results show that Diff-Fastener achieves 99.1% Image AUROC (Area Under the Receiver Operating Characteristic) and 98.9% Pixel AUROC on the railroad fastener dataset, which outperforms the existing models and achieves the best average score on MVTec and VisA datasets. Our research provides new ideas and directions in the field of anomaly detection for rail fasteners.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Diffusion model</kwd>
<kwd>anomaly detection</kwd>
<kwd>unsupervised learning</kwd>
<kwd>rail fastener</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>National Natural Science Foundation of China</funding-source>
<award-id>52272385</award-id>
<award-id>52475085</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Railroad fasteners are key components of the railroad system and play a critical role in maintaining the stability of rail connections and ensuring safety during operation. However, due to long-term use and the influence of external environmental factors, railroad fasteners may experience abnormalities such as missing and damaged parts, which may lead to railroad accidents and operational failures [<xref ref-type="bibr" rid="ref-1">1</xref>]. Therefore, timely and accurate inspection of railroad fasteners is crucial.</p>
<p>However, anomaly detection methods based on supervised learning are limited by the number of samples and perform poorly with unbalanced data. Therefore, unsupervised anomaly detection methods based on unsupervised learning have wider applicability.</p>
<p>Unsupervised anomaly detection models are mainly divided into two types: feature embedding-based models and generative model-based models. PatchCore [<xref ref-type="bibr" rid="ref-2">2</xref>] is an advanced method for anomaly detection that utilizes feature embedding. This model extracts local features from images and uses a memory bank to store the features of normal samples, allowing it to compare the differences between abnormal regions and normal features during the detection process. RD&#x002B;&#x002B; (Reversed Distillation&#x002B;&#x002B;) [<xref ref-type="bibr" rid="ref-3">3</xref>] introduces more complex feature fusion strategies and optimization algorithms, enabling it to capture subtle anomalous information better, making it suitable for high-dimensional and complex anomaly detection tasks. SimpleNet [<xref ref-type="bibr" rid="ref-4">4</xref>] is a simplified feature embedding model designed to enhance computational efficiency by reducing model complexity. FastFlow [<xref ref-type="bibr" rid="ref-5">5</xref>] utilizes flow-based generative models for anomaly detection. DRAEM (Discriminatively trained Reconstruction Anomaly Embedding Model) [<xref ref-type="bibr" rid="ref-6">6</xref>] combines denoising autoencoders with reconstruction methods to detect anomalies by denoising and reconstructing input data. Its innovation lies in its dual reconstruction mechanism, which allows for more accurate identification of anomalous regions while maintaining high-fidelity reconstruction of normal data.</p>
<p>Recently, diffusion modeling has been explored in the field of anomaly detection [<xref ref-type="bibr" rid="ref-7">7</xref>&#x2013;<xref ref-type="bibr" rid="ref-10">10</xref>]. DDPM (Denoising Diffusion Probabilistic Model) [<xref ref-type="bibr" rid="ref-11">11</xref>] demonstrates competitive effectiveness relative to preceding unsupervised and semi-supervised anomaly detection methodologies. It does not require labelled anomalous samples for training, while also exhibiting strong robustness, allowing it to be flexibly applied to different types of data, including images, time series, texts, etc. DDAD (Denoising Diffusion Anomaly Detection) [<xref ref-type="bibr" rid="ref-12">12</xref>] provides a conditional denoising process to generate anomaly-free images that are similar to the target image, improving inference speed while maintaining equivalent anomaly detection performance. Within the realm of healthcare, denoising diffusion models have been employed in the identification of brain tumors [<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-14">14</xref>]. AnoDDPM [<xref ref-type="bibr" rid="ref-15">15</xref>] developed a multi-scale simplex noise diffusion process capable of controlling the target anomaly size. DiAD [<xref ref-type="bibr" rid="ref-16">16</xref>] is a diffusion-based framework containing a semantic bootstrap module and a spatial-aware feature fusion module, which solves the problem of category and semantic loss in Stable Diffusion multi-class anomaly detection. DefectFill [<xref ref-type="bibr" rid="ref-17">17</xref>] fine-tunes the repair diffusion model to produce realistic and high-fidelity defect images with limited reference samples. DiffusionAD [<xref ref-type="bibr" rid="ref-18">18</xref>] consists of a reconstruction sub-network, which uses Residual U-Net to reformulate the reconstruction process as a noise-to-normal paradigm, and a segmentation sub-network, which utilizes the input graphs and their anomaly-free recovery results to predict the pixel-level anomaly scores, and is experimentally shown to outperform the current state-of-the-art (SOTA) methods.</p>
<p>Nevertheless, the enhanced expressiveness and interpretability of the DDPM incur significant computational expenses. Such computational intricacy poses obstacles for anomaly detection endeavors encompassing extensive datasets or continuous data streams. Additionally, one of the key components in diffusion models is the U-Net used for noise prediction, continuous downsampling results in reduced resolution and thus, degradation of fine details. The contributions of this paper can be summarized as follows:
<list list-type="simple">
<list-item><label>(1)</label><p>To the best of our knowledge, this work introduces a diffusion model to the rail fastener anomaly detection task for the first time, named Diff-Fastener. It requires only a small number of normal samples in the training phase, and uses an anomaly synthesis strategy to convert half of the normal samples into anomalous samples online. This feature not only reduces the dependence on the amount of anomaly data but also significantly enhances the model&#x2019;s ability to detect unknown anomalies.</p></list-item>
<list-item><label>(2)</label><p>This work employs one-step denoising to improve the efficiency of image reconstruction, while using a Norm-Guided strategy to alleviate the excessive smoothing problem it brings. The reconstruction results of anomalous regions obtained at relatively large noise scales are utilized to guide the reconstruction at smaller noise scales to improve the accuracy of the reconstruction.</p></list-item>
<list-item><label>(3)</label><p>We propose a module, DACM, used in the middle layer of the reconstructed network to filter and weight the features and increase the receptive field. Sparse-Skip connections are used instead of the original dense skip connections to reduce the computational load of the model and improve scalability, and all activation functions are replaced with GELU (Gaussian Error Linear Unit). Finally, we perform exhaustive experiments on the MVTec, VisA, and railroad fastener datasets.</p></list-item>
</list></p>
</sec>
<sec id="s2">
<label>2</label>
<title>Preliminaries</title>
<sec id="s2_1">
<label>2.1</label>
<title>Diffusion Models</title>
<p>Diffusion models [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-20">20</xref>] represent a category of generative models that draw inspiration from the principles of nonequilibrium thermodynamics. They establish a framework wherein the forward process incrementally introduces random noise into the data, whereas the reverse process systematically generates the targeted data samples from noise.</p>
<sec id="s2_1_1">
<label>2.1.1</label>
<title>Diffusion Process</title>
<p>The diffusion process entails the incremental alteration of the original image <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> into <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> by incrementally adding Gaussian noise, thereby achieving the goal of corrupting the image. This process is also known as the forward (positive) noise addition process, which the formula can represent:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msqrt><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:msqrt><mml:mo>+</mml:mo><mml:msub><mml:mo>&#x03F5;</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:msqrt><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:msqrt><mml:mo>,</mml:mo><mml:msub><mml:mo>&#x03F5;</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x223C;</mml:mo><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>This process transforms the data sample <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> into a noise sample <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, where <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>t</mml:mi></mml:math></inline-formula> is randomly drawn from <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>T</mml:mi><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula>, <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mo>&#x03F5;</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>&#x223C;</mml:mo><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represents Gaussian noise, and <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msubsup><mml:mrow><mml:mo>{</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>}</mml:mo></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> are predefined hyperparameters, known as the noise schedule. These <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> values are crucial as they control the variance of the noise added at each step. A higher <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> implies less noise addition, preserving more of the original image&#x2019;s characteristics at step <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>t</mml:mi></mml:math></inline-formula>, while a lower <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> means more noise is introduced, rapidly transforming the image towards a pure noise state. Typically, <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is set in a decreasing sequence, often following a linear or cosine-based schedule, to ensure a smooth transition from the original image to noise. The choice of the <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> schedule significantly impacts the model&#x2019;s ability to learn the underlying data distribution; a well-designed schedule can facilitate better convergence during training and improve the quality of generated samples.</p>
<p>By iteratively deriving from the above formula, we can obtain the direct transformation formula from <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> to <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> as follows:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:msqrt><mml:msub><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:msqrt><mml:mo>+</mml:mo><mml:msub><mml:mo>&#x03F5;</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:msqrt><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:msqrt><mml:mo>,</mml:mo><mml:msub><mml:mo>&#x03F5;</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x223C;</mml:mo><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>given that <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:msub><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x220F;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, which is a hyperparameter set according to the noise schedule, and <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mo>&#x03F5;</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x223C;</mml:mo><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is also Gaussian noise. Both <xref ref-type="disp-formula" rid="eqn-1">Eqs. (1)</xref> and <xref ref-type="disp-formula" rid="eqn-2">(2)</xref> can describe the forward noise addition process, with the former used for gradually corrupting an image and the latter for corrupting an image in one step.</p>
</sec>
<sec id="s2_1_2">
<label>2.1.2</label>
<title>Reverse Process</title>
<p>The diffusion process encompasses the addition of noise to the image, whereas the reverse process engages in denoising. Knowing the true distribution of each step in the reverse process, denoted as <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mi>q</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, and beginning with random noise <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x223C;</mml:mo><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, a gradual denoising procedure can yield an authentic sample. The parameter <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, related to <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> through <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mrow><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, plays a significant role in the reverse process. A larger <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> implies a more aggressive denoising step, which might introduce more variance in the generated samples but could also lead to faster convergence towards the original data distribution. In the reverse process equation:
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:msqrt><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:msqrt></mml:mfrac><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:msqrt><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:msqrt></mml:mfrac><mml:msub><mml:mo>&#x03F5;</mml:mo><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B2;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mo>&#x223C;</mml:mo><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>given <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mrow><mml:mover><mml:mi>&#x03B2;</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. Since the real noise <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msub><mml:mo>&#x03F5;</mml:mo><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> in the forward process equations is not allowed to be used during the restoration process, the key is to obtain a model <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msub><mml:mo>&#x03F5;</mml:mo><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> that predicts the noise from <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mi>t</mml:mi></mml:math></inline-formula>, where <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mi>&#x03B8;</mml:mi></mml:math></inline-formula> represents the training parameters of the model, and <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mrow><mml:mi>z</mml:mi></mml:mrow><mml:mo>&#x223C;</mml:mo><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is another Gaussian noise used to represent the difference between the prediction and the actual noise. The balance between <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msub><mml:mi>&#x03B1;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> in the reverse process determines the stability and quality of the generated output, with improper choices potentially leading to blurry or inconsistent results.</p>
</sec>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>How to Train</title>
<p>To estimate the distribution <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mi>q</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, the entire training sample is required. We can estimate these distributions using neural networks. Although the derivation behind diffusion models is complex, the optimization goal we ultimately obtain is very straightforward: to make the noise predicted by the network consistent with the actual noise. During the training process, the following loss function is minimized by fitting through a Residual U-Net structure to predict <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mo>&#x03F5;</mml:mo></mml:math></inline-formula>:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mo>&#x223C;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>T</mml:mi><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x223C;</mml:mo><mml:mi>q</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x03F5;</mml:mo><mml:mo>&#x223C;</mml:mo><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>I</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:msup><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:mo>&#x03F5;</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mo>&#x03F5;</mml:mo><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>]</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></disp-formula></p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Preliminaries</title>
<sec id="s3_1">
<label>3.1</label>
<title>Network Structure</title>
<p>The core components of Diff-Fastener are the reconstruction network and the segmentation network. Both the reconstruction network and the segmentation network have U-Net as their core. A workflow of Diff-Fastener is presented in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Workflow diagram of Diff-Fastener. After the original images are input into the model, they are disturbed by Gaussian noise. Through training, the model predicts the noise, and testing is used to learn how to reconstruct the original images. During the inference stage, the diffusion model reconstructs images of abnormal components, and by pixel matching, compares the reconstructed images with the original images to generate accurate anomaly heatmaps</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_66098-fig-1a.tif"/>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_66098-fig-1b.tif"/>
</fig>
<sec id="s3_1_1">
<label>3.1.1</label>
<title>Reconstruction Model</title>
<p>The reconstruction network is designed to predict the noise <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mi>&#x03B5;</mml:mi></mml:math></inline-formula> related to <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mi>t</mml:mi></mml:math></inline-formula>, based on a U-Net-like architecture integrating ResNet [<xref ref-type="bibr" rid="ref-21">21</xref>], PixelCNN&#x002B;&#x002B; [<xref ref-type="bibr" rid="ref-22">22</xref>], and Transformer [<xref ref-type="bibr" rid="ref-23">23</xref>]. In the noise prediction module, the encoder uses ResNet blocks with skip connections for hierarchical feature extraction, the middle layers adopt PixelCNN&#x002B;&#x002B; to model pixel-level spatial dependencies, and the decoder incorporates Transformer layers to capture long-range dependencies. For the module-to-module collaboration, U-Net&#x2019;s skip connections transfer low-level details from the encoder to the decoder, enabling the decoder to combine high-level semantics with fine-grained information. ResNet, PixelCNN&#x002B;&#x002B;, and Transformer components work together to enhance feature representation, refinement and synthesis. The choice of this combination is theoretically justified as U-Net suits detail recovery, ResNet boosts representational power, PixelCNN&#x002B;&#x002B; captures local structure, and Transformer overcomes the limitation of convolutional layers in long-range dependency modeling, achieving a balance between local and global understanding for better noise prediction.</p>
<p>In addition to using a U-Net-like structure as the backbone, each block in the decoder is connected to its corresponding encoder block through skip connections. Early diffusion works [<xref ref-type="bibr" rid="ref-24">24</xref>,<xref ref-type="bibr" rid="ref-25">25</xref>] have demonstrated the importance of skip connections in this architecture. On the one hand, in the encoder-decoder structure, they reduce the information loss during the downsampling process by directly passing the feature maps to the decoder, thus enriching its representation. On the other hand, skip connections accelerate the training process and improve the generation quality by providing additional gradient paths and preserving important features, enabling more efficient and effective diffusion modelling. However, the dense block-wise skip connections have become a bottleneck. A large number of block-wise skip features have to be concatenated to the features and merged along the channel dimension within the corresponding decoder blocks, which consumes excessive computational costs. Such inefficiency hinders the scalability of the model. To address this issue, we propose Sparse-Skip connections, in which the skip connections are applied only every few blocks instead of after each block. This approach improves the performance, and our experiments show that, compared with the densely connected skip design, it indeed yields better results. <xref ref-type="fig" rid="fig-2">Fig. 2</xref> illustrates the differences between the two U-Net structures.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>U-Net with different numbers of skip connections</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_66098-fig-2.tif"/>
</fig>
<p>Initially, the input image <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is corrupted at a random time step <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mi>t</mml:mi></mml:math></inline-formula> during the diffusion process, resulting in <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> as per <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>. As the time step increases, the input image <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> gradually loses its distinctive features&#x2014;the anomalous pixels lose their sharp characteristics and approach an isotropic Gaussian distribution, where <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> represents a normal image or a composite abnormal image. Based on this, the reconstruction network can iteratively obtain the reconstructed anomaly-free <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> through <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref>.</p>
</sec>
<sec id="s3_1_2">
<label>3.1.2</label>
<title>Segmentation Model</title>
<p>The segmentation model, based on U-Net, identifies and classifies each pixel in the image as normal or abnormal, thus achieving precise segmentation of the anomalous areas. It compares the reconstructed image <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> with the original image <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and learns to detect anomalies by learning the commonalities and differences between them.</p>
</sec>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>One-Step Denoising</title>
<p>Although diffusion models possess excellent density estimation capabilities and high sampling quality, classical diffusion models generally require 50&#x2013;1000 iterations, and each iteration corresponds to a round of network inference. Although this multi-step iterative method can gradually refine the reconstruction and ensure high-quality output, the calculation cost is extremely high, resulting in a very slow reasoning speed, which makes it difficult to meet the requirements of real-time application of fastener detection. To address this issue, based on the direct reconstruction using the DDPM [<xref ref-type="bibr" rid="ref-11">11</xref>] theory, we adopt one-step denoising as an alternative to iterative denoising:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:msqrt><mml:msub><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:msqrt></mml:mfrac><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msqrt><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:msqrt><mml:msub><mml:mo>&#x03F5;</mml:mo><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p><inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> is the anomaly-free reconstruction achieved through one-step denoising. Simply put, at any time step <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mi>t</mml:mi></mml:math></inline-formula>, once the diffusion model predicts the noise <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msub><mml:mo>&#x03F5;</mml:mo><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> for <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> through a single-step inference, direct recovery is always effective. By gradually adding noise and employing one-step denoising, it effectively simulates the distribution of normal data while significantly improving inference speed. Compared to <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref>, this direct prediction method is <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:mi>t</mml:mi></mml:math></inline-formula> times faster than iterative prediction.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title> Norm-Guided Paradigm</title>
<p>In one-step denoising, the noise prediction of the model is completed at one time and cannot be corrected many times, which may lead to insufficient reconstruction of the details of the model, especially in the case of large noise scales. If the model&#x2019;s prediction of noise is wrong, this error will directly affect the result of reconstruction and may lead to fastener image distortion, which is more obvious in images containing complex abnormal features, resulting in residual abnormal features.</p>
<p>Experiments have proven that different scales of noise injection are required to repair anomalies of different sizes. Larger abnormal areas require larger-scale noise to disturb, and smaller or non-existent abnormal areas require smaller-scale noise. But large-scale noise may introduce distortion, and some anomalous regions will be left when the diffusion model is used for prediction. Although small-scale noise is difficult to disturb large-area anomalies, it exhibits higher pixel quality and preserves more image details.</p>
<p>To take full advantage of these two noise scales, we use the Norm-Guided paradigm [<xref ref-type="bibr" rid="ref-18">18</xref>] to guide the denoising process of the diffusion model. Specifically, large-scale noise is injected into the original image and denoised to obtain the reconstructed image <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. Then, small-scale noise is injected into the original image and denoised under the guidance of <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> to obtain the final reconstructed image <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. In this way, <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> not only carries out a good reconstruction of anomalies in large areas, but also fully retains the details of the original image. See <xref ref-type="fig" rid="fig-3">Fig. 3</xref> for detailed steps.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Norm-guided paradigm</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_66098-fig-3.tif"/>
</fig>
<p>We divide the random range of <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mi>t</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>T</mml:mi><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> into two parts using <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mi>&#x03C4;</mml:mi></mml:math></inline-formula>, where <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mi>S</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>&#x03C4;</mml:mi><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mi>B</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mi>&#x03C4;</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>&#x03C4;</mml:mi><mml:mo>+</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>T</mml:mi><mml:mo>}</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math></inline-formula> For an input image <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, we first perturb it using two randomly sampled time steps <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mi>S</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mi>B</mml:mi></mml:math></inline-formula> to obtain <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> through <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>. Then, we use the diffusion model to predict the noise of <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> through <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref>, and generate a reconstructed image <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mn>0</mml:mn><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, denoted by <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:mi>n</mml:mi></mml:math></inline-formula>. Then, under the guidance of <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mn>0</mml:mn><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, we use the diffusion model to predict the noise of <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and denoise <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>, bringing <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:msub><mml:mo>&#x03F5;</mml:mo><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> into <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref>, which yields <xref ref-type="disp-formula" rid="eqn-6">Eq. (6)</xref>:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mn>0</mml:mn><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:msqrt><mml:msub><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:msqrt></mml:mfrac><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msqrt><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mover><mml:mi>&#x03B1;</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:msqrt><mml:msub><mml:mo>&#x03F5;</mml:mo><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>finally generate the reconstructed image <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mn>0</mml:mn><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>. <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mn>0</mml:mn><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> performs well in the reconstruction of large abnormal areas and retains the rich details of the original image.</p>
<p>We can see from <xref ref-type="fig" rid="fig-4">Fig. 4</xref> that before using the norm-guided paradigm, the One-Step denoising model lacked a grasp of detailed information, and the reconstructed abnormal fastener with a distorted shape occurred after the displacement was reconstructed. After using the norm-guided paradigm, the reconstructed image showed obvious improvement in detailed information, with a more natural coupling shape and a texture on the rail that was closer to the real state.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>A comparison of the results before and after using the norm-guided paradigm</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_66098-fig-4.tif"/>
</fig>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>DACM</title>
<p>In the face of various situations such as fastener damages of different degrees, complex backgrounds, and harsh acquisition environments, even a powerful network like the residual U-Net network is prone to problems such as unstable performance and reduced generalization ability when dealing with complex and diverse datasets. Therefore, we propose a module named DACM. Experiments have proven that the introduction of DACM can effectively alleviate such problems, especially with remarkable effects in improving accuracy and restoring details.</p>
<p>The DACM module integrates the attention module and dilated convolutions. The former deals with the channel distribution of the feature map and helps the neural network to focus on key pixel regions while ignoring irrelevant parts. The dilated convolution with a kernel size of 3, a dilation rate of 2, and a padding of 2 in the middle layer expands the receptive field without losing spatial resolution to retain more details [<xref ref-type="bibr" rid="ref-26">26</xref>]. It solves the problem of the loss of fastener detail features caused by multiple downsamplings without increasing the number of parameters, provides richer contextual information, and enhances the model&#x2019;s understanding of complex relationships. In addition, the DACM allows for the adjustment of the dilation rate to enable the model to flexibly adapt to feature detection at different granularities, significantly improving performance in complex scenarios. <xref ref-type="fig" rid="fig-5">Fig. 5</xref> shows our DACM.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Structure of DACM. Wherein, &#x201C;G&#x201D; refers to the GELU activation function</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_66098-fig-5.tif"/>
</fig>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Abnormal Fastener Synthesis Strategy</title>
<p>To train the model for detecting abnormal fasteners, in practical applications, we usually lack sufficient labeled samples of abnormal fasteners. Inspired by [<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-28">28</xref>], we adopt an anomaly synthesis strategy to synthesize forged anomaly railway fastener images online, enabling the model to be trained without real anomaly samples. Specifically, our synthesis strategy is to generate areas that resemble actual anomalies by adding inconsistent visual perturbations to the normal fastener images. These forged abnormal areas are defined as &#x201C;out-of-distribution&#x201D; areas that simulate the different types of defects that fasteners can have in the real world.</p>
<p>In addition to the above-mentioned methods, we have replaced all the activation functions with GELU. The overall structure of Diff-Fastener is shown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>, which primarily divides the anomaly detection of rail fasteners into three steps. The first step is data collection and preprocessing, where images of fasteners are collected on a railway image acquisition device and labelled with ground truth on a computer. The second step involves using Diff-Fastener for anomaly detection; the model only requires normal fastener images as input. Initially, the model will transform half of the normal images into anomalous images using an abnormal fastener synthesis strategy, and then, in combination with One-Step denoising and Norm-Guided paradigm, it will detect the fasteners. The final step is data visualization, where the model displays the detection results using heatmaps, making it easier for people to observe the results.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>The overall framework of Diff-Fastener</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_66098-fig-6.tif"/>
</fig>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiments</title>
<sec id="s4_1">
<label>4.1</label>
<title>Datasets</title>
<p>The railway image acquisition device used in this study is shown in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>. The device can adapt to the data collection of tracks with normal gauge, and the collected image field of view includes the rails on both sides and the fasteners. The test route was a certain section of the high-speed ballastless track between Beijing and Shanghai. The basic states of the rail fasteners are mainly divided into four categories: normal, missing, broken, and displaced.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Railway information collection device</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_66098-fig-7.tif"/>
</fig>
<p>After a series of image processing steps, the images collected by the track inspection vehicle are cropped to 512 &#x00D7; 512 pixels. The open-source software LabelMe is used to annotate the ground truth information in the track images.</p>
<p>For the supervised model, we have a total of 1600 fastener images, consisting of 800 normal fastener images and 800 abnormal fastener images. To ensure a more robust evaluation of the model, we adopt <italic>k</italic>-fold cross-validation with <italic>k</italic> &#x003D; 5, i.e., the 1600 images are divided into 5 subsets. In each round of cross-validation, 4 subsets (a total of 1280 images) are used for training, and 1 subset (320 images) is used for validation. After 5 rounds of training and validation, we calculate the average performance metrics to obtain a more comprehensive assessment of the model&#x2019;s performance.</p>
<p>For the unsupervised model, we use 176 images for training. Among these, half of the images (88 images) are automatically turned into abnormal images through the model&#x2019;s abnormal synthesis strategy. Similar to the supervised model, we also use 5-fold cross-validation. The 176 images are divided into 5 subsets. In each iteration, 4 subsets (around 140 images) are used for training, and 1 subset (around 35 images) is used for validation. We summarize the performance of the unsupervised model across the 5 rounds of cross-validation to understand its generalization ability.</p>
<p>The overall segmentation ratio of the training set to the test set is 7:3. In addition, we set aside 200 abnormal fastener images for model inference.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Training &#x0026; Inference</title>
<sec id="s4_2_1">
<label>4.2.1</label>
<title>Training Stage</title>
<p>We trained the reconstruction model and segmentation model together. The reconstruction model learns the entire distribution of a normal sample (<inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>) by minimizing the following loss function:
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msub><mml:mo>&#x03F5;</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mo>&#x03F5;</mml:mo><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msub><mml:mo>&#x03F5;</mml:mo><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>b</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mo>&#x03F5;</mml:mo><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mn>2</mml:mn></mml:mfrac></mml:math></disp-formula></p>
<p>The segmentation model uses the commonalities and differences between <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:msub><mml:mrow><mml:mover><mml:mi>x</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:msub><mml:mn>0</mml:mn><mml:mrow><mml:mi>s</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> to predict pixel-level anomaly scores that are as close as possible to the ground truth. Segmentation loss is defined as:
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>S</mml:mi><mml:mi>m</mml:mi><mml:mi>o</mml:mi><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>M</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:mi>M</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>&#x03B3;</mml:mi><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:mi>M</mml:mi><mml:mo>,</mml:mo><mml:mrow><mml:mover><mml:mi>M</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:mi>M</mml:mi></mml:math></inline-formula> is the truth mask of the input image, and <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mrow><mml:mover><mml:mi>M</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> is the output of the segmentation model. <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:mi>S</mml:mi><mml:mi>m</mml:mi><mml:mi>o</mml:mi><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> were applied simultaneously to reduce oversensitivity to outliers and accurately segment difficult exception examples. <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mi>&#x03B3;</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mo>+</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula> is the hyperparameter that controls the importance of <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>f</mml:mi><mml:mi>o</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. Therefore, the total loss of Diff-Fastener use for combined training is:
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>t</mml:mi><mml:mi>o</mml:mi><mml:mi>t</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>o</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>s</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula></p>
</sec>
<sec id="s4_2_2">
<label>4.2.2</label>
<title>Inference Stage</title>
<p>To achieve higher inference speed and reconstruction quality, we still use one-step norm-guided estimation in the inference stage. After the segmentation model predicts the pixel-level anomaly fraction <inline-formula id="ieqn-84"><mml:math id="mml-ieqn-84"><mml:mrow><mml:mover><mml:mi>M</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula>, we take the average of the first <inline-formula id="ieqn-85"><mml:math id="mml-ieqn-85"><mml:mi>K</mml:mi></mml:math></inline-formula> anomaly pixels in <inline-formula id="ieqn-86"><mml:math id="mml-ieqn-86"><mml:mrow><mml:mover><mml:mi>M</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> as the image-level anomaly fraction. This inference method is hundreds of times faster while maintaining sampling quality.</p>
</sec>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Evaluation Metrics</title>
<p>To evaluate the anomaly detection performance of Diff-Fastener, we use the following metrics: Image AUROC, Pixel AUROC, precision rate and recall rate. Image AUROC is a metric used to evaluate the anomaly detection performance of the entire image and is the most widely used measure for anomaly assessment. It generates an ROC (Receiver Operating Characteristic) curve by calculating the true positive rate and false positive rate at different thresholds and computes the area under this curve. It provides a comprehensive performance evaluation for the entire image, making it suitable for scenarios that require consideration of overall anomalies in the image. Pixel AUROC is a metric used to assess the model&#x2019;s anomaly detection performance at the pixel level. Unlike Image AUROC, Pixel AUROC focuses on the anomaly detection scores for each pixel and calculates the true positive rate and false positive rate at different thresholds. Pixel AUROC offers a detailed evaluation for each pixel, making it suitable for tasks that require precise localization of anomalous regions, such as image segmentation and fine-grained classification.
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-12"><label>(12)</label><mml:math id="mml-eqn-12" display="block"><mml:mrow><mml:mtext>AUROC</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:msubsup><mml:mo>&#x222B;</mml:mo><mml:mrow><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mrow><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mi>R</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mi>R</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-13"><label>(13)</label><mml:math id="mml-eqn-13" display="block"><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-14"><label>(14)</label><mml:math id="mml-eqn-14" display="block"><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-15"><label>(15)</label><mml:math id="mml-eqn-15" display="block"><mml:mrow><mml:mtext>AUPRO</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:msubsup><mml:mo>&#x222B;</mml:mo><mml:mrow><mml:mn>0</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>d</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>Note: <inline-formula id="ieqn-87"><mml:math id="mml-ieqn-87"><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:math></inline-formula> stands for True Positive, <inline-formula id="ieqn-88"><mml:math id="mml-ieqn-88"><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:math></inline-formula> for False Positive, <inline-formula id="ieqn-89"><mml:math id="mml-ieqn-89"><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:math></inline-formula> for True Negative, <inline-formula id="ieqn-90"><mml:math id="mml-ieqn-90"><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:math></inline-formula> for False Negative, <inline-formula id="ieqn-91"><mml:math id="mml-ieqn-91"><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mi>R</mml:mi></mml:math></inline-formula> for True Positive Rate, and <inline-formula id="ieqn-92"><mml:math id="mml-ieqn-92"><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mi>R</mml:mi></mml:math></inline-formula> for False Positive Rate.</p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Experimental Setup</title>
<p>Train and test Diff-Fastener on this dataset. The model adjusts all image inputs in the dataset to a resolution of 256 &#x00D7; 256, with the base channel set to 128, and the attention resolution set at 32, 16, and 8. The number of attention heads is set to 4. Set the number of epochs to 1500 and the batch size to 4, which includes 2 batches of normal samples and 2 batches of abnormal samples synthesized online through the anomaly synthesis strategy. The model employs the Adam optimizer for optimization, with an initial learning rate of <inline-formula id="ieqn-93"><mml:math id="mml-ieqn-93"><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. The research shows that the noise scale <inline-formula id="ieqn-94"><mml:math id="mml-ieqn-94"><mml:mi>t</mml:mi></mml:math></inline-formula> is set to 400 to get better recovery [<xref ref-type="bibr" rid="ref-18">18</xref>]. All experiments were conducted on a GeForce RTX 4090.</p>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Experimental Results and Discussion</title>
<sec id="s4_5_1">
<label>4.5.1</label>
<title>Result Visualization</title>
<p><xref ref-type="fig" rid="fig-8">Fig. 8</xref> shows the visualization of detection results for various types of fastener anomalies. It can be observed that Diff-Fastener does not require prior learning of specific anomaly categories. By simply training with samples of healthy fastener images, it can detect various fastener anomalies and successfully repair them, while maintaining high image fidelity. Among them, the model performs best in detecting two types of anomalies: broken and missing fasteners. It indicates that scenarios where objects transition from absent to present are clearly more suitable for anomaly detection based on generative models.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Visualization of anomalous fastener detection results</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_66098-fig-8.tif"/>
</fig>
</sec>
<sec id="s4_5_2">
<label>4.5.2</label>
<title>Comparative Experiment</title>
<p>By using cross-validation, we have obtained a more reliable assessment of the performance of both the supervised and unsupervised models. The average performance metrics from cross-validation provide a better understanding of the models&#x2019; generalization ability.</p>
<p>(1) Supervised models</p>
<p>This paper trains and tests the current advanced supervised detection model on the rail fastener dataset, and compares the results based on the recall rate and accuracy index at the label-level (i.e., at the whole image level). The comparison results are presented in <xref ref-type="table" rid="table-1">Table 1</xref>, and the visualization of the comparison results can be seen in <xref ref-type="fig" rid="fig-9">Fig. 9</xref>. Diff-Fastener learns fastener features by simply adding noise to and denoising healthy data, it can achieve approximately 99.9% accuracy rates and 99.9% recall rates in the label-level, that is, it can judge whether there is an anomaly in each fastener image almost perfectly. This result is of great significance, as it demonstrates the feasibility of the unsupervised detection model in the field of abnormal fastener detection.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Comparison of anomaly detection performance with advanced supervised models</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th rowspan="2">Models</th>
<th colspan="2">Evaluation metrics (%)</th>
</tr>
<tr>
<th>Recall &#x2191;</th>
<th>Precision &#x2191;</th>
</tr>
</thead>
<tbody>
<tr>
<td>Mask R-CNN</td>
<td>84.5</td>
<td>85.4</td>
</tr>
<tr>
<td>SSD</td>
<td>90.1</td>
<td>86.4</td>
</tr>
<tr>
<td>Faster R-CNN</td>
<td>91.3</td>
<td>88.6</td>
</tr>
<tr>
<td>YOLOX</td>
<td>93.6</td>
<td>91.8</td>
</tr>
<tr>
<td>YOLOv8</td>
<td>96.5</td>
<td>99.9</td>
</tr>
<tr>
<td>RTDETR</td>
<td>99.4</td>
<td>98.7</td>
</tr>
<tr>
<td>RT-DETR-R50</td>
<td>99.8</td>
<td>98.8</td>
</tr>
<tr>
<td>Ours</td>
<td><bold>99.9</bold></td>
<td><bold>99.9</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot><fn id="table-1fn1" fn-type="other"><p>Note: The best results of recall rate and precision rate are highlighted in bold.</p></fn>
</table-wrap-foot>
</table-wrap><fig id="fig-9">
<label>Figure 9</label>
<caption>
<title>Bar chart comparing the label-level P&#x0026;R performance of Diff-Fastener with advanced supervised detection models</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_66098-fig-9.tif"/>
</fig>
<p>(2) Unsupervised models</p>
<p>The MVTec [<xref ref-type="bibr" rid="ref-29">29</xref>] and VisA [<xref ref-type="bibr" rid="ref-30">30</xref>] datasets are essential benchmarks for anomaly detection. MVTec, with its 5000&#x002B; high-resolution images across 15 categories and pixel-precise anomaly annotations, is a gold standard for evaluating industrial inspection algorithms. VisA, offering 9621 normal and 1200 abnormal images in 12 categories, provides a new testbed for diverse industrial scenarios.</p>
<p>We compared the anomaly detection performance of Diff-Fastener with three feature embedding-based methods and three generative model-based methods; the results are presented in <xref ref-type="table" rid="table-2">Table 2</xref>. The visualization of the comparison results can be seen in <xref ref-type="fig" rid="fig-10">Fig. 10</xref>. The results show that for average Image AUROC, our model outperforms the feature embedding-based SOTA by 1.6% and the generative model-based SOTA by 0.3%. For average Pixel AUROC, our model outperforms the feature embedding-based SOTA by 2.9% and the generative model-based SOTA by 0.4%. The performance of Diff-Fastener is significantly better than previous models.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Comparison of anomaly detection performance with advanced unsupervised models</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th></th>
<th></th>
<th colspan="3">Feature embedding&#x002D;based</th>
<th colspan="3">Generative model&#x002D;based</th>
<th></th>
</tr>
<tr>
<th></th>
<th></th>   
<th>PatchCore</th>
<th>RD&#x002B;&#x002B;</th>
<th>SimpleNet</th>
<th>FastFlow</th>
<th>DRAEM</th>
<th>DiffusionAD</th>
<th>Ours</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="3">MVTec</td>
<td>Bottle</td>
<td>99.4/97.3</td>
<td>100/98.6</td>
<td>99.2/97.5</td>
<td>99.3/92.2</td>
<td>98.9/98.0</td>
<td>99.2/99.0</td>
<td>99.4/<bold>99.7</bold></td>
</tr>
<tr>
<td>Cable</td>
<td>99.1/98.0</td>
<td>98.8/97.9</td>
<td>99.8/97.2</td>
<td>91.0/96.9</td>
<td>92.6/93.7</td>
<td>99.1/98.0</td>
<td>99.2/<bold>98.5</bold></td>
</tr>
<tr>
<td>Screw</td>
<td>97.5/98.8</td>
<td>98.0/99.0</td>
<td>98.1/98.9</td>
<td>96.9/98.8</td>
<td>93.5/97.5</td>
<td>98.7/98.9</td>
<td><bold>98.9</bold>/99.0</td>
</tr>
<tr>
<td rowspan="3">VisA</td>
<td>Candle</td>
<td>97.5/94.8</td>
<td>95.7/93.0</td>
<td>98.0/87.9</td>
<td>91.5/85.5</td>
<td>93.2/91.3</td>
<td>98.2/96.8</td>
<td><bold>98.5</bold>/<bold>97.2</bold></td>
</tr>
<tr>
<td>Capsules</td>
<td>80.2/87.5</td>
<td>91.5/95.7</td>
<td>88.8/90.7</td>
<td>69.3/30.2</td>
<td>74.0/82.5</td>
<td>97.5/98.3</td>
<td><bold>97.9</bold>/<bold>98.8</bold></td>
</tr>
<tr>
<td>Fryum</td>
<td>96.0/84.9</td>
<td>94.2/90.3</td>
<td>97.5/86.5</td>
<td>86.9/72.8</td>
<td>96.8/92.8</td>
<td>98.0/95.5</td>
<td><bold>98.4</bold>/95.3</td>
</tr>
<tr>
<td colspan="2">Rail fastener</td>
<td>97.3/98.2</td>
<td>95.2/93.6</td>
<td>98.3/92.8</td>
<td>88.3/85.6</td>
<td>95.4/93.2</td>
<td>98.6/98.2</td>
<td><bold>99.1</bold>/<bold>98.9</bold></td>
</tr>
<tr>
<td colspan="2">Average</td>
<td>95.2/94.2</td>
<td>96.2/95.4</td>
<td>97.1/93.0</td>
<td>89.0/80.2</td>
<td>92.0/92.7</td>
<td>98.4/97.8</td>
<td><bold>98.7</bold>/<bold>98.2</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot><fn id="table-2fn1" fn-type="other"><p>Note: The best results of Image AUROC/Pixel AUROC are highlighted in bold.</p></fn>
</table-wrap-foot>
</table-wrap><fig id="fig-10">
<label>Figure 10</label>
<caption>
<title>Bar chart comparing the average Image AUROC and Pixel AUROC with advanced unsupervised detection models</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_66098-fig-10.tif"/>
</fig>
</sec>
<sec id="s4_5_3">
<label>4.5.3</label>
<title>Ablation Study</title>
<p><xref ref-type="table" rid="table-3">Table 3</xref> shows the results of the ablation study. From the first and second rows of the table, compared to traditional iterative denoising, the inference speed of One-Step denoising is greatly improved, approximately 320 times faster than before. In the third row, the use of the Norm-Guided paradigm enhances the model&#x2019;s ability to preserve image details while effectively removing anomaly regions of varying sizes, achieving more accurate anomaly detection, with a slight decrease in speed. From the fourth to sixth rows, it is evident that the introduction of CBAM (Convolutional Block Attention Module) [<xref ref-type="bibr" rid="ref-31">31</xref>] and dilated convolutions [<xref ref-type="bibr" rid="ref-32">32</xref>] improves the model&#x2019;s semantic awareness while maintaining relatively efficient speed. In addition, sparse skip has enhanced the network performance and detection efficiency. Ablation experiments have demonstrated the effectiveness of the above methods.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>The impact of changes in the denoising method and module on the rail fastener dataset</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th colspan="3">Denoising methods</th>
<th colspan="3">Modules</th>
<th colspan="3">Performance</th>
</tr>
<tr>
<th><bold>Iteration</bold></th>
<th><bold>One-Step</bold></th>
<th><bold>Norm-Guided</bold></th>
<th><bold>CBAM</bold></th>
<th><bold>Dilated Conv</bold></th>
<th><bold>Sparse-Skip</bold></th>
<th><bold>I&#x2191;</bold></th>
<th><bold>P&#x2191;</bold></th>
<th><bold>F&#x2191;</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td>&#x2713;</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
<td>95.2</td>
<td>97.4</td>
<td>0.07</td>
</tr>
<tr>
<td>&#x2717;</td>
<td>&#x2713;</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
<td>95.4</td>
<td>96.5</td>
<td>22.5</td>
</tr>
<tr>
<td>&#x2717;</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
<td>95.9</td>
<td>97.0</td>
<td>18.4</td>
</tr>
<tr>
<td>&#x2717;</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2717;</td>
<td>&#x2717;</td>
<td>96.3</td>
<td>96.9</td>
<td>17.7</td>
</tr>
<tr>
<td>&#x2717;</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2717;</td>
<td>&#x2713;</td>
<td>&#x2717;</td>
<td>96.7</td>
<td>97.1</td>
<td>18.0</td>
</tr>
<tr>
<td>&#x2717;</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2717;</td>
<td>97.0</td>
<td>97.8</td>
<td>17.9</td>
</tr>
<tr>
<td>&#x2717;</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td>&#x2713;</td>
<td><bold>97.3</bold></td>
<td><bold>97.6</bold></td>
<td><bold>24.6</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot><fn id="table-3fn1" fn-type="other"><p>Note: &#x201C;I&#x201D;, &#x201C;P&#x201D;, and &#x201C;F&#x201D; respectively refer to the Image AUROC, Pixel AUROC and FPS (Frames Per Second); The best results are highlighted in bold.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p><xref ref-type="fig" rid="fig-11">Fig. 11</xref> shows the effects of the Norm-Guided paradigm and DACM. <xref ref-type="fig" rid="fig-11">Fig. 11a</xref>&#x2013;<xref ref-type="fig" rid="fig-11">f</xref> represents several different anomalous samples. A indicates a diffusion model structure using only one-step denoising. The introduction of one-step denoising leads to a loss of detail in the reconstruction results, limiting adaptability to different types of anomalies. From the image, it can be seen that the reconstructed semantic edges of the fasteners are unclear, with some even experiencing spatial distortion, and the overall tone of the image appears unnatural. B represents a one-step denoising diffusion model after the introduction of the Norm-Guided paradigm, which allows for more refined reconstruction rather than simple rough recovery. Additionally, the model can effectively handle anomalies in varying-sized areas with two scales of noise, enhancing the model&#x2019;s adaptability. However, there is still an issue of excessive image smoothing. After introducing DACM, the model can emphasize important features while suppressing less significant ones during the reconstruction process. It also captures a broader context in the feature extraction stage, further enhancing the model&#x2019;s reconstruction capability. As shown in <xref ref-type="fig" rid="fig-11">Fig. 11</xref>, our reconstructed images contain richer detail information and are more aligned with real-world situations.</p>
<fig id="fig-11">
<label>Figure 11</label>
<caption>
<title>Ablations of norm-guided paradigm and CBAM &#x0026; Dilated Convolution on rail fastener dataset</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="SDHM_66098-fig-11.tif"/>
</fig>
</sec>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>In this paper, we propose Diff-Fastener, a model that, for the first time, utilizes the generative ability of the diffusion model to detect abnormal fasteners. We employ one-step denoising to alleviate the inefficiency issues associated with traditional iterative denoising methods. The introduction of the Norm-Guided paradigm brings the model&#x2019;s ability for finer reconstruction and better adaptability. DACM further improves the model&#x2019;s feature selection capability, resulting in stronger accuracy in complex scenes. The high efficiency of sparse skip has enhanced the scalability of the model. Experiments have demonstrated that this model has strong effectiveness and practicality, and it possesses practical application value in actual railway engineering scenarios. Meanwhile, we hope this work can reignite interest in denoising-based railway fault diagnosis methods within the context of current research in unsupervised learning.</p>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This research was funded by the National Natural Science Foundation of China, grant number 52272385 and 52475085.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Conceptualization, Peng Sun; methodology, Peng Sun; software, Peng Sun; validation, Peng Sun and Quanyu Long; formal analysis, Peng Sun; investigation, Peng Sun; resources, Dechen Yao and Jianwei Yang; data curation, Peng Sun; writing&#x2014;original draft preparation, Peng Sun; writing&#x2014;review and editing, Peng Sun and Jianwei Yang; visualization, Peng Sun; supervision, Dechen Yao and Jianwei Yang; project administration, Dechen Yao and Jianwei Yang; funding acquisition, Dechen Yao and Jianwei Yang. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>Data not available due to ethical restrictions. Due to the nature of this research, participants of this study did not agree for their data to be shared publicly, so supporting data is not available.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Min</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Jing</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Method for rail surface defect detection based on neural network architecture search</article-title>. <source>Meas Sci Technol</source>. <year>2024</year>;<volume>36</volume>(<issue>1</issue>):<fpage>016027</fpage>. doi:<pub-id pub-id-type="doi">10.1088/1361-6501/ad9048</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Roth</surname> <given-names>K</given-names></string-name>, <string-name><surname>Pemula</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zepeda</surname> <given-names>J</given-names></string-name>, <string-name><surname>Sch&#x00F6;lkopf</surname> <given-names>B</given-names></string-name>, <string-name><surname>Brox</surname> <given-names>T</given-names></string-name>, <string-name><surname>Gehler</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Towards total recall in industrial anomaly detection</article-title>. <comment>arXiv:2106.08265v2. 2021</comment>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Tien</surname> <given-names>TD</given-names></string-name>, <string-name><surname>Nguyen</surname> <given-names>AT</given-names></string-name>, <string-name><surname>Tran</surname> <given-names>NH</given-names></string-name>, <string-name><surname>Huy</surname> <given-names>TD</given-names></string-name>, <string-name><surname>Duong</surname> <given-names>STM</given-names></string-name>, <string-name><surname>Nguyen</surname> <given-names>CDT</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Revisiting reverse distillation for anomaly detection</article-title>. In: <conf-name>Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2023 Jun 17&#x2013;24</conf-name>; <publisher-loc>Vancouver, BC, Canada</publisher-loc>. doi:<pub-id pub-id-type="doi">10.1109/CVPR52729.2023.02348</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>SimpleNet: a simple network for image anomaly detection and localization</article-title>. <comment>arXiv:2303.15140v2. 2023</comment>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Yu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>W</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>R</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>FastFlow: unsupervised anomaly detection and localization via 2D normalizing flows</article-title>. <comment>arXiv:2111.07677v2. 2021</comment>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zavrtanik</surname> <given-names>V</given-names></string-name>, <string-name><surname>Kristan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Sko&#x010D;aj</surname> <given-names>D</given-names></string-name></person-group>. <article-title>DRAEM&#x2014;a discriminatively trained reconstruction embedding for surface anomaly detection</article-title>. <comment>arXiv:2108.07610v2. 2021</comment>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wei</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Review: recent advances for the diffusion model</article-title>. <source>J Phys Conf Ser</source>. <year>2024</year>;<volume>2711</volume>(<issue>1</issue>):<fpage>012005</fpage>. doi:<pub-id pub-id-type="doi">10.1088/1742-6596/2711/1/012005</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Bhosale</surname> <given-names>A</given-names></string-name>, <string-name><surname>Mukherjee</surname> <given-names>S</given-names></string-name>, <string-name><surname>Banerjee</surname> <given-names>B</given-names></string-name>, <string-name><surname>Cuzzolin</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Anomaly detection using diffusion-based methods</article-title>. <comment>arXiv:2412.07539. 2024</comment>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Livernoche</surname> <given-names>V</given-names></string-name>, <string-name><surname>Jain</surname> <given-names>V</given-names></string-name>, <string-name><surname>Hezaveh</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ravanbakhsh</surname> <given-names>S</given-names></string-name></person-group>. <article-title>On diffusion modeling for anomaly detection</article-title>. <comment>arXiv:2305.18593. 2025</comment>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ahsan</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Raman</surname> <given-names>S</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Siddique</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>A comprehensive survey on diffusion models and their applications</article-title>. <comment>arXiv:2408.10207. 2024</comment>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Ho</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jain</surname> <given-names>A</given-names></string-name>, <string-name><surname>Abbeel</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Denoising diffusion probabilistic models</article-title>. <comment>arXiv:2006.11239. 2020</comment>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Mousakhan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Brox</surname> <given-names>T</given-names></string-name>, <string-name><surname>Tayyub</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Anomaly detection with conditioned denoising diffusion models</article-title>. <comment>arXiv:2305.15956. 2023</comment>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wolleb</surname> <given-names>J</given-names></string-name>, <string-name><surname>Bieder</surname> <given-names>F</given-names></string-name>, <string-name><surname>Sandk&#x00FC;hler</surname> <given-names>R</given-names></string-name>, <string-name><surname>Cattin</surname> <given-names>PC</given-names></string-name></person-group>. <article-title>Diffusion models for medical anomaly detection</article-title>. <comment>arXiv:2203.04306. 2022</comment>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Pinaya</surname> <given-names>WHL</given-names></string-name>, <string-name><surname>Graham</surname> <given-names>MS</given-names></string-name>, <string-name><surname>Gray</surname> <given-names>R</given-names></string-name>, <string-name><surname>Costa</surname> <given-names>PFD</given-names></string-name>, <string-name><surname>Tudosiu</surname> <given-names>PD</given-names></string-name>, <string-name><surname>Wright</surname> <given-names>P</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Fast unsupervised brain anomaly detection and segmentation with diffusion models</article-title>. <comment>arXiv:2206.03461. 2022</comment>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Wyatt</surname> <given-names>J</given-names></string-name>, <string-name><surname>Leach</surname> <given-names>A</given-names></string-name>, <string-name><surname>Schmon</surname> <given-names>SM</given-names></string-name>, <string-name><surname>Willcocks</surname> <given-names>CG</given-names></string-name></person-group>. <article-title>AnoDDPM: anomaly detection with denoising diffusion probabilistic models using simplex noise</article-title>. In: <conf-name>Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW); 2022 Jun 19&#x2013;20</conf-name>; <publisher-loc>New Orleans, LA, USA</publisher-loc>. doi:<pub-id pub-id-type="doi">10.1109/CVPRW56347.2022.00080</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>H</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>DiAD: a diffusion-based framework for multi-class anomaly detection</article-title>. <comment>arXiv:2312.06607. 2023</comment>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Song</surname> <given-names>J</given-names></string-name>, <string-name><surname>Park</surname> <given-names>D</given-names></string-name>, <string-name><surname>Baek</surname> <given-names>K</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>S</given-names></string-name>, <string-name><surname>Choi</surname> <given-names>J</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>E</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>DefectFill: realistic defect generation with inpainting diffusion model for visual inspection</article-title>. <comment>arXiv:2503.13985v1. 2025</comment>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>YG</given-names></string-name></person-group>. <article-title>DiffusionAD: norm-guided one-step denoising diffusion for anomaly detection</article-title>. <comment>arXiv:2303.08730. 2023</comment>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Rombach</surname> <given-names>R</given-names></string-name>, <string-name><surname>Blattmann</surname> <given-names>A</given-names></string-name>, <string-name><surname>Lorenz</surname> <given-names>D</given-names></string-name>, <string-name><surname>Esser</surname> <given-names>P</given-names></string-name>, <string-name><surname>Ommer</surname> <given-names>B</given-names></string-name></person-group>. <article-title>High-resolution image synthesis with latent diffusion models</article-title>. <comment>arXiv:2112.10752. 2022</comment>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Song</surname> <given-names>J</given-names></string-name>, <string-name><surname>Meng</surname> <given-names>C</given-names></string-name>, <string-name><surname>Ermon</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Denoising diffusion implicit models</article-title>. <comment>arXiv:2010.02502. 2022</comment>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Deep residual learning for image recognition</article-title>. <comment>arXiv:1512.03385. 2015</comment>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Salimans</surname> <given-names>T</given-names></string-name>, <string-name><surname>Karpathy</surname> <given-names>A</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Kingma</surname> <given-names>DP</given-names></string-name></person-group>. <article-title>PixelCNN&#x002B;&#x002B;: improving the PixelCNN with discretized logistic mixture likelihood and other modifications</article-title>. <comment>arXiv:1701.05517v1. 2017</comment>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Vaswani</surname> <given-names>A</given-names></string-name>, <string-name><surname>Shazeer</surname> <given-names>N</given-names></string-name>, <string-name><surname>Parmar</surname> <given-names>N</given-names></string-name>, <string-name><surname>Uszkoreit</surname> <given-names>J</given-names></string-name>, <string-name><surname>Jones</surname> <given-names>L</given-names></string-name>, <string-name><surname>Gomez</surname> <given-names>AN</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Attention is all you need</article-title>. <comment>arXiv:1706.03762. 2023</comment>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Long</surname> <given-names>C</given-names></string-name>, <string-name><surname>Tan</surname> <given-names>C</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Tan</surname> <given-names>H</given-names></string-name>, <string-name><surname>Duan</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Industrial CT image reconstruction for faster scanning through U-Net&#x002B;&#x002B; with hybrid attention and loss function</article-title>. <source>Nondestruct Test Eval</source>. <year>2024</year>;<volume>39</volume>(<issue>8</issue>):<fpage>2646</fpage>&#x2013;<lpage>65</lpage>. doi:<pub-id pub-id-type="doi">10.1080/10589759.2024.2305329</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shang</surname> <given-names>F</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>D</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>B</given-names></string-name>, <string-name><surname>Han</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhan</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Research on weld seam feature extraction and defect identi-fication technology</article-title>. <source>Nondestruct Test Eval</source>. <year>2024</year>. doi:<pub-id pub-id-type="doi">10.1080/10589759.2024.2405062</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>M</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Pantic</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Dilated convolutions with lateral inhibitions for semantic image segmentation</article-title>. <comment>arXiv:2006.03708. 2022</comment>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>P</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>H</given-names></string-name></person-group>. <article-title>MemSeg: a semi-supervised method for image surface defect detection using differences and commonalities</article-title>. <comment>arXiv:2205.00908. 2022</comment>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>YG</given-names></string-name></person-group>. <article-title>Prototypical residual networks for anomaly detection and localization</article-title>. <comment>arXiv:2212.02031. 2023</comment>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bergmann</surname> <given-names>P</given-names></string-name>, <string-name><surname>Batzner</surname> <given-names>K</given-names></string-name>, <string-name><surname>Fauser</surname> <given-names>M</given-names></string-name>, <string-name><surname>Sattlegger</surname> <given-names>D</given-names></string-name>, <string-name><surname>Steger</surname> <given-names>C</given-names></string-name></person-group>. <article-title>The MVTec anomaly detection dataset: a comprehensive real-world dataset for unsupervised anomaly detection</article-title>. <source>Int J Comput Vis</source>. <year>2021</year>;<volume>129</volume>(<issue>4</issue>):<fpage>1038</fpage>&#x2013;<lpage>59</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11263-020-01400-4</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Jeong</surname> <given-names>J</given-names></string-name>, <string-name><surname>Pemula</surname> <given-names>L</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>D</given-names></string-name>, <string-name><surname>Dabeer</surname> <given-names>O</given-names></string-name></person-group>. <article-title>SPot-the-difference self-supervised pre-training for anomaly detection and segmentation</article-title>. <comment>arXiv:2207.14315. 2022</comment>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Woo</surname> <given-names>S</given-names></string-name>, <string-name><surname>Park</surname> <given-names>J</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>JY</given-names></string-name>, <string-name><surname>Kweon</surname> <given-names>IS</given-names></string-name></person-group>. <article-title>CBAM: convolutional block attention module</article-title>. <comment>arXiv:1807.06521. 2018</comment>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Yu</surname> <given-names>F</given-names></string-name>, <string-name><surname>Koltun</surname> <given-names>V</given-names></string-name></person-group>. <article-title>Multi-scale context aggregation by dilated convolutions</article-title>. <comment>arXiv:1511.07122. 2016</comment>.</mixed-citation></ref>
</ref-list>
</back></article>














