<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">49623</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2024.049623</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>A Tabletop Nano-CT Image Noise Reduction Network Based on 3-Dimensional Axial Attention Mechanism</article-title>
<alt-title alt-title-type="left-running-head">A Tabletop Nano-CT Image Noise Reduction Network Based on 3-Dimensional Axial Attention Mechanism</alt-title>
<alt-title alt-title-type="right-running-head">A Tabletop Nano-CT Image Noise Reduction Network Based on 3-Dimensional Axial Attention Mechanism</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Fu</surname><given-names>Huijuan</given-names></name></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Zhu</surname><given-names>Linlin</given-names></name></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Wang</surname><given-names>Chunhui</given-names></name></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Xi</surname><given-names>Xiaoqi</given-names></name></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Han</surname><given-names>Yu</given-names></name></contrib>
<contrib id="author-6" contrib-type="author">
<name name-style="western"><surname>Li</surname><given-names>Lei</given-names></name></contrib>
<contrib id="author-7" contrib-type="author">
<name name-style="western"><surname>Sun</surname><given-names>Yanmin</given-names></name></contrib>
<contrib id="author-8" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Yan</surname><given-names>Bin</given-names></name><email>ybspace@hotmail.com</email></contrib>
<aff>
<institution>Henan Key Laboratory of Imaging and Intelligent Processing, PLA Strategic Support Force Information Engineering University</institution>, <addr-line>Zhengzhou, 450000</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Bin Yan. Email: <email>ybspace@hotmail.com</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2024</year></pub-date>
<pub-date date-type="pub" publication-format="electronic"><day>18</day><month>7</month><year>2024</year></pub-date>
<volume>80</volume>
<issue>1</issue>
<fpage>1711</fpage>
<lpage>1725</lpage>
<history>
<date date-type="received">
<day>12</day>
<month>1</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>28</day>
<month>4</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2024 Fu et al.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Fu et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_49623.pdf"></self-uri>
<abstract>
<p>Nano-computed tomography (Nano-CT) is an emerging, high-resolution imaging technique. However, due to their low-light properties, tabletop Nano-CT has to be scanned under long exposure conditions, which the scanning process is time-consuming. For 3D reconstruction data, this paper proposed a lightweight 3D noise reduction method for desktop-level Nano-CT called AAD-ResNet (Axial Attention DeNoise ResNet). The network is framed by the U-net structure. The encoder and decoder are incorporated with the proposed 3D axial attention mechanism and residual dense block. Each layer of the residual dense block can directly access the features of the previous layer, which reduces the redundancy of parameters and improves the efficiency of network training. The 3D axial attention mechanism enhances the correlation between 3D information in the training process and captures the long-distance dependence. It can improve the noise reduction effect and avoid the loss of image structure details. Experimental results show that the network can effectively improve the image quality of a 0.1-s exposure scan to a level close to a 3-s exposure, significantly shortening the sample scanning time.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Deep learning</kwd>
<kwd>tabletop Nano-CT</kwd>
<kwd>image denoising</kwd>
<kwd>3D axial attention mechanism</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>National Natural Science Foundation of China</funding-source>
<award-id>62201618</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Computed tomography (CT) is the technique that uses the projection of X-ray through the object and mathematical formulas to reconstruct the three-dimensional internal structure of the object. With the rapid development of X-ray sources and high-resolution detectors, Nano-Computed Tomography (Nano-CT) realizes non-destructive high-resolution imaging of the nanoscale structure of the object being detected [<xref ref-type="bibr" rid="ref-1">1</xref>], which provides an unprecedented opportunity for direct observation of the fine structure of the object and has been widely used in many fields such as material science [<xref ref-type="bibr" rid="ref-2">2</xref>,<xref ref-type="bibr" rid="ref-3">3</xref>], drug development [<xref ref-type="bibr" rid="ref-4">4</xref>], integrated circuits, and so on [<xref ref-type="bibr" rid="ref-5">5</xref>&#x2013;<xref ref-type="bibr" rid="ref-7">7</xref>].</p>
<p>Tabletop Nano-CT (after this referred to as Nano-CT) is exhibits portability characteristics, user-friendly operation, and low maintenance requirements. However, in practical applications, there are still some problems that affect the imaging quality. The X-ray source of Nano-CT adopts the nanoscale focus spot ray source with low-energy characteristics. Compared with micro-CT, the photon intensity of Nano-CT has decreased by 1&#x2013;2 orders of magnitude, and its ray source has small power and weak light intensity. At short exposure times, Nano-CT is in a weak photon condition, and the reduced number of photons received by the detector can lead to noise interference in the projection data [<xref ref-type="bibr" rid="ref-8">8</xref>]. These subtle noise changes will be amplified during the image reconstruction process, resulting in serious noise and artifacts in the reconstructed CT image, which cannot accurately reflect the real structure inside the object, and thus is not conducive to the subsequent analysis of the internal structure of the measured object. Although the noise can be suppressed by increasing the exposure time, this will substantially increase the scanning time, decreasing scanning efficiency. Therefore, reducing the noise level in Nano-CT images is crucial in improving image quality, especially under fast scanning conditions. The achievement of both high efficiency and high-quality imaging in Nano-CT is crucial for advancing the practical application of this technology.</p>
<p>The development of deep neural networks provides a new processing method for image post-processing, and the network realizes the denoising of images through the nonlinear mapping relationship between image features. Thanks to deep models&#x2019; strong linear fitting ability, deep learning has achieved satisfactory performance in medical CT noise reduction [<xref ref-type="bibr" rid="ref-9">9</xref>&#x2013;<xref ref-type="bibr" rid="ref-11">11</xref>]. According to the current research, noise reduction methods for 2D CT images have been relatively abundant and are gradually becoming mature. Although the noise reduction methods for 2D CT images have been relatively mature, the 3D noise reduction of Nano-CT is still a challenging problem. Nano-CT as a three-dimensional imaging technology has many applications in many fields. These application scenarios require high resolution and sharpness of volume data. For example, additive manufacturing of precision components [<xref ref-type="bibr" rid="ref-12">12</xref>], volume measurement [<xref ref-type="bibr" rid="ref-13">13</xref>], and fossil grain characterization [<xref ref-type="bibr" rid="ref-14">14</xref>]. The 3D correlation of Nano-CT exists not only within slices but also between slices. Directly stacking 2D noise reduction slices may cause incoherent volume data [<xref ref-type="bibr" rid="ref-15">15</xref>]. Therefore, it is necessary to consider extracting features from adjacent slices to capture and enhance the structural details of CT images. However, implementing 3D network noise reduction requires storing volume data and numerous network parameters in memory, which requires high memory. Noise reduction of volume data requires considerable time and computational resources. This problem limits further research, development, and applications.</p>
<p>This paper focuses on the problems of a large number of parameters, long training time, and high requirement of computational resources of existing 3D networks. We aim to achieve fast and high-quality Nano-CT imaging by developing efficient 3D image processing algorithms without significantly increasing computational costs. So, we propose a lightweight Nano-CT 3D Axial Attention DeNoise ResNet (AAD-ResNet). The main innovations of this paper are as follows:</p>
<p>(1) We propose an encode-decode noise reduction network based on a three-dimensional axial attention mechanism, which can effectively improve the image quality of a 0.1-s exposure scan to a level close to a 3-s exposure. This significantly shortens the sample scanning time.</p>
<p>(2) We design a 3D axial attention that apply attention independently for each axis, and then correlate the attention of these three axes to realize global attention. This method effectively reduces the number of network parameters while capturing the long-range dependencies among 3D data. The proposed 3D network parameter number can be reduced to 6.64 M, which is suitable for application scenarios with limited computational resources.</p>
<p>(3) Depending on the tabletop Nano-CT device, various samples (including corals and rocks) are scanned to construct the network training dataset. Experimental results show that the proposed method performs significantly better than existing methods in 3D networks. Ablation experiments verified the effectiveness of the components in AAD-ResNet.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>In this section, we review related works including 2D CT image denoising, current 3D network designs and applications, and the application of self-attention mechanism. Recent research has focused on data-driven and model-driven deep learning strategies. Data-driven methods have been used to learn a fully nonlinear mapping from noise images to target images by pairing training data.</p>
<sec id="s2_1">
<label>2.1</label>
<title>2D Image Noise Reduction</title>
<p>There are many studies on improving the image quality of medical CT slices. In 2023, Li et al. [<xref ref-type="bibr" rid="ref-16">16</xref>] proposed a Progressive Circular Convolutional Neural Network (PCCNN) that introduces a noise transfer model. Transfers noise from low-dose CT to normal-dose CT using denoised and noisy CT images generated from unpaired CT images. The denoising model also contains a progressive module that efficiently removes noise through a multistage wavelet transform without sacrificing high-frequency components such as edges and details. In 2023, Ma et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] proposed a Residual Dense Network with Self-Calibrated Convolution for low-dose CT image denoising. This network fully utilizes the hierarchical features of the original image to obtain a reconstructed image with more details.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>3D Network Design and Application</title>
<p>Due to the 3D imaging properties of Nano-CT, the reconstruction results are 3D data; the correlations in the 3D still exist between slices apart from the 2D correlations in the slice. For this reason, extracting the feature from the adjacent slices is very effective for capturing and enhancing fine details.</p>
<p><bold>3D network design:</bold> In 2017, Liu et al. mentioned the importance of the third dimension and applied a 3D convolution over 2D convolutions for the segmentation of Digital Breast Tomosynthesis and CT [<xref ref-type="bibr" rid="ref-18">18</xref>]. In 2023, Liu et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] proposed an attention mechanism based 3D segmentation network, and refer to this algorithm as the Dilated Hierarchical Decoupled Convolution Network with Attention (ADHDC-Net). In 2023, Cao et al. proposed a 3D lightweight segmentation network for multimodal brain tumor MRI images. It added 3D SA attention to skip connections to address the semantic gap problem with spatial and channel aggregation [<xref ref-type="bibr" rid="ref-20">20</xref>].</p>
<p><bold>3D image denoising:</bold> In 2018, You et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] proposed a Structure-Sensitive Multiscale Generative Adversarial Network (SMGAN) that utilizes the information of 3D data to improve image quality, which can effectively preserve the structural and textural information of low-dose CT images and significantly reduce noise and artifacts. In the same year, Shan et al. [<xref ref-type="bibr" rid="ref-22">22</xref>] proposed a Conveying Path-Based Convolutional Encoder-Decoder network used for LDCT noise reduction within the GAN framework. An initial 3D CPCE denoising model is directly obtained by extending a trained 2D CNN. Then, it is fine-tuned to merge 3D spatial information from neighboring slices. The method suppresses image noise and preserves fine image structure. In 2019, Yin et al. [<xref ref-type="bibr" rid="ref-23">23</xref>] proposed a domain-progressive 3D residual convolutional network for LDCT noise reduction, which can reduce CT images from the projection domain to the image domain. In 2020, Li et al. [<xref ref-type="bibr" rid="ref-24">24</xref>] proposed a new 3D self-attention convolutional neural network that utilizes 3D data to capture a wide range of spatial information within and between CT slices. In 2022, Gunduzalp et al. [<xref ref-type="bibr" rid="ref-25">25</xref>] proposed a 3D U-Net based noise reduction method, and the proposed 3D U-NetR architecture was tested on simulated and real chest CT data to achieve better noise reduction. In 2023, Zhou et al. [<xref ref-type="bibr" rid="ref-15">15</xref>] aimed to investigate whether multi-slice input improves the denoising performance compared to single-slice input and whether the 3D network architecture outperforms the 2D version in utilizing multi-slice input.</p>
<p>The existing research on 3D networks provides new ideas and directions to solve the noise reduction of Nano-CT 3D data.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Self-Attention</title>
<p>The self-attention mechanism enables the model to automatically learn the relationships between different positions within a sequence. It dynamically adjusts the level of attention each location pays to other locations by calculating the interrelationships between each location in the sequence and all other locations in the sequence [<xref ref-type="bibr" rid="ref-26">26</xref>]. This mechanism allows the model to capture long-distance dependencies, independent of sequence length.</p>
<p>Self-attention is the core idea of Transformer, and the mechanism of self-attention has become a hot research topic and has been explored in various NLP tasks. For example, Wang et al. [<xref ref-type="bibr" rid="ref-27">27</xref>] proposed a convolution-free Token2Token extended visual Transformer for low-dose CT noise removal. The method applied Transformer for noise suppression processing and enhanced the feature optimization process in the model by introducing inflated convolution, which better removes common boundary artifacts. In 2023, Zhu et al. [<xref ref-type="bibr" rid="ref-28">28</xref>] proposed an LDCT noise and artifact suppressing network based on the Swin Transformer. The noise removal sub-network improves the network&#x2019;s ability to extract relevant image features by using a Swin Transformer with a shift window as an encoder-decoder for global feature fusion. In 2023, Kirillov et al. [<xref ref-type="bibr" rid="ref-29">29</xref>] proposed a robust segment anything model. SAM improves the accuracy and efficiency of image recognition and segmentation by using the Transformer architecture.</p>
<p>Although the self-attention mechanism dramatically improves network performance optimization, it requires high computational resources due to its higher complexity. Currently, there are not many studies on lightweight 3D networks and self-attention mechanisms; later, we introduce in detail the lightweight 3D network proposed in this paper and design a 3D axial attention mechanism to improve noise reduction performance on Nano-CT images.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Method</title>
<p>The reconstructed Nano-CT image mainly contains photon statistical noise [<xref ref-type="bibr" rid="ref-30">30</xref>] and reconstruction noise [<xref ref-type="bibr" rid="ref-31">31</xref>]. In this paper, the short exposure time scan image is assumed to be the approximate sum of the long exposure time scan image and the noise. Then, the Nano-CT reconstructed image can be expressed in a linear algebraic form:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi>&#x03B7;</mml:mi></mml:math></disp-formula>where <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> refers to the short exposure time 3D data, <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> refers to the long exposure time 3D data, and <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>&#x03B7;</mml:mi></mml:math></inline-formula> refers to the noise. The propose Nano-CT noise reduction process in this paper is shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>The propose noise reduction process</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_49623-fig-1.tif"/>
</fig>
<p>First, we scan the samples using the Nano-CT device to obtain training data for long and short exposure times. Subsequently, the FBP algorithm reconstructs the projection data to obtain the reconstructed volume data [<xref ref-type="bibr" rid="ref-32">32</xref>]. Since the 3D volume data occupy high memory, we cropped the 3D volume data. The 512 &#x00D7; 512 &#x00D7; 1024 volume data was cropped into multiple 64 &#x00D7; 64 &#x00D7; 64 block data. Finally, we input these block data into AAD-ResNet to get clean block data.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Network Architecture</title>
<p>Inspired by the successful application of 3D-Unet in medical image segmentation [<xref ref-type="bibr" rid="ref-33">33</xref>], we constructed AAD-ResNet using 3D convolution. <xref ref-type="fig" rid="fig-2">Fig. 2</xref> presents the network architecture, containing encoding and decoding paths, each with three resolution levels. The encoder and decoder contain 3D axial attention mechanisms and residual dense blocks. This network design can efficiently capture contextual information and preserve local details to achieve better performance in Nano-CT noise reduction.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>AAD-ResNet network structure</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_49623-fig-2.tif"/>
</fig>
<p><bold>Three-dimensional axial attention mechanism:</bold> The self-attention mechanism enables associations between a target pixel and any other pixel. In many Natural Language Processing (NLP) applications, Transformer has been demonstrated to encode long-distance dependencies [<xref ref-type="bibr" rid="ref-26">26</xref>]. It is due to the self-attention mechanism that reveals the correlations between the input sequences. This involves considerable computation, especially when the feature map size is large, and can lead to a significant efficiency decrease. To cope with this computational demand while efficiently integrating global information and preserving long-distance dependencies, we propose a three-dimensional axial attention mechanism that improves efficiency by constraining attention on specific axial, as shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Axial attention mechanism schematic</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_49623-fig-3.tif"/>
</fig>
<p>For the computation of the axial attention mechanism, we need to compute Query, Key, and Value for all depths of all columns in each row. These can be realized by linear changes:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>X</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>Q</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>X</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>X</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>X</mml:mi></mml:math></inline-formula> is the input block data with dimension <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>m</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi></mml:math></inline-formula>. <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>Q</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are the weight matrices of Query, Key, Value. In the example of a particular axis, these can be realized by linear changes:
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mi>A</mml:mi><mml:mi>t</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:msub><mml:msubsup><mml:mi>K</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>w</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:msqrt><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:msqrt></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>o</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:msub></mml:math></disp-formula></p>
<p>Typically, the self-attention mechanism needs to computes the relationships between all the pixels in the image when dealing with 3D data, which usually involves a large weight matrix, thus increasing the number of parameters. Pixel position encoding is generally obtained by network training, which also increases the number of parameters of the model. While our proposed 3D axial attention mechanism only compute the relationship between elements in one direction. The computational complexity is reduced by handling rows, columns, and depths and combining their attention separately. It effectively reduces the number of parameters and computation of the model while maintaining the performance of the model. This mechanism is crucial when dealing with large-scale 3D data, as it can improve computational efficiency without sacrificing performance.</p>
<p>For example, suppose the dimension of the input image block is <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>m</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>d</mml:mi></mml:math></inline-formula>, to which we apply the attention mechanism. The traditional self-attention mechanism would use all elements of the input for computation and the computational complexity is <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>O</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mi>m</mml:mi><mml:mi>d</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>n</mml:mi><mml:mi>m</mml:mi><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. In contrast, the axial attention mechanism has a computational complexity of <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>O</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mi>m</mml:mi><mml:mi>d</mml:mi><mml:mo>+</mml:mo><mml:mi>n</mml:mi><mml:mi>m</mml:mi><mml:mi>d</mml:mi><mml:mo>+</mml:mo><mml:mi>n</mml:mi><mml:mi>m</mml:mi><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> by combining the combination of the attention mechanisms on each axis (rows, columns, and depth). After applying the axial attention mechanism, the computational complexity is reduced from <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi>O</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mi>m</mml:mi><mml:mi>d</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>n</mml:mi><mml:mi>m</mml:mi><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> to <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>O</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>n</mml:mi><mml:mi>m</mml:mi><mml:mi>d</mml:mi><mml:mo>+</mml:mo><mml:mi>n</mml:mi><mml:mi>m</mml:mi><mml:mi>d</mml:mi><mml:mo>+</mml:mo><mml:mi>n</mml:mi><mml:mi>m</mml:mi><mml:mi>d</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, which significantly reduces the computational complexity, it is significant when n, m, and d are all large [<xref ref-type="bibr" rid="ref-34">34</xref>].</p>
<p><bold>Residual dense module:</bold> This paper uses residual dense blocks to extract many local properties. Specifically, we use three 3 &#x00D7; 3 &#x00D7; 3 convolutions to construct a third-order dense connection. The residual connection [<xref ref-type="bibr" rid="ref-35">35</xref>] adaptively fuses the previous and current local features to better learn these features. By incorporating residual connection and dense connection [<xref ref-type="bibr" rid="ref-36">36</xref>], the proposed model can learn image features more deeply and improve the noise reduction performance of the network. In addition, since our network applied dense blocks, each layer can directly access the previous layers, reducing the redundancy of parameters and improving the network&#x2019;s training efficiency [<xref ref-type="bibr" rid="ref-37">37</xref>].</p>
<p>The network structure of each section is shown below. Each part of our model uses 3D convolution to enhance the correlation between slices. The specific parameters of the network are shown in <xref ref-type="table" rid="table-1">Table 1</xref>. At feature extraction, we use two convolution blocks with kernel size of 3 &#x00D7; 3 &#x00D7; 3 and the filter is 64. We use convolution with kernel size of 1 &#x00D7; 2 &#x00D7; 2 and stride of 1 &#x00D7; 2 &#x00D7; 2 for down-sampling. Similarly, the transposed convolution with kernel size 1 &#x00D7; 2 &#x00D7; 2 and stride 1 &#x00D7; 2 &#x00D7; 2 is used for up-sampling. We use a 3D axial attention mechanism and residual dense block in the encoder and decoder.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Network configuration of the denoiser</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Module</th>
<th>Layers</th>
<th>Kernel size</th>
<th>Stride</th>
<th>Input size</th>
<th>Output size</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="2">Feature extraction</td>
<td>Conv3d, ReLU</td>
<td>3 &#x00D7; 3 &#x00D7; 3</td>
<td>1 &#x00D7; 1 &#x00D7; 1</td>
<td><inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mn>64</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:math></inline-formula></td>
</tr>
<tr>
<td>Conv3d, ReLU</td>
<td>3 &#x00D7; 3 &#x00D7; 3</td>
<td>1 &#x00D7; 1 &#x00D7; 1</td>
<td><inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mn>64</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mn>64</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:math></inline-formula></td>
</tr>
<tr>
<td>Downsampling</td>
<td>Conv3d</td>
<td>1 &#x00D7; 2 &#x00D7; 2</td>
<td>1 &#x00D7; 2 &#x00D7; 2</td>
<td><inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mn>64</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mn>64</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>H</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>W</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle></mml:math></inline-formula></td>
</tr>
<tr>
<td rowspan="2">Encoder_1</td>
<td>Axial attention_1</td>
<td>3 &#x00D7; 3 &#x00D7; 3</td>
<td>1 &#x00D7; 1 &#x00D7; 1</td>
<td><inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mn>64</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>H</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>W</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mn>64</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>H</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>W</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle></mml:math></inline-formula></td>
</tr>
<tr>
<td>Res-dense block</td>
<td>3 &#x00D7; 3 &#x00D7; 3</td>
<td>1 &#x00D7; 1 &#x00D7; 1</td>
<td><inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mn>64</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>H</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>W</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mn>128</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>H</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>W</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle></mml:math></inline-formula></td>
</tr>
<tr>
<td>Downsampling</td>
<td>Conv3d</td>
<td>1 &#x00D7; 2 &#x00D7; 2</td>
<td>1 &#x00D7; 2 &#x00D7; 2</td>
<td><inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mn>128</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>H</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>W</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mn>128</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>H</mml:mi><mml:mn>4</mml:mn></mml:mfrac></mml:mstyle><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>W</mml:mi><mml:mn>4</mml:mn></mml:mfrac></mml:mstyle></mml:math></inline-formula></td>
</tr>
<tr>
<td rowspan="2">Encoder_2</td>
<td>Axial attention_2</td>
<td>3 &#x00D7; 3 &#x00D7; 3</td>
<td>1 &#x00D7; 1 &#x00D7; 1</td>
<td><inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mn>128</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>H</mml:mi><mml:mn>4</mml:mn></mml:mfrac></mml:mstyle><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>W</mml:mi><mml:mn>4</mml:mn></mml:mfrac></mml:mstyle></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mn>128</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>H</mml:mi><mml:mn>4</mml:mn></mml:mfrac></mml:mstyle><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>W</mml:mi><mml:mn>4</mml:mn></mml:mfrac></mml:mstyle></mml:math></inline-formula></td>
</tr>
<tr>
<td>Res-dense block</td>
<td>3 &#x00D7; 3 &#x00D7; 3</td>
<td>1 &#x00D7; 1 &#x00D7; 1</td>
<td><inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mn>128</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>H</mml:mi><mml:mn>4</mml:mn></mml:mfrac></mml:mstyle><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>W</mml:mi><mml:mn>4</mml:mn></mml:mfrac></mml:mstyle></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mn>256</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>H</mml:mi><mml:mn>4</mml:mn></mml:mfrac></mml:mstyle><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>W</mml:mi><mml:mn>4</mml:mn></mml:mfrac></mml:mstyle></mml:math></inline-formula></td>
</tr>
<tr>
<td>Upsampling</td>
<td>TransposeConv3d</td>
<td>1 &#x00D7; 2 &#x00D7; 2</td>
<td>1 &#x00D7; 2 &#x00D7; 2</td>
<td><inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mn>256</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>H</mml:mi><mml:mn>4</mml:mn></mml:mfrac></mml:mstyle><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>W</mml:mi><mml:mn>4</mml:mn></mml:mfrac></mml:mstyle></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mn>128</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>H</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>W</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle></mml:math></inline-formula></td>
</tr>
<tr>
<td rowspan="2">Decoder_1</td>
<td>Axial attention_3</td>
<td>3 &#x00D7; 3 &#x00D7; 3</td>
<td>1 &#x00D7; 1 &#x00D7; 1</td>
<td><inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mn>128</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>H</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>W</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mn>128</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>H</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>W</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle></mml:math></inline-formula></td>
</tr>
<tr>
<td>Res-dense block</td>
<td>3 &#x00D7; 3 &#x00D7; 3</td>
<td>1 &#x00D7; 1 &#x00D7; 1</td>
<td><inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mn>128</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>H</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>W</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mn>128</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>H</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>W</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle></mml:math></inline-formula></td>
</tr>
<tr>
<td>Upsampling</td>
<td>TransposeConv3d</td>
<td>1 &#x00D7; 2 &#x00D7; 2</td>
<td>1 &#x00D7; 2 &#x00D7; 2</td>
<td><inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mn>128</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>H</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle><mml:mo>&#x00D7;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>W</mml:mi><mml:mn>2</mml:mn></mml:mfrac></mml:mstyle></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mn>64</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>H</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td rowspan="2">Decoder_2</td>
<td>Axial attention_4</td>
<td>3 &#x00D7; 3 &#x00D7; 3</td>
<td>1 &#x00D7; 1 &#x00D7; 1</td>
<td><inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mn>64</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>H</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mn>64</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:math></inline-formula></td>
</tr>
<tr>
<td>Res-dense block</td>
<td>3 &#x00D7; 3 &#x00D7; 3</td>
<td>1 &#x00D7; 1 &#x00D7; 1</td>
<td><inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mn>64</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>H</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mn>64</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>H</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow></mml:math></inline-formula></td>
</tr>
<tr>
<td>Output</td>
<td>Conv3d</td>
<td>3 &#x00D7; 3 &#x00D7; 3</td>
<td>1 &#x00D7; 1 &#x00D7; 1</td>
<td><inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:mn>64</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>H</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow></mml:math></inline-formula></td>
<td><inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>D</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>H</mml:mtext></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:mtext>W</mml:mtext></mml:mrow></mml:math></inline-formula></td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Training Details</title>
<p>The network was trained and tested in the PyTorch framework on the AMAX workstation with an Intel Xeon Gold 5118 CPU and 64 GB of available memory. Used three Graphics Processing Units GeForce RTX 3090Ti. An Adam optimizer with a 1 &#x00D7; 10<sup>&#x2212;4</sup> learning rate was used to minimize the objective function during the network training. In the network parameter training optimization experiments, the batch size was set to 48, the epoch was set to 100, and it took 36 h to complete the training on 14,297 pairs of real datasets.</p>
<p>The proposed method uses deep networks to deal with the problem of Nano-CT image noise, which is independent of the statistical properties of the noise [<xref ref-type="bibr" rid="ref-28">28</xref>]. It attempts to learn from many data pairs to find the mapping relationship between short exposure time CT images and long exposure time. Therefore, Nano-CT denoising can be summarized as the optimal approximation of the objective function, and the denoising process can be expressed as:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mrow><mml:mtext>arg min</mml:mtext></mml:mrow><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:munder><mml:mi>L</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow></mml:math></disp-formula>
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>M</mml:mi><mml:mi>S</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi>m</mml:mi><mml:mi>n</mml:mi><mml:mi>d</mml:mi></mml:mrow></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>a</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>b</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>c</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:munderover><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>c</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mi>g</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>a</mml:mi><mml:mo>,</mml:mo><mml:mi>b</mml:mi><mml:mo>,</mml:mo><mml:mi>c</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></disp-formula>where <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mi>f</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the output of the model, <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:mi>y</mml:mi></mml:math></inline-formula> is the labelled image, and <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>M</mml:mi><mml:mi>S</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the loss function.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiments</title>
<p>In this section, we conduct a series of experiments on the real data to fully evaluate the proposed AAD-ResNet. Detailed information about these experiments is described as follows.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Datasets</title>
<p>To ensure the accuracy and generalization of the model usually requires a large amount of data to train the model. We have accumulated two types of sample data (fossil and coral) based on Nano-CT. We acquire real data under different exposure time scanning conditions. The reconstruction algorithm obtains the 3D volume data to construct the dataset. The scanning parameters are shown in <xref ref-type="table" rid="table-2">Table 2</xref>, the scanning parameter tube voltage is set to 80 kV, and the power of the ray source is set to 3 W. The long exposure time is set to 3 s, and the short exposure time is set to 0.1 s. Different quality CT images are acquired by this acquisition method. The scans that are taken at long exposure times are of high quality and are used as labels for network training. The scans that are taken at short exposure times are of poor quality, contain a lot of noise and artifacts, and are used as training data.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Scanning parameters</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Tube voltage (kV)</th>
<th>Power (W)</th>
<th>Rotation step of round 1 (deg)</th>
<th>Exposure time (s)</th>
</tr>
</thead>
<tbody>
<tr>
<td>80</td>
<td>3</td>
<td>0.225</td>
<td>3.0/0.1</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>In addition, due to the size of the Nano-CT reconstructed volume data, which is 512 &#x00D7; 512 &#x00D7; 1024, the memory consumption is high during processing. Due to the limitation of computer hardware, it is inefficient and infeasible to process complete 3D CT images directly. Therefore, we consider overlapping and cropping the full 3D volume data into block data. This strategy is effective and efficient and can significantly increase the training dataset. In network training, models trained on small image blocks converge faster and use less GPU memory, while models trained on larger image blocks have better image quality and take longer to train. Considering the model convergence, memory usage, model performance, and testing time, the method in this paper selects 64 &#x00D7; 64 &#x00D7; 64 block data for network training. The consecutive 128 layers in the volume data are randomly chosen as the test data (the test data size is 512 &#x00D7; 512 &#x00D7; 128). Each dataset contains 14,297 training data and 2,523 test data.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Evaluation Metrics</title>
<p>The image recovery task widely uses the evaluation metrics Peak Signal Noise Ratio (PSNR) and Structural Similarity (SSIM) to evaluate the method performance. The formulae for PSNR and SSIM are defined as follows:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mi>P</mml:mi><mml:mi>S</mml:mi><mml:mi>N</mml:mi><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mn>10</mml:mn><mml:mo>&#x22C5;</mml:mo><mml:msub><mml:mi>log</mml:mi><mml:mrow><mml:mn>10</mml:mn></mml:mrow></mml:msub><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mi>N</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mrow><mml:mtext>n</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>Y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula></p>
<p><disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mrow><mml:mtext>SSIM</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>2</mml:mn><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>2</mml:mn><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msub><mml:mi>C</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula>where <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mi>x</mml:mi></mml:math></inline-formula> is the output image of the network and <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mi>y</mml:mi></mml:math></inline-formula> is the ground truth, <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> are the mean and variance of the output image, <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msub><mml:mi>&#x03BC;</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:msubsup><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup></mml:math></inline-formula> are the mean and variance of the ground truth, and <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:msub><mml:mi>&#x03C3;</mml:mi><mml:mrow><mml:mi>x</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the variance between the ground truth and the output image. PSNR are generally used to evaluate the deviation of the output image from the ground truth. The larger PSNR means better image quality. The larger SSIM indicates that the output image is visually closer to the ground truth.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Experimental Results and Analysis</title>
<sec id="s5_1">
<label>5.1</label>
<title>Comparison Experiments</title>
<p>To analyze and judge the denoising effect of the methods proposed in this paper, we choose 3D UnetR, ADHDC-Net, and 3D-Unet as the comparison methods. We use their officially released codes as much as possible to make a fairer comparison with the existing methods. We train and test the network according to the training parameters given in the related literature.</p>
<p>First, we show the test results for the fossil data, where we randomly show the xOy and xOz planes of the test data. As shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>, from left to right, long exposure time image, short exposure time image, UnetR, ADHDCNet, 3D-Unet, and the proposed method. From <xref ref-type="fig" rid="fig-4">Fig. 4</xref>, it can be seen that all the methods can reduce the noise in the Nano-CT images with different degrees. <xref ref-type="fig" rid="fig-3">Fig. 3</xref> demonstrates the deviation between the noise reduction results of each method and the label image. All processing methods can suppress the noise in the Nano-CT image. However, each method also loses part of the image information while removing noise, especially the loss of image contour information is more serious. In contrast, the method proposed in this paper avoids this phenomenon as much as possible.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Comparison of the fossil denoising results of actual short-exposure images. From (a) to (f) indicate the long exposure CT images, short exposure CT images, UnetR, ADHDCNet, 3D-Unet, and proposed denoised images</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_49623-fig-4.tif"/>
</fig>
<p>To better identify the low-contrast regions and further compare the results of each method in terms of detailed structural processing effects, we mark the region of interest (ROI) with a red box in <xref ref-type="fig" rid="fig-4">Fig. 4</xref> to observe better and compare the correction effects of each method. More obvious splicing gaps exist in the comparison methods UnetR and ADHDCNet processing results. Our method can effectively improve the image quality of a 0.1-s exposure scan to a level close to a 3-s exposure, preserving the image structure and detail information.</p>

<p>In addition, evaluation metrics are used to evaluate the results of each method quantitatively. As shown in <xref ref-type="table" rid="table-3">Table 3</xref>, the optimal and suboptimal solutions of the quantization results are highlighted in the table, where the optimal solutions are shown in bold and the suboptimal solutions are underlined. The quantitative results are consistent with the visual perception, and the proposed method has the highest PSNR and SSIM among all the test results. Based on the average of the test results, we calculate the percentage change in metric values of the proposed method relative to the other methods. Specifically, the proposed method improves the PSNR and SSIM of the noise reduction results by 1.62% and 0.14% compared to UnetR, 22.96% PSNR and 1.32% SSIM compared to ADHDCNet, and 7.22% PSNR and 0.1% SSIM compared to Unet3D.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Quantitative results (PSNR\SSIM) of the fossil test datasets based on different methods of processing results</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th></th>
<th></th>
<th>Noise</th>
<th>UnetR</th>
<th>ADHDCNet</th>
<th>3D-Unet</th>
<th>AAD-ResNet</th>
</tr>
</thead>
<tbody>
<tr>
<td>PSNR</td>
<td>Slice30</td>
<td>23.52</td>
<td><underline>46.27</underline></td>
<td>38.24</td>
<td>43.85</td>
<td><bold>47.02</bold></td>
</tr>
<tr>
<td></td>
<td>ROI</td>
<td>23.42</td>
<td>40.51</td>
<td>38.66</td>
<td><underline>41.05</underline></td>
<td><bold>41.72</bold></td>
</tr>
<tr>
<td></td>
<td>Average</td>
<td>25.01 &#x00B1; 1.65</td>
<td><underline>47.32 &#x00B1; 2.43</underline></td>
<td>43.38 &#x00B1; 4.09</td>
<td>44.77 &#x00B1; 2.36</td>
<td><bold>47.83 &#x00B1; 2.34</bold></td>
</tr>
<tr>
<td>SSIM</td>
<td>Slice30</td>
<td>0.7796</td>
<td>0.9920</td>
<td>0.9802</td>
<td><underline>0.9924</underline></td>
<td><bold>0.9934</bold></td>
</tr>
<tr>
<td></td>
<td>ROI</td>
<td>0.8221</td>
<td>0.9797</td>
<td>0.9799</td>
<td><bold>0.9819</bold></td>
<td><underline>0.9812</underline></td>
</tr>
<tr>
<td></td>
<td>Average</td>
<td>0.7814 &#x00B1; 0.0023</td>
<td><underline>0.9927 &#x00B1; 0.0014</underline></td>
<td>0.9891 &#x00B1; 0.0045</td>
<td>0.9918 &#x00B1; 0.0006</td>
<td><bold>0.9931 &#x00B1; 0.0005</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Next, we show the denoising results of different methods on the coral data. As above, we show the xOy plane and xOz plane of the coral volume data. Similarly, our method denoises while minimizing the loss of image information. Moreover, we can see from <xref ref-type="fig" rid="fig-5">Fig. 5</xref> that the denoising results of UnetR have serious blurring of image information in the ROIs and more obvious splicing gaps in the xOz plane. ADHDCNet can suppress noise but still produces new artifacts in the low-contrast region. The 3D-Unet method can achieve noise suppression, but the processing results also have blurring, resulting in some image information loss. The proposed method can achieve better noise suppression and retain image details.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Comparison of the coral denoising results of actual short-exposure images. From (a) to (f) indicate the long exposure CT images, short exposure CT images, UnetR, ADHDCNet, 3D-Unet, and proposed denoised images</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_49623-fig-5.tif"/>
</fig>
<p>The quantitative results of each method for coral data denoising are presented in <xref ref-type="table" rid="table-4">Table 4</xref>, using the same approach as in <xref ref-type="table" rid="table-3">Table 3</xref> to highlight the optimal and suboptimal solutions. The quantitative results are consistent with the visual perception, and the AAD-ResNet method proposed in this chapter achieves the highest PSNR and SSIM for both denoising results for the coral test set among all the test results, outperforming the other noise reduction methods in terms of numerical results.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Quantitative results (PSNR\SSIM) of the coral test datasets based on different methods of processing results</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th></th>
<th></th>
<th>Noise</th>
<th>UnetR</th>
<th>ADHDCNet</th>
<th>3D-Unet</th>
<th>AAD-ResNet</th>
</tr>
</thead>
<tbody>
<tr>
<td>PSNR</td>
<td>Slice64</td>
<td>20.43</td>
<td>26.48</td>
<td><underline>44.20</underline></td>
<td>42.84</td>
<td><bold>45.47</bold></td>
</tr>
<tr>
<td></td>
<td>ROI</td>
<td>22.60</td>
<td>24.86</td>
<td>27.4153</td>
<td><underline>39.04</underline></td>
<td><bold>42.76</bold></td>
</tr>
<tr>
<td></td>
<td>Average</td>
<td>21.06 &#x00B1; 1.18</td>
<td>26.49 &#x00B1; 1.48</td>
<td>38.59 &#x00B1; 5.32</td>
<td><underline>43.10 &#x00B1; 1.50</underline></td>
<td><bold>46.72 &#x00B1; 1.46</bold></td>
</tr>
<tr>
<td>SSIM</td>
<td>Slice64</td>
<td>0.6889</td>
<td>0.7795</td>
<td><underline>0.9932</underline></td>
<td>0.9906</td>
<td><bold>0.9943</bold></td>
</tr>
<tr>
<td></td>
<td>ROI</td>
<td>0.8251</td>
<td>0.7997</td>
<td>0.8691</td>
<td><underline>0.9843</underline></td>
<td><bold>0.9882</bold></td>
</tr>
<tr>
<td></td>
<td>Average</td>
<td>0.6870 &#x00B1; 0.0026</td>
<td>0.7864 &#x00B1; 0.0069</td>
<td>0.9772 &#x00B1; 0.0005</td>
<td><underline>0.9913 &#x00B1; 0.0006</underline></td>
<td><bold>0.9951 &#x00B1; 0.0004</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="fig" rid="fig-6">Fig. 6</xref> gives the gray scale distribution of the labeled positions in <xref ref-type="fig" rid="fig-4">Figs. 4</xref> and <xref ref-type="fig" rid="fig-5">5</xref>. The denoising results of the proposed method are closer to the fluctuation of the label image at the edge position, which indicates that the proposed method has good edge preservation ability. The gray level distribution of the rest of the compared methods still has a part of the deviation from the label.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>The grayscale values at the locations marked by the green lines in <xref ref-type="fig" rid="fig-4">Figs. 4</xref> and <xref ref-type="fig" rid="fig-5">5</xref></title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_49623-fig-6.tif"/>
</fig>
<p>In addition, we compare the number of network parameters and test times of these four methods. As shown in <xref ref-type="table" rid="table-5">Table 5</xref>, the number of parameters and test time of the proposed method are closer to that of the UnetR model, and the ADHDCNet model has the least number of parameters, but the test time is slightly longer. The 3D-Unet model has the most significant parameters and the longest test time. Combining the above objective and quantitative analyses, the following conclusion can be drawn: the 3D-Unet model has similar noise reduction performance as our method, but our method has the advantages of small number of parameters and fast test time.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Number of model parameters and testing time for different methods</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>Parameters (M)</th>
<th>Test time (s)</th>
</tr>
</thead>
<tbody>
<tr>
<td>UnetR</td>
<td>5.91</td>
<td>621</td>
</tr>
<tr>
<td>ADHDCNet</td>
<td>1.37</td>
<td>747</td>
</tr>
<tr>
<td>3D-Unet</td>
<td>90.30</td>
<td>1243</td>
</tr>
<tr>
<td>AAD-ResNet</td>
<td>6.64</td>
<td>627</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Ablation Experiments</title>
<p>In this section, we design ablation experiments to demonstrate that each module in the proposed method is helpful for the denoising task by analyzing the noise reduction effect of the 3D axial attention mechanism and residual dense blocks. As shown in <xref ref-type="table" rid="table-6">Table 6</xref>, we validate each module on the coral dataset, while considering the values of the evaluation metrics with the number of network parameters, and verify that each module is helpful for the denoising task through ablation experiments. The network removes the dense residual block, PSNR decreases by 12.94%, and SSIM decreases by 0.7%. Removal of the axial attention block decreases PSNR by 2.61% and SSIM by 0.1%.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Ablation experiments for AAD-ResNet architecture</title>
</caption>
<table frame="hsides">
<colgroup>
<col align="left"/>
<col align="left"/>
<col align="left"/>
<col align="left"/>
</colgroup>
<thead>
<tr>
<th>Module</th>
<th>PSNR (dB)</th>
<th>SSIM</th>
<th>Parameters (M)</th>
</tr>
</thead>
<tbody>
<tr>
<td>Proposed without RDB</td>
<td>40.67 &#x00B1; 1.79</td>
<td><inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mn>0.9875</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>1.8723</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td>1.47</td>
</tr>
<tr>
<td>Proposed without axial attention</td>
<td>45.50 &#x00B1; 1.52</td>
<td><inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mn>0.9941</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>5.5529</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td>3.16</td>
</tr>
<tr>
<td>AAD-ResNet</td>
<td>46.72 &#x00B1; 1.46</td>
<td><inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mn>0.9951</mml:mn><mml:mo>&#x00B1;</mml:mo><mml:mn>3.7316</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td>6.64</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The above results show that the residual dense block can fully utilize the global features, which can effectively improve the ability of the network to remove noise. More features are essential in image recovery to maintain detailed information in the original image.</p>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusion</title>
<p>Tabletop Nano-CT has a wide range of applications as a nondestructive examination tool. However, it is usually necessary to scan under long exposure conditions to obtain high-quality Nano-CT images. It sacrifices imaging efficiency and increases time costs. Since Nano-CT reconstructed images are 3D data, existing methods are often accompanied by high computational costs, and the increased computational demands may limit the practical application of these methods. In order to advance the application of tabletop Nano-CT, this study aims to explore and develop new strategies to improve the quality of Nano-CT imaging under short exposure conditions.</p>
<p>With the development and broad application of deep learning technology, deep learning technology provides a new solution for noise removal of Nano-CT images. In this paper, we proposed a novel residual dense denoising network based on an axial attention mechanism. Experimental results verify the effectiveness of the proposed method. This method effectively suppresses the noise of Nano-CT in short-exposure-time imaging through a lightweight network, thus significantly reducing the scanning time of Nano-CT.</p>
</sec>
</body>
<back>
<ack><p>None.</p>
</ack>
<sec><title>Funding Statement</title>
<p>This work was supported by the National Natural Science Foundation of China (62201618).</p>
</sec>
<sec><title>Author Contributions</title>
<p>Conceptualization, H.F. and L.Z.; methodology, L.Z.; software, H.F.; validation, H.F., C.W. and X.X.; formal analysis, H.F. and C.W.; investigation, H.F.; resources, B.Y.; data curation, Z.C., C.W. and X.X.; writing-original draft preparation, H.F. and Y.H.; writing-review and editing, L.Z., L.L., B.Y. and Y.S.; visualization, H.F., L.Z. and C.W.; supervision, B.Y.; project administration, Y.S.; funding acquisition, Y.H. and L.L. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability"><title>Availability of Data and Materials</title>
<p>The data and the code used for the manuscript are available for researchers on request from the corresponding author.</p>
</sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Sakdinawat</surname></string-name> and <string-name><given-names>D.</given-names> <surname>Attwood</surname></string-name></person-group>, &#x201C;<article-title>Nanoscale X-ray imaging</article-title>,&#x201D; <source>Nat. Photon</source>., vol. <volume>4</volume>, no. <issue>12</issue>, pp. <fpage>840</fpage>&#x2013;<lpage>848</lpage>, <year>Dec. 2010</year>. doi: <pub-id pub-id-type="doi">10.1038/nphoton.2010.267</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Su</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Artificial neural network approach for multiphase segmentation of battery electrode nano-CT images</article-title>,&#x201D; <source>npj Comput. Mater</source>, vol. <volume>8</volume>, no. <issue>1</issue>, pp. <fpage>30</fpage>, <year>Feb. 2022</year>. doi: <pub-id pub-id-type="doi">10.1038/s41524-022-00709-7</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>W. K.</given-names> <surname>Lee</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Ge</surname></string-name></person-group>, &#x201C;<article-title>Sub-10 second fly-scan nano-tomography using machine learning</article-title>,&#x201D; <source>Commun. Mat.</source>, vol. <volume>3</volume>, pp. <fpage>91</fpage>, <year>Dec. 2022</year>. doi: <pub-id pub-id-type="doi">10.1038/s43246-022-00313-8</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E. J.</given-names> <surname>Curry</surname></string-name>, <string-name><given-names>A. D.</given-names> <surname>Henoun</surname></string-name>, <string-name><given-names>A. N.</given-names> <surname>Miller</surname></string-name>, and <string-name><given-names>T. D.</given-names> <surname>Nguyen</surname></string-name></person-group>, &#x201C;<article-title>3D nano- and micro-patterning of biomaterials for controlled drug delivery</article-title>,&#x201D; <source>Ther. Deliv.</source>, vol. <volume>8</volume>, no. <issue>1</issue>, pp. <fpage>15</fpage>&#x2013;<lpage>28</lpage>, <year>Jan. 2017</year>. doi: <pub-id pub-id-type="doi">10.4155/tde-2016-0052</pub-id>; <pub-id pub-id-type="pmid">27982732</pub-id></mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z. H.</given-names> <surname>Levine</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>A tabletop X-ray tomography instrument for nanometer-scale imaging: Reconstructions</article-title>,&#x201D; <source>Microsyst. Nanoeng.</source>, vol. <volume>9</volume>, pp. <fpage>47</fpage>, <year>2023</year>. doi: <pub-id pub-id-type="doi">10.1038/s41378-023-00510-6</pub-id>; <pub-id pub-id-type="pmid">37064166</pub-id></mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Kampschulte</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Nano-computed tomography: Technique and applications</article-title>,&#x201D; <source>Rofo</source>, vol. <volume>188</volume>, no. <issue>2</issue>, pp. <fpage>146</fpage>&#x2013;<lpage>154</lpage>, <year>Feb. 2016</year>; <pub-id pub-id-type="pmid">26815120</pub-id></mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Scharf</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Bridging nano- and microscale X-ray tomography for battery research by leveraging artificial intelligence</article-title>,&#x201D; <source>Nat. Nanotechnol.</source>, vol. <volume>17</volume>, no. <issue>5</issue>, pp. <fpage>446</fpage>&#x2013;<lpage>459</lpage>, <year>May 2022</year>. doi: <pub-id pub-id-type="doi">10.1038/s41565-022-01081-9</pub-id>; <pub-id pub-id-type="pmid">35414116</pub-id></mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Yu</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Phase retrieval in 3D X-ray magnified phase nano CT: Imaging bone tissue at the nanoscale</article-title>,&#x201D; <conf-name>presented at the 2017 IEEE 14th Int. Symp. Biomed. Imaging (ISBI 2017)</conf-name>, <conf-loc>Melbourne, VIC, Australia</conf-loc>, <year>2017</year>, pp. <fpage>56</fpage>&#x2013;<lpage>59</lpage>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Chen</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Low-dose CT with a residual encoder-decoder convolutional neural network</article-title>,&#x201D; <source>IEEE Trans. Med. Imaging</source>, vol. <volume>36</volume>, no. <issue>12</issue>, pp. <fpage>2524</fpage>&#x2013;<lpage>2535</lpage>, <year>Dec. 2017</year>. doi: <pub-id pub-id-type="doi">10.1109/TMI.2017.2715284</pub-id>; <pub-id pub-id-type="pmid">28622671</pub-id></mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Yao</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Chen</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Zhao</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Luo</surname></string-name></person-group>, &#x201C;<article-title>Micro-CT image denoising with an asymmetric perceptual convolutional network</article-title>,&#x201D; <source>Phys. Med. Biol.</source>, vol. <volume>66</volume>, no. <issue>13</issue>, pp. <fpage>135018</fpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.1088/1361-6560/ac0bd2</pub-id>; <pub-id pub-id-type="pmid">34134093</pub-id></mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Ge</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Zhao</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Si</surname></string-name> and <string-name><given-names>D.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>Low-dose CT denoising via sinogram inner-structure transformer</article-title>,&#x201D; <source>IEEE Trans. Med. Imaging</source>, vol. <volume>42</volume>, no. <issue>4</issue>, pp. <fpage>910</fpage>&#x2013;<lpage>921</lpage>, <year>2022</year>. doi: <pub-id pub-id-type="doi">10.1109/TMI.2022.3219856</pub-id>; <pub-id pub-id-type="pmid">36331637</pub-id></mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Shao</surname></string-name>, and <string-name><given-names>L.</given-names> <surname>Li</surname></string-name></person-group>, &#x201C;<article-title>Micro/nano functional devices fabricated by additive manufacturing</article-title>,&#x201D; <source>Prog. Mater. Sci.</source>, vol. <volume>131</volume>, pp. <fpage>101020</fpage>, <year>Jan. 2023</year>. doi: <pub-id pub-id-type="doi">10.1016/j.pmatsci.2022.101020</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Kodama</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Three-dimensional structural measurement and material identification of an all-solid-state lithium-ion battery by X-ray nanotomography and deep learning</article-title>,&#x201D; <source>J. Power Sources Adv.</source>, vol. <volume>8</volume>, pp. <fpage>100048</fpage>, <year>Apr. 2021</year>. doi: <pub-id pub-id-type="doi">10.1016/j.powera.2021.100048</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>M&#x00FC;ller</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Graetz</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Balles</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Stier</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Hanke</surname></string-name> and <string-name><given-names>C.</given-names> <surname>Fella</surname></string-name></person-group>, &#x201C;<article-title>Laboratory-based nano-computed tomography and examples of its application in the field of materials research</article-title>,&#x201D; <source>Crystals</source>, vol. <volume>11</volume>, no. <issue>6</issue>, pp. <fpage>677</fpage>, <year>Jun. 2021</year>. doi: <pub-id pub-id-type="doi">10.3390/cryst11060677</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>N. R.</given-names> <surname>Huber</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Inoue</surname></string-name>, <string-name><given-names>C. H.</given-names> <surname>McCollough</surname></string-name>, and <string-name><given-names>L.</given-names> <surname>Yu</surname></string-name></person-group>, &#x201C;<article-title>Multislice input for 2D and 3D residual convolutional neural network noise reduction in CT</article-title>,&#x201D; <source>J. Med. Imaging</source>, vol. <volume>10</volume>, no. <issue>1</issue>, pp. <fpage>14003</fpage>, <year>Jan. 2023</year>. doi: <pub-id pub-id-type="doi">10.1117/1.JMI.10.1.014003</pub-id>; <pub-id pub-id-type="pmid">36743869</pub-id></mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Q.</given-names> <surname>Li</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Unpaired low-dose computed tomography image denoising using a progressive cyclical convolutional neural network</article-title>,&#x201D; <source>Med. Phys.</source>, vol. <volume>51</volume>, no. <issue>2</issue>, pp. <fpage>1289</fpage>&#x2013;<lpage>1312</lpage>, <year>Feb. 2023</year>; <pub-id pub-id-type="pmid">36841936</pub-id></mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Ma</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>SCRDN: Residual dense network with self-calibrated convolutions for low dose CT image denoising</article-title>,&#x201D; <source>Nuclear Instrum. Methods Phys. Res. Section A: Accelerators, Spectrom., Detectors Associated Equip.</source>, vol. <volume>1045</volume>, pp. <fpage>167625</fpage>, <year>Jan. 2023</year>. doi: <pub-id pub-id-type="doi">10.1016/j.nima.2022.167625</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Liu</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>3D anisotropic hybrid network: Transferring convolutional features from 2D images to 3D anisotropic volumes</article-title>,&#x201D; <conf-name>presented at the Med. Image Comput. Comput. Assisted Interv&#x2013;MICCAI 2018</conf-name>, <conf-loc>Spain</conf-loc>, <year>Sep. 16&#x2013;20, 2018</year>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>G.</given-names> <surname>Huo</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Guan</surname></string-name>, and <string-name><given-names>M. L.</given-names> <surname>Tseng</surname></string-name></person-group>, &#x201C;<article-title>Multiscale lightweight 3D segmentation algorithm with attention mechanism: Brain tumor image segmentation</article-title>,&#x201D; <source>Expert. Syst. Appl.</source>, vol. <volume>214</volume>, pp. <fpage>119166</fpage>, <year>Mar. 2023</year>. doi: <pub-id pub-id-type="doi">10.1016/j.eswa.2022.119166</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Cao</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Zang</surname></string-name>, <string-name><given-names>D.</given-names> <surname>An</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Feng</surname></string-name> and <string-name><given-names>B.</given-names> <surname>Yu</surname></string-name></person-group>, &#x201C;<article-title>MBANet: A 3D convolutional neural network with multi-branch attention for brain tumor segmentation from MRI images</article-title>,&#x201D; <source>Biomed. Signal Process. Control</source>, vol. <volume>80</volume>, pp. <fpage>104296</fpage>, <year>Feb. 2023</year>. doi: <pub-id pub-id-type="doi">10.1016/j.bspc.2022.104296</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>You</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Structurally-sensitive multi-scale deep neural network for low-dose CT denoising</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>6</volume>, pp. <fpage>41839</fpage>&#x2013;<lpage>41855</lpage>, <year>2018</year>. doi: <pub-id pub-id-type="doi">10.1109/ACCESS.2018.2858196</pub-id>; <pub-id pub-id-type="pmid">30906683</pub-id></mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H.</given-names> <surname>Shan</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>3-D convolutional encoder-decoder network for low-dose CT via transfer learning from a 2-D trained network</article-title>,&#x201D; <source>IEEE Trans. Med. Imaging</source>, vol. <volume>37</volume>, no. <issue>6</issue>, pp. <fpage>1522</fpage>&#x2013;<lpage>1534</lpage>, <year>Jun. 2018</year>. doi: <pub-id pub-id-type="doi">10.1109/TMI.2018.2832217</pub-id>; <pub-id pub-id-type="pmid">29870379</pub-id></mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Yin</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Domain progressive 3D residual convolution network to improve low-dose CT imaging</article-title>,&#x201D; <source>IEEE Trans. Med. Imaging</source>, vol. <volume>38</volume>, no. <issue>12</issue>, pp. <fpage>2903</fpage>&#x2013;<lpage>2913</lpage>, <year>Dec. 2019</year>. doi: <pub-id pub-id-type="doi">10.1109/TMI.2019.2917258</pub-id>; <pub-id pub-id-type="pmid">31107644</pub-id></mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Hsu</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Xie</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Cong</surname></string-name>, and <string-name><given-names>W.</given-names> <surname>Gao</surname></string-name></person-group>, &#x201C;<article-title>SACNN: Self-attention convolutional neural network for low-dose CT denoising with self-supervised perceptual loss network</article-title>,&#x201D; <source>IEEE Trans. Med. Imaging.</source>, vol. <volume>39</volume>, no. <issue>7</issue>, pp. <fpage>2289</fpage>&#x2013;<lpage>2301</lpage>, <year>Jul. 2020</year>. doi: <pub-id pub-id-type="doi">10.1109/TMI.2020.2968472</pub-id>; <pub-id pub-id-type="pmid">31985412</pub-id></mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Gunduzalp</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Cengiz</surname></string-name>, <string-name><given-names>M. O.</given-names> <surname>Unal</surname></string-name>, and <string-name><given-names>I.</given-names> <surname>Yildirim</surname></string-name></person-group>, &#x201C;<article-title>3D U-NetR: Low dose computed tomography reconstruction via deep learning and 3 dimensional convolutions</article-title>,&#x201D; <comment>arXiv preprint arXiv:2105.14130</comment>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Vaswani</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Attention is all you need</article-title>,&#x201D; in <source>Advances in Neural Information Processing Systems</source>, vol. <volume>30</volume>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>D.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Fan</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Wang</surname></string-name> and <string-name><given-names>H.</given-names> <surname>Yu</surname></string-name></person-group>, &#x201C;<article-title>CTformer: Convolution-free Token2Token dilated vision transformer for low-dose CT denoising</article-title>,&#x201D; <source>Phys. Med. Biol.</source>, vol. <volume>68</volume>, no. <issue>6</issue>, pp. <fpage>65012</fpage>, <year>Mar. 2023</year>. doi: <pub-id pub-id-type="doi">10.1088/1361-6560/acc000</pub-id>; <pub-id pub-id-type="pmid">36854190</pub-id></mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Zhu</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>STEDNet: Swin transformer-based encoder-decoder network for noise reduction in low-dose CT</article-title>,&#x201D; <source>Med. Phys.</source>, vol. <volume>50</volume>, no. <issue>7</issue>, pp. <fpage>4443</fpage>&#x2013;<lpage>4458</lpage>, <year>Jan. 2023</year>. doi: <pub-id pub-id-type="doi">10.1002/mp.16249</pub-id>; <pub-id pub-id-type="pmid">36708286</pub-id></mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Kirillov</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Segment anything</article-title>,&#x201D; <conf-name>presented at the Int. Conf. Comput. Vis. (ICCV)</conf-name>, <conf-loc>Paris, France</conf-loc>, <year>Oct. 2023</year>, pp. <fpage>4015</fpage>&#x2013;<lpage>4026</lpage>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F. R.</given-names> <surname>Verdun</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Image quality in CT: From physical measurements to model observers</article-title>,&#x201D; <source>Phys. Med.</source>, vol. <volume>31</volume>, no. <issue>8</issue>, pp. <fpage>823</fpage>&#x2013;<lpage>843</lpage>, <year>Dec. 2015</year>. doi: <pub-id pub-id-type="doi">10.1016/j.ejmp.2015.08.007</pub-id>; <pub-id pub-id-type="pmid">26459319</pub-id></mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>A. L.</given-names> <surname>Kwan</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>N. J.</given-names> <surname>Packard</surname></string-name>, and <string-name><given-names>J. M.</given-names> <surname>Boone</surname></string-name></person-group>, &#x201C;<article-title>Noise power properties of a cone-beam CT system for breast cancer detection</article-title>,&#x201D; <source>Med. Phys.</source>, vol. <volume>35</volume>, no. <issue>12</issue>, pp. <fpage>5317</fpage>&#x2013;<lpage>5327</lpage>, <year>2008</year>; <pub-id pub-id-type="pmid">19175091</pub-id></mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Helfen</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Baumbach</surname></string-name>, and <string-name><given-names>H.</given-names> <surname>Suhonen</surname></string-name></person-group>, &#x201C;<article-title>Comparison of image quality in computed laminography and tomography</article-title>,&#x201D; <source>Opt. Express</source>, vol. <volume>20</volume>, no. <issue>2</issue>, pp. <fpage>794</fpage>&#x2013;<lpage>806</lpage>, <year>Jan. 2012</year>. doi: <pub-id pub-id-type="doi">10.1364/OE.20.000794</pub-id>; <pub-id pub-id-type="pmid">22274425</pub-id></mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>&#x00D6;.</given-names> <surname>&#x00C7;i&#x00E7;ek</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Abdulkadir</surname></string-name>, <string-name><given-names>S. S.</given-names> <surname>Lienkamp</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Brox</surname></string-name>, and <string-name><given-names>O.</given-names> <surname>Ronneberger</surname></string-name></person-group>, &#x201C;<article-title>3D U-Net: Learning dense volumetric segmentation from sparse annotation</article-title>,&#x201D; <conf-name>presented at the Med. Image Comput. Comput.-Assisted Interv. (MICCAI 2016)</conf-name>, <conf-loc>Athens, Greece</conf-loc>, <year>Oct. 2016</year>, pp. <fpage>424</fpage>&#x2013;<lpage>432</lpage>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Ho</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Kalchbrenner</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Weissenborn</surname></string-name>, and <string-name><given-names>T.</given-names> <surname>Salimans</surname></string-name></person-group>, &#x201C;<article-title>Axial attention in multidimensional transformers</article-title>,&#x201D; <comment>arXiv preprint arXiv:1912.12180</comment>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>He</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Ren</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Sun</surname></string-name></person-group>, &#x201C;<article-title>Deep residual learning for image recognition</article-title>,&#x201D; <conf-name>presented at the Conf. Comput. Vis. Pattern Recognit. (CVPR)</conf-name>, <conf-loc>Las Vegas, NV, USA</conf-loc>, <year>Jun. 2016</year>, pp. <fpage>770</fpage>&#x2013;<lpage>778</lpage>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>L.</given-names> <surname>van der Maaten</surname></string-name>, and <string-name><given-names>K. Q.</given-names> <surname>Weinberger</surname></string-name></person-group>, &#x201C;<article-title>Densely connected convolutional networks</article-title>,&#x201D; <conf-name>presented at the Conf. Comput. Vis. Pattern Recognit. (CVPR)</conf-name>, <conf-loc>Honolulu, HI, USA</conf-loc>, <year>Jul. 2017</year>, pp. <fpage>2261</fpage>&#x2013;<lpage>2269</lpage>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Tian</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Kong</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Zhong</surname></string-name>, and <string-name><given-names>Y.</given-names> <surname>Fu</surname></string-name></person-group>, &#x201C;<article-title>Residual dense network for image super-resolution</article-title>,&#x201D; <conf-name>presented at the Conf. Comput. Vis. Pattern Recognit. (CVPR)</conf-name>, <conf-loc>Salt Lake City, UT, USA</conf-loc>, <year>Jun. 2018</year>, pp. <fpage>2472</fpage>&#x2013;<lpage>2481</lpage>.</mixed-citation></ref>
</ref-list>
</back></article>