<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">57662</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2024.057662</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Retinexformer&#x002B;: Retinex-Based Dual-Channel Transformer for Low-Light Image Enhancement</article-title>
<alt-title alt-title-type="left-running-head">Retinexformer&#x002B;: Retinex-Based Dual-Channel Transformer for Low-Light Image Enhancement</alt-title>
<alt-title alt-title-type="right-running-head">Retinexformer&#x002B;: Retinex-Based Dual-Channel Transformer for Low-Light Image Enhancement</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Liu</surname><given-names>Song</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Zhang</surname><given-names>Hongying</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><email>zhanghongying@swust.edu.cn</email></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Li</surname><given-names>Xue</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Yang</surname><given-names>Xi</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<aff id="aff-1"><label>1</label><institution>School of Information Engineering, Southwest University of Science and Technology</institution>, <addr-line>Mianyang, 621000</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>Criminal Investigation Department, Sichuan Police College</institution>, <addr-line>Luzhou, 646000</addr-line>, <country>China</country></aff>
<aff id="aff-3"><label>3</label><institution>School of Electronics and Information, Mianyang Polytechnic</institution>, <addr-line>Mianyang, 621000</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Hongying Zhang. Email: <email>zhanghongying@swust.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2025</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>17</day><month>02</month><year>2025</year>
</pub-date>
<volume>82</volume>
<issue>2</issue>
<fpage>1969</fpage>
<lpage>1984</lpage>
<history>
<date date-type="received">
<day>24</day>
<month>8</month>
<year>2024</year>
</date>
<date date-type="accepted">
<day>30</day>
<month>10</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 The Authors.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_57662.pdf"></self-uri>
<abstract>
<p>Enhancing low-light images with color distortion and uneven multi-light source distribution presents challenges. Most advanced methods for low-light image enhancement are based on the Retinex model using deep learning. Retinexformer introduces channel self-attention mechanisms in the IG-MSA. However, it fails to effectively capture long-range spatial dependencies, leaving room for improvement. Based on the Retinexformer deep learning framework, we designed the Retinexformer&#x002B; network. The &#x201C;&#x002B;&#x201D; signifies our advancements in extracting long-range spatial dependencies. We introduced multi-scale dilated convolutions in illumination estimation to expand the receptive field. These convolutions effectively capture the weakening semantic dependency between pixels as distance increases. In illumination restoration, we used Unet&#x002B;&#x002B; with multi-level skip connections to better integrate semantic information at different scales. The designed Illumination Fusion Dual Self-Attention (IF-DSA) module embeds multi-scale dilated convolutions to achieve spatial self-attention. This module captures long-range spatial semantic relationships within acceptable computational complexity. Experimental results on the Low-Light (LOL) dataset show that Retexformer&#x002B; outperforms other State-Of-The-Art (SOTA) methods in both quantitative and qualitative evaluations, with the computational complexity increased to an acceptable 51.63 G FLOPS. On the LOL_v1 dataset, RetinexFormer&#x002B; shows an increase of 1.15 in Peak Signal-to-Noise Ratio (PSNR) and a decrease of 0.39 in Root Mean Square Error (RMSE). On the LOL_v2_real dataset, the PSNR increases by 0.42 and the RMSE decreases by 0.18. Experimental results on the Exdark dataset show that Retexformer&#x002B; can effectively enhance real-scene images and maintain their semantic information.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Low-light image enhancement</kwd>
<kwd>Retinex</kwd>
<kwd>transformer model</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Brightness information is a key indicator of image content representation. Enhancing the brightness of low-light images is an important research direction in the field of computer graphics.</p>
<p>Traditional image processing methods for low-light enhancement include histogram equalization [<xref ref-type="bibr" rid="ref-1">1</xref>&#x2013;<xref ref-type="bibr" rid="ref-5">5</xref>] and gamma correction [<xref ref-type="bibr" rid="ref-6">6</xref>&#x2013;<xref ref-type="bibr" rid="ref-8">8</xref>]. These methods are based on fundamental image principles and are highly interpretable. However, their applicability is limited in complex lighting scenes. According to visual imaging principles, light illuminates different object surfaces, and objects with different materials have distinct reflective properties. These reflected rays ultimately project onto the retina to form an image. The Retinex theory [<xref ref-type="bibr" rid="ref-9">9</xref>] decomposes an image into illumination and reflection components. It enhances low-light images by adjusting the illumination component. This approach provides theoretical support for addressing challenging low-light enhancement problems.</p>
<p>With the development of artificial intelligence technology, especially the wide application of deep learning methods in the field of image processing, new ideas and methods have been provided for low-light image enhancement. Recent research has applied convolutional neural networks [<xref ref-type="bibr" rid="ref-10">10</xref>&#x2013;<xref ref-type="bibr" rid="ref-13">13</xref>] and transformer models [<xref ref-type="bibr" rid="ref-14">14</xref>&#x2013;<xref ref-type="bibr" rid="ref-18">18</xref>] to low-light image enhancement. This research has achieved significant progress. Convolutional networks can effectively capture the regional spatial contextual information of images. They have shown certain efficacy in low-light image enhancement. However, convolutional networks have limitations in capturing long-range dependencies. Transformer-based methods express spatial long-range dependencies by introducing self-attention mechanisms. These methods can better restore lighting details in low-light images. The computational complexity of transformers is typically proportional to the square of the spatial size. This results in slow inference speed and high computational resource consumption.</p>
<p>Based on the summary of the Related Work, the Materials and Methods section provides a detailed description of the principles and framework of the proposed method. The framework consists of two parts: illumination estimation and damage restoration. The illumination estimator utilizes a multi-scale expanded convolution module to extract spatially distant semantic information. It represents the illumination characteristics of different regions and generates a perturbed illumination map. The illumination fusion module (IFU) of the damage restorer adopts the Unet&#x002B;&#x002B; network structure. This algorithm incorporates multi-level skip connections to better fuse semantic information from different levels and reduce information loss. The illumination fusion attention block (IFAB) introduces a spatial attention mechanism to express spatial distant dependencies within an acceptable computational complexity. The key branch of spatial attention employs multi-scale unfold convolutions to extract spatial position features. In the &#x201C;Results&#x201D; section, we first evaluate and validate the advancements of our method in image enhancement from both quantitative and qualitative dimensions using the Low-Light (LOL) dataset. We also compare the computational complexity and parameter count. Next, we demonstrate the effectiveness of our method in enhancing real-scene images using the Exdark dataset and conduct object detection experiments using recommended weights from YOLOv3 to evaluate the preservation of semantic information in enhanced images. In the ablation study, we compare the gains brought by different improvement details to the enhancement effect to support the effectiveness of our method. Finally, we summarize the strengths and weaknesses of our method and propose future research directions.</p>
<p>Extensive experiments show that Retinexformer&#x002B; achieves better quantitative and qualitative results on the LOL dataset [<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-20">20</xref>]. It outperforms other state-of-the-art (SOTA) supervised and unsupervised methods [<xref ref-type="bibr" rid="ref-21">21</xref>&#x2013;<xref ref-type="bibr" rid="ref-25">25</xref>]. Our main innovations are as follows:</p>
<p>1. Proposing a multi-scale dilated convolution structure. This structure expands the receptive field while expressing the weakening semantic dependencies between pixels as the distance increases.</p>
<p>2. Adopting Unet&#x002B;&#x002B; multi-level skip connections in damage restoration. This approach better fuses semantic information from different levels and reduces information loss.</p>
<p>3. Designing a novel multi-scale dilated convolution spatial attention module. This module expresses spatial long-range semantic relationships. It reduces the computational complexity of transformer spatial attention from quadratic to linear with respect to spatial size.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<sec id="s2_1">
<label>2.1</label>
<title>Classical Image Processing Methods</title>
<p>Classical image processing methods include histogram equalization [<xref ref-type="bibr" rid="ref-3">3</xref>&#x2013;<xref ref-type="bibr" rid="ref-5">5</xref>,<xref ref-type="bibr" rid="ref-26">26</xref>,<xref ref-type="bibr" rid="ref-27">27</xref>] and gamma correction [<xref ref-type="bibr" rid="ref-6">6</xref>&#x2013;<xref ref-type="bibr" rid="ref-8">8</xref>]. Cheng et al. [<xref ref-type="bibr" rid="ref-3">3</xref>] proposed a multi-peak generalized histogram equalization method that improves global histogram equalization by using multi-peak histogram equalization combined with local information. Lee et al. [<xref ref-type="bibr" rid="ref-4">4</xref>] proposed a novel contrast enhancement algorithm based on the layered difference representation of 2D histograms to enhance image contrast by amplifying the gray-level differences between adjacent pixels. Wang et al. [<xref ref-type="bibr" rid="ref-6">6</xref>] proposed a new method combining dynamic contrast ratio enhancement and inverse gamma correction for alternating current plasma display panel (AC PDP), and both are realized simultaneously. Huang et al. [<xref ref-type="bibr" rid="ref-7">7</xref>] proposed an automatic transformation technique is presented which improves the brightness of dimmed images via gamma correction and luminance pixel probability distribution and uses temporal information for video enhancement to reduce computational complexity. Rahman et al. [<xref ref-type="bibr" rid="ref-8">8</xref>] proposed an adaptive gamma correction method where parameters are set dynamically based on image information to appropriately enhance the contrast of the image. These methods enhance images by adjusting brightness and contrast. These methods are simple and highly interpretable. However, they do not account for complex real-world lighting scenarios. As a result, enhanced images often lack naturalness and realism.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Traditional Cognition Methods</title>
<p>According to the Retinex theory [<xref ref-type="bibr" rid="ref-9">9</xref>], light illuminates different object surfaces, and objects with different materials have distinct reflective properties. These reflections form images on the retina. The Retinex theory decomposes an image into illumination and reflectance components. It enhances low-light images by adjusting the illumination component [<xref ref-type="bibr" rid="ref-28">28</xref>&#x2013;<xref ref-type="bibr" rid="ref-32">32</xref>]. Fu et al. [<xref ref-type="bibr" rid="ref-28">28</xref>] proposed a fusion-based method for enhancing weakly illuminated images using multiple techniques, including decomposing, deriving, designing, fusing, and compensating to obtain an enhanced image for different weak illumination conditions. Fu et al. [<xref ref-type="bibr" rid="ref-29">29</xref>] proposed a weighted variational model to estimate the reflectance and illumination from an observed image. Guo et al. [<xref ref-type="bibr" rid="ref-30">30</xref>] proposed a simple yet effective low-light image enhancement (LIME) method. Wang et al. [<xref ref-type="bibr" rid="ref-32">32</xref>] made three major contributions including proposing a lightness-order-error measure, a bright-pass filter for image decomposition and a bi-log transformation for mapping illumination. However, these methods require manual setting of illumination priors. Inaccurate priors can lead to artifacts and color distortions in the enhanced results. Additionally, traditional methods often neglect the impact of noise. Simple brightness enhancement can retain and amplify noise.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Deep Learning Methods</title>
<p>In 2017, Lore et al. [<xref ref-type="bibr" rid="ref-33">33</xref>] proposed a deep learning-based low-light image enhancement algorithm. Convolutional neural network (CNN) methods [<xref ref-type="bibr" rid="ref-33">33</xref>&#x2013;<xref ref-type="bibr" rid="ref-37">37</xref>] have been widely applied in low-light image enhancement. Hong et al. [<xref ref-type="bibr" rid="ref-34">34</xref>] proposed a novel unsupervised low-light image enhancement network named LE-GAN based on generative adversarial networks and trained with unpaired low-light images. Lore et al. [<xref ref-type="bibr" rid="ref-33">33</xref>] proposed a deep autoencoder-based approach to identify signal features from low-light images and adaptively brighten images without over-amplifying lighter parts in high dynamic range images. Sharma et al. [<xref ref-type="bibr" rid="ref-35">35</xref>] proposed a method involving estimating the camera response function, decomposing the linearized image into LF and HF feature maps, processing them separately, and combining them to generate an output with increased dynamic range and suppressed light effects. Jiang et al. [<xref ref-type="bibr" rid="ref-36">36</xref>] introduced a highly effective unsupervised generative adversarial network that can be trained without low-light image pairs and generalizes well on real-world test images. Wei et al. [<xref ref-type="bibr" rid="ref-38">38</xref>] and subsequent studies [<xref ref-type="bibr" rid="ref-20">20</xref>,<xref ref-type="bibr" rid="ref-39">39</xref>] combined deep learning methods with the Retinex theory. They adjusted and optimized the Retinex model parameters through convolutional networks. These studies showed improvements over traditional cognitive methods. However, these methods often employ multi-stage training processes. This can be cumbersome. In 2019, Wang et al. [<xref ref-type="bibr" rid="ref-12">12</xref>] proposed a single-stage CNN method to directly predict illumination maps. However, this approach can cause color distortions and noise amplification while enhancing the brightness of low-light images. CNN-based methods still have limitations in expressing long-range spatial dependencies.</p>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Vision Transformer</title>
<p>In 2017, Vaswani et al. [<xref ref-type="bibr" rid="ref-40">40</xref>] presented the Transformer, a neural network architecture relying on self-attention to process sequential data rather than traditional recurrent or convolutional layers. Vision Transformer converts images into sequence data by splitting them into patches and mapping each patch to a vector for processing. Recently, Vision Transformers and their variants have been applied to low-light image enhancement. In 2022, Xu et al. [<xref ref-type="bibr" rid="ref-41">41</xref>] proposed SNR-Net, a CNN-Transformer hybrid model for low-light image enhancement. Due to the computational complexity of Transformers being proportional to the square of the spatial size, SNR-Net integrates a single global Transformer only at the lowest resolution of the U-shaped network. In 2023, Retinexformer [<xref ref-type="bibr" rid="ref-14">14</xref>], based on the Retinex theory, designed a single-stage Transformer network. It concatenates the height and width of the image into HW tokens and performs self-attention calculations in the channel direction. Spatial position information is extracted through two layers of 3 &#x00D7; 3 convolutional networks. Retinexformer achieved SOTA results that year. However, there is still room for improvement in extracting long-range spatial dependency features. The application of deep learning methods has significantly improved low-light image enhancement. However, further research is needed to enhance the extraction of long-range spatial dependency features under acceptable computational complexity.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Materials and Methods</title>
<p><xref ref-type="fig" rid="fig-1">Fig. 1</xref> shows the framework of Retinexformer&#x002B;. Retinexformer&#x002B; consists of an illumination estimator (<xref ref-type="fig" rid="fig-1">Fig. 1a</xref>) and an illumination-fused U-Net (IFU) (<xref ref-type="fig" rid="fig-1">Fig. 1b</xref>). The illumination estimator uses multi-scale dilated convolution. This method extracts image features and expresses pixel semantic dependencies. The damage restorer is designed as an illumination multi-scale fusion U-Net (IFU). Each layer of the IFU uses an illumination-fused attention block (IFAB) to fuse features (<xref ref-type="fig" rid="fig-1">Fig. 1b</xref>).</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Retinexformer&#x002B; consists of an (a) illumination estimator and an (b) illumination-fused U-Net</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57662-fig-1.tif"/>
</fig>
<sec id="s3_1">
<label>3.1</label>
<title>Network Framework Design</title>
<p>The Retinex image enhancement algorithm decomposes a low-light image <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>I</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. It splits it into a reflection component <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>R</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> and an illumination component <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>L</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>L</mml:mi><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, as follows:
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mi>I</mml:mi><mml:mo>=</mml:mo><mml:mi>R</mml:mi><mml:mo>&#x2299;</mml:mo><mml:mi>L</mml:mi></mml:math></disp-formula></p>
<p>The reflection component <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>R</mml:mi></mml:math></inline-formula> expresses the inherent reflective properties of objects. The illumination component <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>L</mml:mi></mml:math></inline-formula> describes the distribution of light in the scene. In real low-light scenarios, image acquisition often includes noise and artifacts. Reflection disturbances are expressed as <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mrow><mml:mover><mml:mi>R</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mn>3</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. The illumination distribution is uneven, with weaker illumination in shadowed areas or multiple light sources in low-light environments. Illumination disturbances are expressed as <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mrow><mml:mover><mml:mi>L</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>L</mml:mi><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. Considering the multi-source light field in real scenes, the illumination component can be decomposed as <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. The illumination disturbance can be decomposed as <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mrow><mml:mover><mml:mi>L</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:mrow><mml:mover><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula>. By introducing reflection disturbances and multi-source light field disturbances, the low-light enhancement function can be expressed as:
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mi>I</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mrow><mml:mover><mml:mi>L</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2299;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>R</mml:mi><mml:mo>+</mml:mo><mml:mrow><mml:mover><mml:mi>R</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2299;</mml:mo><mml:mi>R</mml:mi><mml:mo>+</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mrow><mml:mover><mml:mi>L</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2299;</mml:mo><mml:mi>R</mml:mi><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mrow><mml:mover><mml:mi>L</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2299;</mml:mo><mml:mrow><mml:mover><mml:mi>R</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></disp-formula></p>
<p>Multiplying both sides of the equation by the illumination map <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mover><mml:mi>L</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, we get:
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mover><mml:mi>L</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>&#x2299;</mml:mo><mml:mi>I</mml:mi><mml:mo>=</mml:mo><mml:mi>R</mml:mi><mml:mo>+</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mrow><mml:mover><mml:mi>L</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2299;</mml:mo><mml:mover><mml:mi>L</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>&#x2299;</mml:mo><mml:mi>R</mml:mi><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mrow><mml:mover><mml:mi>L</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2299;</mml:mo><mml:mover><mml:mi>L</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>&#x2299;</mml:mo><mml:mrow><mml:mover><mml:mi>R</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></disp-formula></p>
<p>In this equation, <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mover><mml:mi>L</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>&#x2299;</mml:mo><mml:mi>I</mml:mi></mml:math></inline-formula> is the illuminated image, denoted as <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>u</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. <italic>R</italic> is the normally exposed image. <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mrow><mml:mover><mml:mi>L</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2299;</mml:mo><mml:mover><mml:mi>L</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>&#x2299;</mml:mo><mml:mi>R</mml:mi></mml:math></inline-formula> is the color distortion part after illumination. <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mrow><mml:mover><mml:mi>L</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2299;</mml:mo><mml:mover><mml:mi>L</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>&#x2299;</mml:mo><mml:mrow><mml:mover><mml:mi>R</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow></mml:math></inline-formula> is the noise and artifacts enhanced during illumination.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Illumination Estimator</title>
<p>In the implementation of [<xref ref-type="bibr" rid="ref-14">14</xref>], each pixel of the input image <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mi>I</mml:mi></mml:math></inline-formula> is averaged along the channel dimension to obtain the illumination prior <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. The input image <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mi>I</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>p</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are concatenated into a four-channel input. A 5 &#x00D7; 5 convolution outputs the 40-channel light-up features <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>u</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. A <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> convolution converts the 40 channels into three channels, generating the light-up map <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mover><mml:mi>L</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>L</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. The 40-channel light-up features <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>u</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> describe various illumination characteristics. However, a 5 &#x00D7; 5 convolution kernel is insufficient to extract the spatial distribution features of multiple light sources.</p>
<p>To better extract multi-source distribution features, the convolution network design should consider two aspects. These are expressing long-range spatial semantic relationships and expressing semantic weights between pixels. The number of convolution layers and the convolution kernel size are crucial to achieving these goals. Multi-scale dilated convolution, with different dilation coefficients, can expand the receptive field. This method controls computational complexity and captures multi-scale contextual information. The HDC [<xref ref-type="bibr" rid="ref-37">37</xref>] rule should be followed in designing dilated convolutions. This ensures continuous coverage of the focus area and avoids holes or missing edges. The HDC rules are as follows:
<list list-type="simple">
<list-item><label>1.</label><p>The maximum distance between two non-zero elements in the second layer should be less than the size of the convolution kernel of that layer.</p></list-item>
<list-item><label>2.</label><p>Convolution coefficients should be set in a sawtooth pattern.</p></list-item>
<list-item><label>3.</label><p>The greatest common divisor of the dilation coefficients should not exceed one.</p></list-item>
</list></p>
<p>To balance computational complexity, the convolution kernel size is set to 3. This follows the HDC rule, where the maximum distance between non-zero elements in the second layer is less than 3. We use six layers of convolution with a sawtooth dilation coefficient pattern r &#x003D; 9,3,1,9,3,1. This generates a convolution receptive field of size 53 &#x00D7; 53.</p>
<p>When the dilation coefficient r &#x003D; 9, as shown in <xref ref-type="fig" rid="fig-2">Fig. 2a</xref>. When when r &#x003D; 9,3, as shown in <xref ref-type="fig" rid="fig-2">Fig. 2b</xref>, the maximum distance between non-zero elements is less than 3. When r &#x003D; 9,3,1, as shown in <xref ref-type="fig" rid="fig-2">Fig. 2c</xref>, the convolution is equivalent to a normal convolution with a kernel size of 27 &#x00D7; 27. When r &#x003D; 9,3,1,9, as shown in <xref ref-type="fig" rid="fig-2">Fig. 2d</xref>. When r &#x003D; 9,3,1,9,3, as shown in <xref ref-type="fig" rid="fig-2">Fig. 2e</xref>. When r &#x003D; 9,3,1,9,3,1, as shown in <xref ref-type="fig" rid="fig-2">Fig. 2f</xref>, the semantic weights between pixels decrease with increasing distance from the central pixel.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Multi-scale dilated convolution features</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57662-fig-2.tif"/>
</fig>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Illumination-Fused Unet&#x002B;&#x002B;</title>
<p><xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref> shows that spatial weight disturbances of multi-source illumination on the enhanced image cause noise and color distortion. To address this issue, a three-scale Unet&#x002B;&#x002B; structure is designed, as shown in <xref ref-type="fig" rid="fig-1">Fig. 1b</xref>. This structure uses multi-level skip connections. It can extract and fuse more semantic information at different levels. This reduces information loss and improves segmentation accuracy. It better expresses the spatial distribution features of multi-source illumination.</p>
<p>As shown in <xref ref-type="fig" rid="fig-1">Fig. 1b</xref>, the IFU is designed as a three-layer Unet&#x002B;&#x002B; [<xref ref-type="bibr" rid="ref-42">42</xref>] structure. Unet&#x002B;&#x002B; is a deeply supervised encoder-decoder network. The encoder and decoder sub-networks are connected through a series of nested and dense skip pathways. Compared with the Unet structure, it has enhanced feature extraction ability and better performance. In the encoding part, the illuminated image Ilu undergoes a 3 &#x00D7; 3 convolution with a stride of 2, an IFAB encoding, a 4 &#x00D7; 4 convolution with a stride of 2 for down-sampling, two IFAB encodings, and another 4 &#x00D7; 4 convolution with a stride of 2 to generate multi-scale encoded features <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mfrac><mml:mi>H</mml:mi><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mi>W</mml:mi><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:msup><mml:mn>2</mml:mn><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> where i &#x003D; 0,1,2. Subsequently, <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> undergoes two IFAB encodings. The decoding part features a symmetric up-sampling branch, with up-sampling using 2 &#x00D7; 2 deconv after two IFAB encodings of <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula>, adding skip connections at the same level to reduce information loss during encoding and decoding. Finally, the decoding part outputs the residual image <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, enhancing the illuminated image through <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>e</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>u</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>.</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Illumination-Fused Attention Block</title>
<p><xref ref-type="fig" rid="fig-3">Fig. 3</xref> shows the process flow of obtaining output features by fusing Light-up Feature and Input Feature through the IFAB. <xref ref-type="fig" rid="fig-3">Fig. 3a</xref> shows the structure of the IFAB module, which consists of a Layer Normalization (LN) layer, an Illumination-Fused Dual Self-Attention (IF-DSA) module, and a Feed-Forward Neural Network (FNN). The structure of the IF-DSA module is shown in <xref ref-type="fig" rid="fig-3">Fig. 3b</xref>.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>The Illumination-Fused Attention Block (IFAB) consists of a Layer Normalization (LN) layer, an Illumination-Fused Dual Self-Attention (IF-DSA) module, and a Feed-Forward Neural Network (FNN)</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57662-fig-3.tif"/>
</fig>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Illumination Fusion with Dual Self-Attention</title>
<p>By observing the input image, it is clear that different regions are illuminated by different light sources. Therefore, it is necessary to extract semantic information from different regions of the image to encode the illumination features and achieve brightness enhancement. Retinexformer [<xref ref-type="bibr" rid="ref-14">14</xref>] reshapes the input features <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> into <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mi>X</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>H</mml:mi><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and uses three bias-free fully connected layers to linearly map <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mi>X</mml:mi></mml:math></inline-formula> to <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mi>Q</mml:mi><mml:mo>=</mml:mo><mml:mi>X</mml:mi><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mi>K</mml:mi><mml:mo>=</mml:mo><mml:mi>X</mml:mi><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>K</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>and</mml:mtext></mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>V</mml:mi><mml:mo>=</mml:mo><mml:mi>X</mml:mi><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>H</mml:mi><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> (<inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>Q</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are learnable parameters of the linear mapping layers, and <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mi>T</mml:mi></mml:math></inline-formula> represents matrix transpose). The input light-up feature <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>u</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is reshaped into <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mi>Y</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>H</mml:mi><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, then <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mi>Q</mml:mi><mml:mo>,</mml:mo><mml:mi>K</mml:mi><mml:mo>,</mml:mo><mml:mi>V</mml:mi><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>and</mml:mtext></mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>Y</mml:mi></mml:math></inline-formula> are each split into <italic>k</italic> heads:<inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mi>Q</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>K</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>V</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mi>Y</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>Y</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>Y</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>Y</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>where</mml:mtext></mml:mrow><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>Y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>H</mml:mi><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mi>C</mml:mi><mml:mi>k</mml:mi></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo><mml:mspace width="thinmathspace" /><mml:mrow><mml:mtext>and</mml:mtext></mml:mrow><mml:mspace width="thinmathspace" /><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>k</mml:mi><mml:mo>.</mml:mo></mml:math></inline-formula> The channel self-attention for each head is represented as:
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mi>C</mml:mi><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>n</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>A</mml:mi><mml:mi>t</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>K</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>Y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>Y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2299;</mml:mo><mml:msub><mml:mi>V</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2299;</mml:mo><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:msubsup><mml:mi>K</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mrow><mml:mo>(</mml:mo><mml:mi>&#x03B1;</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x00D7;</mml:mo><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> is a learnable adaptive scaling parameter matrix. The spatial context information is extracted using two 3 &#x00D7; 3 convolutions to obtain the spatial position information <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mi>P</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, which is then connected to the channel multi-head self-attention. Finally, after reconstruction, the output feature <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is obtained.</p>
<p>Retinexformer [<xref ref-type="bibr" rid="ref-14">14</xref>] cleverly compresses the spatial <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:mi>H</mml:mi><mml:mi>W</mml:mi></mml:math></inline-formula> dimensions in the self-attention computation, making the computational complexity of the attention calculation only the square of the single-head channel size. This ensures that the computational complexity remains within an acceptable range while effectively obtaining channel attention to express different light-up features. However, using only two 3 &#x00D7; 3 convolutions to extract spatial semantic information is insufficient to express long-range dependencies in the space. Directly performing self-attention calculations on the spatial dimensions would result in a computational and memory resource overhead proportional to the square of the spatial size <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:math></inline-formula>, which is undeniably substantial. Therefore, expressing long-range dependencies in space while controlling computational complexity is a worthwhile research problem.</p>
<p>Retinexformer&#x002B; reshapes the input features <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> into <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:mi>X</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. It uses bias-free fully connected layers to linearly map <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:mi>X</mml:mi></mml:math></inline-formula> to <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mi>Q</mml:mi><mml:mo>=</mml:mo><mml:mi>X</mml:mi><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mi>V</mml:mi><mml:mo>=</mml:mo><mml:mi>X</mml:mi><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> (<inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>Q</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are learnable parameters of the linear mapping layers, and <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:mi>T</mml:mi></mml:math></inline-formula> represents matrix transpose). After <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mi>X</mml:mi></mml:math></inline-formula> undergoes 3 &#x00D7; 3 convolution operations with dilation coefficients r &#x003D; 9,3,1,9,3,1, <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:msup><mml:mi>K</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is obtained. <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mi>Q</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mi>K</mml:mi></mml:math></inline-formula> are concatenated, followed by two layers of 1 &#x00D7; 1 convolutions for fusion, achieving spatial position self-attention calculation. The input light-up feature <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>l</mml:mi><mml:mi>u</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is reshaped into <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mi>Y</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>H</mml:mi><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, and the spatial self-attention calculation formula is expressed as:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mi>S</mml:mi><mml:mi>p</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>A</mml:mi><mml:mi>t</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>Q</mml:mi><mml:mo>,</mml:mo><mml:mi>K</mml:mi><mml:mo>,</mml:mo><mml:mi>V</mml:mi><mml:mo>,</mml:mo><mml:mi>Y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>Y</mml:mi><mml:mo>&#x2299;</mml:mo><mml:mi>V</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2299;</mml:mo><mml:mi>s</mml:mi><mml:mi>o</mml:mi><mml:mi>f</mml:mi><mml:mi>t</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:mi>x</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>K</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2299;</mml:mo><mml:mi>Q</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x03B1;</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>s</mml:mi></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mi>&#x03B1;</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>s</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is a learnable adaptive scaling parameter matrix. Finally, the channel self-attention in Retinexformer [<xref ref-type="bibr" rid="ref-14">14</xref>] is reconstructed into <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mi>C</mml:mi><mml:mi>h</mml:mi><mml:mi>a</mml:mi><mml:mi>n</mml:mi><mml:mi>n</mml:mi><mml:mi>e</mml:mi><mml:mi>l</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>A</mml:mi><mml:mi>t</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> and connected to the spatial self-attention <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mi>S</mml:mi><mml:mi>p</mml:mi><mml:mi>a</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>A</mml:mi><mml:mi>t</mml:mi><mml:mi>t</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, ultimately reconstructing the output feature <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:msub><mml:mi>F</mml:mi><mml:mrow><mml:mi>o</mml:mi><mml:mi>u</mml:mi><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>C</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>.</p>
</sec>
<sec id="s3_6">
<label>3.6</label>
<title>Complexity Analysis</title>
<p>The computational complexity of the IF-DSA module includes the complexity of convolution operations and pointwise multiplication operations between matrices. The computational complexity for six convolution operations with a kernel size of 3 &#x00D7; 3 is <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:mi>O</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>H</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>W</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mi>C</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>3</mml:mn><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>&#x00D7;</mml:mo><mml:mn>6</mml:mn><mml:mo>=</mml:mo><mml:mn>54</mml:mn><mml:mi>H</mml:mi><mml:mi>W</mml:mi><mml:msup><mml:mi>C</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>. The complexity for pointwise multiplication operations between matrices is <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:mi>O</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>M</mml:mi><mml:mi>u</mml:mi><mml:mi>l</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi><mml:mi>l</mml:mi><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>H</mml:mi><mml:mi>W</mml:mi><mml:mi>C</mml:mi></mml:math></inline-formula>. Therefore, the complexity of the IF-DSA module can be expressed as:
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mi>O</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>I</mml:mi><mml:mi>F</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>D</mml:mi><mml:mi>S</mml:mi><mml:mi>A</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>O</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>C</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mi>v</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>O</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>P</mml:mi><mml:mi>o</mml:mi><mml:mi>i</mml:mi><mml:mi>n</mml:mi><mml:mi>t</mml:mi><mml:mi>w</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>e</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>M</mml:mi><mml:mi>u</mml:mi><mml:mi>l</mml:mi><mml:mi>t</mml:mi><mml:mi>i</mml:mi><mml:mi>p</mml:mi><mml:mi>l</mml:mi><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>54</mml:mn><mml:msup><mml:mi>C</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:mi>C</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>H</mml:mi><mml:mi>W</mml:mi></mml:math></disp-formula></p>
<p>In contrast, the computational complexity of traditional spatial self-attention methods, like the global MSA used in SNR-Net, is:
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mi>O</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>G</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>M</mml:mi><mml:mi>S</mml:mi><mml:mi>A</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mi>C</mml:mi><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mi>H</mml:mi><mml:mi>W</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></disp-formula></p>
<p>Comparing <xref ref-type="disp-formula" rid="eqn-6">Eqs. (6)</xref> and <xref ref-type="disp-formula" rid="eqn-7">(7)</xref> show that the IF-DSA module expresses long-range semantic relationships in space. It reduces the computational complexity of spatial self-attention methods from the square of the spatial size to a linear function of the spatial size.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Results</title>
<sec id="s4_1">
<label>4.1</label>
<title>Datasets and Implementation Details</title>
<p>We conducted experiments on the LOL dataset. The LOL dataset is divided into LOL_v1 [<xref ref-type="bibr" rid="ref-38">38</xref>] and LOL_v2 [<xref ref-type="bibr" rid="ref-19">19</xref>]. LOL_v1 contains 485 pairs of training data. It contains 15 pairs of testing data. LOL_v2 is further divided into LOL_v2_real and LOL_v2_synthetic. LOL_v2_real consists of real-scene photographs, containing 689 pairs of training data and 100 pairs of testing data. LOL_v2_synthetic consists of synthetic data. It contains 900 pairs of training data and 100 pairs of testing data. Each pair of data in all datasets includes one standard reference image. Each pair also includes one corresponding low-light image.</p>
<p>We implemented the Retinexformer&#x002B; model based on PyTorch. Training and testing were conducted on a Linux server equipped with a 3090 24 GB GPU. The system environment included CUDA 11.8, Python 3.7, and PyTorch 1.13. During training, the data size for LOL_v1 and LOL_v2_synthetic was set to 128 &#x00D7; 128, with a batch size of 8. For LOL_v2_real, the data size was 256 &#x00D7; 256, with a batch size of 8. We used random rotation and flipping to augment the training data. The training was performed using the Adam optimizer, with momentum terms of 0.9 and control parameters of 0.999. The aim was to minimize the Mean Absolute Error (MAE) between the enhanced image and the ground truth image. The initial learning rate was set to 2 &#x00D7; 10<sup>&#x2212;4</sup>. It was gradually decreased to 1 &#x00D7; 10<sup>&#x2212;6</sup> using a cosine annealing strategy.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Quantitative Results</title>
<p>In the quantitative evaluation, we used Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index Measure (SSIM), and Root Mean Square Error (RMSE) as metrics. Higher PSNR indicates better image enhancement. Higher SSIM indicates better retention of high-frequency details and structures. Lower RMSE indicates better model performance. As shown in <xref ref-type="table" rid="table-1">Table 1</xref>, our method improves the PSNR by 1.15 on the LOL_v1 dataset. Our method improves the PSNR by 0.42 on the LOL_v2_real dataset. The RMSE decreases by 0.39 on the LOL_v1 dataset. The RMSE decreases by 0.18 on the LOL_ v2_real dataset.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Quantitative comparisons on LOL_v1 [<xref ref-type="bibr" rid="ref-38">38</xref>] and LOL_v2 [<xref ref-type="bibr" rid="ref-19">19</xref>]. (Red data represents the best results, and blue data represents the second-best results. Our Retinexformer&#x002B; algorithm significantly outperforms other algorithms)</title>
</caption>
<table frame="hsides">
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Methods</th>
<th colspan="2" align="center">Complexity </th>
<th colspan="3" align="center">LOL_v1</th>
<th colspan="3" align="center">LOL_v2_real</th>
</tr>
<tr>
<th/>
<th>FLOPS (G)</th>
<th>Params (M)</th>
<th>PSNR&#x2191;</th>
<th>SSIM&#x2191;</th>
<th>RMSE&#x2193;</th>
<th>PSNR&#x2191;</th>
<th>SSIM&#x2191;</th>
<th>RMSE&#x2193;</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>Supervised</bold></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
</tr>
<tr>
<td>MBLLEN [<xref ref-type="bibr" rid="ref-10">10</xref>]</td>
<td>0.45</td>
<td>0.01</td>
<td>17.94</td>
<td>0.70</td>
<td>18.78</td>
<td>15.95</td>
<td>0.70</td>
<td>30.22</td>
</tr>
<tr>
<td>Retntinex-Net [<xref ref-type="bibr" rid="ref-38">38</xref>]</td>
<td>587.47</td>
<td>0.84</td>
<td>17.19</td>
<td>0.59</td>
<td>22.59</td>
<td>16.41</td>
<td>0.64</td>
<td>20.21</td>
</tr>
<tr>
<td>KinD [<xref ref-type="bibr" rid="ref-39">39</xref>]</td>
<td>34.99</td>
<td>8.02</td>
<td>20.35</td>
<td>0.81</td>
<td>14.30</td>
<td>18.07</td>
<td>0.78</td>
<td>18.04</td>
</tr>
<tr>
<td>KinD&#x002B;&#x002B; [<xref ref-type="bibr" rid="ref-20">20</xref>]</td>
<td>35.06</td>
<td>8.27</td>
<td>20.71</td>
<td>0.80</td>
<td>14.34</td>
<td>16.80</td>
<td>0.74</td>
<td>15.64</td>
</tr>
<tr>
<td>MIRNet [<xref ref-type="bibr" rid="ref-13">13</xref>]</td>
<td>785</td>
<td>31.79</td>
<td><styled-content style="color:#0000FF;">24.14</styled-content></td>
<td><styled-content style="color:#0000FF;">0.84</styled-content></td>
<td>12.03</td>
<td>20.36</td>
<td>0.78</td>
<td>14.21</td>
</tr>
<tr>
<td>URetntinex-Net [<xref ref-type="bibr" rid="ref-43">43</xref>]</td>
<td>53.02</td>
<td>0.34</td>
<td>21.45</td>
<td>0.80</td>
<td>13.55</td>
<td>21.55</td>
<td>0.80</td>
<td>14.23</td>
</tr>
<tr>
<td>Retinexformer [<xref ref-type="bibr" rid="ref-44">44</xref>]</td>
<td>15.57</td>
<td>1.61</td>
<td>23.86</td>
<td>0.83</td>
<td>8.30</td>
<td>21.93</td>
<td>0.84</td>
<td>9.56</td>
</tr>
<tr>
<td>RetinexMamba [<xref ref-type="bibr" rid="ref-44">44</xref>]</td>
<td>42.82</td>
<td>4.59</td>
<td>24.03</td>
<td>0.83</td>
<td><styled-content style="color:#0000FF;">8.17</styled-content></td>
<td><styled-content style="color:#0000FF;">22.45</styled-content></td>
<td><styled-content style="color:#0000FF;">0.84</styled-content></td>
<td><styled-content style="color:#0000FF;">9.38</styled-content></td>
</tr>
<tr>
<td>Retinexformer&#x002B;</td>
<td>51.63</td>
<td>3.55</td>
<td><styled-content style="color:#FF0000;">25.29</styled-content></td>
<td><styled-content style="color:#FF0000;">0.84</styled-content></td>
<td><styled-content style="color:#FF0000;">7.78</styled-content></td>
<td><styled-content style="color:#FF0000;">22.87</styled-content></td>
<td><styled-content style="color:#FF0000;">0.84</styled-content></td>
<td><styled-content style="color:#FF0000;">9.20</styled-content></td>
</tr>
<tr>
<td><bold>Unsupervised</bold></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
<td></td>
</tr>
<tr>
<td>Zero-Dce [<xref ref-type="bibr" rid="ref-22">22</xref>]</td>
<td>4.83</td>
<td>0.08</td>
<td>16.76</td>
<td>0.56</td>
<td>34.42</td>
<td>18.06</td>
<td>0.58</td>
<td>29.01</td>
</tr>
<tr>
<td>RUAS [<xref ref-type="bibr" rid="ref-23">23</xref>]</td>
<td>0.83</td>
<td>0.003</td>
<td>16.40</td>
<td>0.50</td>
<td>30.21</td>
<td>16.87</td>
<td>0.51</td>
<td>29.23</td>
</tr>
<tr>
<td>SCI [<xref ref-type="bibr" rid="ref-24">24</xref>]</td>
<td>0.02</td>
<td>0.003</td>
<td>14.86</td>
<td>0.54</td>
<td>24.87</td>
<td>15.34</td>
<td>0.52</td>
<td>27.50</td>
</tr>
<tr>
<td>PairLie [<xref ref-type="bibr" rid="ref-25">25</xref>]</td>
<td>20.81</td>
<td>0.34</td>
<td>19.69</td>
<td>0.71</td>
<td>19.03</td>
<td>19.29</td>
<td>0.68</td>
<td>20.01</td>
</tr>
<tr>
<td>NeRCO [<xref ref-type="bibr" rid="ref-45">45</xref>]</td>
<td>344.53</td>
<td>23.30</td>
<td>19.70</td>
<td>0.77</td>
<td>24.80</td>
<td>19.23</td>
<td>0.67</td>
<td>23.13</td>
</tr>
<tr>
<td>CLIP-LIE [<xref ref-type="bibr" rid="ref-46">46</xref>]</td>
<td>16.96</td>
<td>0.28</td>
<td>17.21</td>
<td>0.59</td>
<td>10.18</td>
<td>17.06</td>
<td>0.59</td>
<td>10.64</td>
</tr>
<tr>
<td>Enlighten-Your-Voice [<xref ref-type="bibr" rid="ref-47">47</xref>]</td>
<td>0.23</td>
<td>0.001</td>
<td>19.73</td>
<td>0.72</td>
<td>10.13</td>
<td>19.34</td>
<td>0.69</td>
<td>10.21</td>
</tr>
</tbody>
</table>
</table-wrap>
<p><xref ref-type="table" rid="table-1">Table 1</xref> compares the evaluation metrics of various supervised and unsupervised state-of-the-art (SOTA) methods. The comparison is based on the publicly available LOL_v1 and LOL_v2_real datasets. The data in the table were directly quoted from these papers.This demonstrates the superior performance of our method.</p>
<p>Our method has FLOPs and Params that are 3.31 times and 2.20 times larger than Retinexformer, and 1.21 times and 0.77 times larger than RetinexMamba. The computational complexity and parameter size are acceptable.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Qualitative Results</title>
<p><xref ref-type="fig" rid="fig-4">Figs. 4</xref> and <xref ref-type="fig" rid="fig-5">5</xref> show the qualitative comparison results of Retinexformer&#x002B; with other SOTA algorithms on the LOL_v1 and LOL_v2_real datasets, respectively. Upon zooming in, it can be observed that the images processed by the Retinex method exhibit significant noise and artifacts. KinD and Uretinex-Net show overexposure and underexposure in different regions. Retinexformer&#x002B; restores the color of the stapler more realistically than the Retinexformer and Retinexmamba methods. Furthermore, to visually demonstrate the effects of our method, we compared the HSV color space images corresponding to the images on the LOL_v1 and LOL_v2_real datasets in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>. It can be observed that the color distribution of the images processed by our method is closer to the Ground Truth images.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Qualitative results on the LOL_v1 dataset, showing that our method effectively controls exposure intensity across different regions, reduces noise, and achieves color reproduction closest to the ground truth image</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57662-fig-4.tif"/>
</fig><fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Qualitative results on the LOL_v2_real dataset, showing that our method effectively controls exposure intensity across different regions, reduces noise, and achieves color reproduction closest to the ground truth image</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57662-fig-5.tif"/>
</fig><fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Image display of the HSV color space</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57662-fig-6.tif"/>
</fig>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Low-Light Object Detection</title>
<p>In order to evaluate the preservation of semantic information by enhancement methods, we followed the evaluation method in References [<xref ref-type="bibr" rid="ref-48">48</xref>,<xref ref-type="bibr" rid="ref-49">49</xref>]. We conducted brightness enhancement on underexposed images from real scenes. Subsequently, we verified the effects using existing object detection models. The Exdark dataset, containing 7363 underexposed images with annotations for 12 object categories, was chosen for the experiment. After applying the brightness enhancement method described in <xref ref-type="table" rid="table-2">Table 2</xref>, we used the recommended weights of the YOLOv3 algorithm to detect the objects. The COCO dataset provides annotations for 80 different object categories found in real scenes. All 12 object categories from Exdark have corresponding types in COCO. We only needed to change the annotations for &#x201C;People&#x201D; and &#x201C;Table&#x201D; to &#x201C;Person&#x201D; and &#x201C;Dining Table&#x201D; in COCO. Finally, we evaluated the performance using the mAP metric.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Target detection results of YOLOv3 on the Exdark dataset after enhancement by different algorithms (Red data represents the best results, and blue data represents the second-best results)</title>
</caption>
<table frame="hsides">
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Methods</th>
<th>Bicycle</th>
<th>Boat</th>
<th>Bottle</th>
<th>Bus</th>
<th>Car</th>
<th>Cat</th>
<th>Chair</th>
<th>Cup</th>
<th>Dog</th>
<th>Motor</th>
<th>Person</th>
<th>Dining table</th>
<th>Mean</th>
</tr>
</thead>
<tbody>
<tr>
<td>Original LLIs</td>
<td>67.59</td>
<td>58.09</td>
<td>58.79</td>
<td>77.29</td>
<td>69.99</td>
<td>53.69</td>
<td>43.99</td>
<td>55.89</td>
<td>60.99</td>
<td>56.09</td>
<td>64.09</td>
<td>38.89</td>
<td>58.78</td>
</tr>
<tr>
<td>Retinex</td>
<td>56.22</td>
<td>49.42</td>
<td>51.72</td>
<td>70.42</td>
<td>62.42</td>
<td>43.62</td>
<td>36.82</td>
<td>48.22</td>
<td>50.12</td>
<td>47.82</td>
<td>52.62</td>
<td>31.12</td>
<td>50.05</td>
</tr>
<tr>
<td>KinD</td>
<td>62.44</td>
<td>56.24</td>
<td>56.04</td>
<td>75.64</td>
<td>67.24</td>
<td>48.64</td>
<td>42.14</td>
<td>53.54</td>
<td>55.54</td>
<td>53.74</td>
<td>60.74</td>
<td>36.24</td>
<td>55.68</td>
</tr>
<tr>
<td>Uretinex-Net</td>
<td>67.83</td>
<td>59.53</td>
<td>60.73</td>
<td>80.33</td>
<td>72.43</td>
<td>54.53</td>
<td>46.13</td>
<td>59.13</td>
<td>64.33</td>
<td>58.23</td>
<td>65.83</td>
<td>41.73</td>
<td>60.90</td>
</tr>
<tr>
<td>Retinexformer</td>
<td><styled-content style="color:#FF0000;">73.86</styled-content></td>
<td><styled-content style="color:#0000FF;">63.96</styled-content></td>
<td>62.66</td>
<td>81.16</td>
<td>72.96</td>
<td><styled-content style="color:#FF0000;">60.66</styled-content></td>
<td>49.06</td>
<td>59.09</td>
<td>63.59</td>
<td><styled-content style="color:#FF0000;">61.16</styled-content></td>
<td><styled-content style="color:#0000FF;">65.89</styled-content></td>
<td>43.06</td>
<td>63.09</td>
</tr>
<tr>
<td>Retinexmamba</td>
<td>72.59</td>
<td>63.19</td>
<td><styled-content style="color:#0000FF;">62.59</styled-content></td>
<td><styled-content style="color:#FF0000;">81.59</styled-content></td>
<td><styled-content style="color:#0000FF;">75.05</styled-content></td>
<td>58.85</td>
<td><styled-content style="color:#FF0000;">50.29</styled-content></td>
<td><styled-content style="color:#0000FF;">60.86</styled-content></td>
<td><styled-content style="color:#0000FF;">65.59</styled-content></td>
<td>60.69</td>
<td>63.16</td>
<td><styled-content style="color:#0000FF;">42.89</styled-content></td>
<td><styled-content style="color:#0000FF;">63.11</styled-content></td>
</tr>
<tr>
<td>Retinexformer&#x002B;</td>
<td><styled-content style="color:#0000FF;">73.65</styled-content></td>
<td><styled-content style="color:#FF0000;">64.15</styled-content></td>
<td><styled-content style="color:#FF0000;">62.75</styled-content></td>
<td><styled-content style="color:#0000FF;">81.15</styled-content></td>
<td><styled-content style="color:#FF0000;">75.19</styled-content></td>
<td><styled-content style="color:#0000FF;">60.19</styled-content></td>
<td><styled-content style="color:#0000FF;">49.85</styled-content></td>
<td><styled-content style="color:#FF0000;">61.25</styled-content></td>
<td><styled-content style="color:#FF0000;">65.76</styled-content></td>
<td><styled-content style="color:#0000FF;">60.75</styled-content></td>
<td><styled-content style="color:#FF0000;">66.15</styled-content></td>
<td><styled-content style="color:#FF0000;">43.25</styled-content></td>
<td><styled-content style="color:#FF0000;">63.67</styled-content></td>
</tr>
</tbody>
</table>
</table-wrap>

<p>The results in <xref ref-type="table" rid="table-2">Table 2</xref> show that images enhanced by our method achieved the highest average precision: 63.67 AP in the YOLOv3 model. This was a 0.56 AP improvement over the second highest. Furthermore, our method also had the best results in detecting categories such as Boat, Bottle, Car, Cup, Dog, Person, and Dining Table. It is worth noting that the detection accuracy for Retinex and KinD enhanced images did not improve.</p>

<p><xref ref-type="fig" rid="fig-7">Fig. 7</xref> shows target detection results of original and enhanced images. Our method suppresses noise and has generalization for real scene image enhancement in Exdark dataset. Using the recommended weights of YOLOv3 to perform target detection on pictures, original image has high detection confidence but missed detections. Retinex-enhanced image reduces confidence and has false detections. The KinD-enhanced image detects a car on the right side, but confidence decreases. Uretinex-Net misdetects car as person. Retinexformer slightly reduces person detection confidence. Retinexmamba significantly reduces person detection confidence. Our method reduces missed and false detection rates while maintaining confidence.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Comparison of object detection in low-light scenes using different methods on the Exdark dataset</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_57662-fig-7.tif"/>
</fig>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Ablation Study</title>
<p>We conducted ablation studies on three datasets: LOL_v1, LOL_v2_real, and LOL_v2_synthetic. We set up four different framework models to verify the contribution of each designed component to the improvement of the algorithm&#x2019;s performance.</p>
<p>&#x201C;Dilated conv&#x201D; indicates using a six-layer dilated convolution structure with dilation rates of r &#x003D; 9,3,1,9,3,1 during the illumination estimation phase. This structure expands the receptive field. It also captures how semantic dependencies between pixels weaken with increasing distance.</p>
<p>&#x201C;Unet&#x002B;&#x002B;&#x201D; indicates using Unet&#x002B;&#x002B; with multi-level skip connections for damage restoration. This approach better integrates semantic information from different levels and reduces information loss.</p>
<p>&#x201C;Conv-transformer&#x201D; indicates adding multi-scale dilated convolution spatial attention to the IF-DSA module. This addition helps express long-range spatial semantic relationships.</p>
<p>&#x201C;Transformer&#x002B;&#x201D; indicates the experimental results after incorporating all components into the network.</p>
<p>Under the same local configuration environment, we tested the PSNR, SSIM, and RMSE values for each framework model. The improvements of each component enhance the algorithm&#x2019;s performance to varying degrees, as shown in <xref ref-type="table" rid="table-3">Table 3</xref>. The data in the table are from the code testing results. Retinexformer&#x002B; performs the best in parameter evaluation metrics. This confirms that our network design is reasonable and effective.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Ablation Study on LOL_v1, LOL_v2_real and LOL_v2_syn (Red data represents the best results)</title>
</caption>
<table frame="hsides">
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Methods</th>
<th colspan="3" align="center">LOL_v1</th>
<th colspan="3" align="center">LOL_v2_real</th>
<th colspan="3" align="center">LOL_v2_syn</th>
</tr>
<tr>
<th/>
<th>PSNR&#x2191;</th>
<th>SSIM&#x2191;</th>
<th>RMSE&#x2193;</th>
<th>PSNR&#x2191;</th>
<th>SSIM&#x2191;</th>
<th>RMSE&#x2193;</th>
<th>PSNR&#x2191;</th>
<th>SSIM&#x2191;</th>
<th>RMSE&#x2193;</th>
</tr>
</thead>
<tbody>
<tr>
<td>Retinexformer [<xref ref-type="bibr" rid="ref-45">45</xref>]</td>
<td>23.86</td>
<td>0.83</td>
<td>8.30</td>
<td>21.93</td>
<td>0.84</td>
<td>9.56</td>
<td>25.37</td>
<td>0.93</td>
<td>8.43</td>
</tr>
<tr>
<td>Dilated conv</td>
<td>23.95</td>
<td>0.83</td>
<td>8.29</td>
<td>22.09</td>
<td>0.85</td>
<td>9.59</td>
<td>25.88</td>
<td>0.93</td>
<td>8.12</td>
</tr>
<tr>
<td>Unet&#x002B;&#x002B;</td>
<td>24.06</td>
<td>0.83</td>
<td>8.33</td>
<td>22.11</td>
<td>0.83</td>
<td>9.60</td>
<td>25.56</td>
<td>0.93</td>
<td>8.28</td>
</tr>
<tr>
<td>Conv-transformer</td>
<td>24.02</td>
<td>0.83</td>
<td>8.34</td>
<td>22.35</td>
<td>0.84</td>
<td>9.37</td>
<td>25.82</td>
<td>0.93</td>
<td>8.13</td>
</tr>
<tr>
<td>Retinexformer&#x002B;</td>
<td><styled-content style="color:#FF0000;">25.29</styled-content></td>
<td><styled-content style="color:#FF0000;">0.84</styled-content></td>
<td><styled-content style="color:#FF0000;">7.78</styled-content></td>
<td><styled-content style="color:#FF0000;">22.87</styled-content></td>
<td><styled-content style="color:#FF0000;">0.84</styled-content></td>
<td><styled-content style="color:#FF0000;">9.20</styled-content></td>
<td><styled-content style="color:#FF0000;">26.07</styled-content></td>
<td><styled-content style="color:#FF0000;">0.93</styled-content></td>
<td><styled-content style="color:#FF0000;">8.00</styled-content></td>
</tr>
</tbody>
</table>
</table-wrap>

</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>This paper introduces Retinexformer&#x002B;, a new architecture designed for low-light image enhancement. This architecture uses a six-layer dilated convolution to capture long-range spatial semantic features. This approach better represents the distribution characteristics of multiple light sources. The damage restorer utilizes Unet&#x002B;&#x002B; with multi-level skip connections. This design effectively integrates semantic information from different levels and reduces information loss. Most importantly, the IF-DSA module includes a novel convolution-based spatial attention module. This module captures long-range spatial dependencies. It reduces the computational complexity of transformer spatial attention. The complexity is reduced from being proportional to the square of the spatial size to being proportional to a multiple of the spatial size.</p>
<p>Based on qualitative and quantitative experimental analyses on the LOL dataset, Retinexformer&#x002B; outperforms current state-of-the-art methods. Experiments on the Exdark dataset verify its generalization for enhancing real-scene images. Detection results with YOLOv3 weights verify that Retinexformer&#x002B; can retain image semantic information while enhancing low-illumination images. Although Retinexformer&#x002B;&#x2019;s computational complexity is acceptable, it increases module parameters and consumes more resources. Future work will focus on improving model performance and reducing resource consumption, called &#x201C;Retinexformer&#x002B;&#x002B;&#x201D;.</p>
</sec>
</body>
<back>
<ack>
<p>The authors also gratefully acknowledge the helpful comments and suggestions of the reviewers, which have improved the presentation.</p>
</ack>
<sec><title>Funding Statement</title>
<p>This work was supported by the Key Laboratory of Forensic Science and Technology at College of Sichuan Province (2023YB04).</p>
</sec>
<sec><title>Author Contributions</title>
<p>Study conception and design: Song Liu, Hongying Zhang; data collection: Song Liu, Xue Li; analysis and interpretation of results: Song Liu, Xi Yang; draft manuscript preparation: Song Liu. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability"><title>Availability of Data and Materials</title>
<p>The data used in this paper can be requested from the corresponding author upon request.</p>
</sec>
<sec><title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement"><title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Abdullah-Al-Wadud</surname></string-name>, <string-name><given-names>M. H.</given-names> <surname>Kabir</surname></string-name>, <string-name><given-names>M. A. A.</given-names> <surname>Dewan</surname></string-name>, and <string-name><given-names>O.</given-names> <surname>Chae</surname></string-name></person-group>, &#x201C;<article-title>A dynamic histogram equalization for image contrast enhancement</article-title>,&#x201D; <source>IEEE Trans. Consum. Electron.</source>, vol. <volume>53</volume>, no. <issue>2</issue>, pp. <fpage>593</fpage>&#x2013;<lpage>600</lpage>, <year>2007</year>. doi: <pub-id pub-id-type="doi">10.1109/TCE.2007.381734</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><collab>Celik</collab> and <string-name><given-names>T.</given-names> <surname>Tjahjadi</surname></string-name></person-group>, &#x201C;<article-title>Contextual and variational contrast enhancement</article-title>,&#x201D; <source>IEEE Trans. Image Process.</source>, vol. <volume>20</volume>, no. <issue>12</issue>, pp. <fpage>3431</fpage>&#x2013;<lpage>3441</lpage>, <year>2011</year>. doi: <pub-id pub-id-type="doi">10.1109/TIP.2011.2157513</pub-id>; <pub-id pub-id-type="pmid">21609884</pub-id></mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>H. -D.</given-names> <surname>Cheng</surname></string-name> and <string-name><given-names>X.</given-names> <surname>Shi</surname></string-name></person-group>, &#x201C;<article-title>A simple and effective histogram equalization approach to image enhancement</article-title>,&#x201D; <source>Digit. Signal Process.</source>, vol. <volume>14</volume>, no. <issue>2</issue>, pp. <fpage>158</fpage>&#x2013;<lpage>170</lpage>, <year>2004</year>. doi: <pub-id pub-id-type="doi">10.1016/j.dsp.2003.07.002</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Lee</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Lee</surname></string-name>, and <string-name><given-names>C. -S.</given-names> <surname>Kim</surname></string-name></person-group>, &#x201C;<article-title>Contrast enhancement based on layered difference representation of 2D histograms</article-title>,&#x201D; <source>IEEE Trans. Image Process.</source>, vol. <volume>22</volume>, no. <issue>12</issue>, pp. <fpage>5372</fpage>&#x2013;<lpage>5384</lpage>, <year>2013</year>. doi: <pub-id pub-id-type="doi">10.1109/TIP.2013.2284059</pub-id>; <pub-id pub-id-type="pmid">24108715</pub-id></mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E. D.</given-names> <surname>Pisano</surname></string-name> <etal>et al</etal></person-group>., &#x201C;<article-title>Contrast limited adaptive histogram equalization image processing to improve the detection of simulated spiculations in dense mammograms</article-title>,&#x201D; <source>J. Digit. Imaging</source>, vol. <volume>11</volume>, no. <issue>4</issue>, pp. <fpage>193</fpage>&#x2013;<lpage>200</lpage>, <year>1998</year>. doi: <pub-id pub-id-type="doi">10.1007/BF03178082</pub-id>; <pub-id pub-id-type="pmid">9848052</pub-id></mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z. -G.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Z. -H.</given-names> <surname>Liang</surname></string-name>, and <string-name><given-names>C. -L.</given-names> <surname>Liu</surname></string-name></person-group>, &#x201C;<article-title>A real-time image processor with combining dynamic contrast ratio enhancement and inverse gamma correction for PDP</article-title>,&#x201D; <source>Displays</source>, vol. <volume>30</volume>, no. <issue>3</issue>, pp. <fpage>133</fpage>&#x2013;<lpage>139</lpage>, <year>2009</year>. doi: <pub-id pub-id-type="doi">10.1016/j.displa.2009.03.006</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. -C.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>F. -C.</given-names> <surname>Cheng</surname></string-name>, and <string-name><given-names>Y. -S.</given-names> <surname>Chiu</surname></string-name></person-group>, &#x201C;<article-title>Efficient contrast enhancement using adaptive gamma correction with weighting distribution</article-title>,&#x201D; <source>IEEE Trans. Image Process.</source>, vol. <volume>22</volume>, no. <issue>3</issue>, pp. <fpage>1032</fpage>&#x2013;<lpage>1041</lpage>, <year>2012</year>. doi: <pub-id pub-id-type="doi">10.1109/TIP.2012.2226047</pub-id>; <pub-id pub-id-type="pmid">23144035</pub-id></mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Rahman</surname></string-name>, <string-name><given-names>M. M.</given-names> <surname>Rahman</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Abdullah-Al-Wadud</surname></string-name>, <string-name><given-names>G. D.</given-names> <surname>Al-Quaderi</surname></string-name>, and <string-name><given-names>M.</given-names> <surname>Shoyaib</surname></string-name></person-group>, &#x201C;<article-title>An adaptive gamma correction for image enhancement</article-title>,&#x201D; <source>EURASIP J. Image Video Process.</source>, vol. <volume>2016</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>13</lpage>, <year>2016</year>. doi: <pub-id pub-id-type="doi">10.1186/s13640-016-0138-1</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E. H.</given-names> <surname>Land</surname></string-name> and <string-name><given-names>J. J.</given-names> <surname>McCann</surname></string-name></person-group>, &#x201C;<article-title>Lightness and retinex theory</article-title>,&#x201D; <source>J. Optical Society. America</source>, vol. <volume>61</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>11</lpage>, <year>1971</year>. doi: <pub-id pub-id-type="doi">10.1364/JOSA.61.000001</pub-id>; <pub-id pub-id-type="pmid">5541571</pub-id></mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>F.</given-names> <surname>Lv</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Lu</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Wu</surname></string-name>, and <string-name><given-names>C.</given-names> <surname>Lim</surname></string-name></person-group>, &#x201C;<article-title>MBLLEN: Low-light image/video enhancement using CNNs</article-title>,&#x201D; in <conf-name>BMVC</conf-name>, <publisher-name>Northumbria University</publisher-name>, <year>2018</year>, vol. <volume>220</volume>, no. <issue>1</issue>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Moran</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Marza</surname></string-name>, <string-name><given-names>S.</given-names> <surname>McDonagh</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Parisot</surname></string-name>, and <string-name><given-names>G.</given-names> <surname>Slabaugh</surname></string-name></person-group>, &#x201C;<article-title>DeepLPF: Deep local parametric filters for image enhancement</article-title>,&#x201D; in <conf-name>Proc. IEEE/CVF Conf. Compu. Vis. Pattern Recognit.</conf-name>, <year>2020</year>, pp. <fpage>12826</fpage>&#x2013;<lpage>12835</lpage>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>Q.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>C. -W.</given-names> <surname>Fu</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Shen</surname></string-name>, <string-name><given-names>W. -S.</given-names> <surname>Zheng</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Jia</surname></string-name></person-group>, &#x201C;<article-title>Underexposed photo enhancement using deep illumination estimation</article-title>,&#x201D; in <conf-name>Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit.</conf-name>, <year>2019</year>, pp. <fpage>6849</fpage>&#x2013;<lpage>6857</lpage>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S. W.</given-names> <surname>Zamir</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Learning enriched features for real image restoration and enhancement</article-title>,&#x201D; in <conf-name>Comput. Vis.&#x2013;ECCV 2020</conf-name>, <publisher-loc>Glasgow, UK</publisher-loc>, <publisher-name>Springer</publisher-name>, <year>Aug. 23&#x2013;28, 2020</year>, pp. <fpage>492</fpage>&#x2013;<lpage>511</lpage>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Cai</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Bian</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Lin</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Timofte</surname></string-name> and <string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>Retinexformer: One-stage retinex-based transformer for low-light image enhancement</article-title>,&#x201D; in <conf-name>Proc. IEEE/CVF Int. Conf. Comput. Vis.</conf-name>, <year>2023</year>, pp. <fpage>12504</fpage>&#x2013;<lpage>12513</lpage>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Yao</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Wu</surname></string-name>, and <string-name><given-names>X.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>Improving depth gradient continuity in transformers: A comparative study on monocular depth estimation with CNN</article-title>,&#x201D; <year>2023</year>, <comment><italic>arXiv:2308.08333</italic></comment>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Zhao</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Gu</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Lu</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Zhu</surname></string-name></person-group>, &#x201C;<article-title>SpA-Former: An effective and lightweight transformer for image shadow removal</article-title>,&#x201D; in <conf-name>2023 Int. Joint Conf. Neural Netw. (IJCNN)</conf-name>, <publisher-name>IEEE</publisher-name>, <year>2023</year>, pp. <fpage>1</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Jiang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Jiang</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Luo</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Gu</surname></string-name></person-group>, &#x201C;<article-title>Star: A structure-aware lightweight transformer for real-time image enhancement</article-title>,&#x201D; in <conf-name>Proc. IEEE/CVF Int. Conf. Comput. Vis.</conf-name>, <year>2021</year>, pp. <fpage>4106</fpage>&#x2013;<lpage>4115</lpage>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>Q.</given-names> <surname>Zhao</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Tang</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Gu</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Zhu</surname></string-name></person-group>, &#x201C;<article-title>Enlighten-anything: When segment anything model meets low-light image enhancement</article-title>,&#x201D; <year>2023</year>, <comment><italic>arXiv:2306.10286</italic></comment>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Wang</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Liu</surname></string-name></person-group>, &#x201C;<article-title>Sparse gradient regularized deep retinex network for robust low-light image enhancement</article-title>,&#x201D; <source>IEEE Trans. Image Process.</source>, vol. <volume>30</volume>, pp. <fpage>2072</fpage>&#x2013;<lpage>2086</lpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.1109/TIP.2021.3050850</pub-id>; <pub-id pub-id-type="pmid">33460379</pub-id></mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Ma</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Liu</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>Beyond brightening low-light images</article-title>,&#x201D; <source>Int. J. Comput. Vis.</source>, vol. <volume>129</volume>, no. <issue>4</issue>, pp. <fpage>1013</fpage>&#x2013;<lpage>1037</lpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.1007/s11263-020-01407-x</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>T.</given-names> <surname>Ma</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Fu</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Zhang</surname></string-name>, and <string-name><given-names>C.</given-names> <surname>Shang</surname></string-name></person-group>, &#x201C;<article-title>RF-Net: Unsupervised low-light image enhancement based on Retinex and exposure fusion</article-title>,&#x201D; <source>Comput. Mater. Contin.</source>, vol. <volume>77</volume>, no. <issue>1</issue>, pp. <fpage>1103</fpage>&#x2013;<lpage>1122</lpage>, <year>2023</year>. doi: <pub-id pub-id-type="doi">10.32604/cmc.2023.042416</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Guo</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Zero-reference deep curve estimation for low-light image enhancement</article-title>,&#x201D; in <conf-name>Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit.</conf-name>, <year>2020</year>, pp. <fpage>1780</fpage>&#x2013;<lpage>1789</lpage>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>R.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Ma</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Fan</surname></string-name>, and <string-name><given-names>Z.</given-names> <surname>Luo</surname></string-name></person-group>, &#x201C;<article-title>Retinex-inspired unrolling with cooperative prior architecture search for low-light image enhancement</article-title>,&#x201D; in <conf-name>Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit.</conf-name>, <year>2021</year>, pp. <fpage>10561</fpage>&#x2013;<lpage>10570</lpage>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>L.</given-names> <surname>Ma</surname></string-name>, <string-name><given-names>T.</given-names> <surname>Ma</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Liu</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Fan</surname></string-name>, and <string-name><given-names>Z.</given-names> <surname>Luo</surname></string-name></person-group>, &#x201C;<article-title>Toward fast, flexible, and robust low-light image enhancement</article-title>,&#x201D; in <conf-name>Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit.</conf-name>, <year>2022</year>, pp. <fpage>5637</fpage>&#x2013;<lpage>5646</lpage>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Fu</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Tu</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Ding</surname></string-name> and <string-name><given-names>K. -K.</given-names> <surname>Ma</surname></string-name></person-group>, &#x201C;<article-title>Learning a simple low-light image enhancer from paired low-light instances</article-title>,&#x201D; in <conf-name>Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit.</conf-name>, <year>2023</year>, pp. <fpage>22252</fpage>&#x2013;<lpage>22261</lpage>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S. M.</given-names> <surname>Pizer</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Adaptive histogram equalization and its variations</article-title>,&#x201D; <source>Comput. Vis., Graph. Image Process.</source>, vol. <volume>39</volume>, no. <issue>3</issue>, pp. <fpage>355</fpage>&#x2013;<lpage>368</lpage>, <year>1987</year>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A. M.</given-names> <surname>Reza</surname></string-name></person-group>, &#x201C;<article-title>Realization of the contrast limited adaptive histogram equalization (CLAHE) for real-time image enhancement</article-title>,&#x201D; <source>J. VLSI Signal Process. Syst. Signal, Image and Video Technol.</source>, vol. <volume>38</volume>, no. <issue>1</issue>, pp. <fpage>35</fpage>&#x2013;<lpage>44</lpage>, <year>2004</year>. doi: <pub-id pub-id-type="doi">10.1023/B:VLSI.0000028532.53893.82</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Fu</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Zeng</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Liao</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Ding</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Paisley</surname></string-name></person-group>, &#x201C;<article-title>A fusion-based enhancing method for weakly illuminated images</article-title>,&#x201D; <source>Signal Process.</source>, vol. <volume>129</volume>, no. <issue>12</issue>, pp. <fpage>82</fpage>&#x2013;<lpage>96</lpage>, <year>2016</year>. doi: <pub-id pub-id-type="doi">10.1016/j.sigpro.2016.05.031</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Fu</surname></string-name>, <string-name><given-names>D.</given-names> <surname>Zeng</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Huang</surname></string-name>, <string-name><given-names>X. -P.</given-names> <surname>Zhang</surname></string-name>, and <string-name><given-names>X.</given-names> <surname>Ding</surname></string-name></person-group>, &#x201C;<article-title>A weighted variational model for simultaneous reflectance and illumination estimation</article-title>,&#x201D; in <conf-name>Proc. IEEE Conf. Comput. Vis. Pattern Recogni.</conf-name>, <year>2016</year>, pp. <fpage>2782</fpage>&#x2013;<lpage>2790</lpage>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Guo</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Li</surname></string-name>, and <string-name><given-names>H.</given-names> <surname>Ling</surname></string-name></person-group>, &#x201C;<article-title>LIME: Low-light image enhancement via illumination map estimation</article-title>,&#x201D; <source>IEEE Trans. Image Process.</source>, vol. <volume>26</volume>, no. <issue>2</issue>, pp. <fpage>982</fpage>&#x2013;<lpage>993</lpage>, <year>2016</year>. doi: <pub-id pub-id-type="doi">10.1109/TIP.2016.2639450</pub-id>; <pub-id pub-id-type="pmid">28113318</pub-id></mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Z. -U.</given-names> <surname>Rahman</surname></string-name>, <string-name><given-names>D. J.</given-names> <surname>Jobson</surname></string-name>, and <string-name><given-names>G. A.</given-names> <surname>Woodell</surname></string-name></person-group>, &#x201C;<article-title>Retinex processing for automatic image enhancement</article-title>,&#x201D; <source>J. Electron. Imaging</source>, vol. <volume>13</volume>, no. <issue>1</issue>, pp. <fpage>100</fpage>&#x2013;<lpage>110</lpage>, <year>2004</year>. doi: <pub-id pub-id-type="doi">10.1117/1.1636183</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Zheng</surname></string-name>, <string-name><given-names>H. -M.</given-names> <surname>Hu</surname></string-name>, and <string-name><given-names>B.</given-names> <surname>Li</surname></string-name></person-group>, &#x201C;<article-title>Naturalness preserved enhancement algorithm for non-uniform illumination images</article-title>,&#x201D; <source>IEEE Trans. Image Process.</source>, vol. <volume>22</volume>, no. <issue>9</issue>, pp. <fpage>3538</fpage>&#x2013;<lpage>3548</lpage>, <year>2013</year>. doi: <pub-id pub-id-type="doi">10.1109/TIP.2013.2261309</pub-id>; <pub-id pub-id-type="pmid">23661319</pub-id></mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K. G.</given-names> <surname>Lore</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Akintayo</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Sarkar</surname></string-name></person-group>, &#x201C;<article-title>LLNet: A deep autoencoder approach to natural low-light image enhancement</article-title>,&#x201D; <source>Pattern Recognit.</source>, vol. <volume>61</volume>, no. <issue>6</issue>, pp. <fpage>650</fpage>&#x2013;<lpage>662</lpage>, <year>2017</year>. doi: <pub-id pub-id-type="doi">10.1016/j.patcog.2016.06.008</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Fu</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Hong</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Chen</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>You</surname></string-name></person-group>, &#x201C;<article-title>LE-GAN: Unsupervised low-light image enhancement network using attention module and identity invariant loss</article-title>,&#x201D; <source>Knowl. Based Syst.</source>, vol. <volume>240</volume>, no. <issue>6</issue>, <year>2022, Art. no. 108010</year>. doi: <pub-id pub-id-type="doi">10.1016/j.knosys.2021.108010</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Sharma</surname></string-name> and <string-name><given-names>R. T.</given-names> <surname>Tan</surname></string-name></person-group>, &#x201C;<article-title>Nighttime visibility enhancement by increasing the dynamic range and suppression of light effects</article-title>,&#x201D; in <conf-name>Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit.</conf-name>, <year>2021</year>, pp. <fpage>11977</fpage>&#x2013;<lpage>11986</lpage>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Jiang</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>EnlightenGAN: Deep light enhancement without paired supervision</article-title>,&#x201D; <source>IEEE Trans. Image Process.</source>, vol. <volume>30</volume>, pp. <fpage>2340</fpage>&#x2013;<lpage>2349</lpage>, <year>2021</year>. doi: <pub-id pub-id-type="doi">10.1109/TIP.2021.3051462</pub-id>; <pub-id pub-id-type="pmid">33481709</pub-id></mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Wang</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Understanding convolution for semantic segmentation</article-title>,&#x201D; in <conf-name>2018 IEEE Winter Conf. Appl. Comput. Vis. (WACV)</conf-name>, <publisher-name>IEEE</publisher-name>, <year>2018</year>, pp. <fpage>1451</fpage>&#x2013;<lpage>1460</lpage>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>C.</given-names> <surname>Wei</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Yang</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Liu</surname></string-name></person-group>, &#x201C;<article-title>Deep retinex decomposition for low-light enhancement</article-title>,&#x201D; <year>2018</year>, <comment><italic>arXiv:1808.04560</italic></comment>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Y.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Zhang</surname></string-name>, and <string-name><given-names>X.</given-names> <surname>Guo</surname></string-name></person-group>, &#x201C;<article-title>Kindling the darkness: A practical low-light image enhancer</article-title>,&#x201D; in <conf-name>Proc. 27th ACM Int. Conf. Multimedia</conf-name>, <year>2019</year>, pp. <fpage>1632</fpage>&#x2013;<lpage>1640</lpage>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Vaswani</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Attention is all you need</article-title>,&#x201D; <source>Adv. Neural Inf. Process Syst.</source>, vol. <volume>30</volume>, <year>2017</year>, pp. <fpage>5998</fpage>&#x2013;<lpage>6008</lpage>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Xu</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>C. -W.</given-names> <surname>Fu</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Jia</surname></string-name></person-group>, &#x201C;<article-title>SNR-aware low-light image enhancement</article-title>,&#x201D; in <conf-name>Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit.</conf-name>, <year>2022</year>, pp. <fpage>17714</fpage>&#x2013;<lpage>17724</lpage>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>M. M.</given-names> <surname>Rahman Siddiquee</surname></string-name>, <string-name><given-names>N.</given-names> <surname>Tajbakhsh</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Liang</surname></string-name></person-group>, &#x201C;<article-title>UNet&#x002B;&#x002B;: A nested u-net architecture for medical image segmentation</article-title>,&#x201D; in <conf-name>Deep Lear. Med. Image Anal. Multimodal Learn. Clinical Decis. Support</conf-name>. <publisher-loc>Granada, Spain</publisher-loc>, <publisher-name>Springer</publisher-name>, <year>Sep. 20, 2018</year>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>W.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Weng</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Zhang</surname></string-name>, <string-name><given-names>X.</given-names> <surname>Wang</surname></string-name>, <string-name><given-names>W.</given-names> <surname>Yang</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Jiang</surname></string-name></person-group>, &#x201C;<article-title>URetinex-Net: Retinex-based deep unfolding network for low-light image enhancement</article-title>,&#x201D; in <conf-name>Proc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit.</conf-name>, <year>2022</year>, pp. <fpage>5901</fpage>&#x2013;<lpage>5910</lpage>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Bai</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Yin</surname></string-name>, and <string-name><given-names>Q.</given-names> <surname>He</surname></string-name></person-group>, &#x201C;<article-title>Retinexmamba: Retinex-based mamba for low-light image enhancement</article-title>,&#x201D; <year>2024</year>, <comment><italic>arXiv:2405.03349</italic></comment>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>S.</given-names> <surname>Yang</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Ding</surname></string-name>, <string-name><given-names>Y.</given-names> <surname>Wu</surname></string-name>, <string-name><given-names>Z.</given-names> <surname>Li</surname></string-name>, and <string-name><given-names>J.</given-names> <surname>Zhang</surname></string-name></person-group>, &#x201C;<article-title>Implicit neural representation for cooperative low-light image enhancement</article-title>,&#x201D; in <conf-name>Proc. IEEE/CVF Int. Conf. Comput. Vis.</conf-name>, <year>2023</year>, pp. <fpage>12918</fpage>&#x2013;<lpage>12927</lpage>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>Z.</given-names> <surname>Liang</surname></string-name>, <string-name><given-names>C.</given-names> <surname>Li</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Zhou</surname></string-name>, <string-name><given-names>R.</given-names> <surname>Feng</surname></string-name>, and <string-name><given-names>C. C.</given-names> <surname>Loy</surname></string-name></person-group>, &#x201C;<article-title>Iterative prompt learning for unsupervised backlit image enhancement</article-title>,&#x201D; in <conf-name>Proc. IEEE/CVF Int. Conf. Comput. Vis.</conf-name>, <year>2023</year>, pp. <fpage>8094</fpage>&#x2013;<lpage>8103</lpage>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>X.</given-names> <surname>Zhang</surname></string-name> <etal>et al.</etal></person-group>, &#x201C;<article-title>Enlighten-your-voice: When multimodal meets zero-shot low-light image enhancement</article-title>,&#x201D; <year>2023</year>, <comment><italic>arXiv:2312.10109</italic></comment>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>R. Al</given-names> <surname>Sobbahi</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Tekli</surname></string-name></person-group>, &#x201C;<article-title>Comparing deep learning models for low-light natural scene image enhancement and their impact on object detection and classification: Overview, empirical evaluation, and challenges</article-title>,&#x201D; <source>Signal Process.: Image Commun.</source>, vol. <volume>109</volume>, no. <issue>12</issue>, <year>2022, Art. no. 116848</year>. doi: <pub-id pub-id-type="doi">10.1016/j.image.2022.116848</pub-id>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Ang</surname></string-name>, <string-name><given-names>W. T.</given-names> <surname>Lim</surname></string-name>, <string-name><given-names>Y. P.</given-names> <surname>Loh</surname></string-name>, and <string-name><given-names>S.</given-names> <surname>Ong</surname></string-name></person-group>, &#x201C;<article-title>Noise-aware zero-reference low-light image enhancement for object detection</article-title>,&#x201D; in <conf-name>2022 Int. Symp. Intell. Signal Process. Commun. Syst. (ISPACS)</conf-name>, <publisher-name>IEEE</publisher-name>, <year>2022</year>, pp. <fpage>1</fpage>&#x2013;<lpage>4</lpage>.</mixed-citation></ref>
</ref-list>
</back></article>