<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">82718</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2026.082718</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Confidence-Regulated Heart Murmur Classification via Joint Representation Learning and Decision Optimization</article-title>
<alt-title alt-title-type="left-running-head">Confidence-Regulated Heart Murmur Classification via Joint Representation Learning and Decision Optimization</alt-title>
<alt-title alt-title-type="right-running-head">Confidence-Regulated Heart Murmur Classification via Joint Representation Learning and Decision Optimization</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author"><contrib-id contrib-id-type="orcid">https://orcid.org/0009-0001-4129-5832</contrib-id>
<name name-style="western"><surname>Chang</surname><given-names>HyeSun</given-names></name></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes"><contrib-id contrib-id-type="orcid">https://orcid.org/0000-0002-4251-7177</contrib-id>
<name name-style="western"><surname>Lee</surname><given-names>Sangjun</given-names></name><email>sangjun@ssu.ac.kr</email></contrib>
<aff id="aff-1">
<institution>Department of AI/SW Convergence, Soongsil University</institution>, <addr-line>Seoul</addr-line>, <country>Republic of Korea</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Sangjun Lee. Email: <email>sangjun@ssu.ac.kr</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>15</day><month>06</month><year>2026</year>
</pub-date>
<volume>88</volume>
<issue>2</issue>
<elocation-id>79</elocation-id>
<history>
<date date-type="received">
<day>21</day>
<month>03</month>
<year>2026</year>
</date>
<date date-type="accepted">
<day>14</day>
<month>05</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_82718.pdf"></self-uri>
<abstract>
<p>Accurate identification of heart murmurs from auscultation recordings is essential for early cardiovascular screening and diagnosis. While deep learning offers strong potential for automated heart murmur classification, existing models often exhibit overconfident, incorrect predictions and limited generalization due to dataset bias and class imbalance. To address these challenges, this study proposes a two-stage confidence-regulated learning framework that jointly optimizes feature representation and decision reliability. Rather than focusing solely on improving classification performance, this work emphasizes enhancing prediction reliability through confidence-aware decision-making. The proposed framework integrates supervised contrastive learning (SCL) to strengthen the discriminative structure of feature embeddings and reward-based optimization (RBO) to regulate prediction confidence under uncertainty. In this framework, a convolutional neural network-based encoder first extracts acoustic representations, and a long short-term memory-based classifier refines the learned embeddings before final prediction. SCL improves intra-class compactness and inter-class separability, while the proposed confidence-regulated mechanism enables the model to adaptively accept or defer predictions based on a dynamically adjusted threshold. This approach allows the model to balance prediction accuracy and decision reliability by reducing overconfident errors in uncertain cases. The proposed method is evaluated on the PhysioNet 2022 heart murmur dataset. Experimental results show that the proposed framework improves the validation Score from 0.8064 to 0.8233, where the Score is defined as the mean of sensitivity and specificity. More importantly, these results demonstrate improved reliability and more balanced decision-making under uncertainty, beyond a marginal increase in aggregate performance. These findings demonstrate that jointly optimizing representation learning and confidence-regulated decision-making provides an effective and clinically relevant approach for robust heart murmur classification.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Heart murmur classification</kwd>
<kwd>phonocardiogram</kwd>
<kwd>supervised contrastive learning</kwd>
<kwd>confidence-aware decision-making</kwd>
<kwd>reinforcement learning</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Korea government (MSIT)</funding-source>
<award-id>IITP-2026-RS-2022-00156360</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Cardiovascular diseases (CVDs) remain one of the leading causes of mortality worldwide, making early detection essential for timely intervention and improved clinical outcomes [<xref ref-type="bibr" rid="ref-1">1</xref>]. Auscultation has long served as a fundamental and non-invasive tool for cardiac assessment; however, its diagnostic accuracy is often influenced by clinician expertise, environmental noise, and the inherent variability of heart sounds [<xref ref-type="bibr" rid="ref-2">2</xref>]. These factors can lead to an inconsistent interpretation of heart murmurs in clinical practice. With the increasing adoption of digital stethoscopes and computer-aided diagnostic technologies, automated analysis of phonocardiogram recordings has emerged as a promising approach to improve the objectivity and accessibility of cardiac screening.</p>
<p>Recent advances in deep learning have significantly enhanced the performance of automated heart sound analysis. By learning representations from time-frequency inputs such as spectrograms, deep learning models can capture both spectral characteristics and temporal dynamics of cardiac cycles [<xref ref-type="bibr" rid="ref-3">3</xref>]. These developments have enabled substantial improvements in heart murmur classification accuracy. Nevertheless, achieving clinically reliable decision support remains challenging. In real-world scenarios, diagnostic models must not only accurately distinguish between murmur categories but also remain robust to data heterogeneity and provide reliable predictions when confronted with uncertain or ambiguous inputs.</p>
<p>Despite their promising performance, existing artificial intelligence-based heart sound classification models exhibit important limitations. A key concern is that such models often produce highly confident yet incorrect predictions, which can be detrimental in safety-critical medical applications. Furthermore, model performance frequently degrades when evaluated on data collected from different populations, devices, or recording conditions, reflecting the effects of dataset bias and limited generalization [<xref ref-type="bibr" rid="ref-4">4</xref>,<xref ref-type="bibr" rid="ref-5">5</xref>]. These challenges indicate that improving classification accuracy alone is insufficient; instead, there is a need for learning frameworks that jointly address feature representation quality and decision reliability.</p>
<p>To address these challenges, this study proposes a learning framework that integrates discriminative representation learning with confidence-aware decision regulation. Rather than focusing solely on improving classification accuracy, the proposed approach emphasizes reliable decision behavior under uncertainty. Specifically, supervised contrastive learning is employed to structure the embedding space and enhance class separability, thereby improving the robustness of learned acoustic representations [<xref ref-type="bibr" rid="ref-6">6</xref>]. In addition, a reward-driven decision mechanism is introduced to regulate predictive confidence and enable adaptive deferral of uncertain cases through an expert query strategy. By incorporating decision cost into the optimization process, the framework encourages more reliable classification behavior when model uncertainty is high [<xref ref-type="bibr" rid="ref-7">7</xref>].</p>
<p>The primary contribution of this study lies in the development of a confidence-regulated learning framework that jointly combines supervised contrastive representation learning with reward-based decision optimization for heart murmur classification. Unlike conventional approaches that focus primarily on improving classification accuracy or feature extraction alone, the proposed method explicitly connects representation refinement and confidence-aware decision regulation within a unified framework. Through this integrated design, the framework aims to improve not only classification performance but also the reliability and clinical applicability of automated heart murmur classification systems.</p>
<p>To provide a clear assessment of the proposed framework, this study focuses on controlled comparisons within a unified experimental setting, thereby isolating the contributions of representation learning and confidence-regulated decision-making.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>Research on automated heart murmur classification has advanced significantly with the development of deep learning techniques, aiming to overcome the subjectivity and variability inherent in traditional auscultation [<xref ref-type="bibr" rid="ref-2">2</xref>,<xref ref-type="bibr" rid="ref-8">8</xref>]. Deep learning-based approaches analyze phonocardiogram (PCG) recordings to automatically extract discriminative features and improve diagnostic accuracy and consistency [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-9">9</xref>&#x2013;<xref ref-type="bibr" rid="ref-11">11</xref>]. This section reviews three key areas relevant to this study: deep learning for PCG analysis, contrastive learning for feature enhancement, and reward-based approaches for confidence-aware decision regulation.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Deep Learning for PCG Analysis and Its Limitations</title>
<p>Deep learning has significantly improved heart murmur classification by enabling automatic feature extraction from heart sound signals [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-9">9</xref>&#x2013;<xref ref-type="bibr" rid="ref-11">11</xref>]. Earlier approaches relied on handcrafted features such as mel-frequency cepstral coefficients (MFCCs) and wavelet transformations. These traditional methods often struggled to generalize across datasets and required domain expertise [<xref ref-type="bibr" rid="ref-12">12</xref>&#x2013;<xref ref-type="bibr" rid="ref-14">14</xref>].</p>
<p>More recent methods leverage Convolutional Neural Networks (CNNs) to learn discriminative spectral representations from PCG spectrograms, while Long Short-Term Memory (LSTM) networks and related architectures capture temporal dependencies in cardiac cycles [<xref ref-type="bibr" rid="ref-3">3</xref>,<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-15">15</xref>]. Recent studies have further explored advanced murmur-classification architectures, including transformer-based and multiscale designs. These models jointly exploit frequency-domain and time-domain characteristics to improve classification performance. However, differences in task formulation and evaluation settings make direct comparisons across studies difficult.</p>
<p>Despite these advancements, several limitations remain. Deep learning models often produce overconfident yet incorrect predictions. Such overconfidence poses a significant risk in clinical applications. These models can also exhibit degraded performance when applied to data from different populations, devices, or recording environments, reflecting the effects of dataset bias and limited generalization [<xref ref-type="bibr" rid="ref-5">5</xref>,<xref ref-type="bibr" rid="ref-16">16</xref>]. The lack of mechanisms to account for prediction uncertainty further limits their reliability in real-world deployment. These challenges highlight the need for approaches that go beyond accuracy and explicitly address decision reliability.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Contrastive Learning for Feature Enhancement</title>
<p>Contrastive learning has emerged as an effective technique for improving representation quality in deep learning models. By encouraging semantically similar samples to be embedded closer together while pushing dissimilar samples apart, contrastive learning structures the feature space and enhances discriminability and generalization [<xref ref-type="bibr" rid="ref-6">6</xref>,<xref ref-type="bibr" rid="ref-17">17</xref>].</p>
<p>Supervised contrastive learning (SCL), which incorporates label information into this process, has demonstrated effectiveness in medical applications, including cardiac sound analysis [<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-18">18</xref>]. By improving class separability in the embedding space, SCL facilitates the detection of subtle acoustic variations in heart sounds, thereby improving classification performance.</p>
<p>However, enhancing feature representations alone does not fully address the problem of unreliable predictions. Models that rely solely on improved feature separability may still exhibit overconfidence in uncertain cases, leading to incorrect yet highly confident outputs in safety-critical settings [<xref ref-type="bibr" rid="ref-10">10</xref>,<xref ref-type="bibr" rid="ref-19">19</xref>]. This limitation motivates the need for complementary mechanisms that regulate decision behavior in addition to improving representation quality.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Reward-Based Approaches for Confidence-Aware Decision Regulation</title>
<p>Beyond representation learning, recent studies have explored decision-level strategies that incorporate feedback signals to improve reliability in medical prediction tasks [<xref ref-type="bibr" rid="ref-7">7</xref>,<xref ref-type="bibr" rid="ref-20">20</xref>]. These approaches introduce adaptive mechanisms that adjust model behavior based on prediction outcomes, enabling more flexible and context-aware decision-making.</p>
<p>In particular, reward-based optimization provides a framework for incorporating decision cost and prediction confidence into the learning process. By associating different outcomes with corresponding utilities, such approaches enable models to balance classification accuracy with the cost of incorrect or uncertain predictions. This perspective is especially relevant in medical applications, where the consequences of incorrect decisions can be significant.</p>
<p>Despite these developments, most existing heart murmur classification models primarily focus on improving classification accuracy and do not explicitly incorporate mechanisms for confidence-aware decision regulation. To the best of our knowledge, reward-based strategies have not been effectively applied to static heart murmur classification tasks to enable adaptive deferral of uncertain predictions. This gap motivates the proposed approach, which integrates discriminative feature learning with confidence-regulated decision optimization in a unified framework.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Methodology</title>
<p>This study proposes a two-stage framework for heart murmur classification that jointly optimizes feature representation and decision reliability. The framework combines supervised contrastive learning (SCL) for structured representation learning with a reward-based optimization (RBO) mechanism for confidence-regulated decision-making.</p>
<sec id="s3_1">
<label>3.1</label>
<title>Dataset</title>
<p>This study utilizes the CirCor DigiScope heart sound dataset from the PhysioNet Challenge 2022 [<xref ref-type="bibr" rid="ref-21">21</xref>,<xref ref-type="bibr" rid="ref-22">22</xref>]. The dataset contains 3163 phonocardiogram (PCG) recordings from 963 pediatric patients across four auscultation sites. The recordings are labeled as Murmur Absent (73.8%), Murmur Present (19.0%), and Unknown (7.2%). Recordings range from 5 to 65 s in duration and were originally sampled at 4000 Hz. The dataset captures substantial physiological variability across patients, auscultation locations, and recording conditions. Murmur annotations are provided by experienced cardiac physiologists based on a combination of auditory assessment and visual inspection of phonocardiogram signals. This annotation process reflects real-world clinical practice, where both acoustic patterns and waveform characteristics are considered during diagnosis. In this study, the task is formulated as a binary classification problem by excluding the Unknown class to reduce label ambiguity. To address class imbalance, a class-weighted loss is applied.</p>
<p>A strict patient-wise split is used. The training set contains 9495 segments from 611 patients (79.7% Absent, 20.3% Present), and the validation set contains 4012 segments from 263 patients (81.0% Absent, 19.0% Present). This patient-wise split prevents data leakage and ensures that the model is evaluated on previously unseen individuals, which is critical for reliable clinical deployment.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Preprocessing</title>
<p>All recordings are resampled to 16 kHz and normalized to the range [&#x2212;1, 1] to ensure a consistent preprocessing pipeline and adequate temporal resolution. Although murmur-related frequency components are primarily below 900 Hz, a higher sampling rate improves waveform representation and supports stable time&#x2013;frequency feature extraction using Mel filter bank representations. This choice preserves diagnostically relevant information while maintaining computational efficiency.</p>
<p>Each recording is segmented into non-overlapping 5-s clips. Segments shorter than 3 s are discarded, while longer residual segments are repeated and truncated. This segmentation strategy is designed to balance temporal coverage and data efficiency. Fixed-length segments allow stable batch processing while preserving sufficient cardiac-cycle information for reliable murmur detection. In particular, each segment typically contains multiple cardiac cycles, allowing the model to learn robust acoustic representations from variable and non-stationary heart sound patterns within a standardized input length. Discarding very short segments prevents the introduction of incomplete or noisy acoustic patterns that may degrade model performance.</p>
<p>Following segmentation, Mel filter bank (Fbank) features are extracted from each audio clip to obtain a time&#x2013;frequency representation of heart sounds. These features are computed using the Kaldi implementation, in which each waveform is divided into overlapping frames using a Hanning window with a 10 ms frame shift. A set of 64 Mel-scaled triangular filters is applied to estimate the energy distribution across frequency bands.</p>
<p>Unlike conventional pipelines that rely on Short-Time Fourier Transform (STFT)-based spectrograms, this approach directly estimates Mel-scale energy features from the waveform. This results in a more compact representation while retaining perceptually relevant frequency information. Prior studies have shown that Mel filter bank representations are robust to noise and effective at capturing subtle acoustic variations in heart sounds [<xref ref-type="bibr" rid="ref-23">23</xref>]. This robustness is particularly important in real-world auscultation scenarios, where recordings may be affected by environmental noise, sensor variability, and patient movement. Overall, the preprocessing pipeline is designed to standardize input representations while preserving clinically relevant acoustic characteristics. This design aims to improve both model stability and generalization performance. An overview of the preprocessing pipeline is illustrated in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. Representative examples of the resulting Fbank features are shown in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Preprocessing pipeline for heart sound recordings. Raw audio signals are normalized and resampled to 16,000 Hz, segmented into fixed-length clips, and transformed into Mel filter bank (Fbank) representations. Segments shorter than 3 s are discarded, while residual segments are repeated and truncated to maintain a consistent input length.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_82718-fig-1.tif"/>
</fig><fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Examples of Mel filter bank representations of heart sounds. (<bold>a</bold>) Murmur Absent sample showing regular cardiac cycles; (<bold>b</bold>) Murmur Present sample exhibiting irregular spectral patterns and additional acoustic components associated with pathological murmurs.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_82718-fig-2.tif"/>
</fig>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Overall Framework</title>
<p>This study presents a structured two-stage training framework for heart murmur classification, comprising a representation learning stage based on supervised contrastive learning (Stage 1) and a confidence-regulated decision-making stage using reward-based optimization (Stage 2). The framework is designed not only to extract discriminative acoustic features from heart sound recordings but also to regulate prediction confidence under uncertainty, which is essential for reliable clinical deployment.</p>
<p>As illustrated in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>, the pipeline begins with preprocessing steps that convert variable-length recordings into standardized time&#x2013;frequency representations. These inputs are then used to learn a structured embedding space with improved intra-class compactness and inter-class separability. In the second stage, the model learns to regulate its own predictions through a confidence-based decision mechanism, enabling it to either make autonomous predictions or defer uncertain cases.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Overall architecture of the proposed framework. The model first learns discriminative feature embeddings using supervised contrastive learning (Stage 1), followed by confidence-regulated classification with reward-based optimization (Stage 2), including an auxiliary value network for decision evaluation.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_82718-fig-3.tif"/>
</fig>
<p>The training follows a sequential design. Stage 1 focuses on constructing a well-separated feature space, while Stage 2 learns decision behavior conditioned on confidence. This separation allows the model to first establish stable representations before introducing decision-level optimization, improving both training stability and interpretability.</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Stage 1: Supervised Contrastive Representation Learning</title>
<p>The first stage of training focuses on learning robust and discriminative feature representations through supervised contrastive learning (SCL). Given an input segment <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msub><mml:mi>X</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>F</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>, where <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>T</mml:mi></mml:math></inline-formula> and <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>F</mml:mi></mml:math></inline-formula> denote the time and frequency dimensions of the Mel filter bank representation, two augmented views <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msubsup><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msubsup><mml:mrow><mml:mover><mml:mi>X</mml:mi><mml:mo stretchy="false">&#x007E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>2</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> are generated using time and frequency masking as proposed in SpecAugment [<xref ref-type="bibr" rid="ref-24">24</xref>]. These augmentations randomly suppress limited regions along the temporal and spectral axes, simulating variability in recording conditions. <xref ref-type="fig" rid="fig-4">Fig. 4</xref> illustrates a comparison between the original Mel filter bank and its SpecAugment-transformed version. This strategy encourages the model to focus on invariant, class-relevant acoustic patterns while improving robustness to noise and localized distortions without systematically eliminating diagnostically relevant murmur information.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Comparison between the original Mel filter bank and the SpecAugment-transformed representation of the same heart sound segment, illustrating localized time and frequency masking applied to partial regions of the input.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_82718-fig-4.tif"/>
</fig>
<p>The augmented views are processed through a shared feature extractor <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, implemented using CNN6 from the Pretrained Audio Neural Networks (PANNs) framework [<xref ref-type="bibr" rid="ref-25">25</xref>]. CNN6 consists of four convolutional blocks, each comprising a 5 <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mo>&#x00D7;</mml:mo></mml:math></inline-formula> 5 convolution, batch normalization, ReLU activation, and pooling operations, enabling hierarchical extraction of time&#x2013;frequency features. A layer-wise summary of the architecture is provided in <xref ref-type="table" rid="table-1">Table 1</xref>. After the final convolutional block, the feature map is aggregated by temporal mean pooling followed by frequency-domain max and mean pooling, and the resulting representations are combined to form the final embedding vector. The network is initialized with pretrained weights from AudioSet, a large-scale audio classification dataset containing over 5000 h of labeled audio across 527 sound classes [<xref ref-type="bibr" rid="ref-26">26</xref>]. Rather than being used as a fixed feature extractor, the pretrained CNN6 encoder is fully fine-tuned during Stage 1 on the target PCG dataset. This design allows the model to benefit from a stable acoustic initialization while adapting its representations to the temporal and spectral characteristics of heart sound signals. Although AudioSet and PCG recordings differ in domain, the pretrained initialization provides general acoustic priors that can be refined through task-specific training.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Layer-wise configuration of the CNN6.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Stage</th>
<th>Layer Type</th>
<th>Configuration</th>
<th>Channels</th>
</tr>
</thead>
<tbody>
<tr>
<td>Input</td>
<td>Mel filter bank</td>
<td>Single-channel input</td>
<td>1</td>
</tr>
<tr>
<td>Block 1</td>
<td>ConvBlock 5 &#x00D7; 5</td>
<td>Conv <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mn>5</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>5</mml:mn></mml:math></inline-formula> &#x002B; BN &#x002B; ReLU &#x002B; AvgPool</td>
<td>64</td>
</tr>
<tr>
<td>Block 2</td>
<td>ConvBlock 5 &#x00D7; 5</td>
<td>Conv <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mn>5</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>5</mml:mn></mml:math></inline-formula> &#x002B; BN &#x002B; ReLU &#x002B; AvgPool</td>
<td>128</td>
</tr>
<tr>
<td>Block 3</td>
<td>ConvBlock 5 &#x00D7; 5</td>
<td>Conv <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mn>5</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>5</mml:mn></mml:math></inline-formula> &#x002B; BN &#x002B; ReLU &#x002B; AvgPool</td>
<td>256</td>
</tr>
<tr>
<td>Block 4</td>
<td>ConvBlock 5 &#x00D7; 5</td>
<td>Conv <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mn>5</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>5</mml:mn></mml:math></inline-formula> &#x002B; BN &#x002B; ReLU &#x002B; AvgPool</td>
<td>512</td>
</tr>
<tr>
<td>Pooling</td>
<td>Aggregation</td>
<td>Mean (time) &#x002B; Max/Mean (freq)</td>
<td>512</td>
</tr>
<tr>
<td>Output</td>
<td>Embedding</td>
<td>Final feature vector</td>
<td>512</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The feature extractor outputs high-dimensional embeddings <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msubsup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msubsup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>2</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>, which are passed through a projection head <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mi>g</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>, modeled as a two-layer multilayer perceptron (MLP) with ReLU activation. The projection head maps the encoder output to a 128-dimensional projected representation for contrastive learning, yielding projected representations <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:msubsup><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msubsup><mml:mi>P</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>2</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>. Supervised contrastive learning is applied to these projected representations, while the encoder output <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mi>Z</mml:mi></mml:math></inline-formula> is retained for downstream classification. Although the contrastive objective is optimized on <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mi>P</mml:mi></mml:math></inline-formula>, the encoder and projection head are trained jointly during Stage 1. Consequently, the discriminative structure encouraged in the projected representations is propagated back to the encoder through gradient updates, improving the quality of the encoder representation <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mi>Z</mml:mi></mml:math></inline-formula>. This design enables contrastive optimization through the projection head while preserving <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mi>Z</mml:mi></mml:math></inline-formula> as the primary representation for subsequent classification.</p>
<p>Supervised contrastive learning is applied to structure the embedding space by leveraging label information to define semantically meaningful positive pairs [<xref ref-type="bibr" rid="ref-27">27</xref>]. This formulation encourages samples from the same class to cluster together while separating samples from different classes. To address class imbalance in the dataset, class weights are incorporated into the supervised contrastive loss, increasing the contribution of underrepresented classes during training.</p>
<p>Inspired by the contrastive repulsion mechanism proposed in [<xref ref-type="bibr" rid="ref-28">28</xref>], an additional repulsion term <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>p</mml:mi><mml:mi>u</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is incorporated to explicitly penalize high similarity between samples from different classes. The repulsion loss is defined as:<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>p</mml:mi><mml:mi>u</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>N</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>N</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:munder><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>m</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mi>N</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>i</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denotes the set of negative pairs from different classes. This term further enforces inter-class separation by reducing similarity among dissimilar samples.</p>
<p>The total Stage 1 objective is defined as:<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>Stage1</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>SCL</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mtext>repulsion</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>p</mml:mi><mml:mi>u</mml:mi><mml:mi>l</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>The balancing coefficient <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">r</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">p</mml:mi><mml:mi mathvariant="normal">u</mml:mi><mml:mi mathvariant="normal">l</mml:mi><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">i</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">n</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> controls the relative contribution of the repulsion term. The combination of class-aware attraction and contrastive repulsion results in a well-structured embedding space with improved intra-class compactness and inter-class separability.</p>
<p>The overall architecture and training flow for this stage are illustrated in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>. The learned representation serves as the foundation for Stage 2, where classification decisions are further refined through confidence-regulated optimization.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Stage 1 pipeline. Two augmented views are processed by a shared feature extractor (CNN6) and then passed through a projection head. The model is trained using supervised contrastive learning with an additional repulsion loss to improve inter-class separation.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_82718-fig-5.tif"/>
</fig>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Stage 2: Confidence-Regulated Decision Optimization</title>
<p>Once feature representations are refined, Stage 2 focuses on classification and confidence-regulated decision-making. A bidirectional LSTM-based classifier is employed to process the encoder-derived embeddings. In this framework, the classifier operates on learned embeddings rather than on raw PCG waveforms. Accordingly, the LSTM serves as a trainable transformation and classification module built on top of the learned representation. This design provides a flexible classifier head while preserving the structured representations learned in Stage 1. The classifier consists of stacked bidirectional LSTM layers followed by a fully connected (FC) layer. The output logits are then transformed into class probabilities using a Softmax function. A dropout layer with a rate of 0.3 is applied to reduce overfitting and improve generalization. The overall architecture and training process are illustrated in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Stage 2 pipeline. Feature embeddings are passed to an LSTM classifier. Predictions are optimized using cross-entropy and reward-based objectives, while a value network is trained separately using mean squared error.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_82718-fig-6.tif"/>
</fig>
<p>The classifier receives input embeddings constructed from the encoder outputs after Stage 1 training. Since two augmented views are generated for each original input during Stage 1, the encoder produces two embeddings per input, denoted as <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msubsup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> and <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msubsup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>2</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>. Using both embeddings directly in Stage 2 would double the number of training instances. To preserve the original number of inputs, the two embeddings are consolidated into a single embedding for each input before classification. Specifically, 80% of the inputs are formed using the averaged embedding <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msub><mml:mrow><mml:mover><mml:mi>Z</mml:mi><mml:mo stretchy="false">&#x00AF;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, while the remaining 20% use a randomly selected embedding from either <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msubsup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula> or <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msubsup><mml:mi>Z</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mn>2</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msubsup></mml:math></inline-formula>. The 80:20 ratio was fixed throughout all experiments and introduced as a heuristic regularization choice to combine representation stability with modest stochastic variation. This design provides a stable representation for most inputs while retaining limited stochastic variation derived from augmentation. The resulting embeddings are processed by the LSTM to produce prediction logits <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, which are compared with the ground truth <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> to compute the classification loss <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> using the standard cross-entropy objective. To address class imbalance, class weights are incorporated into the cross-entropy loss for more balanced learning across classes. The encoder is frozen during Stage 2 to preserve the structured feature representations learned in Stage 1. This prevents changes to the embedding space and ensures consistent feature distributions during classifier training.</p>
<p>To regulate prediction reliability, a confidence threshold <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> is applied to the classifier&#x2019;s Softmax output. If the maximum predicted probability satisfies <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x2265;</mml:mo><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula>, the prediction is accepted; otherwise, the model defers the decision. This confidence-based gating mechanism reduces unreliable predictions and improves decision trustworthiness during training.</p>
<p>To further refine decision behavior under uncertainty, reward-based optimization (RBO) is applied. The classifier output is treated as a categorical distribution over the prediction classes, and the confidence-regulated decision objective is optimized using a stabilized strategy inspired by Proximal Policy Optimization (PPO) [<xref ref-type="bibr" rid="ref-29">29</xref>]. The probability ratio is defined as
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mrow><mml:mtext>old</mml:mtext></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:mi>&#x03B8;</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mrow><mml:msub><mml:mi>&#x03B8;</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">l</mml:mi><mml:mi mathvariant="normal">d</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>a</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> denote the current and previous probabilities assigned to decision <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:msub><mml:mi>a</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> for input <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:msub><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula>, respectively. The advantage is defined as
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>A</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mi>V</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:msub><mml:mi>R</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> is the reward assigned based on the classification outcome, and <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mi>V</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the value estimated by the auxiliary value network. The value network is implemented as a multilayer perceptron (MLP) that takes the encoder-derived embedding as input and outputs a scalar value estimate. It consists of three fully connected layers with dimensions <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mn>512</mml:mn><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>256</mml:mn><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>128</mml:mn><mml:mo stretchy="false">&#x2192;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>, with ReLU activation functions applied after the first two layers. The final output represents the estimated value associated with the input embedding and is used to compute the advantage term in Stage 2. Since heart murmur classification is formulated as a single-step static decision problem, the learning process does not involve sequential state transitions. As a result, the advantage formulation simplifies by excluding future return estimation, allowing direct optimization based on immediate outcomes.</p>
<p>The reward-based decision optimization objective for the classifier is defined as
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>RBO</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mo movablelimits="true" form="prefix">min</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:msub><mml:mi>A</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mtext>&#x00A0;</mml:mtext><mml:mrow><mml:mtext>clip</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>&#x03B5;</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:mi>&#x03B5;</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mi>e</mml:mi></mml:msub><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>entropy</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mi>&#x03B5;</mml:mi></mml:math></inline-formula> is the clipping parameter, <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mi>e</mml:mi></mml:msub></mml:math></inline-formula> controls the contribution of entropy regularization, and <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">n</mml:mi><mml:mi mathvariant="normal">t</mml:mi><mml:mi mathvariant="normal">r</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">p</mml:mi><mml:mi mathvariant="normal">y</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> is defined as
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>entropy</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="double-struck">E</mml:mi></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>&#x0210B;</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>&#x03C0;</mml:mi><mml:mi>&#x03B8;</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo>&#x2223;</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>]</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>The overall classifier objective in Stage 2 is defined as
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>Stage2</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>CE</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>RBO</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:msub><mml:mi>c</mml:mi><mml:mn>1</mml:mn></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:msub><mml:mi>c</mml:mi><mml:mn>2</mml:mn></mml:msub></mml:math></inline-formula> control the balance between supervised classification and confidence-regulated decision optimization.</p>
<p>The value network is optimized separately using a mean squared error loss,
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x02112;</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>value</mml:mtext></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mtext>MSE</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>V</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mi>R</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>which aligns the predicted value with the observed reward signal. This separation improves optimization stability by decoupling confidence-regulated decision optimization from value estimation.</p>
</sec>
<sec id="s3_6">
<label>3.6</label>
<title>Reward Design and Adaptive Threshold</title>
<p>To enable confidence-regulated decision-making, the RBO framework incorporates a dynamically adjusted confidence threshold and a structured reward design.</p>
<p>The confidence threshold <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> is updated using a decreasing sigmoid function over training epochs:<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>&#x03B4;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>epoch</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>epoch</mml:mtext></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mtext>shift</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>This schedule begins with a high, conservative threshold, encouraging frequent deferral of uncertain predictions. As training progresses, the threshold gradually decreases, allowing the model to make more autonomous decisions as its predictive confidence improves. The parameters <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mi>k</mml:mi></mml:math></inline-formula>, <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mrow><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">h</mml:mi><mml:mi mathvariant="normal">i</mml:mi><mml:mi mathvariant="normal">f</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:mrow></mml:math></inline-formula>, and <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:msub><mml:mi>&#x03B4;</mml:mi><mml:mrow><mml:mo movablelimits="true" form="prefix">min</mml:mo></mml:mrow></mml:msub></mml:math></inline-formula> control the decay rate, transition point, and minimum threshold value, respectively.</p>
<p>The reward signal is designed to jointly account for prediction correctness, confidence, and class imbalance:<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msub><mml:mi>R</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mi>R</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mrow><mml:mtext>base</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:mn>0.1</mml:mn><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:msub><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:math></inline-formula> denotes the prediction confidence. The base reward is defined according to prediction correctness and confidence level:<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:msubsup><mml:mi>R</mml:mi><mml:mi>t</mml:mi><mml:mrow><mml:mrow><mml:mtext>base</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mn>0.5</mml:mn><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if&#xA0;</mml:mtext></mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mrow><mml:mtext>&#xA0;and&#xA0;</mml:mtext></mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2265;</mml:mo><mml:mi>&#x03B4;</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if&#xA0;</mml:mtext></mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mrow><mml:mtext>&#xA0;and&#xA0;</mml:mtext></mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x003C;</mml:mo><mml:mi>&#x03B4;</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:mn>2.0</mml:mn><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if&#xA0;</mml:mtext></mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2260;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mrow><mml:mtext>&#xA0;and&#xA0;</mml:mtext></mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2265;</mml:mo><mml:mi>&#x03B4;</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if&#xA0;</mml:mtext></mml:mrow><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x2260;</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mrow><mml:mtext>&#xA0;and&#xA0;</mml:mtext></mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>&#x003C;</mml:mo><mml:mi>&#x03B4;</mml:mi></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Here, <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:msub><mml:mi>w</mml:mi><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula> represents the class-dependent weighting factor used to address class imbalance. The reward structure assigns strong penalties to overconfident incorrect predictions, discouraging unreliable decisions, while rewarding correct predictions with an additional incentive when they are made confidently.</p>
<p>The additional term <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mn>0.1</mml:mn><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> acts as a confidence-based regularizer. This term provides a small incentive for lower-confidence predictions, encouraging the model to remain cautious near the decision boundary and mitigating the tendency toward uniformly overconfident outputs.</p>
<p>This reward formulation guides the model to balance prediction accuracy with confidence, resulting in more reliable and calibrated decision-making in clinically uncertain scenarios.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experiment</title>
<p>To evaluate the effectiveness of the proposed two-stage framework, a comparative analysis is conducted between the baseline model and the proposed method integrating supervised contrastive learning (SCL) and reward-based optimization (RBO). The objective is to assess whether structured representation learning and confidence-regulated decision-making provide complementary benefits in improving both classification performance and decision reliability.</p>
<p>Experiments are performed on the validation split constructed using a strict patient-wise partitioning strategy, ensuring that no patient appears in both training and validation sets. This setup reflects realistic clinical deployment, where models must generalize to previously unseen individuals.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Evaluation Metrics</title>
<p>Model performance is evaluated using sensitivity, specificity, and their arithmetic mean, referred to as the Score. All reported metrics are computed at the segment level on the segmented inputs obtained after the patient-wise train-validation split described above, rather than at the patient level. These metrics are widely used in clinical classification tasks, where both detecting pathological cases and avoiding false alarms are equally critical.</p>
<p>Sensitivity measures the proportion of correctly identified positive cases (murmur present), reflecting the model&#x2019;s ability to detect clinically relevant abnormalities. Specificity measures the proportion of correctly identified negative cases (murmur absent), indicating the model&#x2019;s ability to avoid false positives. Given the inherent class imbalance in the dataset, reliance on a single metric can be misleading. Therefore, the Score, defined as the arithmetic mean of sensitivity and specificity, is used as the primary evaluation criterion to provide a balanced assessment of performance. We note that the official George B. Moody PhysioNet Challenge 2022 adopted a different task formulation and evaluation protocol. In contrast to the official challenge setting, this study formulates murmur classification as a binary task and evaluates performance using sensitivity, specificity, and their mean Score. Accordingly, the reported results are intended to support controlled comparison within the proposed framework rather than direct comparison with official challenge rankings.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Computational Setup and Hyperparameters</title>
<p>The implementation is carried out in Python 3.9 using PyTorch 2.6.0, with Torchaudio 2.6.0 and NumPy 1.24.3. All experiments are conducted on a system equipped with dual NVIDIA GeForce RTX 4090 GPUs (48 GB total VRAM), an Intel&#x00AE; Core&#x2122; i9-13900KF CPU, and 64 GB of RAM, running Ubuntu 22.04.2 LTS. A batch size of 8 is used for training.</p>
<p>The training process follows the two-stage framework described in <xref ref-type="sec" rid="s3">Section 3</xref>. In Stage 1, the CNN6 encoder and projection head are fine-tuned using the Adam optimizer with a learning rate of 0.0005. The total loss combines supervised contrastive learning and the repulsion term, with the balancing coefficient <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:msub><mml:mi>&#x03BB;</mml:mi><mml:mrow><mml:mrow><mml:mi mathvariant="normal">r</mml:mi><mml:mi mathvariant="normal">e</mml:mi><mml:mi mathvariant="normal">p</mml:mi><mml:mi mathvariant="normal">u</mml:mi><mml:mi mathvariant="normal">l</mml:mi><mml:mi mathvariant="normal">s</mml:mi><mml:mi mathvariant="normal">i</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mi mathvariant="normal">n</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula> set to 0.3.</p>
<p>The bidirectional LSTM classifier and the value network are trained using the Adam optimizer with a learning rate of 0.0005. A dropout rate of 0.3 is applied to the LSTM classifier to mitigate overfitting. The total classification loss combines the cross-entropy loss and the reward-based optimization objective, with coefficients <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:msub><mml:mi>c</mml:mi><mml:mn>1</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn>0.3</mml:mn></mml:math></inline-formula> and <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:msub><mml:mi>c</mml:mi><mml:mn>2</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn>0.7</mml:mn></mml:math></inline-formula>, respectively, assigning greater weight to the RBO objective.</p>
<p>The clipping parameter <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mi>&#x03B5;</mml:mi></mml:math></inline-formula> for stabilized updates is set to 0.2. The confidence threshold <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mi>&#x03B4;</mml:mi></mml:math></inline-formula> is dynamically adjusted during training using a decreasing sigmoid function, allowing the model to transition from conservative to more autonomous decision-making behavior as training progresses.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Performance Comparison</title>
<p>The performance comparison between the baseline and the proposed method variants is presented in <xref ref-type="table" rid="table-2">Table 2</xref>. The baseline model consists of a CNN6 encoder and an LSTM classifier trained using standard supervised learning.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Comparative performance of baseline and proposed methods on the validation set in terms of sensitivity, specificity, and score.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Methods</th>
<th>Sensitivity</th>
<th>Specificity</th>
<th>Score</th>
</tr>
</thead>
<tbody>
<tr>
<td>Baseline (CNN6 &#x002B; LSTM)</td>
<td>0.6745</td>
<td>0.9382</td>
<td>0.8064</td>
</tr>
<tr>
<td>Baseline &#x002B; SCL (Feature Learning)</td>
<td>0.6509</td>
<td><bold>0.9563</bold></td>
<td>0.8036</td>
</tr>
<tr>
<td>Baseline &#x002B; RBO (Decision Regulation)</td>
<td>0.7021</td>
<td>0.9317</td>
<td>0.8170</td>
</tr>
<tr>
<td><bold>Proposed Method (SCL &#x002B; RBO)</bold></td>
<td><bold>0.7284</bold></td>
<td>0.9182</td>
<td><bold>0.8233</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The PhysioNet 2022 dataset is a widely used benchmark for heart murmur classification. In this study, the evaluation focuses on controlled comparisons between the baseline and proposed variants to isolate the impact of supervised contrastive learning and confidence-regulated decision optimization. Rather than benchmarking existing methods across varying experimental settings, this design enables a clearer assessment of each component&#x2019;s contribution within a consistent framework.</p>
<p>The baseline model achieves a Score of 0.8064, serving as a reference for comparison. Introducing SCL alone improves feature representation by enhancing inter-class separation, yielding the highest specificity of 0.9563. To further analyze this effect, we visualize the embedding space before and after applying SCL using t-SNE, as shown in <xref ref-type="fig" rid="fig-7">Fig. 7</xref>. The baseline embeddings exhibit substantial overlap between classes, indicating limited discriminative structure. In contrast, the SCL-enhanced embeddings demonstrate improved intra-class compactness and clearer inter-class separation, despite inherent acoustic similarity between heart sound patterns, suggesting that SCL effectively structures the feature space for discriminative learning. This behavior is consistent with the observed reduction in false positives. However, the improvement in specificity does not translate into a higher overall Score, suggesting that representation learning alone is insufficient to optimize decision behavior.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>t-SNE visualization of feature embeddings before and after supervised contrastive learning (SCL). (<bold>a</bold>) Before SCL, the embeddings exhibit substantial overlap between murmur-absent and murmur-present samples, indicating limited discriminative structure. (<bold>b</bold>) After SCL, the embeddings show improved intra-class compactness and clearer inter-class separation, despite the inherent acoustic similarity of heart sound signals.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_82718-fig-7.tif"/>
</fig>
<p>In contrast, applying RBO alone significantly increases sensitivity from 0.6745 to 0.7021, thereby improving the detection of murmur cases. This improvement highlights the effectiveness of confidence-regulated decision optimization in encouraging the model to identify positive cases more aggressively. However, this is accompanied by a slight reduction in specificity, indicating a trade-off between sensitivity and false-positive control.</p>
<p>The proposed method, which integrates both SCL and RBO, achieves the highest Score of 0.8233, representing a relative improvement of 2.1% over the baseline. This result demonstrates that feature representation learning and decision regulation provide complementary benefits. SCL enhances the quality of the learned feature space, while RBO guides the model toward more reliable decision-making under uncertainty. The results confirm that jointly optimizing representation and decision behavior leads to a more balanced and clinically reliable classification system. To further visualize the classification performance of the proposed method, <xref ref-type="fig" rid="fig-8">Fig. 8</xref> presents the confusion matrix and ROC curve on the validation set. The confusion matrix shows strong class-wise performance for both murmur-absent and murmur-present cases, while the ROC curve further confirms the model&#x2019;s discriminative capability across decision thresholds.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Segment-level confusion matrix and ROC curve of the proposed method on the validation set. (<bold>a</bold>) Row-normalized confusion matrix. (<bold>b</bold>) ROC curve with an AUC of 0.878.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_82718-fig-8.tif"/>
</fig>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Discussion</title>
<p>The experimental results demonstrate that jointly optimizing feature representation and confidence-regulated decision-making leads to consistent improvements in heart murmur classification performance. The proposed framework not only enhances overall classification performance but also achieves a better balance between sensitivity and specificity, which is critical in clinical diagnostic settings.</p>
<p><bold>Impact of Feature Representation Learning (SCL):</bold> The <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:mtext>Baseline</mml:mtext><mml:mo>+</mml:mo><mml:mtext>SCL</mml:mtext></mml:math></inline-formula> variant achieved the highest specificity among all methods. This indicates that supervised contrastive learning effectively improves inter-class separability in the embedding space by clustering semantically similar samples while pushing dissimilar ones apart. As a result, the model becomes more conservative in its predictions, reducing false positives and improving reliability in identifying murmur-absent cases. This behavior is particularly desirable in clinical environments, where minimizing unnecessary alarms is important to avoid overdiagnosis and additional testing.</p>
<p>However, the improvement in specificity does not directly translate to a higher overall Score. This suggests that while SCL enhances feature quality, it does not explicitly regulate how these features are used during decision-making. In particular, the increased separation can lead to more conservative decision boundaries, making the model less sensitive to ambiguous or borderline murmur cases, which may limit sensitivity. Consequently, the model remains limited in its ability to adapt its predictions under uncertainty, highlighting the need for an additional mechanism that governs decision behavior.</p>
<p>Although supervised contrastive learning improves representation quality within the labeled dataset used in this study, its generalization benefits may be further enhanced through large-scale unlabeled pretraining. In particular, self-supervised objectives such as masked prediction or related representation learning strategies on larger heart sound datasets could provide a stronger initialization for Stage 1 and further improve the effectiveness of the confidence-regulated decision optimization applied in Stage 2. Exploring such pretraining strategies remains an important direction for future work.</p>
<p><bold>Impact of Confidence-Regulated Decision Optimization (RBO):</bold> The <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:mtext>Baseline</mml:mtext><mml:mo>+</mml:mo><mml:mtext>RBO</mml:mtext></mml:math></inline-formula> variant shows a significant increase in sensitivity, indicating improved detection of murmur-present cases. This improvement arises from the confidence-regulated decision mechanism, which dynamically adjusts the acceptance threshold during training. Initially, the model adopts a conservative strategy by deferring uncertain predictions, effectively reducing overconfident errors. As training progresses and the model becomes more competent, the threshold gradually decreases, allowing the model to make more autonomous predictions.</p>
<p>This adaptive process encourages the model to explore decision boundaries more effectively and reduces the tendency to ignore difficult positive cases. As a result, the model becomes more sensitive to subtle murmur patterns, improving its ability to detect clinically relevant abnormalities. However, this gain in sensitivity is accompanied by a slight reduction in specificity, reflecting the inherent trade-off between detecting positive cases and avoiding false positives.</p>
<p><bold>Complementary Effects of SCL and RBO:</bold> The full proposed method, which combines SCL and RBO, achieves the highest overall Score, demonstrating that the two components provide complementary benefits. SCL enhances the structure and separability of the feature space, ensuring that representations are discriminative and robust. RBO, on the other hand, regulates how these features are utilized during prediction, guiding the model to make more reliable decisions under uncertainty. By integrating these two components, the framework effectively addresses both representation-level and decision-level limitations present in conventional deep learning approaches. The improved sensitivity indicates better detection of murmur cases, while the maintained specificity ensures that false positives remain controlled. This balance is essential for real-world deployment, where both missed diagnoses and false alarms carry significant clinical implications. Furthermore, the proposed framework is not constrained to a specific backbone architecture. Incorporating more advanced feature extraction or classification components is expected to yield further performance improvements, which remains a promising direction for future work.</p>
<p>It should also be noted that the encoder is initialized using AudioSet pretraining, which may introduce a domain mismatch with PCG signals. Although the encoder is fully fine-tuned on the target dataset, a direct comparison with random initialization was beyond the scope of this study and remains an important direction for future investigation.</p>
<p>These findings highlight that improving feature representation alone is insufficient for achieving optimal performance. Instead, incorporating a confidence-aware decision-making mechanism is crucial for developing reliable and clinically applicable models.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion</title>
<p>This study proposes a Joint Feature-Decision Learning Framework that integrates supervised contrastive learning (SCL) and reward-based optimization (RBO) to enhance feature discriminability and improve classification reliability under clinical uncertainty. Experimental results demonstrate that the proposed method outperforms the baseline, achieving a Score of 0.8233. The findings confirm the complementary roles of the two components: SCL improves class separability, leading to higher specificity, while RBO regulates prediction confidence, enhancing sensitivity and enabling more reliable decision-making. By jointly optimizing feature representation and decision behavior, the proposed framework provides a balanced and effective approach for heart murmur classification. This work highlights the importance of integrating structured representation learning with confidence-aware decision mechanisms for developing clinically applicable models. Future work will focus on refining the reward formulation with class-aware penalties, exploring uncertainty-aware value estimation, and extending the framework to more complex settings, such as multi-label and open-set classification.</p>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This work was supported by Innovative Human Resource Development for Local Intellectualization Program through the Institute of Information &#x0026; Communications Technology Planning &#x0026; Evaluation (IITP) grant funded by the Korea government (MSIT) (IITP-2026-RS-2022-00156360).</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm their contributions to the paper as follows: HyeSun Chang: Conceptualization, Methodology, Software, Validation, Formal Analysis, Investigation, Data Curation, Writing&#x2014;Original Draft Preparation, Writing&#x2014;Review and Editing, Visualization. Sangjun Lee: Supervision, Writing&#x2014;Review and Editing, Project Administration, Funding Acquisition. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The data that support the findings of this study are publicly available in the PhysioNet repository (<ext-link ext-link-type="uri" xlink:href="https://physionet.org/">https://physionet.org/</ext-link>), specifically the CirCor DigiScope dataset from the PhysioNet Challenge 2022.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>This study utilizes publicly available anonymized data and does not involve direct human or animal subject interaction. Therefore, ethical approval is not required.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<glossary content-type="abbreviations" id="glossary-1">
<title>Abbreviations</title>
<def-list>
<def-item>
<term>PCG</term>
<def>
<p>Phonocardiogram</p>
</def>
</def-item>
<def-item>
<term>SCL</term>
<def>
<p>Supervised Contrastive Learning</p>
</def>
</def-item>
<def-item>
<term>RBO</term>
<def>
<p>Reward-Based Optimization</p>
</def>
</def-item>
<def-item>
<term>CNN</term>
<def>
<p>Convolutional Neural Network</p>
</def>
</def-item>
<def-item>
<term>LSTM</term>
<def>
<p>Long Short-Term Memory</p>
</def>
</def-item>
<def-item>
<term>GAP</term>
<def>
<p>Global Average Pooling</p>
</def>
</def-item>
<def-item>
<term>FC</term>
<def>
<p>Fully Connected</p>
</def>
</def-item>
<def-item>
<term>CE</term>
<def>
<p>Cross-Entropy</p>
</def>
</def-item>
</def-list>
</glossary>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ameen</surname> <given-names>A</given-names></string-name>, <string-name><surname>Fattoh</surname> <given-names>IE</given-names></string-name>, <string-name><surname>Abd El-Hafeez</surname> <given-names>T</given-names></string-name>, <string-name><surname>Ahmed</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Advances in ECG and PCG-based cardiovascular disease classification: a review of deep learning and machine learning methods</article-title>. <source>J Big Data</source>. <year>2024</year>;<volume>11</volume>(<issue>1</issue>):<fpage>159</fpage>. doi:<pub-id pub-id-type="doi">10.1186/s40537-024-01011-7</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Omarov</surname> <given-names>B</given-names></string-name>, <string-name><surname>Tuimebayev</surname> <given-names>A</given-names></string-name>, <string-name><surname>Abdrakhmanov</surname> <given-names>R</given-names></string-name>, <string-name><surname>Eskarayeva</surname> <given-names>B</given-names></string-name>, <string-name><surname>Sultan</surname> <given-names>D</given-names></string-name>, <string-name><surname>Aidarov</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Digital stethoscope for early detection of heart disease on phonocardiography data</article-title>. <source>Int J Adv Comput Sci Appl</source>. <year>2023</year>;<volume>14</volume>(<issue>9</issue>):<fpage>716</fpage>&#x2013;<lpage>24</lpage>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Kamson</surname> <given-names>AP</given-names></string-name>, <string-name><surname>Crecsilla Lewis</surname> <given-names>M</given-names></string-name>, <string-name><surname>Vishnu Sunil</surname> <given-names>BN</given-names></string-name>, <string-name><surname>Jeevannavar</surname> <given-names>SS</given-names></string-name>, <string-name><surname>Sawant</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ghosh</surname> <given-names>PK</given-names></string-name></person-group>. <article-title>E2E multi-scale CNN with LSTM for murmur detection in PCG or noise identification</article-title>. In: <conf-name>2023 International Conference on Electrical, Communication and Computer Engineering (ICECCE)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2023</year>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Goetz</surname> <given-names>L</given-names></string-name>, <string-name><surname>Seedat</surname> <given-names>N</given-names></string-name>, <string-name><surname>Vandersluis</surname> <given-names>R</given-names></string-name>, <string-name><surname>van der Schaar</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Generalization&#x2014;a key challenge for responsible AI in patient-facing clinical applications</article-title>. <source>npj Digit Med</source>. <year>2024</year>;<volume>7</volume>:<fpage>126</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41746-024-01127-3</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Norori</surname> <given-names>N</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Aellen</surname> <given-names>FM</given-names></string-name>, <string-name><surname>Faraci</surname> <given-names>FD</given-names></string-name>, <string-name><surname>Tzovara</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Addressing bias in big data and AI for health care: a call for open science</article-title>. <source>Patterns</source>. <year>2021</year>;<volume>2</volume>(<issue>10</issue>):<fpage>100347</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.patter.2021.100347</pub-id>; <pub-id pub-id-type="pmid">34693373</pub-id></mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>T</given-names></string-name>, <string-name><surname>Kornblith</surname> <given-names>S</given-names></string-name>, <string-name><surname>Norouzi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Hinton</surname> <given-names>G</given-names></string-name></person-group>. <article-title>A simple framework for contrastive learning of visual representations</article-title>. In: <conf-name>Proceedings of the 37th International Conference on Machine Learning, ICML&#x2019;20; 2020 Jul 12&#x2013;18</conf-name>; <publisher-loc>Vienna, Austria</publisher-loc>. p. <fpage>1597</fpage>&#x2013;<lpage>607</lpage>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gayathri</surname> <given-names>R</given-names></string-name>, <string-name><surname>Sangeetha</surname> <given-names>SKB</given-names></string-name>, <string-name><surname>Mathivanan</surname> <given-names>SK</given-names></string-name>, <string-name><surname>Rajadurai</surname> <given-names>H</given-names></string-name>, <string-name><surname>Benjula Anbu Malar</surname> <given-names>MB</given-names></string-name>, <string-name><surname>Mallik</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Enhancing heart disease prediction with reinforcement learning and data augmentation</article-title>. <source>Syst Soft Comput</source>. <year>2024</year>;<volume>6</volume>:<fpage>200129</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.sasc.2024.200129</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chorba</surname> <given-names>JS</given-names></string-name>, <string-name><surname>Shapiro</surname> <given-names>AM</given-names></string-name>, <string-name><surname>Le</surname> <given-names>L</given-names></string-name>, <string-name><surname>Maidens</surname> <given-names>J</given-names></string-name>, <string-name><surname>Prince</surname> <given-names>J</given-names></string-name>, <string-name><surname>Pham</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Deep learning algorithm for automated cardiac murmur detection via a digital stethoscope platform</article-title>. <source>J Am Heart Assoc</source>. <year>2021</year>;<volume>10</volume>:<fpage>e019905</fpage>. doi:<pub-id pub-id-type="doi">10.1101/2020.04.01.20050518</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Alkhodari</surname> <given-names>M</given-names></string-name>, <string-name><surname>Hadjileontiadis</surname> <given-names>LJ</given-names></string-name>, <string-name><surname>Khandoker</surname> <given-names>AH</given-names></string-name></person-group>. <article-title>Identification of congenital valvular murmurs in young patients using deep learning-based attention transformers and phonocardiograms</article-title>. <source>IEEE J Biomed Health Inform</source>. <year>2024</year>;<volume>28</volume>(<issue>4</issue>):<fpage>1803</fpage>&#x2013;<lpage>14</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jbhi.2024.3357506</pub-id>; <pub-id pub-id-type="pmid">38261492</pub-id></mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Manshadi</surname> <given-names>OD</given-names></string-name>, <string-name><surname>Mihandoost</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Murmur identification and outcome prediction in phonocardiograms using deep features based on Stockwell transform</article-title>. <source>Sci Rep</source>. <year>2024</year>;<volume>14</volume>:<fpage>7592</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41598-024-58274-6</pub-id>; <pub-id pub-id-type="pmid">38555390</pub-id></mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Lu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yip</surname> <given-names>JB</given-names></string-name>, <string-name><surname>Steigleder</surname> <given-names>T</given-names></string-name>, <string-name><surname>Grie&#x00DF;hammer</surname> <given-names>S</given-names></string-name>, <string-name><surname>Heckel</surname> <given-names>M</given-names></string-name>, <string-name><surname>Jami</surname> <given-names>NVSJ</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A lightweight robust approach for automatic heart murmurs and clinical outcomes classification from phonocardiogram recordings</article-title>. In: <conf-name>2022 Computing in Cardiology (CinC); 2022 Sep 4&#x2013;7</conf-name>; <publisher-loc>Tampere, Finland</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>4</lpage>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Noman</surname> <given-names>F</given-names></string-name>, <string-name><surname>Salleh</surname> <given-names>SH</given-names></string-name>, <string-name><surname>Ting</surname> <given-names>CM</given-names></string-name>, <string-name><surname>Samdin</surname> <given-names>SB</given-names></string-name>, <string-name><surname>Ombao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Hussain</surname> <given-names>H</given-names></string-name></person-group>. <article-title>A Markov-switching model approach to heart sound segmentation and classification</article-title>. <source>IEEE J Biomed Health Inform</source>. <year>2020</year>;<volume>24</volume>(<issue>3</issue>):<fpage>705</fpage>&#x2013;<lpage>16</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jbhi.2019.2925036</pub-id>; <pub-id pub-id-type="pmid">31251203</pub-id></mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nogueira</surname> <given-names>DM</given-names></string-name>, <string-name><surname>Ferreira</surname> <given-names>CA</given-names></string-name>, <string-name><surname>Gomes</surname> <given-names>EF</given-names></string-name>, <string-name><surname>Jorge</surname> <given-names>AM</given-names></string-name></person-group>. <article-title>Classifying heart sounds using images of motifs, MFCC and temporal features</article-title>. <source>J Med Syst</source>. <year>2019</year>;<volume>43</volume>:<fpage>168</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s10916-019-1286-5</pub-id>; <pub-id pub-id-type="pmid">31056720</pub-id></mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Khan</surname> <given-names>FA</given-names></string-name>, <string-name><surname>Abid</surname> <given-names>A</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>MS</given-names></string-name></person-group>. <article-title>Automatic heart sound classification from segmented/unsegmented phonocardiogram signals using time and frequency features</article-title>. <source>Physiol Meas</source>. <year>2020</year>;<volume>41</volume>:<fpage>055006</fpage>. doi:<pub-id pub-id-type="doi">10.1088/1361-6579/ab8770</pub-id>; <pub-id pub-id-type="pmid">32259811</pub-id></mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Das</surname> <given-names>S</given-names></string-name>, <string-name><surname>Dandapat</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Heart murmur severity stages classification using Multikernel residual CNN</article-title>. <source>IEEE Sens J</source>. <year>2024</year>;<volume>24</volume>:<fpage>13019</fpage>&#x2013;<lpage>27</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jsen.2024.3373226</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dawood</surname> <given-names>T</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>C</given-names></string-name>, <string-name><surname>Sidhu</surname> <given-names>BS</given-names></string-name>, <string-name><surname>Ruijsink</surname> <given-names>B</given-names></string-name>, <string-name><surname>Gould</surname> <given-names>J</given-names></string-name>, <string-name><surname>Porter</surname> <given-names>B</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Uncertainty aware training to improve deep learning model calibration for classification of cardiac MR images</article-title>. <source>Med Image Anal</source>. <year>2023</year>;<volume>88</volume>:<fpage>102861</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.media.2023.102861</pub-id>; <pub-id pub-id-type="pmid">37327613</pub-id></mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Antoni</surname> <given-names>L</given-names></string-name>, <string-name><surname>Bruoth</surname> <given-names>E</given-names></string-name>, <string-name><surname>Bugata</surname> <given-names>P</given-names></string-name>, <string-name><surname>Bugata</surname> <given-names>P</given-names></string-name>, <string-name><surname>Gajdo&#x0161;</surname> <given-names>D</given-names></string-name>, <string-name><surname>Hud&#x00E1;k</surname> <given-names>D</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Murmur identification using supervised contrastive learning</article-title>. In: <conf-name>2022 Computing in Cardiology (CinC); 2022 Sep 4&#x2013;7</conf-name>; <publisher-loc>Tampere, Finland</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>4</lpage>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Le</surname> <given-names>D</given-names></string-name>, <string-name><surname>Truong</surname> <given-names>S</given-names></string-name>, <string-name><surname>Brijesh</surname> <given-names>P</given-names></string-name>, <string-name><surname>Adjeroh</surname> <given-names>DA</given-names></string-name>, <string-name><surname>Le</surname> <given-names>N</given-names></string-name></person-group>. <article-title>sCL-ST: supervised contrastive learning with semantic transformations for multiple lead ECG arrhythmia classification</article-title>. <source>IEEE J Biomed Health Inform</source>. <year>2023</year>;<volume>27</volume>(<issue>6</issue>):<fpage>2818</fpage>&#x2013;<lpage>28</lpage>. doi:<pub-id pub-id-type="doi">10.1109/jbhi.2023.3246241</pub-id>; <pub-id pub-id-type="pmid">37028019</pub-id></mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kompa</surname> <given-names>B</given-names></string-name>, <string-name><surname>Snoek</surname> <given-names>J</given-names></string-name>, <string-name><surname>Beam</surname> <given-names>AL</given-names></string-name></person-group>. <article-title>Second opinion needed: communicating uncertainty in medical machine learning</article-title>. <source>npj Digit Med</source>. <year>2021</year>;<volume>4</volume>(<issue>1</issue>):<fpage>4</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41746-020-00367-3</pub-id>; <pub-id pub-id-type="pmid">33402680</pub-id></mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jayaraman</surname> <given-names>P</given-names></string-name>, <string-name><surname>Desman</surname> <given-names>J</given-names></string-name>, <string-name><surname>Sabounchi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Nadkarni</surname> <given-names>GN</given-names></string-name>, <string-name><surname>Sakhuja</surname> <given-names>A</given-names></string-name></person-group>. <article-title>A primer on reinforcement learning in medicine for clinicians</article-title>. <source>npj Digit Med</source>. <year>2024</year>;<volume>7</volume>:<fpage>337</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41746-024-01316-0</pub-id>; <pub-id pub-id-type="pmid">39592855</pub-id></mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Reyna</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Kiarashi</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Elola</surname> <given-names>A</given-names></string-name>, <string-name><surname>Oliveira</surname> <given-names>J</given-names></string-name>, <string-name><surname>Renna</surname> <given-names>F</given-names></string-name>, <string-name><surname>Gu</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Heart murmur detection from phonocardiogram recordings: the George B. Moody PhysioNet Challenge 2022</article-title>. <source>PLoS Digit Health</source>. <year>2023</year>;<volume>2</volume>:<fpage>e0000324</fpage>; <pub-id pub-id-type="pmid">37695769</pub-id></mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Oliveira</surname> <given-names>J</given-names></string-name>, <string-name><surname>Renna</surname> <given-names>F</given-names></string-name>, <string-name><surname>Costa</surname> <given-names>PD</given-names></string-name>, <string-name><surname>Nogueira</surname> <given-names>M</given-names></string-name>, <string-name><surname>Oliveira</surname> <given-names>C</given-names></string-name>, <string-name><surname>Ferreira</surname> <given-names>C</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>The CirCor DigiScope dataset: from murmur detection to murmur classification</article-title>. <source>IEEE J Biomed Health Inform</source>. <year>2022</year>;<volume>26</volume>(<issue>6</issue>):<fpage>2524</fpage>&#x2013;<lpage>35</lpage>; <pub-id pub-id-type="pmid">34932490</pub-id></mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Azam</surname> <given-names>FB</given-names></string-name>, <string-name><surname>Ansari</surname> <given-names>MI</given-names></string-name>, <string-name><surname>Nuhash</surname> <given-names>SISK</given-names></string-name>, <string-name><surname>McLane</surname> <given-names>I</given-names></string-name>, <string-name><surname>Hasan</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Cardiac anomaly detection considering an additive noise and convolutional distortion model of heart sound recordings</article-title>. <source>Artif Intell Med</source>. <year>2022</year>;<volume>133</volume>:<fpage>102417</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.artmed.2022.102417</pub-id>; <pub-id pub-id-type="pmid">36328670</pub-id></mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Park</surname> <given-names>DS</given-names></string-name>, <string-name><surname>Chan</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chiu</surname> <given-names>CC</given-names></string-name>, <string-name><surname>Zoph</surname> <given-names>B</given-names></string-name>, <string-name><surname>Cubuk</surname> <given-names>ED</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>SpecAugment: a simple data augmentation method for automatic speech recognition</article-title>. <comment>arXiv:1904.08779. 2019</comment>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kong</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Iqbal</surname> <given-names>T</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Plumbley</surname> <given-names>MD</given-names></string-name></person-group>. <article-title>PANNs: large-scale pretrained audio neural networks for audio pattern recognition</article-title>. <source>IEEE/ACM Trans Audio Speech Lang Process</source>. <year>2020</year>;<volume>28</volume>:<fpage>2880</fpage>&#x2013;<lpage>94</lpage>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Gemmeke</surname> <given-names>JF</given-names></string-name>, <string-name><surname>Ellis</surname> <given-names>DPW</given-names></string-name>, <string-name><surname>Freedman</surname> <given-names>D</given-names></string-name>, <string-name><surname>Jansen</surname> <given-names>A</given-names></string-name>, <string-name><surname>Lawrence</surname> <given-names>W</given-names></string-name>, <string-name><surname>Moore</surname> <given-names>RC</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Audio set: an ontology and human-labeled dataset for audio events</article-title>. In: <conf-name>2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</conf-name>. <publisher-loc>Piscataway, NJ, USA</publisher-loc>: <publisher-name>IEEE</publisher-name>; <year>2017</year>. p. <fpage>776</fpage>&#x2013;<lpage>80</lpage>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Khosla</surname> <given-names>P</given-names></string-name>, <string-name><surname>Teterwak</surname> <given-names>P</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Sarna</surname> <given-names>A</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Isola</surname> <given-names>P</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Supervised contrastive learning</article-title>. <comment>arXiv:2004.11362. 2021</comment>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zheng</surname> <given-names>H</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Contrastive attraction and contrastive repulsion for representation learning</article-title>. <comment>arXiv:2105.03746. 2023</comment>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Schulman</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wolski</surname> <given-names>F</given-names></string-name>, <string-name><surname>Dhariwal</surname> <given-names>P</given-names></string-name>, <string-name><surname>Radford</surname> <given-names>A</given-names></string-name>, <string-name><surname>Klimov</surname> <given-names>O</given-names></string-name></person-group>. <article-title>Proximal policy optimization algorithms</article-title>. <comment>arXiv:1707.06347. 2017</comment>.</mixed-citation></ref>
</ref-list>
</back></article>