<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">81667</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2026.081667</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>A Hybrid Physics-Informed and Data-Driven Feature Framework with Explicit Correlation-Structure Embeddings for Early-Life Prognostics of Lithium-Ion Batteries</article-title>
<alt-title alt-title-type="left-running-head">A Hybrid Physics-Informed and Data-Driven Feature Framework with Explicit Correlation-Structure Embeddings for Early-Life Prognostics of Lithium-Ion Batteries</alt-title>
<alt-title alt-title-type="right-running-head">A Hybrid Physics-Informed and Data-Driven Feature Framework with Explicit Correlation-Structure Embeddings for Early-Life Prognostics of Lithium-Ion Batteries</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Lee</surname><given-names>Kang-Woo</given-names></name></contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Lee</surname><given-names>Dong-Hee</given-names></name><email>dhee@skku.edu</email></contrib>
<contrib id="author-3" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Kwon</surname><given-names>Dae-Il</given-names></name><email>dikwon@skku.edu</email></contrib>
<aff id="aff-1"><institution>Department of Industrial Engineering, Sungkyunkwan University</institution>, <addr-line>Suwon</addr-line>, <country>Republic of Korea</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Authors: Dong-Hee Lee. Email: <email>dhee@skku.edu</email>; Dae-Il Kwon. Email: <email>dikwon@skku.edu</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>15</day><month>06</month><year>2026</year>
</pub-date>
<volume>88</volume>
<issue>2</issue>
<elocation-id>51</elocation-id>
<history>
<date date-type="received">
<day>06</day>
<month>03</month>
<year>2026</year>
</date>
<date date-type="accepted">
<day>22</day>
<month>04</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors. Published by Tech Science Press.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>The Authors</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_81667.pdf"></self-uri>
<abstract>
<p>Early-life cycle-life prediction for lithium-ion batteries&#x2014;estimating end-of-life from initial cycles&#x2014;is valuable for rapid cell screening and battery health management. We investigate whether an explicit correlation-structure descriptor can complement physics-informed <italic>&#x0394;Q</italic>-based indicators and generic early-cycle statistical features on the Severson 124-cell benchmark. We develop a lightweight hybrid framework that combines <italic>&#x0394;Q</italic>-based health indicators, data-driven statistical features, and Laplacian Eigenmaps embeddings derived from a Pearson-correlation feature graph, with XGBoost used as the predictor. Across five feature configurations (<italic>&#x0394;Q Only, &#x0394;Q &#x002B; Statistics, Hybrid Append, VIF &#x002B; Laplacian, and Integrated Laplacian</italic>), we evaluate pointwise regression accuracy using RMSE and <italic>R</italic><sup>2</sup> together with PHM-style error-band measures RA@0.2, PH@0.1, and <italic>&#x03B1;-&#x03BB;</italic>(0.15), computed on implied RUL trajectories induced by the early-life cycle-life estimate. On the Primary test domain, all four non-baseline configurations improved over <italic>&#x0394;Q Only</italic>; Integrated Laplacian achieved the strongest RMSE/<italic>R</italic><sup>2</sup> pair (97.00 cycles, 0.8215), while Hybrid Append remained competitive (102.28 cycles, 0.8016) and improved RA@0.2 and PH@0.1 relative to <italic>&#x0394;Q &#x002B; Statistics</italic>. On the shifted Secondary domain, <italic>&#x0394;Q Only</italic> gave the most favorable RMSE/<italic>R</italic><sup>2</sup> pair (267.82 cycles, 0.2275), whereas <italic>Hybrid Append</italic> and <italic>VIF &#x002B; Laplacian</italic> improved selected error-band metrics. In an additional comparison against PCA, Random Projection, and Truncated SVD conducted at a matched 79-feature scale, with all transforms estimated from the training cells only, the graph-derived embedding remained competitive, but its margin over simpler reductions varied across splits. Taken together, these results support <italic>Hybrid Append</italic> as the main appended-structure configuration in this study, while indicating that the benefit of the correlation-structure descriptor is more visible in selected PHM-style error-band metrics than in uniformly improved pointwise accuracy.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Lithium-ion battery</kwd>
<kwd>early-life cycle life prediction</kwd>
<kwd>remaining useful life (RUL)</kwd>
<kwd>prognostics and health management (PHM)</kwd>
<kwd><italic>&#x0394;Q</italic>-based health indicators</kwd>
<kwd>hybrid feature engineering</kwd>
<kwd>graph Laplacian embedding</kwd>
<kwd>prognostic metrics</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIT)</funding-source>
<award-id>NRF-2022R1C1C1011743</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Lithium-ion batteries are widely deployed in electric vehicles, stationary energy-storage systems, aerospace, and defense applications. As adoption expands, prognostics and health management (PHM)&#x2014;including remaining useful life (RUL) prediction&#x2014;becomes increasingly important for safety, warranty planning, and operational efficiency. Recent reviews emphasize that practical battery PHM algorithms must balance predictive reliability with interpretability and deployment constraints such as limited data and changing operating conditions [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-2">2</xref>].</p>
<p>Within this landscape, early-life prediction aims to estimate the total cycle life of a cell using only a small fraction of the initial cycling data (e.g., the first 5%&#x2013;10% of life). This capability is valuable for cell screening, quality control, and rapid iteration of cell design and fast-charging protocols [<xref ref-type="bibr" rid="ref-3">3</xref>&#x2013;<xref ref-type="bibr" rid="ref-6">6</xref>]. However, it is inherently challenging because early degradation signatures are subtle, and available laboratory datasets are often small. Consequently, robustness and interpretability are as important as raw predictive accuracy.</p>
<p>Prior work on early-life prognostics spans three complementary directions. Physics-informed machine learning (PIML) embeds electrochemical knowledge and degradation constraints into learning pipelines to improve physical consistency [<xref ref-type="bibr" rid="ref-7">7</xref>&#x2013;<xref ref-type="bibr" rid="ref-10">10</xref>]. Data-driven early-life methods leverage engineered early-cycle features or learned representations to predict cycle life from initial measurements [<xref ref-type="bibr" rid="ref-3">3</xref>&#x2013;<xref ref-type="bibr" rid="ref-6">6</xref>]. Graph-based learning has also been explored to capture dependencies among signals, cells, or features, often through deep graph architectures [<xref ref-type="bibr" rid="ref-11">11</xref>&#x2013;<xref ref-type="bibr" rid="ref-15">15</xref>]. While these approaches have advanced the field, practical gaps remain in small-sample early-life settings. High-capacity PIML and deep graph models can be difficult to tune and interpret when data are limited [<xref ref-type="bibr" rid="ref-1">1</xref>], and in many graph-deep-learning frameworks, the graph structure is implicit within the network, limiting direct inspection of which correlations drive predictions [<xref ref-type="bibr" rid="ref-11">11</xref>&#x2013;<xref ref-type="bibr" rid="ref-15">15</xref>]. Moreover, feature choices should align with the intended deployment objective, since protocol-encoding features may inflate apparent accuracy while reducing transferability [<xref ref-type="bibr" rid="ref-16">16</xref>].</p>
<p>Beyond these methodological considerations, a further gap concerns evaluation. Many battery life-prediction studies report only static regression metrics such as RMSE or R<sup>2</sup>, which summarize pointwise error but say little about whether the implied RUL trajectory from an early-life cycle-life estimate stays within practically relevant error bands over time. In PHM, RUL is often assessed through trajectory-oriented measures&#x2014;Relative Accuracy (RA), Prognostic Horizon (PH), and <italic>&#x03B1;-&#x03BB;</italic>&#x2014;that quantify error-band agreement and temporal consistency [<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-18">18</xref>]. These metrics remain less common in early-life battery benchmarks that primarily emphasize RMSE/<italic>R</italic><sup>2</sup> [<xref ref-type="bibr" rid="ref-3">3</xref>&#x2013;<xref ref-type="bibr" rid="ref-6">6</xref>]. Accordingly, we predict total cycle life from initial cycles and additionally evaluate the implied RUL trajectory using PHM-style error-band metrics.</p>
<p>Motivated by these gaps, we propose a lightweight early-life prognostics framework that emphasizes explicit structural representation together with PHM-style evaluation under both in-domain and shifted-domain settings. We use a fixed feature-correlation graph as an external structural descriptor and express it through Laplacian Eigenmaps alongside physics-informed <italic>&#x0394;Q</italic> indicators in a small-sample early-life setting.</p>
<p>Our framework combines three feature groups: <italic>&#x0394;Q</italic>-based health indicators as a physics-informed channel [<xref ref-type="bibr" rid="ref-3">3</xref>], generic early-cycle statistical features, and graph-derived structural embeddings obtained through Laplacian Eigenmaps on a Pearson-correlation feature graph [<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-20">20</xref>]. XGBoost [<xref ref-type="bibr" rid="ref-21">21</xref>] is used as the predictor because it remains robust and parameter-efficient for tabular data in small-sample regimes.</p>
<p>Across five feature configurations (<italic>&#x0394;Q Only</italic>, <italic>&#x0394;Q &#x002B; Statistics</italic>, <italic>Hybrid Append</italic>, <italic>VIF &#x002B; Laplacian</italic>, and <italic>Integrated Laplacian</italic>), we evaluate both regression metrics (RMSE, <italic>R</italic><sup>2</sup>) and PHM-style error-band metrics under an in-domain Primary split and a batch-shifted Secondary split. The analysis emphasizes graph construction confined to the training cells, comparisons performed at matched feature dimensionality, and the contrast between pointwise regression error and threshold-based PHM agreement.</p>
<p>This paper makes three contributions. First, we study a hybrid early-life prognostics pipeline that combines physics-informed <italic>&#x0394;Q</italic> indicators, generic statistical features, and a fixed spectral descriptor derived from a feature-correlation graph. Second, we provide a stepwise feature-design workflow that keeps the physics-informed <italic>&#x0394;Q</italic> channel explicit and separate from the correlation-structure descriptor at the feature-construction level. Third, we evaluate the resulting feature configurations using both standard regression metrics and PHM-style error-band metrics under a within-benchmark batch-shift setting in the Severson benchmark.</p>
<p>The remainder of this paper is organized as follows. <xref ref-type="sec" rid="s2">Section 2</xref> reviews related work on early-life prediction, physics-informed and graph-based learning, and PHM trajectory metrics. <xref ref-type="sec" rid="s3">Section 3</xref> describes the proposed hybrid feature pipeline and experimental protocol. <xref ref-type="sec" rid="s4">Section 4</xref> presents results using both regression and PHM metrics. <xref ref-type="sec" rid="s5">Section 5</xref> concludes with limitations and directions for future research.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Literature Review</title>
<p>This section reviews prior work most relevant to our study and clarifies how our approach is positioned. We focus on battery PHM and early-life cycle-life prediction, physics-informed and graph-based learning with an emphasis on interpretability and small-sample practicality, and PHM trajectory metrics for evaluating prognostic stability. The goal is to connect these threads to our design choice of using graph theory as an explicit, external feature-structure representation in a small-sample early-life setting.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Battery PHM and Early-Life Cycle-Life Prediction</title>
<p>Battery PHM spans degradation-mechanism understanding, state-of-health (SoH) estimation, RUL prediction, and decision-making for maintenance and operation across the battery life cycle. Comprehensive reviews summarize degradation modes, operational factors, and modeling paradigms and emphasize that interpretability and reliability are central to real-world PHM adoption [<xref ref-type="bibr" rid="ref-1">1</xref>].</p>
<p>Early-life prediction targets rapid life estimation using only initial cycles to enable screening and early decisions [<xref ref-type="bibr" rid="ref-3">3</xref>&#x2013;<xref ref-type="bibr" rid="ref-6">6</xref>]. Severson et al. [<xref ref-type="bibr" rid="ref-3">3</xref>] established a widely used benchmark by showing that early-cycle voltage-curve-derived features (notably <italic>&#x0394;Q&#x2212;V</italic>) can be predictive of cycle life before substantial capacity fade. Subsequent studies explored alternative representations and learning models, including image-based encodings with CNNs [<xref ref-type="bibr" rid="ref-4">4</xref>,<xref ref-type="bibr" rid="ref-6">6</xref>], probabilistic regressors such as GPR with degradation-pattern recognition [<xref ref-type="bibr" rid="ref-5">5</xref>], and graphical feature constructions based on <italic>&#x0394;Q&#x2212;V IC</italic>/<italic>&#x0394;Q</italic> curves [<xref ref-type="bibr" rid="ref-22">22</xref>]. Broader-condition work further examined early-life prediction under varying usage conditions and proposed degradation-informed or hierarchical approaches to improve extrapolation [<xref ref-type="bibr" rid="ref-23">23</xref>], while cross-condition deep learning frameworks such as BatLiNet aim to improve reliability across diverse ageing conditions via inter-cell learning mechanisms [<xref ref-type="bibr" rid="ref-24">24</xref>].</p>
<p>These efforts demonstrate the potential of early-cycle data while also highlighting practical trade-offs in small datasets: higher-capacity end-to-end models can be less transparent and more sensitive to modeling choices. In addition, recent perspective work emphasizes that feature selection should match the deployment objective, because protocol-encoding features can bias apparent performance and complicate transferability [<xref ref-type="bibr" rid="ref-16">16</xref>]. Our work adopts a pragmatic stance by combining physics-informed <italic>&#x0394;Q</italic> indicators with generic statistics and an explicit correlation-structure representation to study whether a lightweight tabular pipeline can remain competitive without introducing a complex end-to-end graph learner.</p>
<p>Recent adjacent literature has also explored graph-based battery health modeling and alternative health metrics under related but non-identical tasks, including graph-based SOH estimation from partial discharge segments [<xref ref-type="bibr" rid="ref-25">25</xref>], alternative battery-health metrics such as Health EquiMetrics [<xref ref-type="bibr" rid="ref-26">26</xref>], and matrix-profile-based online knee-onset identification for lithium-ion battery SOH estimation [<xref ref-type="bibr" rid="ref-27">27</xref>]. Because many recent SOH-oriented and alternative battery-health studies use different prediction targets, sensing windows, and evaluation protocols, we discuss them as methodological context rather than as direct numerical baselines for the early-life cycle-life regression task considered here.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Physics-Informed and Graph-Based Learning: Expressiveness vs. Interpretability</title>
<p>Physics-informed machine learning (PIML) integrates physical constraints, electrochemical knowledge, or equivalent-circuit insights into learning pipelines, for example, by embedding governing equations into model structures or training objectives [<xref ref-type="bibr" rid="ref-7">7</xref>&#x2013;<xref ref-type="bibr" rid="ref-10">10</xref>]. PIML can offer improved physical consistency and may enhance generalization when adequate domain information and data support are available. However, PIML often introduces additional modeling layers and expanded hyperparameter spaces, which can be burdensome in small-sample regimes and less accessible for practitioners without deep modeling expertise [<xref ref-type="bibr" rid="ref-1">1</xref>].</p>
<p>Graph-based learning has also received substantial attention for capturing structured dependencies among sensors, cells, or features. Recent approaches include spatio-temporal models with dynamic graphs [<xref ref-type="bibr" rid="ref-11">11</xref>], correlation- or attention-based GNN variants [<xref ref-type="bibr" rid="ref-12">12</xref>&#x2013;<xref ref-type="bibr" rid="ref-14">14</xref>], and graph-attention&#x2013;recurrent hybrids supported by explainable feature fragments [<xref ref-type="bibr" rid="ref-15">15</xref>]. These models can be powerful, but many treat the graph as an internal latent structure whose influence is not always straightforward to interpret directly; moreover, implicit graph learning can raise concerns about tuning effort and reproducibility in small datasets.</p>
<p>In contrast, our study should be understood as a graph-assisted feature-engineering case study. We use graph theory to construct a structurally explicit descriptor at the feature-correlation level. Specifically, we build a feature-by-feature correlation graph using Pearson coefficients [<xref ref-type="bibr" rid="ref-19">19</xref>] and compute low-dimensional Laplacian Eigenmaps embeddings [<xref ref-type="bibr" rid="ref-20">20</xref>] as compact representations of correlation structure. The contribution, therefore, lies in explicit spectral feature engineering rather than in a novel graph-learning architecture.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>PHM Performance Metrics and Evaluation in Small-Sample Settings</title>
<p>Most battery life-prediction studies report static regression metrics such as RMSE, MAE, or <italic>R</italic><sup>2</sup>. While these metrics quantify pointwise accuracy, they do not describe whether the RUL trajectory implied by a predicted cycle life remains inside practically relevant error bands over time. In PHM, RUL is often evaluated as a trajectory, motivating measures that summarize error-band agreement and temporal consistency rather than final pointwise error alone [<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-18">18</xref>].</p>
<p>Saxena et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] formalized offline prognostic metrics including RA, PH, and <italic>&#x03B1;-&#x03BB;</italic>. RA measures the fraction of predictions within a specified tolerance band; PH summarizes how early predictions enter a tighter band; and <italic>&#x03B1;-&#x03BB;</italic> captures agreement within a chosen tolerance window. Sharp [<xref ref-type="bibr" rid="ref-18">18</xref>] further discussed how such measures can be communicated to users with varied backgrounds. Despite their relevance, these trajectory-oriented metrics are still not routinely used in early-life battery benchmarks, where RMSE/<italic>R</italic><sup>2</sup> remain the dominant reporting metrics [<xref ref-type="bibr" rid="ref-3">3</xref>&#x2013;<xref ref-type="bibr" rid="ref-6">6</xref>]. In this work, we therefore report both regression metrics and PHM-style error-band metrics computed on the implied RUL trajectories induced by each model&#x2019;s early-life cycle-life estimate.</p>
<p>Accordingly, our work evaluates models using both regression metrics and PHM-style threshold-based accuracy measures, providing a complementary view of pointwise prediction quality and error-band agreement under batch shift.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Approaches</title>
<p>We describe the problem formulation, the hybrid feature-engineering pipeline, the five feature-set configurations, and the XGBoost [<xref ref-type="bibr" rid="ref-21">21</xref>] modeling and evaluation protocol. The overall workflow is summarized in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Overall pipeline of the proposed early-life cycle-life prediction framework. Early-life signals from the Severson 124-cell dataset are transformed into physics-informed <italic>&#x0394;Q&#x2212;V</italic> health indicators (HIs), data-driven statistical features, and Laplacian features (graph-based Laplacian features; Laplacian Eigenmaps projections) derived from a Pearson-correlation graph constructed on statistical features only. The resulting feature sets are used to train an XGBoost regressor with Bayesian hyperparameter optimization, and model performance is evaluated using RMSE, <italic>R</italic><sup>2</sup>, and PHM-oriented metrics (RA@0.2, PH@0.1, <italic>&#x03B1;-&#x03BB;</italic>(0.15)).</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81667-fig-1.tif"/>
</fig>
<sec id="s3_1">
<label>3.1</label>
<title>Dataset and Problem Formulation</title>
<p>We use the open Severson benchmark [<xref ref-type="bibr" rid="ref-3">3</xref>], which contains 124 commercial A123 APR18650M1A LFP/graphite cells cycled at 30&#x00B0;C under 72 one-step and two-step fast-charging policies. Cycle life is defined as the number of cycles until the cell reaches 80% of nominal capacity [<xref ref-type="bibr" rid="ref-3">3</xref>]. For each cell, the dataset provides cycle-resolved current, terminal voltage, charge/discharge capacity, time, temperature, and internal-resistance-related measurements.</p>
<p>Following Severson et al. [<xref ref-type="bibr" rid="ref-3">3</xref>], we formulate early-life cycle-life prediction as a supervised regression problem. For each cell, we use early-life cycles 1&#x2013;100 as inputs and predict the total cycle life as the target.</p>
<p>For our evaluation, the dataset is organized into three batches (Batch 1&#x2013;3). Batches 1&#x2013;2 form the main development domain and are split into Train (41 cells; odd-indexed cells) and Primary Test (43 cells; even-indexed cells), whereas Batch 3 (40 cells) is retained as the Secondary Test set. Across all three batches, the common protocol backbone is fast charging from 0% to 80% SOC under one of the benchmark charging policies, an internal-resistance measurement at 80% SOC, a uniform 1C CC-CV top-off from 80% to 100% SOC to 3.6 V, and identical 4C CC-CV discharge to 2.0 V [<xref ref-type="bibr" rid="ref-3">3</xref>].</p>
<p>The batches nevertheless differ in their exact rest schedules and in the placement of short pauses around the 80% SOC/internal-resistance/discharge steps. Severson et al. report 1 min and 1 s rests after reaching 80% SOC and after discharge, respectively, for the 12 May 2017 batch; 5 min rests at both locations for the 30 June 2017 batch; and 5 s rests after reaching 80% SOC, after the internal-resistance test, and both before and after discharge for the 12 April 2018 batch [<xref ref-type="bibr" rid="ref-3">3</xref>]. Together, these protocol and distributional differences motivate treating Batch 3 as a protocol-shifted secondary evaluation domain rather than as an in-domain test split. The batch-specific rest/pause schedules and their roles in the evaluation are summarized in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Batch-specific rest/pause schedules in the Severson benchmark and their relevance to the shifted Secondary Test domain.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Batch</th>
<th>Key Protocol Difference</th>
<th>Role In Evaluation</th>
</tr>
</thead>
<tbody>
<tr>
<td>Batch 1 (12 May 2017)</td>
<td>1 min rest after reaching 80% SOC; 1 s rest after discharge</td>
<td>Part of the main in-domain distribution</td>
</tr>
<tr>
<td>Batch 2 (30 June 2017)</td>
<td>5 min rest after reaching 80% SOC and after discharge</td>
<td>Part of the main in-domain distribution, with moderate within-benchmark variation</td>
</tr>
<tr>
<td>Batch 3 (12 April 2018)</td>
<td>5 s rests after reaching 80% SOC, after the internal-resistance test, and before/after discharge</td>
<td>Secondary test domain with protocol shift</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>Because several point features are computed from current, time, and capacity trajectories over cycles 1&#x2013;100, differences in rest timing and pause placement can directly alter the empirical distribution of the statistical-feature channel even when the underlying chemistry is unchanged.</p>
<p>After feature extraction, each cell is represented as a single tabular feature vector. Early-life time series are summarized into physics-informed indicators, generic statistics, and correlation-structure descriptors, yielding a regression task suitable for tree-based gradient boosting.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Hybrid Feature Engineering</title>
<p>Our feature engineering comprises three groups: (i) physics-informed <italic>&#x0394;Q&#x2212;V</italic> health indicators, (ii) data-driven statistical features, and (iii) Laplacian features that encode correlation structure among the statistical features.</p>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>Physics-Informed &#x0394;Q&#x2212;V Health Indicators</title>
<p>We build on the <italic>&#x0394;Q</italic>-based health indicators introduced by Severson et al. [<xref ref-type="bibr" rid="ref-3">3</xref>]. Using early-life charging curves (cycles 10&#x2013;100), we align capacity-voltage curves onto a common voltage grid and compute <italic>&#x0394;Q</italic> between selected cycle pairs within voltage bins, producing <italic>&#x0394;Q</italic> signatures. We then summarize these signatures into seven scalar HIs that capture the magnitude, variability, and characteristic shifts of <italic>&#x0394;Q</italic> along the voltage axis, including statistics computed over degradation-sensitive voltage regions.</p>
<p>These seven HIs are treated as physics-informed descriptors because they arise from physically motivated signal transforms that have previously shown relevance to cycle-life variability in early-life prediction settings [<xref ref-type="bibr" rid="ref-3">3</xref>]. Here, they serve as a dedicated physics-informed feature channel rather than as direct mechanistic state variables.</p>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>Data-Driven Point Features</title>
<p>The second feature group consists of generic statistical features computed directly from early-cycle signals without using <italic>&#x0394;Q</italic> transforms. Here, <italic>Qdlin</italic> denotes the linearly interpolated discharge-capacity signal represented on a common voltage grid. From cycles 1&#x2013;100, we extract 52 statistical features organized into five families: current-derived, time-derived, capacity-derived, differential-capacity&#x2013;derived, and interaction-derived features. Across these families, the features are generated by applying four summary operators&#x2014;mean, standard deviation, minimum, and maximum&#x2014;to raw, transformed, ratio, and interaction signals. A compact family-wise summary is provided in <xref ref-type="table" rid="table-2">Table 2</xref>, and <xref ref-type="table" rid="table-7">Table A1</xref> in <xref ref-type="app" rid="app-1">Appendix A</xref> lists the complete grouped set of the 52 feature names.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Compact family-wise summary of the 52 statistical features used in the data-driven point-feature set, grouped by source-signal family and common aggregation pattern.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Feature Family</th>
<th>Representative Source Signals</th>
<th>Aggregation Pattern</th>
<th>Count</th>
</tr>
</thead>
<tbody>
<tr>
<td>Current-derived</td>
<td><inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>I</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mi>I</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td>mean, std, min, max</td>
<td>8</td>
</tr>
<tr>
<td>Time-derived</td>
<td><inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msup><mml:mi>t</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>t</mml:mi></mml:math></inline-formula></td>
<td>mean, std, min, max</td>
<td>8</td>
</tr>
<tr>
<td>Capacity-derived</td>
<td><inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mrow><mml:mi>Q</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>d</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>l</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>i</mml:mi><mml:mo>&#x2062;</mml:mo><mml:msup><mml:mi>n</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mrow><mml:mi>Q</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>d</mml:mi></mml:mrow><mml:mo>/</mml:mo><mml:mi>t</mml:mi></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mrow><mml:mi>Q</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>d</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>l</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>i</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>n</mml:mi></mml:mrow><mml:mo>/</mml:mo><mml:mrow><mml:mi>Q</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>d</mml:mi></mml:mrow></mml:mrow><mml:mo>,</mml:mo><mml:mspace linebreak="newline"></mml:mspace><mml:mrow><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mi>Q</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>d</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>l</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>i</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:mrow></mml:math></inline-formula></td>
<td>mean, std, min, max</td>
<td>16</td>
</tr>
<tr>
<td>Differential-capacity-derived</td>
<td><inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>Q</mml:mi></mml:mrow><mml:mo>/</mml:mo><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>V</mml:mi></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mrow><mml:mi>I</mml:mi><mml:mo>/</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>Q</mml:mi></mml:mrow><mml:mo>/</mml:mo><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>V</mml:mi></mml:mrow></mml:mrow><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mo>,</mml:mo><mml:mspace linebreak="newline"></mml:mspace><mml:mrow><mml:mrow><mml:mi>Q</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>d</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>l</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>i</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>n</mml:mi></mml:mrow><mml:mo>&#xD7;</mml:mo><mml:mrow><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>Q</mml:mi></mml:mrow><mml:mo>/</mml:mo><mml:mrow><mml:mi>d</mml:mi><mml:mo>&#x2062;</mml:mo><mml:mi>V</mml:mi></mml:mrow></mml:mrow></mml:mrow></mml:math></inline-formula></td>
<td>mean, std, min, max</td>
<td>12</td>
</tr>
<tr>
<td>Interaction-derived</td>
<td><inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>I</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>Q</mml:mi><mml:mi>d</mml:mi><mml:mo>,</mml:mo><mml:mi>Q</mml:mi><mml:mi>d</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>t</mml:mi></mml:math></inline-formula></td>
<td>mean, std, min, max</td>
<td>8</td>
</tr>
<tr>
<td align="center" colspan="3"><bold>Total</bold></td>
<td><bold>52</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>These statistical features complement the <italic>&#x0394;Q</italic> health indicators by capturing broader early-cycle distributional characteristics and transformed-signal patterns present in the raw measurements.</p>
</sec>
<sec id="s3_2_3">
<label>3.2.3</label>
<title>Correlation Graph and Laplacian Features</title>
<p>The third feature group uses an explicit graph-theoretic representation to encode correlation structure among the 52 statistical features. To avoid information leakage, the correlation graph is constructed using the training cells only (41 cells).</p>
<p>Let <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi mathvariant="bold-italic">S</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mn>41</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>52</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> denote the training-set matrix of the 52 statistical features. We compute the feature-by-feature Pearson correlation matrix <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>R</mml:mi><mml:mo>=</mml:mo><mml:mo stretchy="false">[</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula>, where <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the Pearson correlation between statistical features <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> across the 41 training cells.</p>
<p><bold>Visualization (<xref ref-type="fig" rid="fig-2">Fig. 2</xref>).</bold> For visualization, we display <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>R</mml:mi></mml:math></inline-formula> as (a) a heatmap and (b) a correlation network with edges displayed when |<inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>| &#x003E; 0.3; this threshold is used for visualization only.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Correlation structure of the statistical feature set computed from the training cells. (<bold>a</bold>) Pearson correlation heatmap of the 52 statistical features. (<bold>b</bold>) Correlation network visualization with edges shown for |r| &#x003E; 0.3, where nodes represent features and node colors indicate feature categories.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81667-fig-2.tif"/>
</fig>
<p><bold>Graph construction.</bold> For the actual Laplacian feature generation, we define a weighted graph over features using off-diagonal absolute Pearson correlations while excluding self-correlations. Specifically, <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> <bold>&#x003D; |</bold><inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula><bold>|</bold> for <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mi>i</mml:mi><mml:mo>&#x2260;</mml:mo><mml:mi>j</mml:mi></mml:math></inline-formula>, with <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>, and the degree matrix <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mi>D</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mtext>diag</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:munder><mml:msub><mml:mi>W</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula>. The normalized graph Laplacian is
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mi>L</mml:mi><mml:mo>=</mml:mo><mml:mi>I</mml:mi><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mi>D</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mi>W</mml:mi><mml:msup><mml:mi>D</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></disp-formula></p>
<p><bold>Laplacian Eigenmaps.</bold> We apply Laplacian Eigenmaps [<xref ref-type="bibr" rid="ref-20">20</xref>] and extract <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mi>k</mml:mi></mml:math></inline-formula> non-trivial eigenvectors of <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mi>L</mml:mi></mml:math></inline-formula> (excluding the constant eigenvector). <xref ref-type="fig" rid="fig-3">Fig. 3</xref> shows an illustrative 2D embedding (<italic>k</italic> &#x003D; 2).</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Two-dimensional Laplacian Eigenmaps embedding of the 52 statistical-feature nodes constructed from the correlation graph in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>. Features with similar correlation profiles form clusters; colors denote feature categories.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81667-fig-3.tif"/>
</fig>
<p><bold>Cell-level Laplacian features.</bold> Laplacian Eigenmaps provides coordinates for feature nodes. To obtain cell-level features for XGBoost, we project each cell&#x2019;s standardized statistical feature vector onto the selected eigenvector subspace. Before projection, statistical features are z-scored using the training-set mean and standard deviation, and the same transformation is applied to the Primary and Secondary test sets. Let <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mn>52</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> be the selected eigenvectors. For a cell with statistical feature vector <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi mathvariant="bold-italic">s</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mn>52</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, we compute
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mi>g</mml:mi><mml:mo>=</mml:mo><mml:msubsup><mml:mi>U</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>T&#xA0;</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mi>s</mml:mi><mml:mo>&#x2208;</mml:mo><mml:msup><mml:mrow><mml:mi mathvariant="double-struck">R</mml:mi></mml:mrow><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msup></mml:math></disp-formula>and use <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mi>g</mml:mi></mml:math></inline-formula> as the Laplacian feature vector for that cell. The same <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msub><mml:mi>U</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> learned from the training set is applied to the Primary and Secondary test cells.</p>
<p><xref ref-type="fig" rid="fig-2">Figs. 2</xref> and <xref ref-type="fig" rid="fig-3">3</xref> are presented as structural diagnostics of the statistical-feature space, not as direct explanations of predictor behavior. Accordingly, the present study interprets the graph component at the level of feature-space organization rather than feature-attribution to individual predictions.</p>

<p>We evaluated k over a prespecified candidate set and found broadly similar behavior across mid-range embedding dimensions, although the best observed value differed by metric and split. We therefore retain <italic>k</italic> &#x003D; 20 as a practical mid-range operating dimension. The quantitative sensitivity results obtained with graph construction confined to the training cells are shown in <xref ref-type="fig" rid="fig-7">Fig. A1</xref> in <xref ref-type="app" rid="app-2">Appendix B</xref>: over the explored grid, the best observed values occurred at <italic>k</italic> &#x003D; 5 for CV <italic>R</italic><sup>2</sup>, <italic>k</italic> &#x003D; 25 for Primary <italic>R</italic><sup>2</sup>, and <italic>k</italic> &#x003D; 15 for Secondary <italic>R</italic><sup>2</sup>. Importantly, <italic>&#x0394;Q&#x2212;V</italic> HIs are not included in the correlation graph: the graph is built on the statistical features only, and <italic>&#x0394;Q&#x2212;V</italic> HIs remain an independent physics-informed channel. The Integrated Laplacian variant, which recomputes the graph after adding <italic>&#x0394;Q&#x2212;V</italic> HIs, is retained as an explicit ablation in the experiments. <xref ref-type="table" rid="table-3">Table 3</xref> summarizes the three feature groups.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Summary of feature groups used in the proposed hybrid feature engineering framework.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Group</th>
<th>Abbrev.</th>
<th>#Features</th>
<th>Main Variables</th>
<th>Role</th>
</tr>
</thead>
<tbody>
<tr>
<td><italic>&#x0394;Q&#x2212;V</italic> HIs (Physics-informed)</td>
<td><inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mrow><mml:mi mathvariant="normal">&#x0394;</mml:mi></mml:mrow><mml:mi>Q</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>V</mml:mi></mml:math></inline-formula>-HI</td>
<td>7</td>
<td><italic>&#x0394;Q&#x2212;V</italic> signature summaries</td>
<td>Physics-motivated<break/>Signal descriptors</td>
</tr>
<tr>
<td>Data-driven Statistics</td>
<td>Stats</td>
<td>52</td>
<td><inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mi>V</mml:mi><mml:mo>,</mml:mo><mml:mi>I</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi>d</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>T</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:math></inline-formula> (summary &#x002B; trend statistics)</td>
<td>Generic early-life<break/>trend summary</td>
</tr>
<tr>
<td>Laplacian Features</td>
<td>Laplacian</td>
<td>20</td>
<td><inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mi>g</mml:mi><mml:mo>=</mml:mo><mml:msubsup><mml:mi>U</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>T&#xA0;</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mi>s</mml:mi></mml:math></inline-formula> from stats correlation graph</td>
<td>Correlation-structure encoding</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Feature-Set Design and XGBoost Modeling</title>
<p>Using the three feature groups above, we define five feature configurations to isolate the incremental value of adding statistics and correlation structure on top of the <italic>&#x0394;Q&#x2212;V</italic> baseline.
<list list-type="bullet">
<list-item>
<p><bold><italic>&#x0394;Q Only</italic>.</bold> <italic>&#x0394;Q&#x2212;V</italic> HI (7) only. This is closest to the Severson-style <italic>&#x0394;Q&#x2212;V</italic> baseline and serves as a physics-informed reference.</p></list-item>
<list-item>
<p><bold><italic>&#x0394;Q &#x002B; Statistics</italic>.</bold> <italic>&#x0394;Q&#x2212;V</italic> HI (7) plus Stats (52), totaling 59 features. This tests whether generic statistics add complementary information beyond <italic>&#x0394;Q&#x2212;V</italic>.</p></list-item>
<list-item>
<p><bold><italic>Hybrid Append</italic>.</bold> <italic>&#x0394;Q&#x2212;V</italic> HI (7) plus Stats (52) plus Laplacian (20), totaling 79 features. This is the main proposed configuration.</p></list-item>
<list-item>
<p><bold><italic>VIF &#x002B; Laplacian</italic>.</bold> A compact variant in which multicollinear statistical features are pruned using Variance Inflation Factor (VIF), and the remaining statistics are combined with Laplacian (20), totaling 47 features.</p></list-item>
<list-item>
<p><bold><italic>Integrated Laplacian</italic></bold>. A comparison setting in which <italic>&#x0394;Q&#x2212;V</italic> HI and Stats are jointly used to recompute the correlation graph and Laplacian, producing 99 features. This contrasts with Hybrid Append, where <italic>&#x0394;Q&#x2212;V</italic> remains a separate channel.</p></list-item>
</list></p>
<p><bold>Regression model.</bold> We use XGBoost [<xref ref-type="bibr" rid="ref-21">21</xref>] for all feature sets. XGBoost is well-suited to tabular data, captures non-linear interactions, and is relatively parameter-efficient in small-sample regimes. Because tree-based models tolerate heterogeneous feature scales, we combine <italic>&#x0394;Q&#x2212;V</italic> HI, Stats, and Laplacian features without requiring additional normalization for the regressor.</p>
<p>Training and hyperparameter tuning. To enable fair comparison across feature sets, we use a common training protocol. We perform 5-fold cross-validation on the 41 training cells. Key hyperparameters (e.g., tree depth, learning rate, number of estimators, min_child_weight, subsample, colsample_bytree, and <italic>&#x03BB;</italic> regularization) are optimized using Optuna-based hyperparameter optimization, with cross-validation RMSE as the objective. The final model is retrained on the full training set using the selected hyperparameters and evaluated on the Primary and Secondary test sets. All scalers, correlation graphs, Laplacian bases, and PCA/RP/TSVD transforms were fit on the training cells only and then applied unchanged to the Primary and Secondary test splits.</p>
<p><xref ref-type="table" rid="table-4">Table 4</xref> summarizes the feature-set configurations.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Feature-set configurations used in the XGBoost-based early-life prediction experiments.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Feature set</th>
<th>Composition</th>
<th>#Features</th>
<th>Purpose</th>
</tr>
</thead>
<tbody>
<tr>
<td><italic>&#x0394;Q Only</italic></td>
<td><italic>&#x0394;Q-V</italic> HI</td>
<td>7</td>
<td>Severson-style baseline</td>
</tr>
<tr>
<td><italic>&#x0394;Q &#x002B; Statistics</italic></td>
<td><italic>&#x0394;Q-V</italic> HI &#x002B; Stats</td>
<td>59</td>
<td>Physics-informed &#x002B; Generic Statistics</td>
</tr>
<tr>
<td><italic>Hybrid Append</italic></td>
<td><italic>&#x0394;Q-V</italic> HI &#x002B; Stats &#x002B; Laplacian</td>
<td>79</td>
<td>Proposed hybrid configuration</td>
</tr>
<tr>
<td><italic>VIF &#x002B; Laplacian</italic></td>
<td>VIF-pruned Stats &#x002B; Laplacian</td>
<td>47</td>
<td>Compact structure-aware variant</td>
</tr>
<tr>
<td><italic>Integrated Laplacian</italic></td>
<td>(<italic>&#x0394;Q-V</italic> HI &#x002B; Stats) with recomputed Laplacian</td>
<td>99</td>
<td>Ablation: graph over all features</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Evaluation Metrics and PHM-Oriented Assessment</title>
<p>We evaluate each feature configuration using both standard regression metrics and PHM-oriented error-band metrics.</p>
<p><bold>Regression metrics.</bold> We report RMSE and <italic>R</italic><sup>2</sup> for pointwise cycle-life prediction accuracy.</p>
<p><bold>PHM-style error-band metrics.</bold> In addition to RMSE and <italic>R</italic><sup>2</sup>, we assess PHM-style behavior using trajectory-oriented metrics [<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-18">18</xref>]. For each cell, the model outputs a single early-life estimate of total cycle life. From this estimate we construct an implied RUL trajectory by subtracting the cycle index t from the predicted cycle life, while the true RUL is obtained by subtracting t from the observed cycle life. Because the model does not generate sequentially updated predictions, these metrics are applied to the implied trajectory induced by the fixed early-life estimate. We then compute relative error along the trajectory and report:<list list-type="bullet">
<list-item>
<p>RA@0.2: the fraction of time points where the relative error is within &#x00B1;20%;</p></list-item>
<list-item>
<p>PH@0.1: a stricter error-band agreement measure based on a &#x00B1;10% error tolerance;</p></list-item>
<list-item>
<p><italic>&#x03B1;-&#x03BB;</italic>(0.15): a summary score capturing where the relative error remains within a &#x00B1;15% band.</p></list-item>
</list></p>
<p>We report RMSE, <italic>R</italic><sup>2</sup>, and the three PHM metrics for both the Primary and Secondary test sets. This dual evaluation complements pointwise regression accuracy by adding an error-band view of performance under domain shift.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Results</title>
<p>We present quantitative results for the five feature configurations defined in <xref ref-type="sec" rid="s3_3">Section 3.3</xref> under the common training and evaluation protocol. The section first summarizes cross-validation on the Train set, then reports held-out performance on the Primary and Secondary test domains, and finally examines whether the Hybrid Append gain can be reproduced by simpler reduced-dimensional baselines.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Role of Cross-Validation in Model Selection</title>
<p>Cross-validation on the 41-cell training set was used for model selection and hyperparameter tuning. In this small-sample setting, richer feature configurations often appeared promising in CV, but the held-out ordering changed across the Primary and Secondary domains.</p>
<p>Accordingly, the main scientific interpretation in this section is anchored in the held-out test results and the accompanying bootstrap intervals. The additional matched-dimensionality comparisons and the k-sensitivity analysis then provide comparative and robustness evidence for the graph-derived descriptor.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Held-out Test Performance on the Primary and Secondary Domains</title>
<p><xref ref-type="table" rid="table-5">Tables 5</xref> and <xref ref-type="table" rid="table-6">6</xref> report the held-out test results for the five feature configurations. Taken together, they highlight the contrast between in-domain performance on Batches 1&#x2013;2 and shifted-domain generalization on Batch 3. The split difference is substantial not only protocol-wise (<xref ref-type="sec" rid="s3_1">Section 3.1</xref>) but also distributionally: the Primary split has median cycle life 520 cycles, whereas the Secondary split has median 964.5 cycles. The bootstrap intervals reported in <xref ref-type="table" rid="table-5">Tables 5</xref> and <xref ref-type="table" rid="table-6">6</xref> quantify uncertainty for RMSE and <italic>R</italic><sup>2</sup> only; RA, PH, and &#x03B1;-&#x03BB; are reported as point estimates.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Primary test set performance on Batch 1&#x2013;2 (43 cells). RMSE and R<sup>2</sup> values include 95% bootstrap confidence intervals.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Feature Set</th>
<th>RMSE [95% CI]</th>
<th><italic>R</italic><sup>2</sup> [95% CI]</th>
<th>RA@0.2</th>
<th>PH@0.1</th>
<th><italic>&#x03B1;&#x2013;&#x03BB;</italic>(0.15)</th>
</tr>
</thead>
<tbody>
<tr>
<td><italic>&#x0394;Q Only</italic></td>
<td>137.46 [101.29, 176.07]</td>
<td>0.6416 [0.2062, 0.8271]</td>
<td>69.8</td>
<td>46.5</td>
<td>65.1</td>
</tr>
<tr>
<td><italic>&#x0394;Q &#x002B; Statistics</italic></td>
<td>101.03 [74.80, 125.36]</td>
<td>0.8064 [0.6335, 0.8895]</td>
<td>83.7</td>
<td>55.8</td>
<td>74.4</td>
</tr>
<tr>
<td><italic>Hybrid Append</italic></td>
<td>102.28 [71.87, 128.95]</td>
<td>0.8016 [0.6035, 0.8969]</td>
<td>88.4</td>
<td>60.5</td>
<td>74.4</td>
</tr>
<tr>
<td><italic>VIF &#x002B; Laplacian</italic></td>
<td>110.27 [76.22, 144.57]</td>
<td>0.7694 [0.4654, 0.8950]</td>
<td>83.7</td>
<td>60.5</td>
<td>76.7</td>
</tr>
<tr>
<td><italic>Integrated Laplacian</italic></td>
<td>97.00 [71.87, 120.22]</td>
<td>0.8215 [0.6519, 0.8970]</td>
<td>88.4</td>
<td>55.8</td>
<td>74.4</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Secondary test set performance on Batch 3 (40 cells). RMSE and R<sup>2</sup> values include 95% bootstrap confidence intervals.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Feature Set</th>
<th>RMSE [95% CI]</th>
<th>R<sup>2</sup> [95% CI]</th>
<th>RA@0.2</th>
<th>PH@0.1</th>
<th>&#x03B1;&#x2013;&#x03BB;(0.15)</th>
</tr>
</thead>
<tbody>
<tr>
<td><italic>&#x0394;Q Only</italic></td>
<td>267.82 [194.18, 338.27]</td>
<td>0.2275 [&#x2212;0.1015, 0.3851]</td>
<td>62.5</td>
<td>32.5</td>
<td>50.0</td>
</tr>
<tr>
<td><italic>&#x0394;Q &#x002B; Statistics</italic></td>
<td>347.10 [216.05, 455.96]</td>
<td>&#x2212;0.2975 [&#x2212;0.7260, &#x2212;0.0149]</td>
<td>62.5</td>
<td>42.5</td>
<td>60.0</td>
</tr>
<tr>
<td><italic>Hybrid Append</italic></td>
<td>340.26 [213.31, 446.62]</td>
<td>&#x2212;0.2469 [&#x2212;0.6758, 0.0361]</td>
<td>62.5</td>
<td>45.0</td>
<td>60.0</td>
</tr>
<tr>
<td><italic>VIF &#x002B; Laplacian</italic></td>
<td>278.25 [159.31, 375.02]</td>
<td>0.1662 [&#x2212;0.0048, 0.3540]</td>
<td>70.0</td>
<td>40.0</td>
<td>60.0</td>
</tr>
<tr>
<td><italic>Integrated Laplacian</italic></td>
<td>366.51 [234.37, 476.46]</td>
<td>&#x2212;0.4467 [&#x2212;0.9442, &#x2212;0.1349]</td>
<td>62.5</td>
<td>40.0</td>
<td>47.5</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s4_2_1">
<label>4.2.1</label>
<title>Primary Test Set (Batch 1&#x2013;2)</title>
<p><xref ref-type="table" rid="table-5">Table 5</xref> shows that on the Primary test split, all augmented feature configurations improved over <italic>&#x0394;Q</italic> Only. No single method dominated every metric. Integrated Laplacian achieved the lowest RMSE and highest R<sup>2</sup> (97.00 cycles, 0.8215; 95% bootstrap CIs: [71.87, 120.22] and [0.6519, 0.8970]), while Hybrid Append remained competitive in pointwise error (102.28 cycles, 0.8016; [71.87, 128.95] and [0.6035, 0.8969]) and delivered one of the stronger PHM-style profiles (RA@0.2/PH@0.1/<italic>&#x03B1;-&#x03BB;</italic>(0.15) &#x003D; 88.4/60.5/74.4). Relative to <italic>&#x0394;Q</italic> &#x002B; Statistics, Hybrid Append preserved very similar RMSE/R<sup>2</sup> while improving RA@0.2 and PH@0.1, suggesting that the appended structural channel was more visible in PHM-style error-band metrics than in pointwise error alone. Although the Integrated Laplacian achieved the strongest Primary RMSE/<italic>R</italic><sup>2</sup> pair, the Hybrid Append design remains important because it preserves the explicit <italic>&#x0394;Q</italic> branch while adding a competitive structural descriptor. In this sense, the Primary split supports Hybrid Append as a practically well-balanced appended design rather than as a uniformly dominant configuration.</p>

</sec>
<sec id="s4_2_2">
<label>4.2.2</label>
<title>Secondary Test Set (Batch 3, Domain Shift)</title>
<p><xref ref-type="table" rid="table-6">Table 6</xref> shows that on the Secondary test split, <italic>&#x0394;Q</italic> only gave the most favorable RMSE/<italic>R</italic><sup>2</sup> pair among the tested configurations, with RMSE/<italic>R</italic><sup>2</sup> of 267.82/0.2275 (95% bootstrap CIs: [194.18, 338.27] and [&#x2212;0.1015, 0.3851]). VIF &#x002B; Laplacian was the next most competitive shifted-domain configuration in <italic>R</italic><sup>2</sup> (0.1662) and achieved the highest RA@0.2 (70.0). Hybrid Append improved over <italic>&#x0394;Q</italic> &#x002B; Statistics and Integrated Laplacian in both RMSE and <italic>R</italic><sup>2</sup> and reached the highest PH@0.1 (45.0), but it did not exceed <italic>&#x0394;Q</italic> Only on RMSE or R<sup>2</sup>. Under the shifted Batch 3 domain, the sparse <italic>&#x0394;Q</italic>-only baseline remained the most robust pointwise regressor, whereas Hybrid Append and VIF &#x002B; Laplacian improved selected PHM threshold metrics. This contrast is informative for the proposed design: while the strongest shifted-domain RMSE/<italic>R</italic><sup>2</sup> remained with the sparse physics-informed baseline, Hybrid Append retained value as the study&#x2019;s principal hybrid configuration by preserving the explicit <italic>&#x0394;Q</italic> channel and improving selected error-band metrics under shift.</p>

<p>The integrated-graph ablation is also revealing: although Integrated Laplacian was strong on the Primary split, it generalized poorly to Batch 3, which favors keeping <italic>&#x0394;Q</italic> as a separate physics-informed branch rather than absorbing it into the graph.</p>
</sec>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Comparison against Simpler Reduced-Dimensional Baselines</title>
<p>To examine whether the Hybrid Append behavior could be reproduced by generic dimensionality reduction alone, we conducted an additional comparison against three non-graph alternatives at a matched 79-feature scale. These alternatives were obtained by appending 20 features from PCA, Gaussian Random Projection, and Truncated SVD to the 59-feature <italic>&#x0394;Q</italic> &#x002B; Statistics baseline, as shown in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>. This supplementary analysis helps interpret the appended structural descriptor alongside the main five-configuration held-out results.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Comparison at a matched 79-feature scale against simpler reduced-dimensional baselines. Primary and Secondary test <italic>R</italic><sup>2</sup> are shown for <italic>&#x0394;Q</italic> &#x002B; Statistics (59 features), Hybrid Append (79 features), and three 79-feature alternatives obtained by appending 20 features from PCA, Random Projection, and Truncated SVD. All learned transforms were estimated from the training cells only.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81667-fig-4.tif"/>
</fig>
<p>All compared variants were evaluated under a common train/Primary/Secondary split and XGBoost protocol, and every learned transform was fit on the training cells only before being applied unchanged to the Primary and Secondary test splits.</p>
<p>In this comparison, with all learned transforms estimated from the training cells and then applied unchanged to the test splits, Hybrid Append achieved CV <italic>R</italic><sup>2</sup> 0.7497, Primary <italic>R</italic><sup>2</sup> 0.8185, and Secondary <italic>R</italic><sup>2</sup> &#x2212;0.3205. The best matched-dimensionality non-graph baseline reached Primary <italic>R</italic><sup>2</sup> 0.8107 and Secondary <italic>R</italic><sup>2</sup> &#x2212;0.1955, while &#x002B;RandomProjection20 achieved the highest CV <italic>R</italic><sup>2</sup>. The comparison is therefore mixed rather than uniform: the graph-derived embedding remained competitive, but its margin over simpler reductions at the same feature dimensionality depended on the evaluation split. These results instead position Hybrid Append as the study&#x2019;s principal hybrid configuration, whose added value is best understood as complementary and split-dependent.</p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>PHM-Oriented Analysis of Implied RUL Trajectories</title>
<p><xref ref-type="fig" rid="fig-5">Fig. 5</xref> compares RA@0.2, PH@0.1, and <italic>&#x03B1;-&#x03BB;</italic> (0.15) across feature sets on both test domains. These PHM-style error-band metrics highlight differences that are not fully captured by RMSE and <italic>R</italic><sup>2</sup> alone. For compact visualization, <xref ref-type="fig" rid="fig-5">Fig. 5</xref> displays the scores on a 0&#x2013;1 scale, whereas <xref ref-type="table" rid="table-5">Tables 5</xref> and <xref ref-type="table" rid="table-6">6</xref> report the corresponding values in percentage form. Because each model produces a single early-life cycle-life estimate per cell, the scores are computed on the implied RUL trajectory defined in <xref ref-type="sec" rid="s3_4">Section 3.4</xref>.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>PHM-oriented error-band metrics on the implied RUL trajectories for the Primary and Secondary test sets: (<bold>a</bold>) RA@0.2 (&#x00B1;20% tolerance), (<bold>b</bold>) PH@0.1 (&#x00B1;10% tolerance), and (<bold>c</bold>) <italic>&#x03B1;-&#x03BB;</italic> (0.15) (&#x00B1;15% tolerance) for <italic>&#x0394;Q Only, &#x0394;Q &#x002B; Statistics, Hybrid Append, VIF &#x002B; Laplacian, and Integrated Laplacian</italic>. Scores are shown as proportions on a 0&#x2013;1 scale; the dashed line indicates the ideal score (1.0).</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81667-fig-5.tif"/>
</fig>
<p>On the Primary split, Hybrid Append and Integrated Laplacian tie for the highest RA@0.2 (88.4%), Hybrid Append and VIF &#x002B; Laplacian tie for the highest PH@0.1 (60.5%), and VIF &#x002B; Laplacian achieves the highest <italic>&#x03B1;-&#x03BB;</italic>(0.15) (76.7%). On the Secondary split, VIF &#x002B; Laplacian reaches the highest RA@0.2 (70.0), Hybrid Append reaches the highest PH@0.1 (45.0), and <italic>&#x0394;Q</italic> &#x002B; Statistics, Hybrid Append, and VIF &#x002B; Laplacian tie on <italic>&#x03B1;-&#x03BB;</italic>(0.15) (60.0).</p>
<p>Viewed together with the held-out RMSE/<italic>R</italic><sup>2</sup> results, the PHM-style metrics suggest that the contribution of the appended structural channel is not expressed identically across evaluation criteria. Thus, no single configuration dominates across all criteria; instead, the preferred feature set depends on whether the emphasis is on pointwise regression error or stricter error-band agreement. In particular, the Hybrid Append configuration is most naturally interpreted as improving selected error-band behavior while maintaining competitive pointwise accuracy on the in-domain split and improving PH@0.1 relative to <italic>&#x0394;Q</italic> &#x002B; Statistics on both test domains (60.5 vs. 55.8 on Primary; 45.0 vs. 42.5 on Secondary).</p>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Visualization of Prediction Distributions</title>
<p>To complement the aggregate metrics, <xref ref-type="fig" rid="fig-6">Fig. 6</xref> plots observed vs. predicted cycle life for the Primary and Secondary test cells across the five feature configurations. The Primary-domain points cluster much closer to the identity line than the Secondary points, consistent with the protocol differences summarized in <xref ref-type="table" rid="table-1">Table 1</xref> and the shifted-domain performance reported in <xref ref-type="table" rid="table-6">Table 6</xref>.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Observed vs. predicted cycle life for the five feature configurations: (<bold>a</bold>) <italic>&#x0394;Q Only</italic>, (<bold>b</bold>) <italic>&#x0394;Q</italic> &#x002B; <italic>Statistics</italic>, (<bold>c</bold>) <italic>Hybrid Append</italic>, (<bold>d</bold>) <italic>VIF &#x002B; Laplacian</italic>, and (<bold>e</bold>) <italic>Integrated Laplacian</italic>. Red squares and green triangles denote Primary Test and Secondary Test cells, respectively. The dashed line indicates perfect prediction, and shaded bands represent &#x00B1;10% and &#x00B1;20% relative-error regions.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81667-fig-6.tif"/>
</fig>
<p>Among the augmented feature sets, the Primary scatter for <italic>&#x0394;Q</italic> &#x002B; Statistics, Hybrid Append, and Integrated Laplacian is comparatively tight, consistent with the stronger in-domain RMSE/<italic>R</italic><sup>2</sup> values in <xref ref-type="table" rid="table-5">Table 5</xref>. The broader Secondary dispersion of the integrated variant mirrors its weaker shifted-domain behavior, while Hybrid Append remains visually close to <italic>&#x0394;Q</italic> &#x002B; Statistics in pointwise scatter but achieves slightly stronger PH@0.1. Viewed together with <xref ref-type="table" rid="table-5">Tables 5</xref> and <xref ref-type="table" rid="table-6">6</xref> and <xref ref-type="fig" rid="fig-5">Fig. 5</xref>, these plots indicate that added feature channels help in-domain, whereas shifted-domain behavior remains more metric-dependent.</p>

</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusions and Future Work</title>
<p>This paper examined early-life cycle-life prediction for lithium-ion batteries in a small-sample setting, focusing on whether an explicit correlation-structure descriptor can complement physics-informed <italic>&#x0394;Q</italic> indicators and generic early-cycle statistical features. The graph component was implemented as a fixed spectral feature-engineering step derived from a feature-correlation graph.</p>
<p>Across the five main feature configurations, all augmented feature sets improved over <italic>&#x0394;Q Only</italic> on the in-domain Primary split, but no single configuration dominated every metric. Integrated Laplacian achieved the strongest Primary RMSE/<italic>R</italic><sup>2</sup> pair, whereas Hybrid Append remained competitive in pointwise accuracy and delivered one of the stronger PHM-style profiles. Relative to <italic>&#x0394;Q &#x002B; Statistics</italic>, Hybrid Append preserved similar Primary RMSE/<italic>R</italic><sup>2</sup> while improving RA@0.2 and PH@0.1, which highlights the practical value of the appended structural channel.</p>
<p>Under the shifted Batch 3 domain, the sparse <italic>&#x0394;Q-only</italic> baseline remained the most robust pointwise regressor, while Hybrid Append and <italic>VIF &#x002B; Laplacian</italic> improved selected PHM threshold metrics. Taken together, these findings support Hybrid Append as the study&#x2019;s principal hybrid configuration: it preserves the explicit <italic>&#x0394;Q</italic> branch, remains competitive on the Primary split, and provides a more favorable appended design than the integrated-graph alternative when error-band behavior under shift is also considered.</p>
<p>The additional comparison at a matched 79-feature scale against PCA, Random Projection, and Truncated SVD further refines this interpretation. In those comparisons, all learned transforms were estimated from the training cells and then applied unchanged to the test splits. The graph-derived embedding remained competitive, but its margin over simpler reductions varied by split. The graph component is therefore best understood not as a universal substitute for generic dimensionality reduction, but as a complementary structural descriptor within the Hybrid Append design.</p>
<p>Overall, this work contributes (i) a small-sample case study of combining physics-informed indicators, generic statistical features, and graph-derived structural descriptors for early-life battery prognostics, (ii) a feature-design workflow that keeps the <italic>&#x0394;Q</italic> channel explicit while adding an explicit correlation-structure descriptor, and (iii) a joint evaluation using regression and PHM-style error-band metrics under a within-benchmark batch-shift setting. In practical terms, the present evidence supports using the appended correlation-structure descriptor as a supplementary design option when PHM-style tolerance behavior is of interest, while pointwise shifted-domain accuracy should still be benchmarked against sparse physics-informed baselines.</p>
<p>This work also has limitations. Experiments were conducted on a single public benchmark (the Severson 124-cell LFP/graphite dataset) with a specific early-life window (cycles 1&#x2013;100) and a batch-defined domain shift. Both the structural descriptor and the PHM-style evaluation remain intentionally simple at this stage: the graph descriptor is based on Pearson correlation, and the PHM-style metrics are computed from the implied RUL trajectory induced by a single early-life estimate per cell rather than from sequentially updated online predictions. The present results therefore establish within-benchmark behavior for this early-life setting.</p>
<p>Future work will extend the framework in several directions. First, we will evaluate robustness across additional chemistries, protocols, and datasets. Second, we will investigate alternative correlation-structure constructions, such as sparsified graphs, partial-correlation, or regularized dependence measures, to improve structural stability under protocol changes. Third, we will incorporate uncertainty-aware prediction and assess PHM performance with explicit uncertainty quantification. Finally, we will examine model-selection and evaluation strategies for sequential or online prediction settings while maintaining interpretability and small-sample practicality.</p>
</sec>
</body>
<back>
<ack>
<p>None.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This research was funded by the National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIT) (No. NRF-2022R1C1C1011743).</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: study conception and design: Kang-Woo Lee; data collection: Kang-Woo Lee; analysis and interpretation of results: Kang-Woo Lee; draft manuscript preparation: Kang-Woo Lee; supervision and manuscript review: Dong-Hee Lee and Dae-Il Kwon. All authors reviewed and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The data supporting the findings of this study are publicly available from the Severson lithium-ion battery dataset at <ext-link ext-link-type="uri" xlink:href="https://data.matr.io/1/">https://data.matr.io/1/</ext-link>.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest.</p>
</sec>
<app-group id="appg-1">
<app id="app-1">
<title>Appendix A</title>
<table-wrap id="table-7">
<label>Table A1 </label>
<caption>
<title>Complete grouped list of the 52 statistical features used in the data-driven point-feature set.</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Feature Family</th>
<th>Exact Feature Names</th>
<th>Count</th>
</tr>
</thead>
<tbody>
<tr>
<td>Current-derived</td>
<td><italic>I_mean, I_std, I_min, I_max, I</italic><sup><italic>2</italic></sup><italic>_mean, I</italic><sup><italic>2</italic></sup><italic>_std, I</italic><sup><italic>2</italic></sup><italic>_min, I</italic><sup><italic>2</italic></sup><italic>_max</italic></td>
<td>8</td>
</tr>
<tr>
<td>Time-derived</td>
<td><italic>t</italic><sup><italic>2</italic></sup><italic>_mean, t</italic><sup><italic>2</italic></sup><italic>_std, t</italic><sup><italic>2</italic></sup><italic>_min, t</italic><sup><italic>2</italic></sup><italic>_max, logt_mean, logt_std, logt_min, logt_max</italic></td>
<td>8</td>
</tr>
<tr>
<td>Capacity-derived</td>
<td><italic>Qdlin</italic><sup><italic>2</italic></sup><italic>_mean, Qdlin</italic><sup><italic>2</italic></sup><italic>_std, Qdlin</italic><sup><italic>2</italic></sup><italic>_min, Qdlin</italic><sup><italic>2</italic></sup><italic>_max, Qd/t_mean, Qd/t_std, Qd/t_min, Qd/t_max, Qdlin/Qd_mean, Qdlin/Qd_std, Qdlin/Qd_min, Qdlin/Qd_max, logQdlin_mean, logQdlin_std, logQdlin_min, logQdlin_max</italic></td>
<td>16</td>
</tr>
<tr>
<td>Differential-capacity&#x2013;derived</td>
<td><italic>(dQ/dV)</italic><sup><italic>2</italic></sup><italic>_mean, (dQ/dV)</italic><sup><italic>2</italic></sup><italic>_std, (dQ/dV)</italic><sup><italic>2</italic></sup><italic>_min, (dQ/dV)</italic><sup><italic>2</italic></sup><italic>_max, I/(dQ/dV)_mean, I/(dQ/dV)_std, I/(dQ/dV)_min, I/(dQ/dV)_max, Qdlin &#x00D7; dQ/dV_mean, Qdlin &#x00D7; dQ/dV_std, Qdlin &#x00D7; dQ/dV_min, Qdlin &#x00D7; dQ/dV_max</italic></td>
<td>12</td>
</tr>
<tr>
<td>Interaction-derived</td>
<td><italic>I &#x00D7; Qd_mean, I &#x00D7; Qd_std, I &#x00D7; Qd_min, I &#x00D7; Qd_max, Qd &#x00D7; t_mean, Qd &#x00D7; t_std, Qd &#x00D7; t_min, Qd &#x00D7; t_max</italic></td>
<td>8</td>
</tr>
<tr><td></td>
<td><bold>Total</bold></td>
<td><bold>52</bold></td>
</tr>
</tbody>
</table>
</table-wrap>
</app>
<app id="app-2">
<title>Appendix B</title>
<fig id="fig-7">
<label>Figure A1 </label>
<caption>
<title>Train-only sensitivity analysis of the Laplacian embedding dimension <italic>k</italic>. CV, primary, and secondary performance are shown across multiple values of <italic>k</italic> using <italic>R</italic><sup>2</sup> (left) and RMSE (right). Mid-range values perform comparably, but the best observed values differ by metric and split; <italic>k</italic> &#x003D; 20 is therefore presented as a practical mid-range operating point rather than as a unique optimum.</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_81667-fig-7.tif"/>
</fig>
</app>
</app-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Che</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>X</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>J</given-names></string-name>, <string-name><surname>Teodorescu</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Health prognostics for lithium-ion batteries: mechanisms, methods, and prospects</article-title>. <source>Energy Environ Sci</source>. <year>2023</year>;<volume>16</volume>(<issue>2</issue>):<fpage>338</fpage>&#x2013;<lpage>71</lpage>. doi:<pub-id pub-id-type="doi">10.1039/d2ee03019e</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Thelen</surname> <given-names>A</given-names></string-name>, <string-name><surname>Huan</surname> <given-names>X</given-names></string-name>, <string-name><surname>Paulson</surname> <given-names>N</given-names></string-name>, <string-name><surname>Onori</surname> <given-names>S</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Probabilistic machine learning for battery health diagnostics and prognostics&#x2014;review and perspectives</article-title>. <source>npj Mater Sustain</source>. <year>2024</year>;<volume>2</volume>(<issue>1</issue>):<fpage>14</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s44296-024-00011-1</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Severson</surname> <given-names>KA</given-names></string-name>, <string-name><surname>Attia</surname> <given-names>PM</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>N</given-names></string-name>, <string-name><surname>Perkins</surname> <given-names>N</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>Z</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Data-driven prediction of battery cycle life before capacity degradation</article-title>. <source>Nat Energy</source>. <year>2019</year>;<volume>4</volume>(<issue>5</issue>):<fpage>383</fpage>&#x2013;<lpage>91</lpage>. doi:<pub-id pub-id-type="doi">10.1038/s41560-019-0356-8</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Early lifetime prediction of lithium-ion batteries based on classical image encoding methods</article-title>. <source>Energy</source>. <year>2025</year>;<volume>336</volume>(<issue>2</issue>):<fpage>138372</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.energy.2025.138372</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Fu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>B</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>X</given-names></string-name>, <string-name><surname>Dai</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Early remaining useful life prediction for lithium-ion batteries using a Gaussian process regression model based on degradation pattern recognition</article-title>. <source>Batteries</source>. <year>2025</year>;<volume>11</volume>(<issue>6</issue>):<fpage>221</fpage>. doi:<pub-id pub-id-type="doi">10.3390/batteries11060221</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Ultra-early prediction of lithium-ion battery cycle life based on assembled capacity curve extracted from a single cycle</article-title>. <source>J Power Sources</source>. <year>2025</year>;<volume>640</volume>(<issue>3</issue>):<fpage>236620</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jpowsour.2025.236620</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nascimento</surname> <given-names>RG</given-names></string-name>, <string-name><surname>Viana</surname> <given-names>FAC</given-names></string-name>, <string-name><surname>Corbetta</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kulkarni</surname> <given-names>CS</given-names></string-name></person-group>. <article-title>A framework for Li-ion battery prognosis based on hybrid Bayesian physics-informed neural networks</article-title>. <source>Sci Rep</source>. <year>2023</year>;<volume>13</volume>(<issue>1</issue>):<fpage>13856</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41598-023-33018-0</pub-id>; <pub-id pub-id-type="pmid">37620364</pub-id></mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ma</surname> <given-names>L</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Accurate and efficient remaining useful life prediction of batteries enabled by physics-informed machine learning</article-title>. <source>J Energy Chem</source>. <year>2024</year>;<volume>91</volume>:<fpage>512</fpage>&#x2013;<lpage>21</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jechem.2023.12.043</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shi</surname> <given-names>J</given-names></string-name>, <string-name><surname>Rivera</surname> <given-names>A</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Battery health management using physics-informed machine learning: online degradation modeling and remaining useful life prediction</article-title>. <source>Mech Syst Signal Process</source>. <year>2022</year>;<volume>179</volume>(<issue>7</issue>):<fpage>109347</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ymssp.2022.109347</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>F</given-names></string-name>, <string-name><surname>Zhai</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Di</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Physics-informed neural network for lithium-ion battery degradation stable modeling and prognosis</article-title>. <source>Nat Commun</source>. <year>2024</year>;<volume>15</volume>(<issue>1</issue>):<fpage>4332</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41467-024-48779-z</pub-id>; <pub-id pub-id-type="pmid">38773131</pub-id></mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>DGL-STFA: predicting lithium-ion battery health with dynamic graph learning and spatial-temporal fusion attention</article-title>. <source>Energy AI</source>. <year>2025</year>;<volume>19</volume>(<issue>7</issue>):<fpage>100462</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.egyai.2024.100462</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>E</given-names></string-name></person-group>. <article-title>Correlation based-graph neural network for health prognosis of non-fully charged and discharged lithium-ion batteries</article-title>. <source>J Power Sources</source>. <year>2025</year>;<volume>629</volume>(<issue>2</issue>):<fpage>235984</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jpowsour.2024.235984</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wei</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Prediction of state of health and remaining useful life of lithium-ion battery using graph convolutional network with dual attention mechanisms</article-title>. <source>Reliab Eng Syst Saf</source>. <year>2023</year>;<volume>230</volume>(<issue>1</issue>):<fpage>108947</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ress.2022.108947</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>J</given-names></string-name></person-group>. <article-title>State of health estimation for batteries based on a dynamic graph pruning neural network with a self-attention mechanism</article-title>. <source>Energies</source>. <year>2025</year>;<volume>18</volume>(<issue>20</issue>):<fpage>5333</fpage>. doi:<pub-id pub-id-type="doi">10.3390/en18205333</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Luan</surname> <given-names>W</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>B</given-names></string-name></person-group>. <article-title>State of health for lithium-ion batteries based on explainable feature fragments via graph attention network and bi-directional gated recurrent unit</article-title>. <source>Sensors</source>. <year>2025</year>;<volume>25</volume>(<issue>19</issue>):<fpage>5953</fpage>. doi:<pub-id pub-id-type="doi">10.3390/s25195953</pub-id>; <pub-id pub-id-type="pmid">41094776</pub-id></mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Geslin</surname> <given-names>A</given-names></string-name>, <string-name><surname>van Vlijmen</surname> <given-names>B</given-names></string-name>, <string-name><surname>Cui</surname> <given-names>X</given-names></string-name>, <string-name><surname>Bhargava</surname> <given-names>A</given-names></string-name>, <string-name><surname>Asinger</surname> <given-names>PA</given-names></string-name>, <string-name><surname>Braatz</surname> <given-names>RD</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Selecting the appropriate features in battery lifetime predictions</article-title>. <source>Joule</source>. <year>2023</year>;<volume>7</volume>(<issue>9</issue>):<fpage>1956</fpage>&#x2013;<lpage>65</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.joule.2023.07.021</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Saxena</surname> <given-names>A</given-names></string-name>, <string-name><surname>Celaya</surname> <given-names>J</given-names></string-name>, <string-name><surname>Saha</surname> <given-names>B</given-names></string-name>, <string-name><surname>Saha</surname> <given-names>S</given-names></string-name>, <string-name><surname>Goebel</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Metrics for offline evaluation of prognostic performance</article-title>. <source>Int J Progn Health Manag</source>. <year>2010</year>;<volume>1</volume>(<issue>1</issue>):<fpage>4</fpage>&#x2013;<lpage>23</lpage>. doi:<pub-id pub-id-type="doi">10.36001/ijphm.2010.v1i1.1336</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sharp</surname> <given-names>ME</given-names></string-name></person-group>. <article-title>Simple metrics for evaluating and conveying prognostic model performance to users with varied backgrounds</article-title>. <source>Annu Conf PHM Soc</source>. <year>2013</year>;<volume>5</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>10</lpage>. doi:<pub-id pub-id-type="doi">10.36001/phmconf.2013.v5i1.2317</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lee Rodgers</surname> <given-names>J</given-names></string-name>, <string-name><surname>Nicewander</surname> <given-names>WA</given-names></string-name></person-group>. <article-title>Thirteen ways to look at the correlation coefficient</article-title>. <source>Am Stat</source>. <year>1988</year>;<volume>42</volume>(<issue>1</issue>):<fpage>59</fpage>&#x2013;<lpage>66</lpage>. doi:<pub-id pub-id-type="doi">10.1080/00031305.1988.10475524</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Belkin</surname> <given-names>M</given-names></string-name>, <string-name><surname>Niyogi</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Laplacian eigenmaps for dimensionality reduction and data representation</article-title>. <source>Neural Comput</source>. <year>2003</year>;<volume>15</volume>(<issue>6</issue>):<fpage>1373</fpage>&#x2013;<lpage>96</lpage>. doi:<pub-id pub-id-type="doi">10.1162/089976603321780317</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>T</given-names></string-name>, <string-name><surname>Guestrin</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Xgboost: a scalable tree boosting system</article-title>. In: <conf-name>Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; 2016 Aug 13&#x2013;17; San Francisco, CA, USA</conf-name>. p. <fpage>785</fpage>&#x2013;<lpage>94</lpage>. doi:<pub-id pub-id-type="doi">10.1145/2939672.2939785</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>N</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chai</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Early prediction of battery lifetime based on graphical features and convolutional neural networks</article-title>. <source>Appl Energy</source>. <year>2024</year>;<volume>353</volume>:<fpage>122048</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.apenergy.2023.122048</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>T</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Thelen</surname> <given-names>A</given-names></string-name>, <string-name><surname>Howey</surname> <given-names>DA</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Predicting battery lifetime under varying usage conditions from early aging data</article-title>. <source>Cell Rep Phys Sci</source>. <year>2024</year>;<volume>5</volume>(<issue>4</issue>):<fpage>101891</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.xcrp.2024.101891</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zheng</surname> <given-names>S</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Gui</surname> <given-names>X</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>W</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Battery lifetime prediction across diverse ageing conditions with inter-cell deep learning</article-title>. <source>Nat Mach Intell</source>. <year>2025</year>;<volume>7</volume>(<issue>2</issue>):<fpage>270</fpage>&#x2013;<lpage>7</lpage>. doi:<pub-id pub-id-type="doi">10.1038/s42256-024-00972-x</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>KQ</given-names></string-name>, <string-name><surname>Qin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yuen</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Graph neural network-based lithium-ion battery state of health estimation using partial discharging curve</article-title>. <source>J Energy Storage</source>. <year>2024</year>;<volume>100</volume>(<issue>11</issue>):<fpage>113502</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.est.2024.113502</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>KQ</given-names></string-name>, <string-name><surname>Qin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Bhattacharjee</surname> <given-names>A</given-names></string-name>, <string-name><surname>Tushar</surname> <given-names>W</given-names></string-name>, <string-name><surname>Chong</surname> <given-names>ACC</given-names></string-name>, <string-name><surname>Yuen</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Health EquiMetrics for battery assessment&#x2014;a novel metric for end-of-life evaluation beyond state-of-health</article-title>. <source>J Energy Storage</source>. <year>2026</year>;<volume>147</volume>(<issue>2</issue>):<fpage>120064</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.est.2025.120064</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhou</surname> <given-names>KQ</given-names></string-name>, <string-name><surname>Qin</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Yuen</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Lithium-ion battery state of health estimation by matrix profile empowered online knee onset identification</article-title>. <source>IEEE Trans Transp Electrif</source>. <year>2024</year>;<volume>10</volume>(<issue>1</issue>):<fpage>1935</fpage>&#x2013;<lpage>46</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TTE.2023.3265981</pub-id>; <pub-id pub-id-type="pmid">25079929</pub-id></mixed-citation></ref>
</ref-list>
</back></article>