<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">75957</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2026.075957</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Multi-Algorithm Machine Learning Framework for Predicting Crystal Structures of Lithium Manganese Silicate Cathodes Using DFT Data</article-title>
<alt-title alt-title-type="left-running-head">Multi-Algorithm Machine Learning Framework for Predicting Crystal Structures of Lithium Manganese Silicate Cathodes Using DFT Data</alt-title>
<alt-title alt-title-type="right-running-head">Multi-Algorithm Machine Learning Framework for Predicting Crystal Structures of Lithium Manganese Silicate Cathodes Using DFT Data</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Ishtiaq</surname><given-names>Muhammad</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Lee</surname><given-names>Yeon-Ju</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Bhavani</surname><given-names>Annabathini Geetha</given-names></name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<contrib id="author-4" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Kang</surname><given-names>Sung-Gyu</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><email>s.kang@gnu.ac.kr</email></contrib>
<contrib id="author-5" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Reddy</surname><given-names>Nagireddy Gari Subba</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><email>nsreddy@gnu.ac.kr</email></contrib>
<aff id="aff-1"><label>1</label><institution>Department of Materials Engineering and Convergence Technology, Gyeongsang National University, 501 Jinju-Daero</institution>, <addr-line>Jinju, 52828</addr-line>, <country>Republic of Korea</country></aff>
<aff id="aff-2"><label>2</label><institution>School of Materials Science and Engineering, Engineering Research Institute, Gyeongsang National University, 501 Jinju-Daero</institution>, <addr-line>Jinju, 52828</addr-line>, <country>Republic of Korea</country></aff>
<aff id="aff-3"><label>3</label><institution>Department of Chemistry, SRM Institute of Science and Technology, Delhi-NCR Campus, Delhi-Meerut Road</institution>, <addr-line>Modinagar, Ghaziabad, 201204, Uttar Pradesh</addr-line>, <country>India</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Authors: Sung-Gyu Kang. Email: <email>s.kang@gnu.ac.kr</email>; Nagireddy Gari Subba Reddy. Email: <email>nsreddy@gnu.ac.kr</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2026</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>10</day><month>2</month><year>2026</year>
</pub-date>
<volume>87</volume>
<issue>1</issue>
<elocation-id>21</elocation-id>
<history>
<date date-type="received">
<day>11</day>
<month>11</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>05</day>
<month>01</month>
<year>2026</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2026 The Authors.</copyright-statement>
<copyright-year>2026</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_75957.pdf"></self-uri>
<abstract>
<p>Lithium manganese silicate (Li-Mn-Si-O) cathodes are key components of lithium-ion batteries, and their physical and mechanical properties are strongly influenced by their underlying crystal structures. In this study, a range of machine learning (ML) algorithms were developed and compared to predict the crystal systems of Li-Mn-Si-O cathode materials using density functional theory (DFT) data obtained from the Materials Project database. The dataset comprised 211 compositions characterized by key descriptors, including formation energy, energy above the hull, bandgap, atomic site number, density, and unit cell volume. These features were utilized to classify the materials into monoclinic (0) and triclinic (1) crystal systems. A comprehensive comparison of various classification algorithms including Decision Tree, Random Forest, XGBoost, Support Vector Machine, k-Nearest Neighbor, Stochastic Gradient Descent, Gaussian Na&#x00EF;ve Bayes, Gaussian Process, and Artificial Neural Network (ANN) was conducted. Among these, the optimized ANN architecture (6&#x2013;14-14-14-1) exhibited the highest predictive performance, achieving an accuracy of 95.3%, a Matthews correlation coefficient (MCC) of 0.894, and an F-score of 0.963, demonstrating excellent consistency with DFT-predicted crystal structures. Meanwhile, Random Forest and Gaussian Process models also exhibited reliable and consistent predictive capability, indicating their potential as complementary approaches, particularly when data are limited or computational efficiency is required. This comparative framework provides valuable insights into model selection for crystal system classification in complex cathode materials.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Machine learning</kwd>
<kwd>crystal structure</kwd>
<kwd>classification</kwd>
<kwd>cathode materials: batteries</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>National Research Foundation of Korea (NRF) grant funded by the Ministry of Education</funding-source>
<award-id>RS-2023-00301974</award-id>
</award-group>
<award-group id="awg2">
<funding-source>Glocal University 30 Project fund of Gyeongsang National University</funding-source>
<award-id>2025</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>The ever-increasing global population and the ease of transportation have necessitated continuous advancements in transportation systems, particularly in the automotive sector. The growing demand for electric vehicles has intensified efforts to develop smart batteries capable of operating for extended periods with minimal charging times [<xref ref-type="bibr" rid="ref-1">1</xref>]. Among various energy storage technologies, lithium-ion batteries have attracted significant research interest due to their superior performance characteristics [<xref ref-type="bibr" rid="ref-2">2</xref>]. In this context, lithium manganese silicate Li-M-Si-O (M&#x003D;Fe, Mn, and Co) cathodes have emerged as promising candidates for next-generation lithium-ion batteries [<xref ref-type="bibr" rid="ref-3">3</xref>]. The silicate family is known to be rich in polyforms such as Li<sub>2</sub>FeSiO<sub>4</sub>, Li<sub>2</sub>CoSiO<sub>4</sub>, and Li<sub>2</sub>MnSiO<sub>4</sub>. Among them, Li<sub>2</sub>MnSiO<sub>4</sub> has a higher capacity than commercial batteries due to its oxidation potential (4.2 and 4.4 V) [<xref ref-type="bibr" rid="ref-4">4</xref>]. The polymorphs of Li<sub>2</sub>MnSiO<sub>4</sub> are orthorhombic (Pmnb and Pmn21), monoclinic (P21/n) [<xref ref-type="bibr" rid="ref-5">5</xref>]. However, when polymorphs undergo crystal structure degradation due to instability during the delithiation process, a significant capacity loss occurs during subsequent cycling. To solve the problem, we need to predict the main factor of the stable crystal system of Li<sub>2</sub>MnSiO<sub>4</sub> in delithiated state [<xref ref-type="bibr" rid="ref-6">6</xref>]. Extensive experimental efforts have been devoted to investigating the crystal structures of cathode and anode materials used in batteries. For example, Luo et al. [<xref ref-type="bibr" rid="ref-7">7</xref>] examined the structural characteristics of Li ion battery electrode materials using neutron diffraction, while Nowakowski et al. [<xref ref-type="bibr" rid="ref-8">8</xref>] studied the influence of crystallographic orientation on Li metal anodes. Although such studies provide valuable insights, they require substantial time, specialized expertise, advanced instrumentation, and high-purity materials, which make experimental exploration both resource intensive and costly.</p>
<p>The Materials Project [<xref ref-type="bibr" rid="ref-9">9</xref>] provides an open, web-based database that enables the calculation of physical and chemical properties of both known and predicted materials using density functional theory (DFT). Researchers can access valuable data related to cathode materials; however, the extensive datasets can sometimes be confusing or misleading. Therefore, there is a need for specialized algorithms capable of accurately identifying complex, multifaceted correlations that are difficult to capture using traditional statistical methods. Machine learning (ML) methods have been extensively employed to predict various structures and properties of materials in the field of materials science and engineering [<xref ref-type="bibr" rid="ref-10">10</xref>,<xref ref-type="bibr" rid="ref-11">11</xref>]. Various studies have utilized predictive models to estimate the discharging capacities [<xref ref-type="bibr" rid="ref-12">12</xref>], and health state of Li-ion batteries [<xref ref-type="bibr" rid="ref-13">13</xref>]. Wang and Jiang demonstrated the successful prediction of battery life cycles even in the presence of incomplete data [<xref ref-type="bibr" rid="ref-14">14</xref>]. Zhang et al. also predicted the battery lifespan through a feature construction-based approach [<xref ref-type="bibr" rid="ref-15">15</xref>]. Since many material properties are strongly influenced by crystal structure [<xref ref-type="bibr" rid="ref-16">16</xref>], accurate prediction becomes challenging when different crystal systems exhibit similar characteristics. Prosini employed the K-nearest neighbors (K-NN) to predict the crystal group of lithium manganese oxides [<xref ref-type="bibr" rid="ref-17">17</xref>]. Overlaps in unit cell volumes, bond angles, and energy levels can make distinguishing between structures difficult. These similarities often lead to uncertainties and can reduce the reliability of conventional classification methods. For example, small differences in formation energy (E<sub>f</sub>), density (&#x03C1;), or bandgap (E<sub>g</sub>) may cause a monoclinic structure to be interpreted as orthorhombic. Such inaccuracies in identifying the crystal structure of cathode materials can ultimately compromise their performance in practical applications.</p>
<p>A previous ML based study [<xref ref-type="bibr" rid="ref-11">11</xref>] used the Materials Project database to predict crystal systems, but it employed only five ML algorithms. In contrast, our work extends this approach by implementing and systematically comparing nine different ML methods. These models include Decision Tree (DT), Random Forest (RF), Extreme Gradient Boosting (XGBoost), Support Vector Machine (SVM) classifier, k-Nearest Neighbors (k-NN) classifier, Stochastic Gradient Descent (SGD), Gaussian Na&#x00EF;ve Bayes (GNB), Gaussian Process (GP), and Artificial Neural Network (ANN). These models were chosen to provide a comprehensive comparison across a diverse range of algorithmic families, including tree-based methods (DT, RF, XGBoost), distance-based learning (k-NN), margin-based classification (SVM), probabilistic approaches (GNB, GP), linear optimization (SGD), and deep learning (ANN). This diversity allows us to evaluate how different learning paradigms handle the nonlinear and complex relationships present in the DFT-derived features of Li-Mn-Si-O cathode materials. This broader evaluation provides a more comprehensive assessment of predictive performance and significantly enhances the reliability and generalizability of the results, which constitutes a key novelty of the present study.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Multi-Algorithm ML Frameworks</title>
<sec id="s2_1">
<label>2.1</label>
<title>Brief Notes for Various ML Frameworks Employed in This Study</title>
<sec id="s2_1_1">
<label>2.1.1</label>
<title>Random Forest (RF)</title>
<p>RF is a ML algorithm that integrates the results of multiple decision trees built from randomly selected subsets of training data. Each tree is generated using unique random vector, denoted as &#x0398;<sub>k</sub>, which is independent of the random vectors used for previous trees (&#x0398;<sub>1</sub>,&#x2026;, &#x0398;<sub>k&#x2212;1</sub>). Using these random parameters and corresponding training subsets, each tree produces an individual classifier h(x, &#x0398;<sub>k</sub>), where x represents the input vector. The randomization process typically involves selecting random integer indices corresponding to features or samples, ensuring diversity among trees. The overall prediction of the RF is obtained by aggregating the outputs of all trees, which improves predictive accuracy and mitigates overfittings. The character and dimensionality of &#x0398; depend on its use in tree construction [<xref ref-type="bibr" rid="ref-18">18</xref>].</p>
</sec>
<sec id="s2_1_2">
<label>2.1.2</label>
<title>Decision Tree (DT)</title>
<p>In a DT algorithm, the dataset is recursively partitioned into smaller subsets based on specific feature values. At each node, the algorithm evaluates all variable attributes to determine the most effective feature and threshold for splitting, typically using criteria such as information gain, Gini impurity, or entropy reduction. This ensures that each division maximizes class homogeneity within the resulting subsets. The splitting process continues iteratively for each child node, forming a hierarchical tree structure where internal nodes represent decision rules and leaf nodes correspond to final class labels. The recursive partitioning terminates when all data points within a node belong to a single class or when no further meaningful division can be made. This step-by-step segregation allows DTs to capture nonlinear relationships and provide transparent, interpretable decision boundaries [<xref ref-type="bibr" rid="ref-19">19</xref>].</p>
</sec>
<sec id="s2_1_3">
<label>2.1.3</label>
<title>Extreme Gradient Boosting (XGBoost)</title>
<p>The XGBoost algorithm, developed by Chen and Guestrin, is an advanced implementation of the Gradient Boosting framework optimized for classification and regression tasks [<xref ref-type="bibr" rid="ref-20">20</xref>]. It combines multiple weak learners, typically decision trees, into a strong predictive model through iterative boosting. XGBoost enhances generalization by incorporating regularization terms in its objective function, thereby minimizing overfitting while maintaining computational efficiency. During training, parallelized feature processing accelerates model optimization. Each successive learner is trained on the residuals of the previous iteration, progressively improving model accuracy. The final output is obtained by aggregating the predictions from all individual learners, as expressed in <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>.
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msubsup><mml:mi>f</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>+</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <italic>f</italic><sub><italic>t</italic></sub>(<italic>x</italic><sub><italic>i</italic></sub>) is the learner at step <italic>t</italic>, <italic>f</italic><sub><italic>i</italic></sub><sup>(<italic>t</italic>)</sup> and <italic>f</italic><sub><italic>i</italic></sub><sup>(<italic>t</italic>&#x2212;1)</sup> are the predictions at steps <italic>t</italic> and <italic>t</italic> &#x2212; 1, and <italic>x</italic><sub><italic>i</italic></sub> is the input variable. To mitigate overfitting while maintaining computational efficiency, the XGBoost algorithm formulates an analytical objective function (<xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>) that quantifies the model&#x2019;s performance or &#x201C;goodness&#x201D; based on both predictive accuracy and regularization terms.
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mi>O</mml:mi><mml:mi>b</mml:mi><mml:msup><mml:mi>j</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>t</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover><mml:mi>l</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></disp-formula>where <italic>l</italic> is the loss function, <italic>n</italic> is the number of observations used, and &#x03A9; is the regularization term, and defined by the relation given in <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref>.
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mrow><mml:mi mathvariant="normal">&#x03A9;</mml:mi></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>f</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x03B3;</mml:mi><mml:mi>T</mml:mi><mml:mo>+</mml:mo><mml:mi>&#x03BB;</mml:mi><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mi>&#x03C9;</mml:mi><mml:msup><mml:mo fence="false" stretchy="false">&#x2016;</mml:mo><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></disp-formula>where &#x03C9; is the vector of scores in the leaves, &#x03BB; is the regularization parameter, and &#x03B3; is the minimum loss needed to further partition the leaf node. The detailed information and computation procedures of the XGBoost algorithm can be found in Chen and Guestrin [<xref ref-type="bibr" rid="ref-20">20</xref>].</p>
</sec>
<sec id="s2_1_4">
<label>2.1.4</label>
<title>Nearest Neighbors Classifier Method</title>
<p>Among supervised learning techniques, the K-NN algorithm is widely recognized for its reliable performance without requiring assumptions about the underlying data distribution. It operates by comparing a new data point with labeled examples from the training set and assigning the class based on the majority label among its &#x2018;<italic>k&#x2019;</italic> nearest neighbors. Typically, <italic>k</italic> is chosen as a small, odd number (e.g., 1, 3, or 5) to prevent ties, while higher <italic>k</italic> values can minimize the impact of noise. The optimal <italic>k</italic> is generally determined using cross-validation to balance bias and variance [<xref ref-type="bibr" rid="ref-21">21</xref>].</p>
</sec>
<sec id="s2_1_5">
<label>2.1.5</label>
<title>Stochastic Gradient Descent (SGD)</title>
<p>SGD has been recognized for its respectable status and fast computation when the learning data is huge. For the scattered data, this technique is known for its scaling capability to a huge number of features and samples. SGD is an efficient algorithm because of its linear complexity. Let Q be the matrix having a size (a,b), then the cost of training the system is O(ia&#x03B4;), where i is the number of iterations and &#x03B4; is the average of the number of non-zero attributes over all the samples in the dataset [<xref ref-type="bibr" rid="ref-22">22</xref>].</p>
</sec>
<sec id="s2_1_6">
<label>2.1.6</label>
<title>Gaussian Process (GP)</title>
<p>The training dataset consists of <italic>N</italic> observations, denoted as <italic>D</italic> &#x003D; {(<italic>x</italic><sub><italic>i</italic></sub>, <italic>y</italic><sub><italic>i</italic></sub>)|<italic>i</italic> &#x003D; 1, ..., <italic>N</italic>}, where <italic>x</italic> represents the input and <italic>y</italic> the corresponding output. The objective is to learn an underlying function <italic>f</italic> that can predict the output for an unseen input <italic>x</italic>&#x002A;. Since multiple functions may fit the data, Gaussian Process (GP) regression introduces a probabilistic framework that assigns likelihoods to possible functions based on their ability to model the data. A prior distribution encodes initial assumptions about the function&#x2019;s mean, variance, and smoothness, the latter being governed by a covariance function (kernel). By combining the prior with observed data, a posterior distribution is obtained, enabling both predictions and uncertainty estimates for new inputs. Owing to its Bayesian nature, the GP model continually improves as more data become available [<xref ref-type="bibr" rid="ref-23">23</xref>].</p>
</sec>
<sec id="s2_1_7">
<label>2.1.7</label>
<title>Gaussian Na&#x00EF;ve Bayes (GNB)</title>
<p>A NB classifier calculates the probability of a given instance belonging to a certain class. Given an instance <italic>X</italic> described by its feature vector (<italic>x</italic><sub>1</sub>, &#x2026;, <italic>x</italic><sub><italic>n</italic></sub>) and a class target <italic>y</italic>, the conditional probability <italic>P</italic>(<italic>y</italic>|<italic>X</italic>) can be expressed as a product of simpler probabilities using the Naive independence assumption according to Bayes&#x2019; theorem represented by <xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref>.
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mrow><mml:mtext>P</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mi>X</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>P</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi>y</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>X</mml:mi><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>X</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>y</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:munderover><mml:mo>&#x220F;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:mi>X</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>Here, the target <italic>y</italic> may have two values, where <italic>y</italic> &#x003D; 1 means a hot spot residue and <italic>y</italic> &#x003D; 0 represents a non-hot spot residue. <italic>X</italic> for one residue (one instance) is a feature vector with the same size for describing its characteristics using high-frequency modes generated by GNM. For example, <italic>X</italic> is equal to a vector composed of <italic>i</italic>th component u<italic>ki</italic> for <italic>i</italic>th residue in a sequence when only one high frequency mode u<italic>k</italic> is used. If three high-frequency modes, denoted by u<sub>1</sub>, u<sub>2</sub>, and u<sub>3</sub>, are taken into account, the vector <italic>X</italic> will be (u<sub>1<italic>i</italic></sub>, u<sub>2<italic>i</italic></sub>, u<sub>3<italic>i</italic></sub>) for residue <italic>i</italic> in a protein sequence. Moreover, if a window size of 3 with respect to the residue <italic>i</italic> is adopted, <italic>X</italic> becomes (u<sub>1<italic>i</italic></sub>&#x2212;1, u<sub>1<italic>i</italic></sub>, u<sub>1<italic>i</italic>&#x002B;1</sub>, u<sub>2<italic>i</italic>&#x2212;1</sub>, u<sub>2<italic>i</italic></sub>, u<sub>2<italic>i</italic>&#x002B;1</sub>, u<sub>3<italic>i</italic>&#x2212;1</sub>, u<sub>3<italic>i</italic></sub>, u<sub>3<italic>i</italic>&#x002B;1</sub>). Since (<italic>X</italic>) is constant for a given instance, the following rule is adopted to classify the instance whose class is unknown [<xref ref-type="bibr" rid="ref-24">24</xref>].</p>
</sec>
<sec id="s2_1_8">
<label>2.1.8</label>
<title>Support Vector Machines (SVM)</title>
<p>SVM methods find the maximum margin hyperplane w<sup>T</sup>&#x03C6;(x<sub>i</sub>) &#x002B; b that separates the positive datapoints from the negative datapoints [<xref ref-type="bibr" rid="ref-25">25</xref>]. Where <italic>w</italic> is the normal vector to the hyperplane, <italic>x</italic><sub><italic>i</italic></sub> is the training dataset, and &#x03C6;(<italic>x</italic><sub><italic>i</italic></sub>) maps the training data to the feature. The optimization problem can be formulated by Minimize <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mi>w</mml:mi><mml:mo>+</mml:mo><mml:mi>C</mml:mi><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:munder><mml:msub><mml:mi>&#x03BE;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, Subject to:
<disp-formula id="ueqn-5"><mml:math id="mml-ueqn-5" display="block"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>(</mml:mo><mml:msup><mml:mi>w</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mi>&#x03C6;</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2265;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>&#x03BE;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>;</mml:mo><mml:msub><mml:mi>&#x03BE;</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2265;</mml:mo><mml:mn>0</mml:mn></mml:math></disp-formula>where C &#x003E; 0 is the parameter that controls the trade-off between the training errors and the model complexity, <italic>&#x03BE;</italic><sub><italic>i</italic></sub> are slack variables used to achieve a soft margin, and &#x03C6; is a non-linear mapping from an input space into a feature space. By introducing the Lagrange multiplier <italic>a</italic><sub><italic>i</italic></sub>, a corresponding dual problem can be derived by following the quadratic programming (QP) problem, maximize <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:munder><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:munder><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, subject to
<disp-formula id="ueqn-6"><mml:math id="mml-ueqn-6" display="block"><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:munder><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>0</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>a</mml:mi><mml:mrow><mml:mtext>&#xA0;i&#xA0;</mml:mtext></mml:mrow><mml:mo>&#x2264;</mml:mo><mml:mrow><mml:mtext>C</mml:mtext></mml:mrow><mml:mo>,</mml:mo><mml:mrow><mml:mtext>I</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mrow><mml:mtext>N</mml:mtext></mml:mrow><mml:mo>,</mml:mo></mml:math></disp-formula>where <italic>k</italic> is a kernel function <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>k</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo>=&#x003C;&#x22A2;</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>X</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>&#x03D5;</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mo>&#x003E;&#x22A3;</mml:mo></mml:math></inline-formula>, e.g., radial basis function (RBF) kernel <italic>k</italic>(<italic>x</italic><sub><italic>i</italic></sub>,<italic>x</italic><sub><italic>j</italic></sub>) &#x003D; exp(&#x2016;<italic>x</italic><sub><italic>i</italic></sub> &#x2212; <italic>x</italic><sub><italic>j</italic></sub>&#x2016;2)/2&#x03C3;2). Once the dual QP problem is solved, the resulting decision function at any test data point x is as follows: <italic>f</italic>(<italic>x</italic>) &#x003D; w<sup>T</sup>&#x03C6;(x) &#x002B; b &#x003D; <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mi>k</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo>+</mml:mo><mml:mi>b</mml:mi><mml:mo>=</mml:mo><mml:munder><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mi>S</mml:mi><mml:mi>V</mml:mi></mml:mrow></mml:munder><mml:msub><mml:mi>a</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mi>k</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mi>x</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>+</mml:mo><mml:mi>b</mml:mi></mml:math></inline-formula></p>
<p>Only those data points for which ai is nonzero are referred to as support vectors, and they define the decision function. In the test phase, we estimate the class of the test datapoint <italic>x</italic> based on the sign(<italic>f</italic>(<italic>x</italic>)). Since P(X) is constant for a given instance, the following rule is adopted to classify the instance whose class is unknown, as given in <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref>.
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mo>=</mml:mo><mml:mi>arg</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mo movablelimits="true" form="prefix">max</mml:mo><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mrow><mml:mtext>P</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>y</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:msubsup><mml:mo movablelimits="false">&#x220F;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:msubsup><mml:mi>P</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="false">|</mml:mo></mml:mrow><mml:mtext>y</mml:mtext><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula></p>
</sec>
<sec id="s2_1_9">
<label>2.1.9</label>
<title>Artificial Neural Network (ANN)</title>
<p>An ANN model, which is based on multilayer perceptrons, consists of input, hidden, and output layers in computational systems. The input layer has neurons for obtaining multiple inputs. Each input is multiplied by its weight, which can be summarized as a neuron of a hidden layer. The neurons in the hidden layer use the transmission function to generate new values, and these new values are multiplied again by the weight for the output layer. The model is trained as a backpropagation algorithm and feed-forward using the sigmoid function as an activation function. ANN model has five sequentially optimized factors (Neurons, Hidden Layer, Learning Rate, Momentum terms, and Iterations) [<xref ref-type="bibr" rid="ref-26">26</xref>].</p>
<p>The summary of strengths, limitations and key characteristics of these models are given in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Summary of the ML models employed in this study</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>ML method</th>
<th>Strength</th>
<th>Limitations</th>
<th>Key characteristics</th>
</tr>
</thead>
<tbody>
<tr>
<td>Decision tree</td>
<td>Fast training</td>
<td>Not suitable for small dataset</td>
<td>Rule-based hierarchical splits</td>
</tr>
<tr>
<td>Random forest</td>
<td>Good accuracy and less overfitting</td>
<td>More computations involved</td>
<td>Ensemble of multiple decision trees</td>
</tr>
<tr>
<td>XGBoost</td>
<td>Efficient boosting</td>
<td>Sensitive to noise</td>
<td>Gradient-boosted decision tree ensemble</td>
</tr>
<tr>
<td>SVM classifier</td>
<td>Suitable for small dataset</td>
<td>Required kernel section</td>
<td>Maximizes margin between classes</td>
</tr>
<tr>
<td>Nearest neighbors classifier</td>
<td>Very simple</td>
<td>Slow for large datasets</td>
<td>Distance-based classification</td>
</tr>
<tr>
<td>SGD classifier</td>
<td>Simple implementation</td>
<td>Sensitive to learning rate</td>
<td>Linear classifier optimized via SGD</td>
</tr>
<tr>
<td>Gaussian naive bayes</td>
<td>Very fast</td>
<td>Limited with correlated features</td>
<td>Bayes theorem with Gaussian likelihood</td>
</tr>
<tr>
<td>Gaussian process</td>
<td>High accuracy</td>
<td>More computations involved</td>
<td>More computations involved</td>
</tr>
<tr>
<td>ANN</td>
<td>Captures complex nonlinear patterns</td>
<td>Needs careful tuning</td>
<td>Multilayer architecture</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Materials and Methods</title>
<sec id="s3_1">
<label>3.1</label>
<title>Workflow for ML&#x2013;Based Frameworks for Prediction of Crystal Structures</title>
<p><xref ref-type="fig" rid="fig-1">Fig. 1</xref> presents the workflow employed in this study to predict the crystal structure of lithium manganese silicate cathodes using data sourced from the Materials Project. The dataset consisted of 211 DFT-computed entries containing the selected input features, while the output crystal structure was encoded as a binary label, with 0 representing monoclinic and 1 representing triclinic structures.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Workflow of the present study for predicting the crystal structure of lithium manganese silicate cathodes using Materials Project data</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_75957-fig-1.tif"/>
</fig>
<p>The data were divided into 169 training samples and 42 testing samples, and nine ML algorithms were applied to develop predictive models. This 80:20 division follows a widely accepted practice in machine-learning studies, providing a balanced compromise between model training and unbiased evaluation. The split was generated through random partitioning to avoid sampling bias and ensure that the model performance reflects true generalization to unseen data. Model performance was evaluated using multiple assessment metrics including accuracy (ACC), Matthew&#x2019;s correlation coefficient (MCC), recall (RCC), precision (PRE), F-score (F), negative predictive value (NPV) using <xref ref-type="disp-formula" rid="eqn-6">Eqs. (6)</xref>&#x2013;<xref ref-type="disp-formula" rid="eqn-11">(11)</xref>. Based on overall prediction accuracy, the best-performing model was identified and subsequently subjected to detailed optimization and analysis of its architecture and predictive behavior.
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mrow><mml:mtext>ACC&#xA0;</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mrow><mml:mtext>MCC&#xA0;</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>&#x2217;</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>&#x2217;</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow><mml:msqrt><mml:mrow><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:msqrt></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mrow><mml:mtext>Recall&#xA0;</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>RCC</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mrow><mml:mtext>Precision&#xA0;</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>PRE</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mtext>&#xA0;score&#xA0;</mml:mtext></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mrow><mml:mtext>F</mml:mtext></mml:mrow><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mo>&#x2217;</mml:mo><mml:mi>R</mml:mi><mml:mi>C</mml:mi><mml:mi>C</mml:mi><mml:mo>&#x2217;</mml:mo><mml:mi>P</mml:mi><mml:mi>R</mml:mi><mml:mi>E</mml:mi></mml:mrow><mml:mrow><mml:mi>R</mml:mi><mml:mi>C</mml:mi><mml:mi>C</mml:mi><mml:mo>+</mml:mo><mml:mi>P</mml:mi><mml:mi>R</mml:mi><mml:mi>E</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-11"><label>(11)</label><mml:math id="mml-eqn-11" display="block"><mml:mrow><mml:mtext>NPV&#xA0;</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>Here, TP &#x003D; true positive, TN &#x003D; true negative, FP &#x003D; false positve, FN &#x003D; false negative.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Dataset Description and Preprocessing</title>
<p>The entire dataset was obtained from the Materials Project Database [<xref ref-type="bibr" rid="ref-9">9</xref>], which provides DFT computed properties of 211 cathode materials with Li-Si-(Mn)-O compositions. All DFT calculations and structural optimizations were performed using the VASP software within the Materials Project framework. The exchange&#x2013;correlation potentials were treated using the generalized gradient approximation (GGA) or GGA &#x002B; U, and the DFT energies for the Li-Si-(Mn)-O systems were generated through a high-throughput computational workflow. The initial DFT calculations containing positions of atoms and lattice parameters of crystals can be based on available data from inorganic crystal structure database [<xref ref-type="bibr" rid="ref-27">27</xref>].</p>
<p>The dataset includes the E<sub>f</sub>, E<sub>h</sub>, E<sub>g</sub>, number of sites (N<sub>s</sub>), density (&#x03C1;), the volume of the unit cell (V), and crystal structure of each electrode. The available dataset is divided into 80:20 as training and testing datasets. To avoid any bias in the training process, each method-based model is trained 100 times, and the models are stored. The data inputs are the chemical formula, space group, E<sub>f</sub> (eV) E<sub>h</sub> (eV), E<sub>g</sub> (eV), N<sub>s</sub>, &#x03C1; (g.cm<sup>&#x2212;3</sup>) and V (A<sup>3</sup>). The output is the crystal structures of Li-Si-(Mn)-O cathode materials that are monoclinic (0) or triclinic (1). <xref ref-type="fig" rid="fig-2">Fig. 2</xref> presents the pair plot of the properties of Li-Si-(Mn)-O cathode materials in the dataset. The diagonal elements illustrate the distribution of individual features, while the off-diagonal plots show pairwise relationships between variables. The symmetry along the diagonal reflects similar distributions across parameters, and the axes appear mirrored due to the pairwise plotting arrangement.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Pair plot of the properties of the Li-M-Si-O (M&#x003D;Fe, Mn, and Co) cathode materials from Materials Project about the relative input parameter</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_75957-fig-2.tif"/>
</fig>
<p>No clear correlation is observed between the selected features and the resulting crystal system, highlighting that simple linear or direct relationships are insufficient to describe the underlying structure&#x2013;property interactions. This lack of explicit trends underscores the necessity of employing advanced modelling techniques capable of capturing complex, nonlinear dependencies within the data. Therefore, ML-based modelling becomes essential for reliably predicting the crystal structures of lithium manganese silicate cathodes, where multiple compositional and structural factors interact in a non-trivial manner.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Results and Discussion</title>
<sec id="s4_1">
<label>4.1</label>
<title>Prediction Performance of ML Models</title>
<p>A comparative assessment of each of the nine ML algorithms was conducted to evaluate their performance using various statistical indicators, including ACC, RCC, PRE, specificity, NPV, and F-score (<xref ref-type="table" rid="table-2">Table 2</xref>). The results revealed substantial variability in the generalization capability of different models. The DT model exhibited good fitting during training (accuracy &#x003D; 0.887) but demonstrated a sharp decline in testing accuracy (0.639), indicating overfitting. In contrast, the ensemble-based models, RF and XGBoost, showed superior performance, achieving nearly perfect training accuracies (1.000 and 0.993, respectively) and strong testing accuracies (0.803 and 0.721, respectively). Notably, RF outperformed all other models in terms of balanced accuracy, recall, and precision on the test set (accuracy &#x003D; 0.803, recall &#x003D; 0.786, precision &#x003D; 0.917, F-score &#x003D; 0.846), highlighting its robustness and effective handling of complex nonlinear relationships. The SVM classifier demonstrated moderate predictive ability (testing accuracy &#x003D; 0.688), providing a stable but not outstanding performance. The Nearest Neighbors Classifier yielded lower testing accuracy (0.606), which may be attributed to its sensitivity to noise and local data variations.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Summary of the evaluation of training data and testing data accuracy using the various ML methods</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>ML Method</th>
<th>Train/Test</th>
<th>ACC</th>
<th>RCC</th>
<th>PRE</th>
<th>Specificity</th>
<th>NPV</th>
<th>F-Score</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="2">Decision tree</td>
<td>Training</td>
<td>0.887</td>
<td>0.858</td>
<td>1</td>
<td>1</td>
<td>0.638</td>
<td>0.924</td>
</tr>
<tr>
<td>Testing</td>
<td>0.639</td>
<td>0.630</td>
<td>0.944</td>
<td>0.714</td>
<td>0.200</td>
<td>0.756</td>
</tr>
<tr>
<td rowspan="2">Random forest</td>
<td>Training</td>
<td>1</td>
<td>1</td>
<td>1</td>
<td>1</td>
<td>1</td>
<td>1</td>
</tr>
<tr>
<td>Testing</td>
<td>0.803</td>
<td>0.786</td>
<td>0.917</td>
<td>0.842</td>
<td>0.64</td>
<td>0.846</td>
</tr>
<tr>
<td rowspan="2">XGBoost</td>
<td>Training</td>
<td>0.993</td>
<td>0.990</td>
<td>1</td>
<td>1</td>
<td>0.979</td>
<td>0.995</td>
</tr>
<tr>
<td>Testing</td>
<td>0.721</td>
<td>0.731</td>
<td>0.833</td>
<td>0.7</td>
<td>0.56</td>
<td>0.779</td>
</tr>
<tr>
<td rowspan="2">SVM classifier</td>
<td>Training</td>
<td>0.92</td>
<td>0.925</td>
<td>0.961</td>
<td>0.907</td>
<td>0.829</td>
<td>0.942</td>
</tr>
<tr>
<td>Testing</td>
<td>0.688</td>
<td>0.729</td>
<td>0.75</td>
<td>0.625</td>
<td>0.6</td>
<td>0.739</td>
</tr>
<tr>
<td rowspan="2">Nearest neighbors classifier</td>
<td>Training</td>
<td>0.827</td>
<td>0.841</td>
<td>0.922</td>
<td>0.783</td>
<td>0.617</td>
<td>0.879</td>
</tr>
<tr>
<td>Testing</td>
<td>0.606</td>
<td>0.636</td>
<td>0.778</td>
<td>0.529</td>
<td>0.36</td>
<td>0.7</td>
</tr>
<tr>
<td rowspan="2">SGD Classifier</td>
<td>Training</td>
<td>0.413</td>
<td>0.683</td>
<td>0.279</td>
<td>0.312</td>
<td>0.723</td>
<td>0.389</td>
</tr>
<tr>
<td>Testing</td>
<td>0.443</td>
<td>0.583</td>
<td>0.194</td>
<td>0.408</td>
<td>0.800</td>
<td>0.292</td>
</tr>
<tr>
<td rowspan="2">Gaussian naive bayes</td>
<td>Training</td>
<td>0.647</td>
<td>0.731</td>
<td>0.767</td>
<td>0.429</td>
<td>0.383</td>
<td>0.749</td>
</tr>
<tr>
<td>Testing</td>
<td>0.656</td>
<td>0.683</td>
<td>0.778</td>
<td>0.600</td>
<td>0.480</td>
<td>0.727</td>
</tr>
<tr>
<td rowspan="2">Gaussian process</td>
<td>Training</td>
<td>0.993</td>
<td>0.990</td>
<td>1.000</td>
<td>1.000</td>
<td>0.979</td>
<td>0.995</td>
</tr>
<tr>
<td>Testing</td>
<td>0.705</td>
<td>0.737</td>
<td>0.778</td>
<td>0.652</td>
<td>0.600</td>
<td>0.757</td>
</tr>
<tr>
<td rowspan="2">ANN</td>
<td>Training</td>
<td>1.000</td>
<td>1.000</td>
<td>1.000</td>
<td>1.000</td>
<td>1.000</td>
<td>1.000</td>
</tr>
<tr>
<td>Testing</td>
<td>0.761</td>
<td>0.447</td>
<td>0.800</td>
<td>0.857</td>
<td>0.827</td>
<td>0.906</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The SGD classifier exhibited the weakest performance across all metrics (testing accuracy &#x003D; 0.443, F-score &#x003D; 0.292), indicating poor convergence in nonlinear feature spaces. Among the probabilistic approaches, Gaussian NB and GP classifiers achieved comparable results, with the latter showing slightly higher predictive balance (testing accuracy &#x003D; 0.705, F-score &#x003D; 0.757). The ANN achieved perfect training accuracy (1.000) and demonstrated strong generalization on the testing dataset (accuracy &#x003D; 0.761, F-score &#x003D; 0.906). Although minor overfitting was observed, the ANN effectively captured intricate nonlinear dependencies, outperforming most conventional algorithms in terms of overall predictive reliability.</p>
<p>These findings suggest the ANN model achieved the highest prediction accuracy and demonstrated the strongest capability to learn the complex, nonlinear interdependencies among the input features. The remaining eight models were optimized using standard and widely accepted hyperparameter-tuning procedures (e.g., grid search, cross-validation, and built-in optimization routines), and their performance showed relatively low sensitivity to tuning variations. Therefore, an extensive architectural explanation was not required for them. In contrast, the ANN contains multiple architecture-dependent parameters such as the number of layers, neurons per layer, activation functions, learning rate, and momentum terms, and its performance was highly sensitive to these choices. To ensure transparency, fairness, and reproducibility, the detailed ANN architecture, optimization strategy, and training behavior will be provided in the coming sections.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Data Splitting for ANN Model</title>
<p><xref ref-type="fig" rid="fig-3">Fig. 3</xref> shows the data classed for training and testing data in monoclinic and triclinic crystal structures. In the monoclinic crystal structure, 111 training data and 28 data out of 139 data were investigated, and the training data were not classified, with four of the testing data being unclassified. Also, 58 training data and 14 testing data out of the total 72 data were investigated in the triclinic crystal structure, and the triclinic crystal structure also showed unclassified data in the training data, and the testing data showed unclassified data in six testing data.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>The misclassified training data and testing data in monoclinic and triclinic crystal systems</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_75957-fig-3.tif"/>
</fig>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Optimum ANN Model Architecture</title>
<p>The ANN model was trained using the training dataset with systematic optimization of key hyperparameters. The core ANN algorithm was implemented in the C programming language for computational efficiency, while a user-friendly graphical user interface (GUI) was developed in Java to facilitate model execution, parameter adjustment, and visualization of results. The number of hidden layers were varied from 1 to 3, the number of neurons per layer from 1 to 30, the momentum coefficient from 0.1 to 1.0, the learning rate from 0.1 to 1.0, and the number of training iterations from 5000 to 70,000. The corresponding changes in model behavior under these settings are illustrated in <xref ref-type="fig" rid="fig-4">Fig. 4</xref>. The optimized ANN architecture consists of three hidden layers with fourteen neurons in each layer (<xref ref-type="fig" rid="fig-4">Fig. 4a</xref>). A momentum value of 0.3 (<xref ref-type="fig" rid="fig-4">Fig. 4b</xref>), a learning rate of 0.6 (<xref ref-type="fig" rid="fig-4">Fig. 4c</xref>), and 20,000 training iterations (<xref ref-type="fig" rid="fig-4">Fig. 4d</xref>) yielded the best performance, achieving a prediction accuracy of 94.31%. Further refinement through iteration tuning demonstrated that the highest accuracy of 95.26% was reached at 20,000 iterations. These results clearly demonstrate the significant impact of hyperparameter selection on the ANN&#x2019;s ability to accurately predict the crystal system of Li-Mn-Si-O materials.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>The line graphs show the accuracy of different hyperparameters for the ANN</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_75957-fig-4.tif"/>
</fig>
<p>All data, including the chemical formula, space group, experimental, and ANN predicted crystal structure, are presented in the <xref ref-type="table" rid="table-3">Table 3.</xref> The four unclassified data are monoclinic crystal structures in the experimental crystal system and represent triclinic crystal structures in the ANN model. The six unclassified data represent the triclinic crystal structure in the experimental and the monoclinic crystal structure in the ANN model. The commonality of these is that the Space group is P1.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Data for misclassified Lithium manganese silicate cathodes from the dataset. Bold compositions of cathode materials were misclassified</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Sr. No.</th>
<th>Formula</th>
<th>Space group</th>
<th>Experimental</th>
<th>ANN predicted</th>
</tr>
</thead>
<tbody>
<tr>
<td>1</td>
<td>Li<sub>4</sub>Fe<sub>3</sub>(SiO<sub>4</sub>)<sub>3</sub></td>
<td>P1</td>
<td>Triclinic</td>
<td>Triclinic</td>
</tr>
<tr>
<td>2</td>
<td>Li<sub>2</sub>Fe(Si<sub>2</sub>O<sub>5</sub>)<sub>3</sub></td>
<td>P1</td>
<td>Triclinic</td>
<td>Triclinic</td>
</tr>
<tr>
<td>3</td>
<td><bold>Li</bold><sub><bold>16</bold></sub><bold>Fe</bold><sub><bold>4</bold></sub><bold>SiO</bold><sub><bold>16</bold></sub></td>
<td><bold>P1</bold></td>
<td><bold>Triclinic</bold></td>
<td><bold>Monoclinic</bold></td>
</tr>
<tr>
<td>4</td>
<td><bold>LiFe</bold><sub><bold>2</bold></sub><bold>(SiO</bold><sub><bold>4</bold></sub><bold>)</bold><sub><bold>2</bold></sub></td>
<td><bold>P1</bold></td>
<td><bold>Triclinic</bold></td>
<td><bold>Monoclinic</bold></td>
</tr>
<tr>
<td>5</td>
<td><bold>Li</bold><sub><bold>7</bold></sub><bold>Fe</bold><sub><bold>7</bold></sub><bold>SiO</bold><sub><bold>16</bold></sub></td>
<td><bold>P1</bold></td>
<td><bold>Triclinic</bold></td>
<td><bold>Monoclinic</bold></td>
</tr>
<tr>
<td>6</td>
<td><bold>Li</bold><sub><bold>2</bold></sub><bold>Co</bold><sub><bold>3</bold></sub><bold>(SiO</bold><sub><bold>4</bold></sub><bold>)</bold><sub><bold>2</bold></sub></td>
<td><bold>P1</bold></td>
<td><bold>Triclinic</bold></td>
<td><bold>Monoclinic</bold></td>
</tr>
<tr>
<td>7</td>
<td>Li<sub>3</sub>Co<sub>2</sub>(SiO<sub>4</sub>)<sub>2</sub></td>
<td>P1</td>
<td>Triclinic</td>
<td>Triclinic</td>
</tr>
<tr>
<td>8</td>
<td>Li<sub>3</sub>Co<sub>2</sub>(SiO<sub>4</sub>)<sub>2</sub></td>
<td>P1</td>
<td>Triclinic</td>
<td>Triclinic</td>
</tr>
<tr>
<td>9</td>
<td>Li<sub>2</sub>Co(Si<sub>2</sub>O<sub>5</sub>)<sub>2</sub></td>
<td>P1</td>
<td>Triclinic</td>
<td>Triclinic</td>
</tr>
<tr>
<td>10</td>
<td><bold>Li</bold><sub><bold>6</bold></sub><bold>Co(SiO</bold><sub><bold>4</bold></sub><bold>)</bold><sub><bold>2</bold></sub></td>
<td><bold>P1</bold></td>
<td><bold>Triclinic</bold></td>
<td><bold>Monoclinic</bold></td>
</tr>
<tr>
<td>11</td>
<td>LiCo<sub>3</sub>(SiO<sub>4</sub>)<sub>2</sub></td>
<td>P1</td>
<td>Triclinic</td>
<td>Triclinic</td>
</tr>
<tr>
<td>12</td>
<td>Li<sub>5</sub>Co<sub>4</sub>(Si<sub>3</sub>O<sub>10</sub>)<sub>2</sub></td>
<td>P1</td>
<td>Triclinic</td>
<td>Triclinic</td>
</tr>
<tr>
<td>13</td>
<td><bold>LiCoSiO</bold><sub><bold>4</bold></sub></td>
<td><bold>P1</bold></td>
<td><bold>Triclinic</bold></td>
<td><bold>Monoclinic</bold></td>
</tr>
<tr>
<td>14</td>
<td>Li<sub>3</sub>Co<sub>2</sub>(SiO<sub>4</sub>)<sub>2</sub></td>
<td>P1</td>
<td>Triclinic</td>
<td>Triclinic</td>
</tr>
<tr>
<td>15</td>
<td>Li<sub>2</sub>MnSiO<sub>4</sub></td>
<td>Pc</td>
<td>Monoclinic</td>
<td>Monoclinic</td>
</tr>
<tr>
<td>16</td>
<td>Li<sub>2</sub>MnSiO<sub>4</sub></td>
<td>P21/c</td>
<td>Monoclinic</td>
<td>Monoclinic</td>
</tr>
<tr>
<td>17</td>
<td>Li<sub>4</sub>MnSi<sub>2</sub>O<sub>7</sub></td>
<td>Cc</td>
<td>Monoclinic</td>
<td>Monoclinic</td>
</tr>
<tr>
<td>18</td>
<td>Li<sub>4</sub>Mn<sub>2</sub>Si<sub>3</sub>O<sub>10</sub></td>
<td>C2/c</td>
<td>Monoclinic</td>
<td>Monoclinic</td>
</tr>
<tr>
<td>19</td>
<td><bold>Li</bold><sub><bold>2</bold></sub><bold>Mn</bold><sub><bold>3</bold></sub><bold>Si</bold><sub><bold>3</bold></sub><bold>O</bold><sub><bold>10</bold></sub></td>
<td><bold>C2/c</bold></td>
<td><bold>Monoclinic</bold></td>
<td><bold>Triclinic</bold></td>
</tr>
<tr>
<td>20</td>
<td>Li<sub>4</sub>MnSi<sub>2</sub>O<sub>7</sub></td>
<td>C2</td>
<td>Monoclinic</td>
<td>Monoclinic</td>
</tr>
<tr>
<td>21</td>
<td>LiMnSiO<sub>4</sub></td>
<td>P21</td>
<td>Monoclinic</td>
<td>Monoclinic</td>
</tr>
<tr>
<td>22</td>
<td>Li<sub>2</sub>MnSiO<sub>4</sub></td>
<td>P21/c</td>
<td>Monoclinic</td>
<td>Monoclinic</td>
</tr>
<tr>
<td>23</td>
<td><bold>LiMn(SiO</bold><sub><bold>3</bold></sub><bold>)</bold><sub><bold>2</bold></sub></td>
<td><bold>C2/c</bold></td>
<td><bold>Monoclinic</bold></td>
<td><bold>Triclinic</bold></td>
</tr>
<tr>
<td>24</td>
<td>Li<sub>2</sub>Mn(SiO<sub>3</sub>)<sub>2</sub></td>
<td>Cc</td>
<td>Monoclinic</td>
<td>Monoclinic</td>
</tr>
<tr>
<td>25</td>
<td>Li<sub>2</sub>MnSiO<sub>4</sub></td>
<td>P21/c</td>
<td>Monoclinic</td>
<td>Monoclinic</td>
</tr>
<tr>
<td>26</td>
<td>Li<sub>2</sub>Mn(SiO<sub>3</sub>)<sub>2</sub></td>
<td>C2/c</td>
<td>Monoclinic</td>
<td>Monoclinic</td>
</tr>
<tr>
<td>27</td>
<td><bold>Li</bold><sub><bold>2</bold></sub><bold>Mn</bold><sub><bold>2</bold></sub><bold>Si</bold><sub><bold>2</bold></sub><bold>O</bold><sub><bold>7</bold></sub></td>
<td><bold>P21/c</bold></td>
<td><bold>Monoclinic</bold></td>
<td><bold>Triclinic</bold></td>
</tr>
<tr>
<td>28</td>
<td>Li<sub>10</sub>Mn(SiO<sub>5</sub>)<sub>2</sub></td>
<td>C2/m</td>
<td>Monoclinic</td>
<td>Monoclinic</td>
</tr>
<tr>
<td>29</td>
<td>Li<sub>3</sub>MnSi<sub>2</sub>O<sub>7</sub></td>
<td>P21</td>
<td>Monoclinic</td>
<td>Monoclinic</td>
</tr>
<tr>
<td>30</td>
<td>Li<sub>5</sub>Mn(SiO<sub>4</sub>)<sub>2</sub></td>
<td>C2</td>
<td>Monoclinic</td>
<td>Monoclinic</td>
</tr>
<tr>
<td>31</td>
<td>Li<sub>2</sub>Mn(Si<sub>2</sub>O<sub>5</sub>)<sub>2</sub></td>
<td>P21/c</td>
<td>Monoclinic</td>
<td>Monoclinic</td>
</tr>
<tr>
<td>32</td>
<td>Li<sub>2</sub>Mn<sub>2</sub>Si<sub>3</sub>O<sub>10</sub></td>
<td>Cc</td>
<td>Monoclinic</td>
<td>Monoclinic</td>
</tr>
<tr>
<td>33</td>
<td>Li<sub>2</sub>Mn<sub>2</sub>(SiO<sub>3</sub>)<sub>3</sub></td>
<td>P21/c</td>
<td>Monoclinic</td>
<td>Monoclinic</td>
</tr>
<tr>
<td>34</td>
<td>LiMn(SiO<sub>3</sub>)<sub>2</sub></td>
<td>C2/c</td>
<td>Monoclinic</td>
<td>Monoclinic</td>
</tr>
<tr>
<td>35</td>
<td>Li<sub>2</sub>MnSi<sub>3</sub>O<sub>8</sub></td>
<td>P21</td>
<td>Monoclinic</td>
<td>Monoclinic</td>
</tr>
<tr>
<td>36</td>
<td>Li<sub>3</sub>Mn<sub>2</sub>(SiO<sub>4</sub>)<sub>2</sub></td>
<td>P21</td>
<td>Monoclinic</td>
<td>Monoclinic</td>
</tr>
<tr>
<td>37</td>
<td>Li<sub>4</sub>Mn(SiO<sub>3</sub>)<sub>3</sub></td>
<td>C2</td>
<td>Monoclinic</td>
<td>Monoclinic</td>
</tr>
<tr>
<td>38</td>
<td>Li<sub>2</sub>MnSi<sub>3</sub>O<sub>8</sub></td>
<td>P21</td>
<td>Monoclinic</td>
<td>Monoclinic</td>
</tr>
<tr>
<td>39</td>
<td><bold>Li</bold><sub><bold>2</bold></sub><bold>Mn(SiO</bold><sub><bold>3</bold></sub><bold>)</bold><sub><bold>2</bold></sub></td>
<td><bold>C2</bold></td>
<td><bold>Monoclinic</bold></td>
<td><bold>Triclinic</bold></td>
</tr>
<tr>
<td>40</td>
<td>LiMn<sub>2</sub>Si<sub>2</sub>O<sub>7</sub></td>
<td>Cc</td>
<td>Monoclinic</td>
<td>Monoclinic</td>
</tr>
<tr>
<td>41</td>
<td>Li<sub>3</sub>Mn<sub>2</sub>(SiO<sub>4</sub>)<sub>2</sub></td>
<td>Pc</td>
<td>Monoclinic</td>
<td>Monoclinic</td>
</tr>
<tr>
<td>42</td>
<td>Li<sub>3</sub>Mn<sub>2</sub>(SiO<sub>4</sub>)<sub>2</sub></td>
<td>Pc</td>
<td>Monoclinic</td>
<td>Monoclinic</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Evaluation of the Confusion Matrix</title>
<p>A confusion matrix between the ANN model and the DFT calculated data predicts all data in the dataset as either positive or negative. This classification produces four outcomes.</p>
<p>TP values are accurate positive predictions, FP values are incorrect positive predictions, TN shows accurate negative predictions, and FN represents incorrect negative predictions. The TP value represents the data number when the ANN model prediction and the DFT prediction are both monoclinic, and the TN values represent the data number when both the ANN model and the DFT prediction are Triclinic. <xref ref-type="table" rid="table-4">Table 4</xref> shows the evaluation of the confusion matrix. The total data sets are 211. The number of TP, FP, FN, and TN values show 135, 6, 4, and 66. To ensure the reliability of the classification results and to assess the performance of the ANN model, its predictions were compared with the corresponding DFT-calculated data. The model&#x2019;s performance, evaluated using several statistical metrics, is presented in <xref ref-type="fig" rid="fig-5">Fig. 5</xref>.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Confusion matrix between the ANN model and the DFT calculated data</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/> </colgroup>
<thead>
<tr>
<th>Data</th>
<th>No. of Data</th>
<th>DFT (0)</th>
<th>DFT (1)</th>
<th>Total</th>
</tr>
</thead>
<tbody>
<tr>
<td rowspan="3">Test</td>
<td>ANN (0)</td>
<td>TP (111)</td>
<td>FP (0)</td>
<td>111</td>
</tr>
<tr>
<td>ANN (1)</td>
<td>FN (0)</td>
<td>TN (58)</td>
<td>58</td>
</tr>
<tr>
<td>Total</td>
<td>111</td>
<td>58</td>
<td>169</td>
</tr>
<tr>
<td rowspan="3">Train</td>
<td>ANN (0)</td>
<td>TP (24)</td>
<td>FP (4)</td>
<td>28</td>
</tr>
<tr>
<td>ANN (1)</td>
<td>FN (6)</td>
<td>TN (8)</td>
<td>14</td>
</tr>
<tr>
<td>Total</td>
<td>30</td>
<td>12</td>
<td>42/61</td>
</tr>
<tr>
<td rowspan="3">All Data</td>
<td>ANN (0)</td>
<td>TP (135)</td>
<td>FP (6)</td>
<td>141</td>
</tr>
<tr>
<td>ANN (1)</td>
<td>FN (4)</td>
<td>TN (66)</td>
<td>70</td>
</tr>
<tr>
<td>Total</td>
<td>129</td>
<td>72</td>
<td>211</td>
</tr>
</tbody>
</table>
</table-wrap><fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>The spider plot showing the performance of ANN by different evaluation matrices</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_75957-fig-5.tif"/>
</fig>
<p>The ACC value of the optimized ANN model was found to be 0.953, indicating a high level of predictive reliability. The MCC, which evaluates the balance between under- and over-predictions&#x2014;where MCC &#x003D; 1 represents a perfect prediction and MCC &#x003D; 0 corresponds to a random assignment&#x2014;was 0.894, signifying strong consistency between predicted and actual classifications. The precision (PRE), representing the ratio of correctly predicted positive cases to all predicted positives, was 0.957. The F-score, defined as the harmonic mean of precision and recall (ideal value &#x003D; 1), was 0.963, further confirming the model&#x2019;s strong performance. The NPV, which measures the ratio of correctly predicted negatives to total predicted negatives, was 0.943, demonstrating that the model effectively distinguishes between the two crystal systems.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusions</title>
<p>In this study, various machine learning (ML) algorithms were developed and compared for predicting the crystal system of lithium manganese silicate (Li-Mn-Si-O) cathode materials using density functional theory (DFT) data from the Materials Project database. The dataset contained 211 compositions with key features such as formation energy, energy above the hull, bandgap, number of atomic sites, density, and unit cell volume. These descriptors were used to classify the crystal system into monoclinic (0) and triclinic (1) phases.</p>
<p>The comparative analysis of multiple classification techniques&#x2014;Decision Tree, Random Forest, XGBoost, Support Vector Machine, Nearest Neighbor Classifier, Stochastic Gradient Descent, Gaussian Na&#x00EF;ve Bayes, Gaussian Process, and Artificial Neural Network (ANN)&#x2014;revealed that the ANN model exhibited the highest predictive performance. The optimized ANN architecture (6&#x2013;14-14-14-1) achieved an accuracy of 95.3%, a Matthews correlation coefficient (MCC) of 0.894, and an F-score of 0.963, indicating strong consistency between DFT-predicted and ANN-classified crystal systems. Random Forest and Gaussian Process models also showed high accuracies (0.803 and 0.705, respectively) and served as robust complementary approaches, particularly when data are limited or computational efficiency is required.</p>
<p>This study establishes a reliable ML-based framework for classifying lithium manganese silicate crystal structures, providing a solid foundation for future generative work. Although the present focus is classification, the developed model and insights will guide our next phase, where we aim to extend the approach toward predicting and generating new crystal structures.</p>
</sec>
<sec id="s6">
<label>6</label>
<title>Limitations and Future Work</title>
<p>The present study is limited by the size and scope of the available dataset, which restricts the application of advanced validation strategies and additional thermodynamic analyses such as convex-hull stability mapping. In addition, the current framework is focused on accurate crystal-structure classification rather than generative prediction of new structures, which represents an important next step for real-world materials discovery. Future work will focus on expanding the dataset, incorporating comprehensive phase-stability information, and extending the model toward generative and predictive capabilities, complemented by thermodynamic calculations and experimental validation to further strengthen the robustness and generality of the proposed approach.</p>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This work was supported by the Learning &#x0026; Academic Research Institution for Master&#x2019;s, PhD students, and Postdocs LAMP Program of the National Research Foundation of Korea (NRF) grant funded by the Ministry of Education (No. RS-2023-00301974). This work was also supported by the Glocal University 30 Project fund of Gyeongsang National University in 2025.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>Conceptualization, Muhammad Ishtiaq and Nagireddy Gari Subba Reddy; methodology, Muhammad Ishtiaq and Yeon-Ju Lee; software, Annabathini Geetha Bhavani and Nagireddy Gari Subba Reddy; validation, Annabathini Geetha Bhavani and Yeon-Ju Lee; formal analysis, Yeon-Ju Lee and Muhammad Ishtiaq, investigation, Muhammad Ishtiaq and Nagireddy Gari Subba Reddy; resources, Sung-Gyu Kang; data curation, Annabathini Geetha Bhavani and Yeon-Ju Lee; writing&#x2014;original draft preparation, Muhammad Ishtiaq; writing&#x2014;review and editing, Sung-Gyu Kang and Nagireddy Gari Subba Reddy; supervision, Nagireddy Gari Subba Reddy and Sung-Gyu Kang. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The authors confirm that the data supporting the findings of this study are available within the article.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<glossary content-type="abbreviations" id="glossary-1">
<title>Abbreviations</title>
<def-list>
<def-item>
<term>ANN</term>
<def>
<p>Artificial Neural Network</p>
</def>
</def-item>
<def-item>
<term>DFT</term>
<def>
<p>Density Functional Theory</p>
</def>
</def-item>
<def-item>
<term>MCC</term>
<def>
<p>Matthews correlation coefficient</p>
</def>
</def-item>
<def-item>
<term>NPV</term>
<def>
<p>Negative Predictive Value</p>
</def>
</def-item>
<def-item>
<term>FP</term>
<def>
<p>False Positives</p>
</def>
</def-item>
<def-item>
<term>TP</term>
<def>
<p>True Positives</p>
</def>
</def-item>
<def-item>
<term>TN</term>
<def>
<p>True Negatives</p>
</def>
</def-item>
<def-item>
<term>FN</term>
<def>
<p>False Negatives</p>
</def>
</def-item>
</def-list>
</glossary>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Roy</surname> <given-names>H</given-names></string-name>, <string-name><surname>Roy</surname> <given-names>BN</given-names></string-name>, <string-name><surname>Hasanuzzaman</surname> <given-names>M</given-names></string-name>, <string-name><surname>Islam</surname> <given-names>MS</given-names></string-name>, <string-name><surname>Abdel-Khalik</surname> <given-names>AS</given-names></string-name>, <string-name><surname>Hamad</surname> <given-names>MS</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Global advancements and current challenges of electric vehicle batteries and their prospects: a comprehensive review</article-title>. <source>Sustainability</source>. <year>2022</year>;<volume>14</volume>(<issue>24</issue>):<fpage>16684</fpage>. doi:<pub-id pub-id-type="doi">10.3390/su142416684</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ali Ijaz Malik</surname> <given-names>M</given-names></string-name>, <string-name><surname>Kalam</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Ikram</surname> <given-names>A</given-names></string-name>, <string-name><surname>Zeeshan</surname> <given-names>S</given-names></string-name>, <string-name><surname>Raza Zahidi</surname> <given-names>SQ</given-names></string-name></person-group>. <article-title>Energy transition towards electric vehicle technology: recent advancements</article-title>. <source>Energy Rep</source>. <year>2025</year>;<volume>13</volume>:<fpage>2958</fpage>&#x2013;<lpage>96</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.egyr.2025.02.029</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>N</given-names></string-name>, <string-name><surname>Yin</surname> <given-names>J</given-names></string-name>, <string-name><surname>Li</surname> <given-names>H</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Cui</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>W</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Recent advance in Mn-based Li-rich cathode materials: oxygen release mechanism and its solution strategies based on electronic structure perspective, spanning from commercial liquid batteries to all-solid-state batteries</article-title>. <source>Next Mater</source>. <year>2025</year>;<volume>6</volume>:<fpage>100408</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.nxmate.2024.100408</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lee</surname> <given-names>S</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>W</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>SH</given-names></string-name>, <string-name><surname>Joo</surname> <given-names>SH</given-names></string-name>, <string-name><surname>Nam</surname> <given-names>G</given-names></string-name>, <string-name><surname>Oh</surname> <given-names>P</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Oxygen vacancy diffusion and condensation in lithium-ion battery cathode materials</article-title>. <source>Angew Chem Int Ed</source>. <year>2019</year>;<volume>58</volume>(<issue>31</issue>):<fpage>10478</fpage>&#x2013;<lpage>85</lpage>. doi:<pub-id pub-id-type="doi">10.1002/anie.201904469</pub-id>; <pub-id pub-id-type="pmid">31119837</pub-id></mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Arroyo-de Dompablo</surname> <given-names>ME</given-names></string-name>, <string-name><surname>Armand</surname> <given-names>M</given-names></string-name>, <string-name><surname>Tarascon</surname> <given-names>JM</given-names></string-name>, <string-name><surname>Amador</surname> <given-names>U</given-names></string-name></person-group>. <article-title>On-demand design of polyoxianionic cathode materials based on electronegativity correlations: an exploration of the Li<sub>2</sub>MSiO<sub>4</sub> system (M&#x003D;Fe, Mn, Co, Ni)</article-title>. <source>Electrochem Commun</source>. <year>2006</year>;<volume>8</volume>(<issue>8</issue>):<fpage>1292</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.elecom.2006.06.003</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Duncan</surname> <given-names>H</given-names></string-name>, <string-name><surname>Kondamreddy</surname> <given-names>A</given-names></string-name>, <string-name><surname>Mercier</surname> <given-names>PHJ</given-names></string-name>, <string-name><surname>Le Page</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Abu-Lebdeh</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Couillard</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Novel <italic>Pn</italic> polymorph for Li<sub>2</sub>MnSiO<sub>4</sub> and its electrochemical activity as a cathode material in Li-ion batteries</article-title>. <source>Chem Mater</source>. <year>2011</year>;<volume>23</volume>(<issue>24</issue>):<fpage>5446</fpage>&#x2013;<lpage>56</lpage>. doi:<pub-id pub-id-type="doi">10.1021/cm202793j</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Luo</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>X</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>T</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Exploring the structural properties of cathode and anode materials in Li-ion battery via neutron diffraction technique</article-title>. <source>Chin J Struct Chem</source>. <year>2023</year>;<volume>42</volume>(<issue>5</issue>):<fpage>100032</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.cjsc.2023.100032</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nowakowski</surname> <given-names>P</given-names></string-name>, <string-name><surname>Bonifacio</surname> <given-names>C</given-names></string-name>, <string-name><surname>Ray</surname> <given-names>M</given-names></string-name>, <string-name><surname>Fischione</surname> <given-names>P</given-names></string-name></person-group>. <article-title>The crystal orientation of Li metal anodes: a better understanding of lithium-ion solid-state batteries</article-title>. <source>Microsc Microanal</source>. <year>2024</year>;<volume>30</volume>:<fpage>ozae044.878</fpage>. doi:<pub-id pub-id-type="doi">10.1093/mam/ozae044.878</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jain</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ong</surname> <given-names>SP</given-names></string-name>, <string-name><surname>Hautier</surname> <given-names>G</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>W</given-names></string-name>, <string-name><surname>Richards</surname> <given-names>WD</given-names></string-name>, <string-name><surname>Dacek</surname> <given-names>S</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Commentary: the materials project: a materials genome approach to accelerating materials innovation</article-title>. <source>APL Mater</source>. <year>2013</year>;<volume>1</volume>:<fpage>011002</fpage>. doi:<pub-id pub-id-type="doi">10.1063/1.4812323</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ren</surname> <given-names>D</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Han</surname> <given-names>S</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>W</given-names></string-name></person-group>. <article-title>Harmonizing physical and deep learning modeling: a computationally efficient and interpretable approach for property prediction</article-title>. <source>Scr Mater</source>. <year>2025</year>;<volume>255</volume>:<fpage>116350</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.scriptamat.2024.116350</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shandiz</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Gauvin</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Application of machine learning methods for the prediction of crystal system of cathode materials in lithium-ion batteries</article-title>. <source>Comput Mater Sci</source>. <year>2016</year>;<volume>117</volume>:<fpage>270</fpage>&#x2013;<lpage>8</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.commatsci.2016.02.021</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Fearn</surname> <given-names>T</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Choy</surname> <given-names>KL</given-names></string-name></person-group>. <article-title>Machine-learning approach for predicting the discharging capacities of doped lithium nickel-cobalt&#x2013;manganese cathode materials in Li-ion batteries</article-title>. <source>ACS Cent Sci</source>. <year>2021</year>;<volume>7</volume>(<issue>9</issue>):<fpage>1551</fpage>&#x2013;<lpage>60</lpage>. doi:<pub-id pub-id-type="doi">10.1021/acscentsci.1c00611</pub-id>; <pub-id pub-id-type="pmid">34584957</pub-id></mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ng</surname> <given-names>MF</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Conduit</surname> <given-names>GJ</given-names></string-name>, <string-name><surname>Seh</surname> <given-names>ZW</given-names></string-name></person-group>. <article-title>Predicting the state of charge and health of batteries using data-driven machine learning</article-title>. <source>Nat Mach Intell</source>. <year>2020</year>;<volume>2</volume>(<issue>3</issue>):<fpage>161</fpage>&#x2013;<lpage>70</lpage>. doi:<pub-id pub-id-type="doi">10.1038/s42256-020-0156-7</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Attention mechanism-based neural network for prediction of battery cycle life in the presence of missing data</article-title>. <source>Batteries</source>. <year>2024</year>;<volume>10</volume>(<issue>7</issue>):<fpage>229</fpage>. doi:<pub-id pub-id-type="doi">10.3390/batteries10070229</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Feng</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>M</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>R</given-names></string-name></person-group>. <article-title><italic>In-situ</italic> battery life prognostics amid mixed operation conditions using physics-driven machine learning</article-title>. <source>J Power Sources</source>. <year>2023</year>;<volume>577</volume>:<fpage>233246</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jpowsour.2023.233246</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Longo</surname> <given-names>RC</given-names></string-name>, <string-name><surname>Xiong</surname> <given-names>K</given-names></string-name>, <string-name><surname>Santosh</surname> <given-names>KC</given-names></string-name>, <string-name><surname>Cho</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Crystal structure and multicomponent effects in Tetrahedral Silicate Cathode Materials for Rechargeable Li-ion Batteries</article-title>. <source>Electrochim Acta</source>. <year>2014</year>;<volume>121</volume>:<fpage>434</fpage>&#x2013;<lpage>42</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.electacta.2013.12.104</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Prosini</surname> <given-names>PP</given-names></string-name></person-group>. <article-title>Crystal group prediction for lithiated manganese oxides using machine learning</article-title>. <source>Batteries</source>. <year>2023</year>;<volume>9</volume>(<issue>2</issue>):<fpage>112</fpage>. doi:<pub-id pub-id-type="doi">10.3390/batteries9020112</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sun</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Li</surname> <given-names>P</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>X</given-names></string-name></person-group>. <article-title>An improved random forest based on the classification accuracy and correlation measurement of decision trees</article-title>. <source>Expert Syst Appl</source>. <year>2024</year>;<volume>237</volume>:<fpage>121549</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.eswa.2023.121549</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Rokach</surname> <given-names>L</given-names></string-name>, <string-name><surname>Maimon</surname> <given-names>O</given-names></string-name></person-group>. <chapter-title>Decision trees</chapter-title>. In: <person-group person-group-type="editor"><string-name><surname>Maimon</surname> <given-names>O</given-names></string-name>, <string-name><surname>Rokach</surname> <given-names>L</given-names></string-name></person-group>, editors. <source>Data mining and knowledge discovery handbook</source>. <publisher-loc>Boston, MA, USA</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2006</year>. p. <fpage>165</fpage>&#x2013;<lpage>92</lpage>. doi:<pub-id pub-id-type="doi">10.1007/0-387-25465-x_9</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>T</given-names></string-name>, <string-name><surname>Guestrin</surname> <given-names>C</given-names></string-name></person-group>. <article-title>XGBoost: a scalable tree boosting system</article-title>. In: <conf-name>Proceedings of the 22nd ACM Sigkdd International Conference on Knowledge Discovery and Data Mining; 2016 Aug 13&#x2013;17</conf-name>; <publisher-loc>San Francisco, CA, USA</publisher-loc>. doi:<pub-id pub-id-type="doi">10.1145/2939672.2939785</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Halder</surname> <given-names>RK</given-names></string-name>, <string-name><surname>Uddin</surname> <given-names>MN</given-names></string-name>, <string-name><surname>Uddin</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Aryal</surname> <given-names>S</given-names></string-name>, <string-name><surname>Khraisat</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Enhancing K-nearest neighbor algorithm: a comprehensive review and performance analysis of modifications</article-title>. <source>J Big Data</source>. <year>2024</year>;<volume>11</volume>(<issue>1</issue>):<fpage>113</fpage>. doi:<pub-id pub-id-type="doi">10.1186/s40537-024-00973-y</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tian</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Recent advances in stochastic gradient descent in deep learning</article-title>. <source>Mathematics</source>. <year>2023</year>;<volume>11</volume>(<issue>3</issue>):<fpage>682</fpage>. doi:<pub-id pub-id-type="doi">10.3390/math11030682</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Deringer</surname> <given-names>VL</given-names></string-name>, <string-name><surname>Bart&#x00F3;k</surname> <given-names>AP</given-names></string-name>, <string-name><surname>Bernstein</surname> <given-names>N</given-names></string-name>, <string-name><surname>Wilkins</surname> <given-names>DM</given-names></string-name>, <string-name><surname>Ceriotti</surname> <given-names>M</given-names></string-name>, <string-name><surname>Cs&#x00E1;nyi</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Gaussian process regression for materials and molecules</article-title>. <source>Chem Rev</source>. <year>2021</year>;<volume>121</volume>(<issue>16</issue>):<fpage>10073</fpage>&#x2013;<lpage>141</lpage>. doi:<pub-id pub-id-type="doi">10.1021/acs.chemrev.1c00022</pub-id>; <pub-id pub-id-type="pmid">34398616</pub-id></mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Peretz</surname> <given-names>O</given-names></string-name>, <string-name><surname>Koren</surname> <given-names>M</given-names></string-name>, <string-name><surname>Koren</surname> <given-names>O</given-names></string-name></person-group>. <article-title>Naive Bayes classifier-An ensemble procedure for recall and precision enrichment</article-title>. <source>Eng Appl Artif Intell</source>. <year>2024</year>;<volume>136</volume>:<fpage>108972</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.engappai.2024.108972</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Cortes</surname> <given-names>C</given-names></string-name>, <string-name><surname>Vapnik</surname> <given-names>V</given-names></string-name></person-group>. <article-title>Support-vector networks</article-title>. <source>Mach Learn</source>. <year>1995</year>;<volume>20</volume>(<issue>3</issue>):<fpage>273</fpage>&#x2013;<lpage>97</lpage>. doi:<pub-id pub-id-type="doi">10.1007/BF00994018</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>SC</given-names></string-name></person-group>. <chapter-title>Artificial neural network</chapter-title>. In: <source>Interdisciplinary computing in Java programming</source>. <publisher-loc>Boston, MA, USA</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2003</year>. p. <fpage>81</fpage>&#x2013;<lpage>100</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-1-4615-0377-4</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hinuma</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Hayashi</surname> <given-names>H</given-names></string-name>, <string-name><surname>Kumagai</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Tanaka</surname> <given-names>I</given-names></string-name>, <string-name><surname>Oba</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Comparison of approximations in density functional theory calculations: energetics and structure of binary oxides</article-title>. <source>Phys Rev B</source>. <year>2017</year>;<volume>96</volume>(<issue>9</issue>):<fpage>094102</fpage>. doi:<pub-id pub-id-type="doi">10.1103/physrevb.96.094102</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>