<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="review-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">67427</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2025.067427</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Review</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Large Language Model-Driven Knowledge Discovery for Designing Advanced Micro/Nano Electrocatalyst Materials</article-title>
<alt-title alt-title-type="left-running-head">Large Language Model-Driven Knowledge Discovery for Designing Advanced Micro/Nano Electrocatalyst Materials</alt-title>
<alt-title alt-title-type="right-running-head">Large Language Model-Driven Knowledge Discovery for Designing Advanced Micro/Nano Electrocatalyst Materials</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Shen</surname><given-names>Ying</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Zhao</surname><given-names>Shichao</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Lv</surname><given-names>Yanfei</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Chen</surname><given-names>Fei</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-5" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Fu</surname><given-names>Li</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><email>fuli@hdu.edu.cn</email></contrib>
<contrib id="author-6" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Karimi-Maleh</surname><given-names>Hassan</given-names></name><xref ref-type="aff" rid="aff-2">2</xref><email>hassan@wmu.edu.cn</email></contrib>
<aff id="aff-1"><label>1</label><institution>College of Materials and Environmental Engineering, Hangzhou Dianzi University</institution>, <addr-line>Hangzhou, 310018</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>The Quzhou Affiliated Hospital of Wenzhou Medical University, Quzhou People&#x2019;s Hospital</institution>, <addr-line>Quzhou, 324000</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Authors: Li Fu. Email: <email>fuli@hdu.edu.cn</email>; Hassan Karimi-Maleh. Email: <email>hassan@wmu.edu.cn</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2025</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>03</day><month>07</month><year>2025</year>
</pub-date>
<volume>84</volume>
<issue>2</issue>
<fpage>1921</fpage>
<lpage>1950</lpage>
<history>
<date date-type="received">
<day>03</day>
<month>5</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>12</day>
<month>6</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 The Authors.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_67427.pdf"></self-uri>
<abstract>
<p>This review presents a comprehensive and forward-looking analysis of how Large Language Models (LLMs) are transforming knowledge discovery in the rational design of advanced micro/nano electrocatalyst materials. Electrocatalysis is central to sustainable energy and environmental technologies, but traditional catalyst discovery is often hindered by high complexity, fragmented knowledge, and inefficiencies. LLMs, particularly those based on Transformer architectures, offer unprecedented capabilities in extracting, synthesizing, and generating scientific knowledge from vast unstructured textual corpora. This work provides the first structured synthesis of how LLMs have been leveraged across various electrocatalysis tasks, including automated information extraction from literature, text-based property prediction, hypothesis generation, synthesis planning, and knowledge graph construction. We comparatively analyze leading LLMs and domain-specific frameworks (e.g., CatBERTa, CataLM, CatGPT) in terms of methodology, application scope, performance metrics, and limitations. Through curated case studies across key electrocatalytic reactions&#x2014;HER, OER, ORR, and CO<sub>2</sub>RR&#x2014;we highlight emerging trends such as the growing use of embedding-based prediction, retrieval-augmented generation, and fine-tuned scientific LLMs. The review also identifies persistent challenges, including data heterogeneity, hallucination risks, lack of standard benchmarks, and limited multimodal integration. Importantly, we articulate future research directions, such as the development of multimodal and physics-informed MatSci-LLMs, enhanced interpretability tools, and the integration of LLMs with self-driving laboratories for autonomous discovery. By consolidating fragmented advances and outlining a unified research roadmap, this review provides valuable guidance for both materials scientists and AI practitioners seeking to accelerate catalyst innovation through large language model technologies.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Large language models</kwd>
<kwd>electrocatalysis</kwd>
<kwd>nanomaterials</kwd>
<kwd>knowledge discovery</kwd>
<kwd>materials design</kwd>
<kwd>artificial intelligence</kwd>
<kwd>natural language processing</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Advanced electrocatalyst materials, particularly those engineered at the micro- and nano-scale, stand at the forefront of solutions addressing critical global challenges in sustainable energy and environmental stewardship [<xref ref-type="bibr" rid="ref-1">1</xref>]. Their importance is underscored by their central role in enabling key technologies such as fuel cells for clean power generation [<xref ref-type="bibr" rid="ref-2">2</xref>], water electrolyzers for green hydrogen production [<xref ref-type="bibr" rid="ref-3">3</xref>], systems for converting greenhouse gases like CO<sub>2</sub> into valuable chemicals via CO<sub>2</sub> reduction [<xref ref-type="bibr" rid="ref-4">4</xref>], and biomass/waste valorization through electrochemical upgrading of compounds such as waste glycerol [<xref ref-type="bibr" rid="ref-5">5</xref>], bio-oil [<xref ref-type="bibr" rid="ref-6">6</xref>], and waste cooking oil [<xref ref-type="bibr" rid="ref-7">7</xref>] into value-added fuels and chemicals [<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-9">9</xref>]. The efficiency, selectivity, and durability of these electrochemical processes are fundamentally dictated by the performance of the electrocatalyst employed. Consequently, the development of catalysts exhibiting superior activity, enhanced selectivity towards desired products, and robust stability under demanding operational conditions remains a paramount objective in materials science and chemical engineering research. Micro/nano structuring improves catalytic performance by increasing surface area and exposing more active sites, such as edges and defects. Despite decades of research, the traditional path to discovering and optimizing electrocatalysts is often arduous and inefficient [<xref ref-type="bibr" rid="ref-10">10</xref>&#x2013;<xref ref-type="bibr" rid="ref-12">12</xref>]. Traditional discovery often depends on chemical intuition, time-consuming experiments, and DFT-based simulations [<xref ref-type="bibr" rid="ref-13">13</xref>&#x2013;<xref ref-type="bibr" rid="ref-15">15</xref>]. These approaches face significant hurdles. The combinatorial chemical space encompassing potential catalyst compositions, structures, and morphologies is astronomically vast, rendering exhaustive exploration practically impossible [<xref ref-type="bibr" rid="ref-12">12</xref>,<xref ref-type="bibr" rid="ref-14">14</xref>,<xref ref-type="bibr" rid="ref-16">16</xref>]. Furthermore, the relationship between a material&#x2019;s atomic-level structure, its electronic properties, and its ultimate catalytic performance is exceptionally complex and often non-intuitive, making rational design challenging [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-12">12</xref>,<xref ref-type="bibr" rid="ref-17">17</xref>]. Although fundamental studies using model systems are insightful, they often face the &#x201C;materials gap&#x201D; and &#x201C;pressure gap&#x201D;. This means findings from idealized conditions, such as single crystals under ultrahigh vacuum, may not apply to real-world nanoparticle catalysts. The complex structure of nanocatalysts often exceeds the ability of current tools to identify active sites or explain reaction mechanisms.</p>
<p>Recognizing these limitations, the scientific community is increasingly embracing a data-driven paradigm, often referred to as the &#x201C;fourth paradigm of science&#x201D;, to accelerate materials discovery [<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-18">18</xref>&#x2013;<xref ref-type="bibr" rid="ref-20">20</xref>]. This approach leverages the power of Artificial Intelligence (AI) and Machine Learning (ML) techniques to analyze large datasets, uncover hidden patterns, predict material properties, and guide experimental efforts [<xref ref-type="bibr" rid="ref-10">10</xref>,<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-20">20</xref>&#x2013;<xref ref-type="bibr" rid="ref-23">23</xref>]. The exponential growth in both computational power and the availability of materials data, generated through high-throughput simulations and automated experiments, has fueled this transition. ML models have demonstrated success in predicting catalyst properties and screening candidate materials based on structured datasets derived from simulations or experiments [<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-20">20</xref>,<xref ref-type="bibr" rid="ref-23">23</xref>]. These traditional ML approaches typically rely on numerical descriptors or graph-based representations derived from well-structured input data, such as atomic coordinates or engineered features. While powerful, their utility is often limited by the availability of high-quality, labeled datasets and their inability to fully leverage the rich but unstructured scientific knowledge embedded in textual sources.</p>
<p>In contrast, LLMs represent a fundamentally different class of AI systems. Built on Transformer architectures and trained on vast corpora of unstructured text, LLMs are capable of understanding and generating human language, making them especially suited for processing scientific literature. Unlike conventional ML models that operate primarily on structured numerical input, LLMs can perform tasks such as text mining, summarization, question answering, and entity-relation extraction directly from natural language. These abilities enable LLMs to unlock and synthesize latent scientific insights from previously untapped textual resources, offering a powerful complement to existing data-driven catalyst design strategies [<xref ref-type="bibr" rid="ref-24">24</xref>]. Transformer architecture and trained on massive corpora of scientific literature, LLMs excel at tasks such as text mining, summarization, question answering, and entity-relation extraction. These capabilities are particularly relevant to catalysis research, which is characterized by extensive unstructured knowledge dispersed across publications, patents, and technical reports [<xref ref-type="bibr" rid="ref-25">25</xref>]. Recent work has demonstrated the potential of LLMs to extract synthesis protocols and performance metrics from catalyst literature [<xref ref-type="bibr" rid="ref-26">26</xref>], build structured databases from text, and even assist in predicting catalyst properties based on natural language descriptions. For instance, LLM-based models such as SciBERT and CatBERTa have been used to extract materials and adsorption energy data with high accuracy, and domain-specific models like CataLM have been fine-tuned specifically for electrocatalytic materials. Furthermore, generative models such as CatGPT have been trained to propose novel catalyst compositions and reaction mechanisms, enabling hypothesis generation at scale. These developments illustrate how LLMs are being increasingly integrated into catalyst design pipelines&#x2014;either as standalone tools for literature analysis or as components in multi-agent or closed-loop discovery systems. By processing the language of science itself, LLMs are beginning to serve not only as data extraction engines, but as active collaborators in the generation of knowledge and the acceleration of catalyst innovation. This is vital because most scientific knowledge exists in unstructured text, not in structured databases [<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-27">27</xref>&#x2013;<xref ref-type="bibr" rid="ref-29">29</xref>]. Traditional ML methods typically struggle to directly utilize this vast textual resource, often requiring laborious manual data extraction or sophisticated, domain-specific Natural Language Processing (NLP) pipelines. LLMs, however, offer a more direct route to harnessing this knowledge, acting potentially as &#x201C;second brains&#x201D; for researchers by processing and synthesizing information from literature at an unprecedented scale [<xref ref-type="bibr" rid="ref-29">29</xref>].</p>
<p>This review focuses specifically on the application of LLMs to drive knowledge discovery for the purpose of designing advanced micro/nano electrocatalyst materials. <xref ref-type="fig" rid="fig-1">Fig. 1</xref> provides a schematic roadmap outlining the core roles LLMs play in this context. LLMs assist in several key tasks: extracting synthesis protocols and performance data from the literature, predicting catalytic properties using semantic understanding, generating hypothetical structures or mechanisms, and integrating knowledge via summarization or knowledge graphs. By capturing the entire knowledge discovery cycle&#x2014;ranging from data extraction to hypothesis formulation&#x2014;the figure visually reinforces the manuscript&#x2019;s central thesis that LLMs are emerging as comprehensive enablers in catalyst design.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Schematic overview of how LLMs can be applied for knowledge discovery in electrocatalyst design</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67427-fig-1.tif"/>
</fig>
</sec>
<sec id="s2">
<label>2</label>
<title>Foundational Concepts: Advanced Electrocatalysis and LLMs</title>
<p>A comprehensive understanding of the LLM-driven approach to electrocatalyst design necessitates familiarity with the fundamental aspects of both the target materials and the AI tools being employed. This section provides foundational concepts for advanced micro/nano electrocatalyst materials and LLMs, setting the stage for subsequent discussions on their intersection.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Advanced Micro/Nano Electrocatalyst Materials</title>
<p>Electrocatalysis is fundamentally the process by which the rate of an electrochemical reaction occurring at an electrode-electrolyte interface is enhanced through the action of a catalytic material. These reactions are central to numerous energy conversion and storage technologies, including fuel cells, electrolyzers, and metal-air batteries, as well as environmental applications like CO<sub>2</sub> conversion and pollutant degradation [<xref ref-type="bibr" rid="ref-30">30</xref>]. Advanced electrocatalysts are engineered with unique structural or compositional features, often at the micro/nanoscale (1&#x2013;100 nm). Nanostructuring offers multiple benefits. It increases the surface area available to reactants, introduces quantum confinement and modified electronic properties, and exposes low-coordination sites like edges and corners that enhance catalytic activity [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-31">31</xref>]. These benefits lead to higher activity, better selectivity, and improved atom efficiency, making nano-electrocatalysts valuable for energy and environmental applications.</p>
<p>The field encompasses a diverse array of material types, many of which are being explored using LLM-driven approaches. Recent years have seen substantial advances in the rational design of micro/nano electrocatalysts that exhibit improved intrinsic activity, selectivity, and stability. For example, single-atom catalysts (SACs) with isolated metal atoms anchored on nitrogen-doped carbon matrices have demonstrated record-low overpotentials for HER and OER [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-32">32</xref>]. Transition metal dichalcogenides (TMDs), especially MoS<sub>2</sub> and WS<sub>2</sub>, have been modified through phase engineering and defect creation to enhance ORR and CO<sub>2</sub>RR performance [<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-30">30</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>]. High-entropy alloys (HEAs) have emerged as tunable electrocatalyst platforms with compositional flexibility, showing promising results for multi-functional catalysis [<xref ref-type="bibr" rid="ref-12">12</xref>,<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>,<xref ref-type="bibr" rid="ref-35">35</xref>]. Metal-organic frameworks (MOFs) and their derivatives are now widely investigated for their hierarchical porosity and tunable active sites, often serving as precursors for metal-N-C catalysts. Porous and hollow nanostructures, such as yolk-shell and nanoframe morphologies, are increasingly favored for their ability to enhance mass transport and facilitate intermediate desorption. <xref ref-type="table" rid="table-1">Table 1</xref> synthesizes recent advances across various categories of micro/nano electrocatalyst materials, providing a comparative overview of their unique structural features, target reactions, representative materials, and performance benchmarks. Metal nanoparticles, often based on noble metals like Pt, Pd, or Au, or non-noble transition metals, can be used either unsupported or dispersed on conductive supports like carbon materials, oxides, or nitrides [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-31">31</xref>]. Their catalytic properties are highly sensitive to size, shape, and composition. One-dimensional (1D) nanostructures, such as nanowires, nanorods, and nanotubes, offer directed electron transport pathways and high surface areas [<xref ref-type="bibr" rid="ref-15">15</xref>]. Two-dimensional (2D) materials, including graphene and its analogues, TMDs, MXenes, and the more recently explored metallenes (ultrathin metal nanosheets), provide extremely high surface exposure and unique electronic properties stemming from their reduced dimensionality [<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-30">30</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>]. High-entropy materials (HEMs), including HEAs and oxides, contain five or more main elements in nearly equal ratios. Their disordered atomic structures produce unique synergistic effects and a vast compositional space for catalyst design [<xref ref-type="bibr" rid="ref-12">12</xref>,<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>,<xref ref-type="bibr" rid="ref-35">35</xref>]. SACs represent the ultimate limit of atom efficiency, where individual metal atoms are dispersed on a support material, providing well-defined active sites and potentially unique catalytic behavior distinct from nanoparticles [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-32">32</xref>]. Other important classes include nanostructured metal oxides, sulfides, phosphides, and (oxy)hydroxides, which are often investigated for reactions like oxygen evolution [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-36">36</xref>], as well as porous materials like MOFs that can serve as catalyst precursors or platforms [<xref ref-type="bibr" rid="ref-30">30</xref>,<xref ref-type="bibr" rid="ref-37">37</xref>].</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Summary of recent advances in micro/nano electrocatalyst materials</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center">Catalyst type</th>
<th align="center">Key features</th>
<th align="center">Target reaction (s)</th>
<th align="center">Representative materials</th>
<th align="center">Performance highlights</th>
<th align="center">Refs.</th>
</tr>
</thead>
<tbody>
<tr>
<td>Single-Atom Catalysts (SACs)</td>
<td>Atomically dispersed metals on supports</td>
<td>HER, OER, ORR, CO<sub>2</sub>RR</td>
<td>Fe&#x2013;N&#x2013;C, Co&#x2013;N&#x2013;C</td>
<td>Overpotential &#x003C; 50 mV (HER); High TOF</td>
<td>[<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>]</td>
</tr>
<tr>
<td>Transition Metal Dichalcogenides (TMDs)</td>
<td>Layered 2D structures, defect engineering</td>
<td>HER, ORR, CO<sub>2</sub>RR</td>
<td>MoS<sub>2</sub>, WS<sub>2</sub></td>
<td>HER onset potential &#x007E;70 mV</td>
<td>[<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-36">36</xref>,<xref ref-type="bibr" rid="ref-39">39</xref>]</td>
</tr>
<tr>
<td>High-Entropy Alloys (HEAs)</td>
<td>Multicomponent equiatomic alloys</td>
<td>HER, ORR</td>
<td>PtPdNiFeCo, AgPdRu</td>
<td>Adjustable binding energies, low &#x03B7;<sub>10</sub>&#x002A;</td>
<td>[<xref ref-type="bibr" rid="ref-12">12</xref>,<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-40">40</xref>,<xref ref-type="bibr" rid="ref-41">41</xref>]</td>
</tr>
<tr>
<td>MOFs/Derived carbon composites</td>
<td>Hierarchical porosity, tunable active sites</td>
<td>CO<sub>2</sub>RR, OER</td>
<td>ZIF-8 derived Co&#x2013;N&#x2013;C</td>
<td>FE(CO) &#x003E; 90% (CO<sub>2</sub>RR)</td>
<td>[<xref ref-type="bibr" rid="ref-36">36</xref>,<xref ref-type="bibr" rid="ref-42">42</xref>]</td>
</tr>
<tr>
<td>Nanoframes &#x0026; hollow structures</td>
<td>Enhanced surface exposure, mass transport</td>
<td>OER, ORR</td>
<td>PtNi nanoframes, Co<sub>3</sub>O<sub>4</sub> hollow spheres</td>
<td>High stability, &#x03B7;<sub>10</sub> &#x003C; 300 mV (OER)</td>
<td>[<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-43">43</xref>]</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-1fn1" fn-type="other">
<p>Note: &#x03B7;<sub>10</sub>&#x002A; refers to the overpotential required to achieve a current density of 10 mA&#x00B7;cm<sup>&#x2212;2</sup>, which is a standard metric for assessing electrocatalytic activity. Lower &#x03B7;<sub>10</sub> values indicate higher catalytic efficiency, as less energy is required to drive the target electrochemical reaction at a practical current density.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>Evaluating the effectiveness of these diverse electrocatalysts requires standardized key performance indicators (KPIs). Catalyst activity is often measured by overpotential (&#x03B7;), the extra voltage needed to reach 10 mA cm<sup>&#x2212;2</sup> current density. Lower overpotential signifies higher activity. Other key metrics include onset potential, current density, Tafel slope for mechanism insight, and TOF per active site, though defining active sites can be difficult [<xref ref-type="bibr" rid="ref-16">16</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>]. Selectivity, crucial for reactions producing multiple products like CO<sub>2</sub>RR, is measured by the Faradaic efficiency (FE), representing the percentage of electrons contributing to the formation of the desired product [<xref ref-type="bibr" rid="ref-39">39</xref>,<xref ref-type="bibr" rid="ref-40">40</xref>]. Stability reflects a catalyst&#x2019;s ability to retain performance and structure during long-term use, often tracked by potential or current changes [<xref ref-type="bibr" rid="ref-38">38</xref>]. These KPIs are essential for comparing different materials and guiding the design process.</p>
<p>These advanced electrocatalysts are indispensable for driving key electrochemical reactions central to sustainable technologies. Prominent examples include the hydrogen evolution reaction (HER) and oxygen evolution reaction (OER) involved in water splitting for hydrogen production; the oxygen reduction reaction (ORR) and hydrogen oxidation reaction (HOR) critical for fuel cells; the carbon dioxide reduction reaction (CO<sub>2</sub>RR) for converting CO<sub>2</sub> into fuels and chemicals; the Nitrogen Reduction Reaction (NRR) for ammonia synthesis; and the oxidation of fuels like ethanol or methanol in direct fuel cells [<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-41">41</xref>,<xref ref-type="bibr" rid="ref-43">43</xref>]. However, the rational design of these high-performance micro/nano electrocatalysts remains a formidable task due to several intrinsic challenges [<xref ref-type="bibr" rid="ref-42">42</xref>]. The sheer complexity of these systems, involving the intricate interplay between composition, size, shape, surface structure, defects, and support interactions, makes predicting behavior difficult [<xref ref-type="bibr" rid="ref-11">11</xref>,<xref ref-type="bibr" rid="ref-12">12</xref>]. Establishing clear structure-property relationships that link these nanoscale features to macroscopic catalytic performance (activity, selectivity, stability) is often elusive [<xref ref-type="bibr" rid="ref-12">12</xref>,<xref ref-type="bibr" rid="ref-17">17</xref>]. The Sabatier principle states that effective catalysis requires moderate binding: too weak fails to activate, too strong poisons the surface [<xref ref-type="bibr" rid="ref-1">1</xref>]. Finding materials that strike this delicate balance for complex multi-step reactions is challenging. Precise synthesis control to reproducibly fabricate nanomaterials with desired structural attributes remains a significant hurdle. Characterizing these materials under real conditions is difficult, limiting understanding of their behavior and true active sites [<xref ref-type="bibr" rid="ref-38">38</xref>]. These challenges collectively highlight the need for advanced tools, such as LLMs, capable of navigating complexity, extracting knowledge from vast datasets, and accelerating the design cycle.</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title> LLMs for Scientific Discovery</title>
<p>LLMs represent a significant advancement in artificial intelligence, characterized by their massive scale (billions of parameters) and their ability to process and generate human-like text [<xref ref-type="bibr" rid="ref-24">24</xref>]. Their development has been fueled by breakthroughs in neural network architectures, particularly the Transformer, increased computational power, and the availability of vast amounts of training data.</p>
<p>The core principles underlying most modern LLMs involve several key components. The transformer architecture is fundamental, replacing the sequential processing of older recurrent neural networks (RNNs) with parallel processing enabled by attention mechanisms. Self-attention allows the model to weigh the importance of different words or tokens within an input sequence relative to each other, capturing contextual relationships and long-range dependencies effectively [<xref ref-type="bibr" rid="ref-24">24</xref>]. Variations like sparse attention, flash attention, and multi-query attention have been developed to improve efficiency and handle longer sequences. LLMs typically undergo a two-stage training process. First is pre-training, where the model learns general language understanding and world knowledge by training on enormous, diverse text corpora (often terabytes of data) using self-supervised objectives like predicting the next word in a sequence or filling in masked words. This phase imbues the model with broad linguistic competence. Second is fine-tuning, where the pre-trained model is further trained on smaller, more specific datasets tailored to particular tasks or domains. This adaptation stage includes instruction-tuning (training on instruction-response pairs to follow commands) and alignment-tuning (often using reinforcement learning from human feedback, RLHF) to make the model more helpful, honest, and harmless [<xref ref-type="bibr" rid="ref-44">44</xref>]. A remarkable aspect of LLMs is their emergent abilities&#x2014;capabilities like complex reasoning, planning, few-shot learning (performing tasks with only a few examples), and zero-shot learning (performing tasks without examples) that appear as model scale increases and were not explicitly programmed [<xref ref-type="bibr" rid="ref-45">45</xref>].</p>
<p>These underlying principles grant LLMs a suite of capabilities highly relevant to scientific knowledge discovery, particularly from the vast body of scientific literature. IE is a key capability, allowing LLMs to identify and extract specific pieces of information from unstructured text, such as experimental parameters, material compositions, property values, and synthesis steps [<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-24">24</xref>,<xref ref-type="bibr" rid="ref-29">29</xref>,<xref ref-type="bibr" rid="ref-46">46</xref>,<xref ref-type="bibr" rid="ref-47">47</xref>]. This includes tasks like Named Entity Recognition (NER) to identify relevant terms and Relation Extraction to understand how they are connected [<xref ref-type="bibr" rid="ref-19">19</xref>]. Knowledge synthesis and summarization enable LLMs to condense information from lengthy articles or multiple sources, generate literature reviews, and identify key findings, helping researchers grapple with information overload. Question answering (Q&#x0026;A) allows users to pose specific questions and receive answers based on the LLM&#x2019;s internal knowledge or provided documents. Techniques like retrieval-augmented generation (RAG), where the LLM retrieves relevant information from an external database or corpus before generating an answer, enhance factual grounding and accuracy [<xref ref-type="bibr" rid="ref-25">25</xref>,<xref ref-type="bibr" rid="ref-48">48</xref>]. LLMs can also contribute to hypothesis generation by analyzing existing literature to identify knowledge gaps, inconsistencies, or potential correlations that might suggest new research avenues or material candidates [<xref ref-type="bibr" rid="ref-49">49</xref>,<xref ref-type="bibr" rid="ref-50">50</xref>]. Furthermore, many LLMs possess strong code generation abilities, capable of writing code snippets for data analysis, simulation setup, or even controlling automated laboratory equipment based on natural language prompts [<xref ref-type="bibr" rid="ref-51">51</xref>,<xref ref-type="bibr" rid="ref-52">52</xref>]. Finally, their core text generation capability can assist researchers in drafting manuscripts, reports, or documentation [<xref ref-type="bibr" rid="ref-47">47</xref>,<xref ref-type="bibr" rid="ref-53">53</xref>]. <xref ref-type="fig" rid="fig-2">Fig. 2</xref> serves as a conceptual framework illustrating the multilayered architecture of LLMs&#x2014;from transformer-based attention mechanisms to fine-tuning strategies&#x2014;and maps these capabilities to practical functions in scientific discovery. The figure highlights key tasks such as named entity recognition, summarization, question answering, and code generation, all of which are crucial for mining and synthesizing information from scientific literature. This visual contextualizes how LLMs act as both knowledge extractors and generators, thereby facilitating end-to-end integration into research workflows, including literature analysis, database construction, and experimental planning.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Schematic illustration of how LLMs, through their core architecture and training processes, enable key capabilities such as IE, summarization, and code generation, thereby supporting various stages of scientific knowledge discovery, from literature analysis to hypothesis formulation and manuscript preparation</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67427-fig-2.tif"/>
</fig>
<p>Collectively, these capabilities position LLMs as powerful tools for a new knowledge discovery paradigm in materials science and other scientific fields. Traditional knowledge discovery often relies on structured databases, which cover only a fraction of published knowledge, or manual literature surveys, which are slow and limited in scope. LLMs offer the potential to bridge this gap by directly processing and interpreting the primary source of scientific knowledge&#x2014;the unstructured text of publications&#x2014;at scale. By extracting, synthesizing, and reasoning over this vast textual information, LLMs can potentially accelerate the identification of structure-property relationships, suggest novel materials, optimize experimental procedures, and ultimately speed up the cycle of scientific discovery in complex fields like electrocatalyst design.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>LLM Methodologies in Micro/Nano Electrocatalyst Research</title>
<p>The application of LLMs in the domain of micro/nano electrocatalyst research is rapidly evolving, with various methodologies being developed and tested to leverage their unique capabilities. These methodologies broadly fall into categories focused on extracting existing knowledge, predicting properties based on that knowledge, generating new hypotheses or plans, and synthesizing information across the literature.</p>
<sec id="s3_1">
<label>3.1</label>
<title>IE from Electrocatalysis Literature</title>
<p>In the pursuit of accelerating electrocatalyst discovery, one of the most formidable challenges lies in the accessibility of materials data. Unlike structured databases, a significant proportion of valuable knowledge remains embedded in the unstructured text of scientific publications, patents, and supplementary materials. IE using LLMs has emerged as a transformative solution to this problem. By converting free-text content into structured, machine-readable formats, LLMs enable the creation of comprehensive datasets that capture key experimental and material parameters&#x2014;thereby forming the foundation for downstream tasks such as property prediction, hypothesis generation, and synthesis planning [<xref ref-type="bibr" rid="ref-54">54</xref>,<xref ref-type="bibr" rid="ref-55">55</xref>]. The overarching goal is to create comprehensive, machine-readable datasets detailing materials, synthesis methods, experimental conditions, properties, and performance metrics from the vast electrocatalysis literature [<xref ref-type="bibr" rid="ref-54">54</xref>].</p>
<p>NER and relation extraction are fundamental tasks in this process. NER involves identifying specific entities within the text, such as catalyst materials (e.g., &#x2018;Pt nanoparticles&#x2019;, &#x2018;Cu-based alloys&#x2019;, &#x2018;perovskite oxides&#x2019;), precursors, reaction products (e.g., &#x2018;H<sub>2</sub>&#x2019;, &#x2018;CO&#x2019;, &#x2018;CH<sub>4</sub>&#x2019;), performance metrics (e.g., FE, overpotential, current density, Tafel slope), experimental parameters (e.g., temperature, pressure, electrolyte composition, potential range), synthesis operations (e.g., &#x2018;calcination&#x2019;, &#x2018;electrodeposition&#x2019;), and characterization techniques [<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-40">40</xref>,<xref ref-type="bibr" rid="ref-54">54</xref>,<xref ref-type="bibr" rid="ref-55">55</xref>]. Relation Extraction then identifies the semantic links between these entities, for example, connecting a specific catalyst material to its measured FE for a particular product under defined conditions.</p>
<p>The effectiveness of IE in electrocatalysis relies heavily on the capabilities of the underlying LLMs. Early efforts employed general-purpose models such as GPT-o1 and GPT-4o, which&#x2014;when guided by carefully designed prompts&#x2014;proved capable of extracting synthesis parameters and performance metrics with high accuracy from large corpora [<xref ref-type="bibr" rid="ref-37">37</xref>,<xref ref-type="bibr" rid="ref-46">46</xref>,<xref ref-type="bibr" rid="ref-54">54</xref>]. Examples include SciBERT [<xref ref-type="bibr" rid="ref-40">40</xref>], MatSciBERT [<xref ref-type="bibr" rid="ref-44">44</xref>,<xref ref-type="bibr" rid="ref-54">54</xref>], and CataLM [<xref ref-type="bibr" rid="ref-44">44</xref>], which is specifically tailored for electrocatalytic materials using the Vicuna-13B model as a base and trained on domain literature and expert annotations [<xref ref-type="bibr" rid="ref-44">44</xref>]. Comparative studies have shown promising results; for instance, GPT-4 outperformed a rule-based method (ChemDataExtractor) in extracting band gap information from materials science literature, demonstrating better handling of complex material names and interdependency resolution, although weaknesses in hallucination and identifying specific value types were noted [<xref ref-type="bibr" rid="ref-46">46</xref>]. In the context of CO<sub>2</sub> reduction electrocatalysis, a framework combining BERT embeddings with a BiLSTM-CRF architecture showed strong performance in recognizing key entities [<xref ref-type="bibr" rid="ref-56">56</xref>].</p>
<p>To enhance accuracy and mitigate issues like hallucination, various techniques are employed. Prompt engineering involves carefully crafting the input query to elicit the desired structured output from the LLM [<xref ref-type="bibr" rid="ref-37">37</xref>,<xref ref-type="bibr" rid="ref-46">46</xref>,<xref ref-type="bibr" rid="ref-54">54</xref>]. Fine-tuning adapts the model&#x2019;s parameters using domain-specific labeled data [<xref ref-type="bibr" rid="ref-44">44</xref>]. RAG approaches, often utilizing vector databases built from literature embeddings (e.g., using SciBERT), retrieve relevant context before generation, grounding the LLM&#x2019;s output and improving factual accuracy [<xref ref-type="bibr" rid="ref-25">25</xref>,<xref ref-type="bibr" rid="ref-54">54</xref>]. The ChatExtract method employs a conversational approach, using follow-up prompts to verify extracted data and introduce uncertainty checks, achieving high precision and recall for materials data extraction with models like GPT-4 [<xref ref-type="bibr" rid="ref-36">36</xref>].</p>
<p>These extraction efforts are enabling the creation of valuable resources. Several projects have focused on building specialized corpora, such as the benchmark and extended corpora for the CO<sub>2</sub> electrocatalytic reduction process, containing thousands of manually verified or automatically extracted records detailing materials, methods, products, efficiencies, and conditions [<xref ref-type="bibr" rid="ref-54">54</xref>]. Such corpora serve as vital training and evaluation datasets for future NLP models in the field.</p>
<p>Beyond simple entity and property extraction, LLMs are being used to parse complex experimental synthesis procedures. This involves identifying the sequence of actions (e.g., mixing, heating, washing, annealing), associated parameters (temperature, time, concentration, atmosphere), starting materials, and target products described in narrative text [<xref ref-type="bibr" rid="ref-54">54</xref>,<xref ref-type="bibr" rid="ref-57">57</xref>,<xref ref-type="bibr" rid="ref-58">58</xref>]. Sequence-to-sequence Transformer models have been developed to convert these unstructured descriptions into structured &#x201C;action sequences&#x201D;, providing a machine-readable representation of the synthesis protocol that can be used for analysis, comparison, or even guiding automated synthesis platforms [<xref ref-type="bibr" rid="ref-54">54</xref>]. Examples include extracting synthesis parameters for inorganic materials [<xref ref-type="bibr" rid="ref-59">59</xref>] and MOFs [<xref ref-type="bibr" rid="ref-37">37</xref>].</p>
<p>While most current efforts focus on text, the challenge of extracting information from multimodal sources like tables and figures is recognized [<xref ref-type="bibr" rid="ref-27">27</xref>]. Early work includes extracting data from tables in conjunction with text, with specialized models like MaTableGPT emerging [<xref ref-type="bibr" rid="ref-58">58</xref>]. The future likely involves multimodal LLMs capable of interpreting images (plots, micrographs), tables, and text simultaneously [<xref ref-type="bibr" rid="ref-25">25</xref>,<xref ref-type="bibr" rid="ref-42">42</xref>]. LLM-driven IE is not simply a technical exercise in data extraction; it is a foundational enabler for the entire AI-driven catalyst discovery pipeline. From generating structured corpora and parsing synthesis protocols to integrating text with knowledge graphs and multimodal inputs, IE continues to evolve into a pivotal capability that underpins knowledge accessibility in electrocatalysis research. The narrative arc of these efforts demonstrates a clear trajectory: from isolated data extraction to the construction of intelligent, interconnected research ecosystems.</p>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Property Prediction and Structure-Property Relationships</title>
<p>A central goal in materials science is to predict the properties of a material based on its composition and structure. LLMs are introducing novel approaches to this challenge in electrocatalysis, moving beyond traditional methods that rely solely on explicit structural inputs or computationally expensive simulations [<xref ref-type="bibr" rid="ref-60">60</xref>]. These new methodologies leverage the ability of LLMs to understand and process textual information or to generate meaningful numerical representations (embeddings) from scientific language.</p>
<p>One promising direction involves text-based property prediction, where models predict properties directly from natural language descriptions of the material or catalytic system. CatBERTa, a model based on the RoBERTa Transformer encoder, exemplifies this approach [<xref ref-type="bibr" rid="ref-61">61</xref>]. CatBERTa, a domain-specific adaptation of the RoBERTa Transformer encoder, exemplifies the use of text-based property prediction in electrocatalysis. By encoding structured textual representations of adsorption systems (e.g., catalyst surface facets, adsorbates, site types), it enables semantic learning without needing explicit atomic coordinates. The model demonstrated comparable or superior performance to conventional GNNs like CGCNN on benchmark datasets such as OC20. A particularly impactful finding was CatBERTa&#x2019;s ability to minimize systematic error when evaluating energy differences among chemically similar systems, making it useful for high-throughput screening. However, its reliance on well-formatted text input can limit applicability in cases where such structured language representations are unavailable. It takes human-interpretable text strings as input, which can include information like the chemical symbols of the adsorbate and bulk catalyst, the crystal facet&#x2019;s Miller index, the type of adsorption site, and potential atomic properties. CatBERTa is fine-tuned to predict adsorption energies, a key descriptor for catalytic activity [<xref ref-type="bibr" rid="ref-61">61</xref>]. <xref ref-type="fig" rid="fig-3">Fig. 3</xref> illustrates the architecture and operational flow of CatBERTa, a domain-specific LLM fine-tuned for predicting adsorption energies in catalysis. The conversion of structural descriptors into textual inputs enables the model to derive embeddings using a Transformer encoder, which is then processed via a regression head to output predicted energy values. This framework is particularly powerful because it leverages semantic representations of material features, enabling effective prediction even in cases where atomic coordinates may be incomplete or unavailable. The visual aids in understanding the interpretability advantage and prediction mechanics of text-based LLM models in materials science. Its performance has been shown to be comparable to some established Graph Neural Networks (GNNs) like CGCNN and SchNet, particularly on certain subsets of data, achieving Mean Absolute Errors (MAE) in the range of 0.35&#x2013;0.82 eV depending on the dataset and input features. A notable strength of CatBERTa is its ability to significantly cancel systematic errors when predicting energy differences between chemically similar systems, outperforming GNNs in this aspect. This suggests its utility in comparative studies and screening. Other work has also explored using text descriptions generated by tools like Robocrystallographer as input for transformer models (BERT, MatBERT) to classify materials based on properties like formation energy and band gap, achieving high accuracy [<xref ref-type="bibr" rid="ref-60">60</xref>].</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Overview of CatBERTa: (<bold>a</bold>) Transformation of structural data into a textual format; The structural data undergoes conversion into two types of textual inputs: strings and descriptions; (<bold>b</bold>) Visualization of the fine-tuning process. The embedding from the special token &#x201C;&#x003C;s&#x003E;&#x201D; is input to the regression head, comprising a linear layer and an activation layer; (<bold>c</bold>) Illustration of the Transformer encoder and a multihead attention mechanism [<xref ref-type="bibr" rid="ref-61">61</xref>]</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67427-fig-3.tif"/>
</fig>
<p>Another approach utilizes embedding-based prediction. This involves representing materials, properties, or concepts as dense numerical vectors (embeddings) learned by NLP models trained on large text corpora. The relationships between these vectors can then be exploited for prediction. For instance, one study trained a Word2Vec model on materials science abstracts and used the cosine similarity between material composition vectors and property term vectors (e.g., &#x2018;conductivity&#x2019;, &#x2018;dielectric&#x2019;) as objectives for Pareto optimization [<xref ref-type="bibr" rid="ref-14">14</xref>]. This allowed the prediction and screening of candidate electrocatalyst compositions (e.g., Ag-Pd-Pt, Ag-Pd-Ru) for specific reactions like HER, OER, and ORR, purely based on latent knowledge extracted from text, with predictions matching experimental activity trends well [<xref ref-type="bibr" rid="ref-14">14</xref>]. Similarly, another approach combined word embeddings derived from literature (capturing semantic information) with graph embeddings derived from a constructed knowledge graph (capturing structural relationships) [<xref ref-type="bibr" rid="ref-40">40</xref>]. This hybrid embedding was fed into a deep learning model to successfully predict the FE of Cu-based catalysts for CO<sub>2</sub> reduction, demonstrating the power of integrating textual semantics with structured knowledge (<xref ref-type="fig" rid="fig-4">Fig. 4</xref>) [<xref ref-type="bibr" rid="ref-40">40</xref>].</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>(<bold>a</bold>) (Top) Stacked histograms of various Cu-based electrocatalysts in articles published in the last dozen years. (Bottom) Stacked histograms of the percentage of Cu-based electrocatalysts in articles normalized by year; (<bold>b</bold>) Overall representation of materials, products, and methods for CO<sub>2</sub> reduction. The size of the balls indicates the number of corresponding papers; (<bold>c</bold>) Alluvial plot showing the development of Cu-based alloy electrocatalysts across the last 30 years [<xref ref-type="bibr" rid="ref-40">40</xref>]</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67427-fig-4.tif"/>
</fig>
<p>A potential advantage of text-based LLM approaches is interpretability. Unlike many complex ML models often termed &#x201C;black boxes&#x201D;, text-based models offer avenues for understanding the prediction process. The attention mechanisms inherent in Transformer models like CatBERTa can be analyzed to reveal which parts of the input text (specific words or tokens, such as those related to the adsorbate or interacting atoms) the model focuses on when making a prediction [<xref ref-type="bibr" rid="ref-61">61</xref>]. This provides insights into the features the model deems important. Furthermore, using human-readable text descriptions as the primary input representation itself enhances interpretability, allowing researchers to more easily relate the model&#x2019;s input to their domain knowledge [<xref ref-type="bibr" rid="ref-60">60</xref>]. This contrasts with GNNs where interpretation often relies on analyzing learned graph features or using post-hoc methods like SHAP or LIME [<xref ref-type="bibr" rid="ref-17">17</xref>]. When compared with GNNs and traditional ML, LLM-based property prediction offers distinct characteristics. GNNs excel at capturing fine-grained structural information when precise atomic coordinates are available [<xref ref-type="bibr" rid="ref-61">61</xref>,<xref ref-type="bibr" rid="ref-62">62</xref>]. Traditional descriptor-based ML relies on carefully engineered features derived from physical or chemical principles [<xref ref-type="bibr" rid="ref-22">22</xref>,<xref ref-type="bibr" rid="ref-23">23</xref>]. LLMs, particularly text-based ones, operate on a different modality, leveraging semantic understanding and contextual information from language. This allows them to potentially incorporate qualitative knowledge, handle incomplete structural information, and utilize human-interpretable inputs. However, they might lack the explicit structural resolution of GNNs, leading to trade-offs in accuracy depending on the specific task and available data [<xref ref-type="bibr" rid="ref-61">61</xref>]. The choice between these approaches depends on the nature of the available data (structured coordinates vs. text descriptions), the importance of interpretability vs. raw predictive power, and the specific property being predicted. LLM-based methods provide a complementary pathway, enriching the toolkit for understanding and predicting electrocatalyst behavior.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Hypothesis Generation and Synthesis Planning</title>
<p>Beyond analyzing existing knowledge, LLMs are increasingly being explored for their generative capabilities&#x2014;their potential to propose novel materials, structures, or experimental plans, thereby acting as creative partners in the scientific discovery process [<xref ref-type="bibr" rid="ref-10">10</xref>,<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-48">48</xref>&#x2013;<xref ref-type="bibr" rid="ref-50">50</xref>,<xref ref-type="bibr" rid="ref-58">58</xref>,<xref ref-type="bibr" rid="ref-63">63</xref>]. This represents a significant shift towards using AI not just for prediction but also for exploration and innovation in electrocatalyst design.</p>
<p>One application is the generation of novel material compositions or structures. Inspired by the success of LLMs in generating text and code, researchers are training them on representations of materials to suggest new candidates. CatGPT, for example, is a GPT-based model trained on millions of catalyst structures (represented as tokenized strings of lattice parameters, atomic symbols, and coordinates) from the Open Catalyst database [<xref ref-type="bibr" rid="ref-64">64</xref>,<xref ref-type="bibr" rid="ref-65">65</xref>]. It can autoregressively generate string representations of new inorganic catalyst structures, including surface and adsorbate atoms. Fine-tuning CatGPT on a smaller dataset specific to the two-electron oxygen reduction reaction (2e-ORR) enabled the discovery of five novel and promising 2e-ORR catalyst candidates, demonstrating the potential for targeted discovery. Ensuring the chemical and structural validity of generated structures remains a challenge, requiring specialized validity metrics or post-generation filtering [<xref ref-type="bibr" rid="ref-64">64</xref>,<xref ref-type="bibr" rid="ref-65">65</xref>]. Other work explores generating crystal structures using LLMs trained on crystallographic information file (CIF) formats [<xref ref-type="bibr" rid="ref-66">66</xref>].</p>
<p>Frameworks are also being developed to leverage LLMs for generating scientific hypotheses. MOOSE-Chem is a multi-agent system that uses LLMs to tackle the complex task of hypothesis generation in chemistry [<xref ref-type="bibr" rid="ref-49">49</xref>]. MOOSE-Chem is a novel multi-agent LLM system developed to simulate the multi-step process of hypothesis generation in chemistry. It decomposes this process into document retrieval, knowledge synthesis, and hypothesis formulation. Notably, MOOSE-Chem was shown to independently rediscover hypotheses similar to those in high-impact 2024 papers, based on literature only available up to 2023. This demonstrates its potential to uncover latent connections in existing literature. However, its effectiveness heavily depends on the quality of the retrieval module and the diversity of the source documents. It breaks the process down into stages: retrieving relevant &#x201C;inspiration&#x201D; papers from literature based on a background question, synthesizing novel hypotheses by combining the background and inspirations and evaluating the quality of the generated hypotheses. Using LLMs trained on data up to 2023, MOOSE-Chem was able to rediscover hypotheses from high-impact chemistry papers published in 2024 with high similarity, showcasing the potential of LLMs to autonomously generate scientifically valid and novel ideas. Another approach, LLM-Feynman, combines LLM reasoning with optimization techniques to discover interpretable scientific formulae and theories from data, successfully rediscovering physics formulae and deriving accurate formulae for materials properties like synthesizability and ionic conductivity [<xref ref-type="bibr" rid="ref-50">50</xref>].</p>
<p>LLMs are also proving adept at synthesis planning and optimization, moving beyond simply extracting procedures to actively proposing or refining them. The ChemCrow agent, powered by GPT-4 and equipped with 18 expert-designed tools, demonstrated the ability to autonomously plan and conceptually execute the synthesis of organic molecules, including organocatalysts [<xref ref-type="bibr" rid="ref-10">10</xref>]. In the realm of inorganic nanomaterials, a framework was developed using fine-tuned open-source LLMs for the synthesis of quantum dots (QDs) [<xref ref-type="bibr" rid="ref-59">59</xref>,<xref ref-type="bibr" rid="ref-67">67</xref>,<xref ref-type="bibr" rid="ref-68">68</xref>]. A dedicated QD synthesis planner framework was recently introduced, combining a fine-tuned LLM for protocol generation with a property prediction module. This system takes as input the desired properties (e.g., emission wavelength, particle size) and a masked base protocol, and outputs optimized synthetic recipes. The generated protocols underwent computational validation, expert review, and experimental testing. Impressively, 3 of 6 proposed protocols succeeded in advancing the multi-objective Pareto front. This study illustrated the feasibility of integrating LLMs into closed-loop design cycles, though challenges remain in ensuring generalizability and robustness across materials domains. This system integrates a protocol generation model (taking target properties and a masked reference protocol as input) and a property prediction model. Generated protocols are validated computationally, assessed for novelty, evaluated by humans, and finally tested experimentally. This LLM-driven approach successfully generated QD synthesis protocols that improved target properties and updated the Pareto front for multi-objective optimization [<xref ref-type="bibr" rid="ref-68">68</xref>]. LLMs have also been used to suggest element libraries for exploring HEAs electrocatalysts for the ORR, providing a starting point for high-throughput experimental screening [<xref ref-type="bibr" rid="ref-34">34</xref>]. Furthermore, LLMs can predict suitable precursors for synthesizing target inorganic materials based on patterns learned from the literature [<xref ref-type="bibr" rid="ref-59">59</xref>] and assist in designing continuous flow microreactor systems [<xref ref-type="bibr" rid="ref-69">69</xref>].</p>
<p>Finally, LLMs can provide design recommendations or suggest specific experimental strategies. The CataLM model, for instance, was evaluated on a control method recommendation task for electrocatalytic materials, leveraging its fine-tuned knowledge base [<xref ref-type="bibr" rid="ref-44">44</xref>]. In a study on electrochemical C-H oxidation, LLMs were prompted to generate Python code that iteratively optimized reaction conditions to improve yields, demonstrating a collaborative human-AI approach to synthesis optimization (<xref ref-type="fig" rid="fig-5">Fig. 5</xref>) [<xref ref-type="bibr" rid="ref-51">51</xref>]. These examples highlight the potential of LLMs to not only generate static plans but also participate in dynamic optimization loops.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>(<bold>A</bold>) Design and assembly of the 24-well electrochemical platform and the schematic overview of the electrochemical C(sp<sup>3</sup>)&#x2212;H oxidation process using the electrocatalyst; (<bold>B</bold>) Semantic literature analysis for reaction data mining using a language model with natural language prompts. Performance is evaluated by comparing the ground truth with LLM-assigned labels and examining the impact of prompt quality; (<bold>C</bold>) Overview of training data preparation and machine learning models used for predicting electrochemical C&#x2212;H oxidation reactivity (Task 1) and selectivity (Task 2). Models with different architectures are evaluated for accuracy and AUC [<xref ref-type="bibr" rid="ref-51">51</xref>]</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67427-fig-5a.tif"/>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67427-fig-5b.tif"/>
</fig>
<p>The application of LLMs in hypothesis generation and synthesis planning is still nascent but holds immense promise for accelerating the creative and planning aspects of electrocatalyst research. Key challenges include ensuring the scientific validity and synthesizability of generated outputs and effectively integrating these AI-driven suggestions with experimental validation and refinement.</p>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Knowledge Synthesis and Integration</title>
<p>The exponential growth of scientific literature presents a significant challenge for researchers seeking to stay abreast of developments and synthesize existing knowledge [<xref ref-type="bibr" rid="ref-29">29</xref>]. LLMs, with their advanced natural language understanding and generation capabilities, offer powerful tools to manage this information deluge, integrate findings from disparate sources, and present synthesized knowledge in accessible formats [<xref ref-type="bibr" rid="ref-52">52</xref>].</p>
<p>One direct application is automated literature review and summarization. LLMs can process large numbers of research papers on a specific topic (e.g., propane dehydrogenation catalysts, CO<sub>2</sub> reduction catalysts) and generate comprehensive summaries or structured reviews [<xref ref-type="bibr" rid="ref-29">29</xref>,<xref ref-type="bibr" rid="ref-70">70</xref>]. These automated systems can significantly reduce the cognitive load on researchers and accelerate the process of understanding the state of the art. However, ensuring the factual accuracy and proper citation integrity of these generated reviews is critical. Multi-tier quality control strategies, potentially involving validation against multiple LLMs or expert verification, are necessary to mitigate the risk of hallucinations&#x2014;the generation of plausible but false information&#x2014;which is a known issue with current LLMs, especially in specialized domains. Studies have shown that with careful quality control, hallucination risks in generated reviews can be reduced significantly [<xref ref-type="bibr" rid="ref-29">29</xref>].</p>
<p>LLMs can also facilitate knowledge integration through knowledge graphs (KGs). KGs provide a structured way to represent entities (like materials, properties, and reactions) and the relationships between them [<xref ref-type="bibr" rid="ref-40">40</xref>]. LLMs can be employed in the construction of these KGs by automatically extracting entities and relationships from text. Furthermore, LLMs can interact with existing KGs, allowing researchers to query this structured knowledge base using natural language [<xref ref-type="bibr" rid="ref-25">25</xref>]. The synergy between the semantic understanding of LLMs and the structured representation of KGs is particularly powerful. For example, combining word embeddings (capturing textual context) with graph embeddings (capturing relational structure from a KG of Cu-based CO<sub>2</sub>RR catalysts) led to improved prediction of FE, demonstrating how LLMs can leverage both unstructured and structured knowledge sources.</p>
<p>By processing vast amounts of extracted data or interacting with KGs, LLMs can assist in trend analysis and insight generation. They can help identify historical developments in catalyst research, pinpoint emerging materials or synthesis techniques, and uncover correlations between different factors influencing catalytic performance (e.g., relationships between catalyst composition, regulation methods, and product selectivity in CO<sub>2</sub>RR) [<xref ref-type="bibr" rid="ref-40">40</xref>,<xref ref-type="bibr" rid="ref-70">70</xref>]. This ability to synthesize information across numerous studies can potentially reveal hidden connections or patterns that might be missed through manual review [<xref ref-type="bibr" rid="ref-19">19</xref>].</p>
<p>The development of intelligent Q&#x0026;A systems tailored for electrocatalysis is another promising direction. These systems, often built using LLMs fine-tuned on domain literature or grounded using RAG techniques with extracted data or KGs, allow researchers to ask specific questions and receive informative answers supported by evidence from the literature [<xref ref-type="bibr" rid="ref-25">25</xref>,<xref ref-type="bibr" rid="ref-37">37</xref>]. An example is the development of a data-grounded chatbot for answering questions about MOF synthesis procedures [<xref ref-type="bibr" rid="ref-37">37</xref>].</p>
<p>In essence, LLMs are emerging as crucial tools for knowledge management and synthesis in the rapidly expanding field of electrocatalysis. By automating literature reviews, facilitating the construction and use of knowledge graphs, enabling trend analysis, and powering intelligent Q&#x0026;A systems, they help researchers navigate, understand, and build upon the collective knowledge embedded within scientific publications, ultimately fostering a more integrated and accelerated discovery process.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Applications, Performance Analysis, and Comparative Insights</title>
<sec id="s4_1">
<label>4.1</label>
<title>Case Studies across Electrocatalytic Reactions and Materials</title>
<p>The versatility of LLMs is reflected in their application across the major electrocatalytic reactions critical for energy and environmental technologies. For the HER, which is fundamental to water splitting for hydrogen fuel production, LLMs and related NLP/ML techniques are being used to accelerate the discovery of efficient, low-cost catalysts, often aiming to replace expensive platinum-group metals [<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-71">71</xref>]. ML frameworks have been developed to screen large numbers of candidate alloys, identifying promising high-performance materials like AgPd, whose potential was subsequently verified experimentally and computationally under realistic conditions (<xref ref-type="fig" rid="fig-6">Fig. 6</xref>) [<xref ref-type="bibr" rid="ref-72">72</xref>]. Research specifically targets low-dimensional materials like nanoparticles, nanotubes, and nanosheets, leveraging ML to predict their HER performance based on various descriptors [<xref ref-type="bibr" rid="ref-15">15</xref>]. Text mining combined with Word2Vec embeddings and Pareto optimization has been applied to predict HER-active candidate compositions based on their textual similarity to relevant properties like conductivity [<xref ref-type="bibr" rid="ref-14">14</xref>]. LLMs are also envisioned to provide direct design guidance for HER catalysts [<xref ref-type="bibr" rid="ref-42">42</xref>]. High-throughput experimental methods combined with data-driven strategies, potentially guided by LLM-suggested element libraries, are accelerating the discovery of HER-active HEAs [<xref ref-type="bibr" rid="ref-73">73</xref>].</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>An ML framework for high-throughput screening of electrocatalysts. In the left of the section &#x201C;Constructing Adsorption Database&#x201D;: Adsorption sites for binary alloys, including ontop, bridge, and hollow sites, denoted by black star, red &#x201C;&#x002B;&#x201D;, and blue &#x201C;&#x00D7;&#x201D;, respectively [<xref ref-type="bibr" rid="ref-72">72</xref>]</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67427-fig-6.tif"/>
</fig>
<p>The OER, the typically sluggish anodic counterpart to HER in water splitting, is another key target. LLMs are being applied for predictive analytics, for example, in the context of (oxy)hydroxide-based OER catalysts [<xref ref-type="bibr" rid="ref-36">36</xref>]. While details are emerging, this suggests LLMs are used for tasks like extracting performance data from literature or predicting activity based on compositional or textual features for these materials. ML, often coupled with high-throughput experiments or simulations, is actively used to discover and optimize OER catalysts, including perovskite oxides, where active learning approaches have identified compositions with exceptionally low overpotentials. Text mining and embedding-based approaches, similar to those used for HER, are also applied to predict promising OER candidate compositions [<xref ref-type="bibr" rid="ref-14">14</xref>]. LLMs are expected to play a role in providing design guidance for OER catalysts as well [<xref ref-type="bibr" rid="ref-42">42</xref>]. Interpretable ML models are being developed to unify activity prediction across multiple reactions, including OER, using intrinsic material properties [<xref ref-type="bibr" rid="ref-39">39</xref>].</p>
<p>The ORR is crucial for fuel cells and metal-air batteries. LLM-based generative models like CatGPT have been fine-tuned specifically for discovering catalysts for the selective two-electron ORR pathway, which produces H<sub>2</sub>O<sub>2</sub>, a valuable chemical [<xref ref-type="bibr" rid="ref-64">64</xref>,<xref ref-type="bibr" rid="ref-65">65</xref>]. The model successfully generated novel candidate structures validated by further analysis. In another approach, an LLM provided an initial element library to guide the high-throughput experimental discovery of Pt-based quinary HEAs for ORR [<xref ref-type="bibr" rid="ref-34">34</xref>]. This involved microscale precursor printing and rapid screening with scanning electrochemical cell microscopy (SECCM), demonstrating a powerful synergy between LLM guidance and automated experimentation. Text mining and embedding-based methods are also used for predicting ORR candidates [<xref ref-type="bibr" rid="ref-14">14</xref>], and LLMs are anticipated to offer design guidance [<xref ref-type="bibr" rid="ref-42">42</xref>]. ML models are also being developed to optimize HEA compositions for ORR activity [<xref ref-type="bibr" rid="ref-35">35</xref>].</p>
<p>The CO<sub>2</sub>RR aims to convert CO<sub>2</sub> into valuable fuels and feedstocks, mitigating greenhouse gas emissions. This area has seen significant LLM application, particularly in knowledge extraction and synthesis. LLM-enhanced methods have been used to create large corpora detailing CO<sub>2</sub>RR electrocatalysts (beyond just Cu-based systems) and their synthesis procedures by extracting information like materials, products, Faradaic efficiencies (FE), and experimental conditions from thousands of papers [<xref ref-type="bibr" rid="ref-54">54</xref>]. NLP tools have been used to analyze trends in non-Cu catalysts from literature, identifying emerging materials like perovskites and bismuth oxyhalides [<xref ref-type="bibr" rid="ref-70">70</xref>]. A notable study constructed a knowledge graph specifically for Cu-based CO<sub>2</sub>RR catalysts using a SciBERT-based framework [<xref ref-type="bibr" rid="ref-40">40</xref>]. This KG not only visualized development trends but was also used, by combining word and graph embeddings, to predict the FE for specific products, showcasing the integration of structured and unstructured knowledge. LLMs are also expected to provide design guidance for CO<sub>2</sub>RR catalysts [<xref ref-type="bibr" rid="ref-42">42</xref>].</p>
<p>Beyond these specific reactions, LLMs are being applied to broader tasks in catalyst research relevant to electrocatalysis. This includes text mining and prediction for MOF synthesis using ChatGPT, LLM-driven synthesis planning for QDs which can have electrocatalytic applications [<xref ref-type="bibr" rid="ref-68">68</xref>], IE for single-atom heterogeneous catalysts [<xref ref-type="bibr" rid="ref-58">58</xref>], and assisting in the exploration of electrochemical C-H oxidation reactions through literature mining and code generation for ML model training [<xref ref-type="bibr" rid="ref-51">51</xref>]. These case studies illustrate the diverse ways LLMs are beginning to impact the field, from large-scale data aggregation to targeted material prediction and synthesis planning across various important electrocatalytic systems. <xref ref-type="table" rid="table-2">Table 2</xref> categorizes the application of LLMs across major electrocatalytic reactions such as HER, OER, ORR, and CO<sub>2</sub>RR, organizing them by task type (e.g., property prediction, synthesis planning, knowledge extraction) and highlighting specific use cases and insights. For example, in the case of CO<sub>2</sub>RR, LLMs have been utilized for corpus creation, trend analysis, and FE prediction using hybrid embedding techniques. For HER and OER, LLM-guided screening has facilitated the identification of promising alloy compositions. This table crystallizes how LLMs function as modular tools across the catalyst discovery landscape, tailored to the nuances of different electrochemical reactions.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Overview of LLM applications in specific electrocatalytic reactions</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center">Reaction</th>
<th align="center">LLM task</th>
<th align="center">Specific examples/Refs.</th>
<th align="center">Key insights/Findings for reaction</th>
</tr>
</thead>
<tbody>
<tr>
<td>HER</td>
<td>Property prediction/Design</td>
<td>ML framework for alloy discovery (AgPd) [<xref ref-type="bibr" rid="ref-72">72</xref>]; ML for low-dim catalysts [<xref ref-type="bibr" rid="ref-15">15</xref>]; Text mining/Word2Vec prediction [<xref ref-type="bibr" rid="ref-14">14</xref>]; LLM design guidance [<xref ref-type="bibr" rid="ref-42">42</xref>]; HEA discovery w/HT expts [<xref ref-type="bibr" rid="ref-73">73</xref>]</td>
<td>Acceleration of alloy screening; Prediction based on text similarity; Guidance for catalyst design</td>
</tr>
<tr>
<td>OER</td>
<td>Property prediction/Design</td>
<td>LLM predictive analytics for (oxy)hydroxides [<xref ref-type="bibr" rid="ref-36">36</xref>]; ML/AI for perovskites [<xref ref-type="bibr" rid="ref-36">36</xref>]; Text mining/Word2Vec prediction [<xref ref-type="bibr" rid="ref-14">14</xref>]; LLM design guidance [<xref ref-type="bibr" rid="ref-42">42</xref>]; Interpretable ML for activity [<xref ref-type="bibr" rid="ref-39">39</xref>]</td>
<td>Prediction for specific material classes; Text-based candidate screening; Design guidance</td>
</tr>
<tr>
<td rowspan="3">ORR</td>
<td>Structure generation/Design</td>
<td>CatGPT fine-tuned for 2e-ORR catalyst discovery [<xref ref-type="bibr" rid="ref-65">65</xref>]</td>
<td>Discovery of novel 2e-ORR candidates</td>
</tr>
<tr>
<td>Element selection/Design</td>
<td>LLM generating element library for Pt-HEA discovery &#x002B; HT expts [<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td>LLM-guided high-throughput discovery workflow</td>
</tr>
<tr>
<td>Property prediction/Design</td>
<td>Text mining/Word2Vec prediction [<xref ref-type="bibr" rid="ref-14">14</xref>]; LLM design guidance [<xref ref-type="bibr" rid="ref-42">42</xref>]; ML optimization for HEAs [<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>Text-based candidate screening; Design guidance</td>
</tr>
<tr>
<td rowspan="3">CO<sub>2</sub>RR</td>
<td>IE/Corpus creation</td>
<td>LLM-enhanced corpus creation (general CO<sub>2</sub>RR) [<xref ref-type="bibr" rid="ref-54">54</xref>]; NLP review of non-Cu catalysts [<xref ref-type="bibr" rid="ref-70">70</xref>]</td>
<td>Large-scale data aggregation; Trend analysis beyond Cu</td>
</tr>
<tr>
<td>KG Construction/FE Prediction</td>
<td>KG for Cu-based catalysts (SciBERT); FE prediction via word &#x002B; graph embeddings [<xref ref-type="bibr" rid="ref-40">40</xref>]</td>
<td>Structured knowledge representation; Prediction integrating semantics &#x0026; structure</td>
</tr>
<tr>
<td>Design guidance</td>
<td>LLM design guidance [<xref ref-type="bibr" rid="ref-42">42</xref>]</td>
<td>Guidance for catalyst design</td>
</tr>
<tr>
<td rowspan="3">General/Other</td>
<td>Synthesis planning</td>
<td>LLM for QD synthesis [<xref ref-type="bibr" rid="ref-68">68</xref>]; ChatGPT for MOF synthesis [<xref ref-type="bibr" rid="ref-37">37</xref>]</td>
<td>Optimization of nanomaterial synthesis protocols</td>
</tr>
<tr>
<td>IE</td>
<td>LLM for MOF synthesis [<xref ref-type="bibr" rid="ref-37">37</xref>]; LLM for single-atom catalysts [<xref ref-type="bibr" rid="ref-58">58</xref>]; LLM for C-H oxidation lit. mining [<xref ref-type="bibr" rid="ref-51">51</xref>]</td>
<td>Data extraction for specific catalyst types/reactions</td>
</tr>
<tr>
<td>Property prediction</td>
<td>CatBERTa for adsorption energy [<xref ref-type="bibr" rid="ref-61">61</xref>]</td>
<td>Text-based prediction of fundamental catalytic properties</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Comparative Analysis of LLM Approaches</title>
<p>As LLMs become more integrated into electrocatalysis research, understanding the nuances between different models, training strategies, and application frameworks is crucial for selecting and developing effective tools.</p>
<p>Regarding model architectures, the field utilizes both BERT-based encoders (like SciBERT, RoBERTa used in CatBERTa, MatBERT) primarily for understanding and classification tasks, and GPT-style decoders or encoder-decoders for generative tasks like text generation, hypothesis generation, or synthesis planning [<xref ref-type="bibr" rid="ref-40">40</xref>,<xref ref-type="bibr" rid="ref-44">44</xref>,<xref ref-type="bibr" rid="ref-46">46</xref>,<xref ref-type="bibr" rid="ref-61">61</xref>,<xref ref-type="bibr" rid="ref-68">68</xref>]. BERT-based models excel at extracting semantic meaning and have shown strong performance in NER and property prediction from text [<xref ref-type="bibr" rid="ref-40">40</xref>,<xref ref-type="bibr" rid="ref-60">60</xref>,<xref ref-type="bibr" rid="ref-61">61</xref>]. GPT-based models, with their strong generative capabilities, are increasingly used for proposing synthesis routes, generating novel structures, or acting as conversational agents [<xref ref-type="bibr" rid="ref-37">37</xref>,<xref ref-type="bibr" rid="ref-49">49</xref>,<xref ref-type="bibr" rid="ref-65">65</xref>,<xref ref-type="bibr" rid="ref-68">68</xref>]. The choice of architecture often depends on the primary goal, whether it is analyzing existing text or generating new information.</p>
<p>Training strategies significantly impact performance. While general-purpose LLMs like GPT-4 can perform reasonably well on some tasks using sophisticated prompt engineering [<xref ref-type="bibr" rid="ref-37">37</xref>,<xref ref-type="bibr" rid="ref-46">46</xref>], studies consistently show that domain-specific pre-training or fine-tuning yields substantial improvements for specialized scientific tasks [<xref ref-type="bibr" rid="ref-10">10</xref>,<xref ref-type="bibr" rid="ref-44">44</xref>,<xref ref-type="bibr" rid="ref-54">54</xref>,<xref ref-type="bibr" rid="ref-60">60</xref>]. Models like MatBERT (pre-trained on materials science literature) [<xref ref-type="bibr" rid="ref-60">60</xref>], CatBERTa (fine-tuned RoBERTa for catalyst energy prediction) [<xref ref-type="bibr" rid="ref-61">61</xref>], and CataLM (Vicuna fine-tuned on electrocatalysis literature and expert data) [<xref ref-type="bibr" rid="ref-44">44</xref>] demonstrate enhanced understanding of domain terminology and concepts, leading to higher accuracy in tasks like NER, property prediction, and recommendation [<xref ref-type="bibr" rid="ref-44">44</xref>,<xref ref-type="bibr" rid="ref-60">60</xref>,<xref ref-type="bibr" rid="ref-61">61</xref>]. Parameter-Efficient Fine-Tuning (PEFT) techniques are also being adopted to adapt large models to specific tasks more efficiently, as seen in the QD synthesis planning framework [<xref ref-type="bibr" rid="ref-68">68</xref>]. The trade-off lies between the versatility of large general models and the specialized accuracy of fine-tuned models, with the latter often preferred for demanding scientific applications where domain knowledge is critical.</p>
<p>The type of input data processed by the LLM also differentiates approaches. Many applications focus purely on textual input, leveraging LLMs&#x2019; core strength in NLP [<xref ref-type="bibr" rid="ref-60">60</xref>,<xref ref-type="bibr" rid="ref-61">61</xref>]. CatBERTa predicts energy from textual descriptions [<xref ref-type="bibr" rid="ref-61">61</xref>], and MOOSE-Chem generates hypotheses from background questions and literature text [<xref ref-type="bibr" rid="ref-49">49</xref>]. Other methods integrate structured data or embeddings. The Cu-CO<sub>2</sub>RR KG study combined word embeddings from text with graph embeddings from the structured KG for FE prediction [<xref ref-type="bibr" rid="ref-40">40</xref>]. Word2Vec embeddings derived from abstracts were used as numerical inputs for Pareto optimization [<xref ref-type="bibr" rid="ref-14">14</xref>]. The future direction clearly points towards multimodal inputs, integrating text with figures, tables, and potentially experimental data streams, although this remains a significant developmental challenge [<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-42">42</xref>].</p>
<p>Finally, LLMs are being applied both as standalone tools and as components within larger integrated workflows. Standalone applications might involve using an LLM directly for Q&#x0026;A, summarization, or prediction based on a prompt [<xref ref-type="bibr" rid="ref-37">37</xref>,<xref ref-type="bibr" rid="ref-61">61</xref>]. Integrated approaches are becoming more common and powerful. Examples include using LLM-based IE to populate databases that feed into downstream ML models [<xref ref-type="bibr" rid="ref-56">56</xref>], employing LLMs to generate code for ML model training or optimization [<xref ref-type="bibr" rid="ref-51">51</xref>], combining LLMs with knowledge graphs [<xref ref-type="bibr" rid="ref-40">40</xref>], and embedding LLMs within autonomous laboratory frameworks to guide experiments [<xref ref-type="bibr" rid="ref-34">34</xref>]. These integrated systems leverage the LLM&#x2019;s language and reasoning capabilities while connecting them to other computational tools or physical experiments, amplifying their impact. The trend suggests a move towards more sophisticated, integrated systems where LLMs act as orchestrators or intelligent interfaces within broader scientific discovery pipelines. <xref ref-type="table" rid="table-3">Table 3</xref> presents a systematic comparison of various LLM frameworks applied to IE tasks within electrocatalysis. It categorizes each model by the specific task (e.g., band gap extraction, synthesis parameter identification), the source and size of input data, and the resulting output format. Additionally, the table evaluates each model&#x2019;s effectiveness and limitations based on published benchmarks and practical deployment. For instance, GPT-4 demonstrates strong performance in general extraction tasks but suffers from hallucination and high computational costs. In contrast, domain-specific models like CataLM and SciBERT-BiLSTM-CRF offer improved accuracy in materials entity recognition within defined contexts (e.g., CO<sub>2</sub>RR), albeit at the cost of requiring domain-specific training data. This comparative analysis highlights the trade-offs between generalizability and domain accuracy in current LLM-powered IE pipelines.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Overview of LLMs applied to electrocatalyst information extraction</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center">Model/Framework</th>
<th align="center">Task</th>
<th align="center">Input data (Source, Size)</th>
<th align="center">Output format</th>
<th align="center">Comparative evaluation: effectiveness, efficiency, and limitations</th>
<th align="center">Refs.</th>
</tr>
</thead>
<tbody>
<tr>
<td>GPT-4</td>
<td>Band gap extraction</td>
<td>415 random materials science articles (manual eval)</td>
<td>Band gap values, materials</td>
<td>Effectiveness: 88% accuracy. Efficiency: High GPU demand (e.g., A100). Limitations: Prone to hallucination, general-domain model.</td>
<td>[<xref ref-type="bibr" rid="ref-46">46</xref>]</td>
</tr>
<tr>
<td>ChatGPT (GPT-3.5/4)</td>
<td>MOF synthesis extraction</td>
<td>Peer-reviewed MOF articles (&#x007E;800 MOFs)</td>
<td>26k&#x002B; synthesis parameters (structured)</td>
<td>Effectiveness: F1 90%&#x2013;99%. Efficiency: Fast inference with engineered prompts. Limitations: Context-length sensitivity, depends on prompt clarity.</td>
<td>[<xref ref-type="bibr" rid="ref-37">37</xref>]</td>
</tr>
<tr>
<td>SciBERT &#x002B; BiLSTM-CRF/BERT-BiLSTM-CRF</td>
<td>CO<sub>2</sub>RR entity recognition</td>
<td>835 publications (benchmark); 372 full texts (extended)</td>
<td>Entities (Material, Method, Product, FE, etc.)</td>
<td>Effectiveness: Micro-F1 &#x007E;82%. Efficiency: Moderate training cost. Limitations: Requires annotated corpora, limited to CO<sub>2</sub>RR context.</td>
<td>[<xref ref-type="bibr" rid="ref-56">56</xref>]</td>
</tr>
<tr>
<td>LLMs (General) &#x002B; NLP Techniques</td>
<td>CO<sub>2</sub>RR corpus creation &#x0026; synthesis extraction</td>
<td>5941 documents (metadata); 2776 full texts</td>
<td>Benchmark corpus (7k records); Extended corpora (77k &#x002B; 30k records); Action sequences</td>
<td>Effectiveness: Comprehensive corpus created. Efficiency: Costly pretraining &#x002B; fine-tuning. Limitations: Requires large, domain-specific data.</td>
<td>[<xref ref-type="bibr" rid="ref-54">54</xref>]</td>
</tr>
<tr>
<td>ChatExtract (Conversational LLM, e.g., GPT-4)</td>
<td>General materials data extraction</td>
<td>Research papers (test sets mentioned)</td>
<td>Extracted data points</td>
<td>Effectiveness: &#x007E;90% precision/recall. Efficiency: Lightweight conversational method. Limitations: Hallucination possible, context-sensitive.</td>
<td>[<xref ref-type="bibr" rid="ref-36">36</xref>]</td>
</tr>
<tr>
<td>NLP Tools (General)</td>
<td>CO<sub>2</sub>RR literature review (non-Cu)</td>
<td>7292 published articles</td>
<td>Trends, emerging materials (perovskites, Bi-oxyhalides), common elements, electrolytes</td>
<td>Effectiveness: Qualitative trend mapping. Efficiency: Text-only mining; low compute need. Limitations: Lacks structured extraction.</td>
<td>[<xref ref-type="bibr" rid="ref-70">70</xref>]</td>
</tr>
<tr>
<td>CataLM (Fine-tuned Vicuna-13B)</td>
<td>Entity extraction (Electrocatalysis)</td>
<td>Domain literature, expert annotations</td>
<td>Entities (Material, Method, Product, FE, etc.)</td>
<td>Effectiveness: High domain-specific accuracy. Efficiency: PEFT applied for tuning. Limitations: Performance drops outside domain.</td>
<td>[<xref ref-type="bibr" rid="ref-44">44</xref>]</td>
</tr>
<tr>
<td>Transformer (ACE)</td>
<td>Information extraction (SACs)</td>
<td>Literature</td>
<td>Structured data</td>
<td>Effectiveness: Focused and relevant SAC data. Efficiency: Narrow application. Limitations: Unclear generalizability, lacks metric benchmarks.</td>
<td>[<xref ref-type="bibr" rid="ref-58">58</xref>]</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Performance Evaluation and Benchmarking</title>
<p>Assessing the performance of LLMs in the specialized domain of electrocatalysis requires appropriate metrics and rigorous evaluation, although standardized benchmarks are still largely lacking [<xref ref-type="bibr" rid="ref-16">16</xref>,<xref ref-type="bibr" rid="ref-22">22</xref>]. The metrics used vary depending on the specific task being performed. For IE tasks like NER and relation extraction, standard NLP metrics such as Precision, Recall, and F1-score are commonly employed [<xref ref-type="bibr" rid="ref-36">36</xref>,<xref ref-type="bibr" rid="ref-37">37</xref>]. These metrics quantify the accuracy of identifying entities and relationships compared to a ground truth (often manually annotated data). For instance, the ChatGPT-based MOF synthesis extraction achieved F1 scores of 90%&#x2013;99% using careful prompt engineering, while the BERT-BiLSTM-CRF model for CO<sub>2</sub>RR entity recognition reached micro-average F1 scores of around 82% [<xref ref-type="bibr" rid="ref-56">56</xref>]. The ChatExtract method reported precision and recall close to 90% for materials data extraction.</p>
<p>In property prediction, common regression metrics are used when predicting continuous values like adsorption energy or FE. MAE is frequently reported, indicating the average absolute difference between predicted and true values [<xref ref-type="bibr" rid="ref-60">60</xref>,<xref ref-type="bibr" rid="ref-61">61</xref>,<xref ref-type="bibr" rid="ref-72">72</xref>]. CatBERTa reported MAEs between 0.35 and 0.75 eV for adsorption energy prediction, depending on the dataset and input features [<xref ref-type="bibr" rid="ref-61">61</xref>]. The R<sup>2</sup> is also used to measure the proportion of variance explained by the model [<xref ref-type="bibr" rid="ref-50">50</xref>]. For classification tasks (e.g., predicting synthesizability or classifying materials based on properties), accuracy and metrics like the Matthews correlation coefficient are used [<xref ref-type="bibr" rid="ref-60">60</xref>].</p>
<p>Evaluating generative and planning tasks requires different approaches. For hypothesis generation, metrics might include similarity scores to known valid hypotheses or expert evaluation of novelty and feasibility [<xref ref-type="bibr" rid="ref-49">49</xref>]. For structure generation, validity checks (e.g., detecting overlapping atoms, and ensuring chemical sensibility) are crucial, alongside assessing the novelty and predicted properties of the generated structures [<xref ref-type="bibr" rid="ref-64">64</xref>]. For synthesis planning (e.g., LLM-driven QD synthesis), success can be measured by the rate at which generated protocols lead to successful experiments, improve target properties, or advance the Pareto front in multi-objective optimization [<xref ref-type="bibr" rid="ref-68">68</xref>]. The QD synthesis work reported that 3 out of 6 LLM-generated protocols updated the Pareto front.</p>
<p>Beyond accuracy metrics, efficiency is also a key evaluation criterion. This can involve measuring the reduction in time required for tasks like literature review (e.g., seconds per article for automated review generation [<xref ref-type="bibr" rid="ref-29">29</xref>]) or the potential reduction in computational or experimental cost achieved by using LLM predictions to guide efforts [<xref ref-type="bibr" rid="ref-56">56</xref>].</p>
<p>Despite these reported successes, a significant benchmarking gap exists. There is a lack of standardized, publicly available benchmark datasets and evaluation protocols specifically designed for testing LLMs on various tasks within electrocatalysis or even the broader field of materials science [<xref ref-type="bibr" rid="ref-16">16</xref>,<xref ref-type="bibr" rid="ref-22">22</xref>]. This makes it challenging to directly compare the performance of different models and methodologies developed in separate studies. Establishing such benchmarks would be crucial for driving progress and ensuring rigorous assessment of new LLM approaches. Current evaluations often rely on specific internal datasets or comparisons against limited baselines, highlighting the need for community-wide efforts in developing standardized evaluation frameworks. <xref ref-type="table" rid="table-4">Table 4</xref> summarizes the predictive performance of different LLM-based frameworks in estimating key electrocatalytic properties such as adsorption energy and FE. The models are organized by input type (e.g., textual descriptions, word or graph embeddings), predicted properties, target catalyst systems, and quantitative performance metrics (e.g., MAE, accuracy). For example, CatBERTa&#x2014;trained on structured textual inputs like composition and surface features&#x2014;achieved an MAE of 0.35&#x2013;0.75 eV for adsorption energy predictions, making it a compelling alternative to more computationally demanding Graph Neural Networks. Other models, such as those using hybrid word-graph embeddings, demonstrated strong performance in FE prediction tasks for Cu-based CO<sub>2</sub>RR catalysts. The table emphasizes how LLMs can complement or even rival traditional methods by enabling property prediction from semantically rich, yet structurally limited, inputs. These insights are particularly relevant for scenarios where experimental or atomic-scale data is sparse or unavailable. <xref ref-type="table" rid="table-5">Table 5</xref> highlights key frameworks and models that harness the generative and planning capabilities of LLMs within the realm of electrocatalyst research. It details each framework&#x2019;s core task (e.g., hypothesis generation, structure proposal, synthesis optimization), material focus, the LLM&#x2019;s specific role (e.g., retriever, generator, predictor), and outcomes or validation methods. For instance, MOOSE-Chem uses a multi-agent architecture for autonomous hypothesis formation in chemistry, while CatGPT excels at proposing valid catalyst structures for 2e-ORR through fine-tuned autoregressive generation. Other systems like the LLM-Feynman framework demonstrate symbolic regression abilities, rediscovering known physical laws, while QD-focused models successfully improved experimental synthesis outcomes using Pareto optimization. This table illustrates the growing sophistication of LLM-enabled systems capable of creativity, planning, and real-world lab guidance, bridging the gap between literature mining and experimental realization.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Performance comparison of LLMs for electrocatalyst property prediction</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col/>
</colgroup>
<thead>
<tr>
<th align="center">Model/Framework</th>
<th align="center">Input type</th>
<th align="center">Predicted property</th>
<th align="center">Material system/Reaction</th>
<th align="center">Performance metric/Value</th>
<th align="center">Comparative evaluation: effectiveness, efficiency, and limitations</th>
<th>Refs.</th>
</tr>
</thead>
<tbody>
<tr>
<td>CatBERTa (RoBERTa-based)</td>
<td>Textual description (Composition, Structure features)</td>
<td>Adsorption energy</td>
<td>General catalysts (OC20 dataset)</td>
<td>MAE: 0.75 eV (100k data); MAE: 0.35 eV (high-accuracy subset)</td>
<td>Effectiveness: MAE 0.35&#x2013;0.75 eV. Efficiency: Requires moderate fine-tuning. Limitations: Less precise than GNNs for local atomic interactions.</td>
<td>[<xref ref-type="bibr" rid="ref-61">61</xref>]</td>
</tr>
<tr>
<td>Word2Vec &#x002B; Pareto Optimization</td>
<td>Word embeddings (from Abstracts)</td>
<td>Similarity to &#x2018;Conductivity&#x2019;/&#x2018;Dielectric&#x2019; (proxy for activity)</td>
<td>Candidate Compositions (e.g., AgPdPt, AgPdRu)/HER, OER, ORR</td>
<td>Predicted high-performing compositions matched experimental trends</td>
<td>Effectiveness: High match with experiment. Efficiency: Lightweight; corpus-based. Limitations: Relies on textual semantic similarity, lacks explicit structural context.</td>
<td>[<xref ref-type="bibr" rid="ref-14">14</xref>]</td>
</tr>
<tr>
<td>LLM (SciBERT) &#x002B; KG Embeddings &#x002B; DL</td>
<td>Word embeddings &#x002B; Graph Embeddings</td>
<td>FE</td>
<td>Cu-based Catalysts/CO2RR</td>
<td>Demonstrated FE prediction capability</td>
<td>Effectiveness: Good FE prediction. Efficiency: Requires graph construction. Limitations: Dependent on quality of both text and KG data.</td>
<td>[<xref ref-type="bibr" rid="ref-40">40</xref>]</td>
</tr>
<tr>
<td>Transformer (MatBERT)</td>
<td>Textual Description (Robocrystallographer)</td>
<td>Energy above hull, Band gap, SLME, Spin-orbit spillage, Magnetic moment (Classification)</td>
<td>General inorganic materials</td>
<td>Accuracy &#x003E; GNNs (ALIGNN) for 4/5 properties; High accuracy even with small datasets</td>
<td>Effectiveness: High classification accuracy. Efficiency: Pretrained model with lightweight fine-tuning. Limitations: Performance depends on Robocrystallographer input quality.</td>
<td>[<xref ref-type="bibr" rid="ref-60">60</xref>]</td>
</tr>
</tbody>
</table>
</table-wrap><table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>LLM-driven hypothesis generation and synthesis planning examples</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col/>
</colgroup>
<thead>
<tr>
<th align="center">Framework/Model</th>
<th align="center">Task</th>
<th align="center">Material focus</th>
<th align="center">LLM role</th>
<th align="center">Key outcome/Validation</th>
<th align="center">Comparative evaluation: effectiveness, efficiency, and limitations</th>
<th>Refs.</th>
</tr>
</thead>
<tbody>
<tr>
<td>MOOSE-Chem</td>
<td>Hypothesis generation</td>
<td>General chemistry</td>
<td>Retrieval, synthesis, ranking (Multi-agent)</td>
<td>Rediscovered novel hypotheses from 2024 papers with high similarity</td>
<td>Effectiveness: High novelty similarity. Efficiency: Multi-agent but scalable. Limitations: Performance depends on retriever quality.</td>
<td>[<xref ref-type="bibr" rid="ref-49">49</xref>]</td>
</tr>
<tr>
<td>CatGPT (GPT-based)</td>
<td>Structure generation</td>
<td>Inorganic catalysts (OC20)/2e-ORR Catalysts</td>
<td>Generator</td>
<td>Generated valid catalyst structures; Discovered 5 novel 2e-ORR candidates after fine-tuning</td>
<td>Effectiveness: Discovered 5 new candidates. Efficiency: Heavy compute during generation. Limitations: Validity filtering needed post-generation.</td>
<td>[<xref ref-type="bibr" rid="ref-65">65</xref>]</td>
</tr>
<tr>
<td>LLM-Feynman</td>
<td>Formula/Theory discovery</td>
<td>Physics, materials science (Synthesizability, ionic conductivity, bandgap)</td>
<td>Symbolic regression, interpreter</td>
<td>Rediscovered &#x003E;90% physics formulae; Derived interpretable formulae for materials properties (Acc &#x003E; 90%, R<sup>2</sup> &#x003E; 0.8)</td>
<td>Effectiveness: &#x003E;90% formula rediscovery. Efficiency: Symbolic regression efficient. Limitations: May not scale to highly complex systems.</td>
<td>[<xref ref-type="bibr" rid="ref-50">50</xref>]</td>
</tr>
<tr>
<td>LLM-driven framework (Fine-tuned Open LLM &#x002B; PEFT)</td>
<td>Synthesis planning/Optimization</td>
<td>Quantum Dots (QDs)</td>
<td>Generator, Predictor</td>
<td>Generated 6 protocols; 3 updated Pareto front; All improved &#x2265; 1 property; Validated experimentally</td>
<td>Effectiveness: All generated protocols improved some targets. Efficiency: PEFT used. Limitations: Dependent on accurate property predictor.</td>
<td>[<xref ref-type="bibr" rid="ref-59">59</xref>]</td>
</tr>
<tr>
<td>LLM (unspecified)</td>
<td>Element selection</td>
<td>Pt-based high-entropy alloys (HEAs)/ORR</td>
<td>Generator (Element library)</td>
<td>Provided element library guiding high-throughput experimental screening (SECCM)</td>
<td>Effectiveness: Enabled targeted HT screening. Efficiency: Fast suggestion generation. Limitations: No accuracy benchmark reported.</td>
<td>[<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
</tr>
<tr>
<td>LLM (unspecified) &#x002B; Code Generation</td>
<td>Synthesis optimization</td>
<td>Electrochemical C-H Oxidation</td>
<td>Code generator (for ML optimization)</td>
<td>Iteratively improved reaction yields based on prompts; Optimized conditions for 8 drug-like substrates</td>
<td>Effectiveness: Improved multi-objective outcomes. Efficiency: Code auto-generation fast. Limitations: Prompt dependency and generalizability.</td>
<td>[<xref ref-type="bibr" rid="ref-51">51</xref>]</td>
</tr>
<tr>
<td>ChemCrow (GPT-4 &#x002B; Tools)</td>
<td>Synthesis planning</td>
<td>Organic molecules (incl. Organocatalysts)</td>
<td>Planner, tool-user (Agent)</td>
<td>Autonomously planned and executed syntheses</td>
<td>Effectiveness: Demonstrated successful planning. Efficiency: Combined with external tools. Limitations: Hallucination and chaining logic errors possible.</td>
<td>[<xref ref-type="bibr" rid="ref-10">10</xref>]</td>
</tr>
<tr>
<td>CataLM (Fine-tuned Vicuna-13B)</td>
<td>Control method recommendation</td>
<td>Electrocatalytic Materials</td>
<td>Recommender</td>
<td>Validated on recommendation task using domain knowledge</td>
<td>Effectiveness: Relevant method suggestions. Efficiency: Fast inference. Limitations: Domain-specific tuning required.</td>
<td>[<xref ref-type="bibr" rid="ref-44">44</xref>]</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Challenges, Limitations, and Future Outlook</title>
<p>While the application of LLMs in electrocatalyst research holds significant promise, the field faces substantial challenges and limitations that must be addressed to realize its full potential. Concurrently, exciting future directions are emerging, pointing towards more powerful and integrated AI-driven discovery workflows.</p>
<sec id="s5_1">
<label>5.1</label>
<title>Current Challenges and Limitations</title>
<p>Several key hurdles currently impede the widespread and reliable application of LLMs in electrocatalyst design. A fundamental issue lies with data. Despite the vastness of scientific literature, accessing high-quality, comprehensive, and standardized data remains difficult [<xref ref-type="bibr" rid="ref-10">10</xref>,<xref ref-type="bibr" rid="ref-16">16</xref>,<xref ref-type="bibr" rid="ref-22">22</xref>]. LLM training requires large datasets, but electrocatalysis data can be sparse for specific materials or reactions, heterogeneous due to varying experimental conditions and reporting standards, and potentially biased due to the tendency to underreport negative results [<xref ref-type="bibr" rid="ref-13">13</xref>,<xref ref-type="bibr" rid="ref-22">22</xref>,<xref ref-type="bibr" rid="ref-27">27</xref>]. Extracting complete information is further complicated by knowledge fragmented across multiple publications and supplementary materials, which current LLMs struggle to synthesize effectively [<xref ref-type="bibr" rid="ref-27">27</xref>].</p>
<p>The interpretability and explainability of LLMs pose another significant challenge. Many state-of-the-art models function as &#x201C;black boxes&#x201D;, making it difficult to understand why they make a particular prediction or suggestion [<xref ref-type="bibr" rid="ref-22">22</xref>]. This lack of transparency hinders scientific understanding, trust in the model&#x2019;s output, and the ability to extract generalizable design principles [<xref ref-type="bibr" rid="ref-10">10</xref>]. While techniques like attention analysis or post-hoc explanations offer some insight [<xref ref-type="bibr" rid="ref-61">61</xref>], achieving true mechanistic understanding from LLM outputs remains an open research area.</p>
<p>Hallucinations and factual inaccuracy represent a critical barrier to the reliable use of LLMs in science [<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-46">46</xref>,<xref ref-type="bibr" rid="ref-50">50</xref>,<xref ref-type="bibr" rid="ref-54">54</xref>]. LLMs can generate text that sounds scientifically plausible but is factually incorrect, unsubstantiated, or physically inconsistent. This risk is potentially amplified in specialized domains like electrocatalysis where the training data might be less comprehensive compared to general text [<xref ref-type="bibr" rid="ref-29">29</xref>]. Rigorous validation and mitigation strategies, such as RAG, careful prompting, and expert verification, are essential but add complexity to the workflow [<xref ref-type="bibr" rid="ref-29">29</xref>,<xref ref-type="bibr" rid="ref-36">36</xref>].</p>
<p>Current LLMs also exhibit limitations in domain knowledge grounding and reasoning [<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-50">50</xref>]. While they possess vast general knowledge, their understanding of fundamental materials science and electrochemistry principles can be shallow. They struggle with complex numerical reasoning, unit conversions, understanding intricate chemical notations (e.g., varied formulae, crystallographic information like CIF files), and applying core concepts like crystal symmetry or reaction stoichiometry correctly [<xref ref-type="bibr" rid="ref-27">27</xref>]. This limits their ability to perform deep, physically grounded reasoning. The predominantly text-based nature of most current LLMs restricts their ability to process multimodal data. Scientific publications in materials science heavily rely on figures (micrographs, diffraction patterns, performance plots), complex tables, and chemical structure diagrams to convey crucial information. LLMs that cannot interpret these visual or tabular formats miss a significant portion of the available knowledge.</p>
<p>Finally, even with advanced AI predictions, the experimental validation bottleneck persists [<xref ref-type="bibr" rid="ref-10">10</xref>,<xref ref-type="bibr" rid="ref-27">27</xref>]. Any catalyst candidate or synthesis plan proposed by an LLM must ultimately be tested in the laboratory, a process that remains resource-intensive and time-consuming. LLMs primarily accelerate the <italic>in silico</italic> stages of discovery and design.</p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Future Research Directions</title>
<p>Addressing the current limitations and harnessing the full potential of LLMs in electrocatalysis necessitates focused research efforts along several key directions.</p>
<p>A critical need is the development of specialized Materials Science LLMs (MatSci-LLMs). This involves moving beyond general-purpose models towards architectures pre-trained or extensively fine-tuned on vast corpora of materials science literature, textbooks, and databases. Such models would possess deeper domain knowledge, better understand specialized terminology and notations (including chemical formulas and crystallographic data), and exhibit improved reasoning capabilities grounded in materials science principles. Examples like MatSciBERT, CataLM, and CatBERTa represent early steps in this direction [<xref ref-type="bibr" rid="ref-44">44</xref>,<xref ref-type="bibr" rid="ref-61">61</xref>].</p>
<p>The development of multimodal LLMs is arguably one of the most crucial future directions. Models capable of seamlessly integrating and reasoning over information from text, tables, figures (plots, images, schematics), chemical structures, and potentially experimental data streams (e.g., spectra) will unlock a much larger fraction of scientific knowledge and enable more holistic analysis [<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-42">42</xref>].</p>
<p>Improving reasoning capabilities and grounding LLM outputs in fundamental physical and chemical laws is essential for generating scientifically valid and reliable predictions or hypotheses. Integrating LLMs with physics-informed AI principles, where physical constraints or knowledge are incorporated into the model architecture or training process, holds significant promise [<xref ref-type="bibr" rid="ref-10">10</xref>,<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-22">22</xref>].</p>
<p>Enhancing interpretability and explainability (XAI) remains paramount for building trust and extracting scientific insights [<xref ref-type="bibr" rid="ref-22">22</xref>,<xref ref-type="bibr" rid="ref-60">60</xref>,<xref ref-type="bibr" rid="ref-74">74</xref>]. Future work should focus on developing robust XAI techniques tailored for LLMs in scientific domains, allowing researchers to understand the basis of model predictions and identify potential failure modes.</p>
<p>Continued research into mitigating hallucinations is vital for scientific applications where factual accuracy is non-negotiable. This includes refining techniques like RAG, developing better-prompting strategies, improving fine-tuning methods focused on factuality, incorporating self-evaluation mechanisms within LLMs, and utilizing multi-agent frameworks for cross-validation [<xref ref-type="bibr" rid="ref-49">49</xref>].</p>
<p>The community needs to establish standardized datasets and benchmarks for evaluating LLMs on various electrocatalysis-related tasks [<xref ref-type="bibr" rid="ref-16">16</xref>,<xref ref-type="bibr" rid="ref-22">22</xref>]. This will enable objective comparison of different models and methodologies, track progress more effectively, and identify areas needing further improvement.</p>
<p>A highly promising future direction is the integration of LLMs with autonomous experimentation platforms or Self-Driving Laboratories (SDLs) [<xref ref-type="bibr" rid="ref-75">75</xref>]. In such closed-loop systems, LLMs could analyze previous results, consult literature knowledge, propose the next set of experiments, generate the necessary code to control robotic hardware, interpret incoming data, and iteratively refine the search for optimal catalysts or synthesis conditions. This synergy between AI-driven decision-making and automated execution has the potential to dramatically accelerate the pace of discovery.</p>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusion</title>
<p>The integration of LLMs into the field of electrocatalysis represents a nascent but rapidly advancing frontier with the potential to reshape materials discovery and design. This review has highlighted the diverse methodologies being employed, spanning automated information extraction from the vast scientific literature, novel approaches to property prediction based on textual data and embeddings, the generation of hypotheses for new materials and synthesis routes, and the synthesis of knowledge scattered across countless publications. Case studies across critical reactions like HER, OER, ORR, and CO2RR demonstrate tangible progress, with LLMs contributing to the creation of valuable datasets, the prediction of catalytic performance, and even the suggestion of novel catalyst candidates and optimized synthesis protocols. Models specifically adapted to the materials science domain, such as CatBERTa and CataLM, alongside innovative frameworks like MOOSE-Chem and LLM-driven synthesis planners, showcase the growing sophistication of these tools. However, significant challenges remain. Issues surrounding data availability and quality, the inherent &#x201C;black-box&#x201D; nature and potential for hallucinations in LLMs, limitations in deep scientific reasoning and multimodal data processing, and the persistent need for experimental validation must be rigorously addressed. Overcoming these hurdles will require concerted efforts in developing domain-specific and multimodal models, enhancing interpretability and factual grounding, establishing standardized benchmarks, and fostering collaborative research practices. The fusion of LLMs with physics-informed AI, their incorporation into autonomous experimental workflows within self-driving laboratories, and their role as sophisticated collaborators alongside human researchers promise to significantly accelerate the pace of innovation. By effectively harnessing the ability of LLMs to process, synthesize, and generate knowledge from the ever-expanding body of scientific literature, we can anticipate a future where the rational design and discovery of advanced micro/nano electrocatalyst materials&#x2014;crucial components for a sustainable energy landscape&#x2014;is achieved with unprecedented speed and efficiency.</p>
</sec>
</body>
<back>
<ack>
<p>Not applicable.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>The authors received no specific funding for this study.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Conceptualization, Hassan Karimi-Maleh and Li Fu; methodology, Ying Shen and Shichao Zhao; software, Ying Shen and Fei Chen; validation, Ying Shen, Shichao Zhao and Yanfei Lv; formal analysis, Ying Shen and Yanfei Lv; investigation, Ying Shen and Fei Chen; resources, Fei Chen and Li Fu; data curation, Ying Shen and Fei Chen; writing&#x2014;original draft preparation, Ying Shen and Yanfei Lv; writing&#x2014;review and editing, Hassan Karimi-Maleh; visualization, Ying Shen and Fei Chen; supervision, Li Fu and Hassan Karimi-Maleh; project administration, Li Fu and Hassan Karimi-Maleh. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The authors confirm that the data supporting the findings of this study are available within the article.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Basyooni-M Kabatas</surname> <given-names>MA</given-names></string-name></person-group>. <article-title>A comprehensive review on electrocatalytic applications of 2D metallenes</article-title>. <source>Nanomaterials</source>. <year>2023</year>;<volume>13</volume>(<issue>22</issue>):<fpage>2966</fpage>. doi:<pub-id pub-id-type="doi">10.3390/nano13222966</pub-id>; <pub-id pub-id-type="pmid">37999320</pub-id></mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lucas</surname> <given-names>FWS</given-names></string-name>, <string-name><surname>Grim</surname> <given-names>RG</given-names></string-name>, <string-name><surname>Tacey</surname> <given-names>SA</given-names></string-name>, <string-name><surname>Downes</surname> <given-names>CA</given-names></string-name>, <string-name><surname>Hasse</surname> <given-names>J</given-names></string-name>, <string-name><surname>Roman</surname> <given-names>AM</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Electrochemical routes for the valorization of biomass-derived feedstocks: from chemistry to application</article-title>. <source>ACS Energy Lett</source>. <year>2021</year>;<volume>67</volume>(<issue>11</issue>):<fpage>1205</fpage>&#x2013;<lpage>70</lpage>. doi:<pub-id pub-id-type="doi">10.1021/acsenergylett.0c02692</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>T&#x00FC;ys&#x00FC;z</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Alkaline water electrolysis for green hydrogen production</article-title>. <source>Acc Chem Res</source>. <year>2024</year>;<volume>57</volume>(<issue>4</issue>):<fpage>558</fpage>&#x2013;<lpage>67</lpage>. doi:<pub-id pub-id-type="doi">10.1021/acs.accounts.3c00709</pub-id>; <pub-id pub-id-type="pmid">38335244</pub-id></mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Fan</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Xia</surname> <given-names>R</given-names></string-name>, <string-name><surname>Meyer</surname> <given-names>TJ</given-names></string-name></person-group>. <article-title>CO<sub>2</sub> reduction: from homogeneous to heterogeneous electrocatalysis</article-title>. <source>Acc Chem Res</source>. <year>2020</year>;<volume>53</volume>(<issue>1</issue>):<fpage>255</fpage>&#x2013;<lpage>64</lpage>. doi:<pub-id pub-id-type="doi">10.1021/acs.accounts.9b00496</pub-id>; <pub-id pub-id-type="pmid">31913013</pub-id></mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Braun</surname> <given-names>M</given-names></string-name>, <string-name><surname>Santana</surname> <given-names>CS</given-names></string-name>, <string-name><surname>Garcia</surname> <given-names>AC</given-names></string-name>, <string-name><surname>Andronescu</surname> <given-names>C</given-names></string-name></person-group>. <article-title>From waste to value&#x2014;glycerol electrooxidation for energy conversion and chemical production</article-title>. <source>Curr Opin Green Sustain Chem</source>. <year>2023</year>;<volume>41</volume>:<fpage>100829</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.cogsc.2023.100829</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Page</surname> <given-names>JR</given-names></string-name>, <string-name><surname>Manfredi</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Bliznakov</surname> <given-names>S</given-names></string-name>, <string-name><surname>Valla</surname> <given-names>JA</given-names></string-name></person-group>. <article-title>Recent progress in electrochemical upgrading of bio-oil model compounds and bio-oils to renewable fuels and platform chemicals</article-title>. <source>Materials</source>. <year>2023</year>;<volume>16</volume>(<issue>1</issue>):<fpage>394</fpage>. doi:<pub-id pub-id-type="doi">10.3390/ma16010394</pub-id>; <pub-id pub-id-type="pmid">36614733</pub-id></mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kumar</surname> <given-names>A</given-names></string-name>, <string-name><surname>Bhayana</surname> <given-names>S</given-names></string-name>, <string-name><surname>Singh</surname> <given-names>PK</given-names></string-name>, <string-name><surname>Tripathi</surname> <given-names>AD</given-names></string-name>, <string-name><surname>Paul</surname> <given-names>V</given-names></string-name>, <string-name><surname>Balodi</surname> <given-names>V</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Valorization of used cooking oil: challenges, current developments, life cycle assessment and future prospects</article-title>. <source>Discov Sustain</source>. <year>2025</year>;<volume>6</volume>(<issue>1</issue>):<fpage>119</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s43621-025-00905-7</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Garedew</surname> <given-names>M</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>F</given-names></string-name>, <string-name><surname>Song</surname> <given-names>B</given-names></string-name>, <string-name><surname>DeWinter</surname> <given-names>TM</given-names></string-name>, <string-name><surname>Jackson</surname> <given-names>JE</given-names></string-name>, <string-name><surname>Saffron</surname> <given-names>CM</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Greener routes to biomass waste valorization: lignin transformation through electrocatalysis for renewable chemicals and fuels production</article-title>. <source>ChemSusChem</source>. <year>2020</year>;<volume>13</volume>(<issue>17</issue>):<fpage>4214</fpage>&#x2013;<lpage>37</lpage>. doi:<pub-id pub-id-type="doi">10.1002/cssc.202000987</pub-id>; <pub-id pub-id-type="pmid">32460408</pub-id></mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ram&#x00ED;rez</surname> <given-names>&#x00C1;.</given-names></string-name>, <string-name><surname>Mu&#x00F1;oz-Morales</surname> <given-names>M</given-names></string-name>, <string-name><surname>Fern&#x00E1;ndez-Morales</surname> <given-names>FJ</given-names></string-name>, <string-name><surname>Llanos</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Valorization of polluted biomass waste for manufacturing sustainable cathode materials for the production of hydrogen peroxide</article-title>. <source>Electrochim Acta</source>. <year>2023</year>;<volume>456</volume>:<fpage>142383</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.electacta.2023.142383</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>L</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Salim</surname> <given-names>F</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>AI-empowered catalyst discovery: a survey from classical machine learning approaches to large language models</article-title>. <comment>arXiv:2502.13626. 2025</comment>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xia</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Campbell</surname> <given-names>CT</given-names></string-name>, <string-name><surname>Roldan Cuenya</surname> <given-names>B</given-names></string-name>, <string-name><surname>Mavrikakis</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Introduction: advanced materials and methods for catalysis and electrocatalysis by transition metals</article-title>. <source>Chem Rev</source>. <year>2021</year>;<volume>121</volume>(<issue>2</issue>):<fpage>563</fpage>&#x2013;<lpage>6</lpage>. doi:<pub-id pub-id-type="doi">10.1021/acs.chemrev.0c01269</pub-id>; <pub-id pub-id-type="pmid">33499607</pub-id></mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Huo</surname> <given-names>W</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Dominguez-Gutierrez</surname> <given-names>FJ</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>K</given-names></string-name>, <string-name><surname>Kurpaska</surname> <given-names>&#x0141;</given-names></string-name>, <string-name><surname>Fang</surname> <given-names>F</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>High-entropy materials for electrocatalytic applications: a review of first principles modeling and simulations</article-title>. <source>Mater Res Lett</source>. <year>2023</year>;<volume>11</volume>(<issue>9</issue>):<fpage>713</fpage>&#x2013;<lpage>32</lpage>. doi:<pub-id pub-id-type="doi">10.1080/21663831.2023.2224397</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>Z</given-names></string-name>, <string-name><surname>He</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Recent advances and applications of machine learning in electrocatalysis</article-title>. <source>J Mater Inf</source>. <year>2023</year>;<volume>3</volume>(<issue>3</issue>):<fpage>18</fpage>. doi:<pub-id pub-id-type="doi">10.20517/jmi.2023.23</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Stricker</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Electrocatalyst discovery through text mining and multi-objective optimization</article-title>. <comment>arXiv:2502.20860. 2025</comment>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhao</surname> <given-names>G</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Jena</surname> <given-names>P</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>XC</given-names></string-name></person-group>. <article-title>Machine learning-assisted low-dimensional electrocatalysts design for energy conversion</article-title>. <source>Chem Rev</source>. <year>2023</year>;<volume>123</volume>(<issue>3</issue>):<fpage>1398</fpage>&#x2013;<lpage>454</lpage>. doi:<pub-id pub-id-type="doi">10.1021/acs.chemrev.2c00297</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ding</surname> <given-names>R</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Bando</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Unlocking the potential: machine learning applications in electrocatalyst design for electrochemical hydrogen energy transformation</article-title>. <source>Chem Soc Rev</source>. <year>2024</year>;<volume>53</volume>(<issue>23</issue>):<fpage>11390</fpage>&#x2013;<lpage>461</lpage>. doi:<pub-id pub-id-type="doi">10.1039/d4cs00844h</pub-id>; <pub-id pub-id-type="pmid">39382108</pub-id></mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>M</given-names></string-name>, <string-name><surname>Cheng</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>M</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Interpretable physics-informed machine learning approaches to accelerate electrocatalyst development</article-title>. <source>J Mater Inf</source>. <year>2025</year>;<volume>5</volume>(<issue>2</issue>):<fpage>15</fpage>. doi:<pub-id pub-id-type="doi">10.20517/jmi.2024.67</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Choudhary</surname> <given-names>K</given-names></string-name>, <string-name><surname>DeCost</surname> <given-names>B</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>C</given-names></string-name>, <string-name><surname>Jain</surname> <given-names>A</given-names></string-name>, <string-name><surname>Tavazza</surname> <given-names>F</given-names></string-name>, <string-name><surname>Cohn</surname> <given-names>R</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Recent advances and applications of deep learning methods in materials science</article-title>. <source>npj Comput Mater</source>. <year>2022</year>;<volume>8</volume>(<issue>1</issue>):<fpage>59</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41524-022-00734-6</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jiang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>S</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Lookman</surname> <given-names>T</given-names></string-name>, <string-name><surname>Su</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Applications of natural language processing and large language models in materials discovery</article-title>. <source>npj Comput Mater</source>. <year>2025</year>;<volume>11</volume>(<issue>1</issue>):<fpage>79</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41524-025-01554-0</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shetty</surname> <given-names>V</given-names></string-name>, <string-name><surname>Shabari Shedthi</surname> <given-names>B</given-names></string-name>, <string-name><surname>Shashishekar</surname> <given-names>C</given-names></string-name></person-group>. <article-title>Application and challenges of machine learning techniques in mining engineering and material science</article-title>. <source>J Mines Met Fuels</source>. <year>2023</year>:<fpage>1989</fpage>&#x2013;<lpage>2000</lpage>. doi:<pub-id pub-id-type="doi">10.18311/jmmf/2023/36099</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yin</surname> <given-names>G</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Li</surname> <given-names>T</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>C</given-names></string-name>, <string-name><surname>Jia</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Machine learning-assisted high-throughput screening for electrocatalytic hydrogen evolution reaction</article-title>. <source>Molecules</source>. <year>2025</year>;<volume>30</volume>(<issue>4</issue>):<fpage>759</fpage>. doi:<pub-id pub-id-type="doi">10.3390/molecules30040759</pub-id>; <pub-id pub-id-type="pmid">40005070</pub-id></mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mai</surname> <given-names>H</given-names></string-name>, <string-name><surname>Le</surname> <given-names>TC</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>D</given-names></string-name>, <string-name><surname>Winkler</surname> <given-names>DA</given-names></string-name>, <string-name><surname>Caruso</surname> <given-names>RA</given-names></string-name></person-group>. <article-title>Machine learning for electrocatalyst and photocatalyst design and discovery</article-title>. <source>Chem Rev</source>. <year>2022</year>;<volume>122</volume>(<issue>16</issue>):<fpage>13478</fpage>&#x2013;<lpage>515</lpage>. doi:<pub-id pub-id-type="doi">10.1021/acs.chemrev.2c00061</pub-id>; <pub-id pub-id-type="pmid">35862246</pub-id></mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Nie</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Application of machine learning in material synthesis and property prediction</article-title>. <source>Materials</source>. <year>2023</year>;<volume>16</volume>(<issue>17</issue>):<fpage>5977</fpage>. doi:<pub-id pub-id-type="doi">10.3390/ma16175977</pub-id>; <pub-id pub-id-type="pmid">37687675</pub-id></mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Naveed</surname> <given-names>H</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>AU</given-names></string-name>, <string-name><surname>Qiu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Saqib</surname> <given-names>M</given-names></string-name>, <string-name><surname>Anwar</surname> <given-names>S</given-names></string-name>, <string-name><surname>Usman</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A comprehensive overview of large language models</article-title>. <comment>arXiv:2307.06435. 2023</comment>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Chandrasekhar</surname> <given-names>A</given-names></string-name>, <string-name><surname>Farimani</surname> <given-names>OB</given-names></string-name>, <string-name><surname>Ajenifujah</surname> <given-names>OT</given-names></string-name>, <string-name><surname>Ock</surname> <given-names>J</given-names></string-name>, <string-name><surname>Farimani</surname> <given-names>AB</given-names></string-name></person-group>. <article-title>NANOGPT: a query-driven large language model retrieval-augmented generation system for nanotechnology research</article-title>. <comment>arXiv:2502.20541. 2025</comment>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Polak</surname> <given-names>MP</given-names></string-name>, <string-name><surname>Morgan</surname> <given-names>D</given-names></string-name></person-group>. <article-title>Extracting accurate materials data from research papers with conversational language models and prompt engineering</article-title>. <source>Nat Commun</source>. <year>2024</year>;<volume>15</volume>(<issue>1</issue>):<fpage>1569</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41467-024-45914-8</pub-id>; <pub-id pub-id-type="pmid">38383556</pub-id></mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Miret</surname> <given-names>S</given-names></string-name>, <string-name><surname>Anoop Krishnan</surname> <given-names>NM</given-names></string-name></person-group>. <article-title>Are LLMs ready for real-world materials discovery?</article-title> <comment>arXiv:2402.05200. 2024</comment>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Foppiano</surname> <given-names>L</given-names></string-name>, <string-name><surname>Lambard</surname> <given-names>G</given-names></string-name>, <string-name><surname>Amagasa</surname> <given-names>T</given-names></string-name>, <string-name><surname>Ishii</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Mining experimental data from materials science literature with large language models: an evaluation study</article-title>. <source>Sci Technol Adv Mater Meth</source>. <year>2024</year>;<volume>4</volume>(<issue>1</issue>):<fpage>2356506</fpage>. doi:<pub-id pub-id-type="doi">10.1080/27660400.2024.2356506</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ma</surname> <given-names>X</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>D</given-names></string-name>, <string-name><surname>Li</surname> <given-names>L</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>X</given-names></string-name>, <string-name><surname>Chang</surname> <given-names>X</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Automated review generation method based on large language models</article-title>. <comment>arXiv:2407.20906. 2024</comment>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Choi</surname> <given-names>J</given-names></string-name>, <string-name><surname>Im</surname> <given-names>S</given-names></string-name>, <string-name><surname>Choi</surname> <given-names>J</given-names></string-name>, <string-name><surname>Surendran</surname> <given-names>S</given-names></string-name>, <string-name><surname>Moon</surname> <given-names>DJ</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>JY</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Recent advances in 2D structured materials with defect-exploiting design strategies for electrocatalysis of nitrate to ammonia</article-title>. <source>Energy Mater</source>. <year>2024</year>;<volume>4</volume>(<issue>2</issue>):<fpage>400020</fpage>. doi:<pub-id pub-id-type="doi">10.20517/energymater.2023.67</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Omeiza</surname> <given-names>LA</given-names></string-name>, <string-name><surname>Abdalla</surname> <given-names>AM</given-names></string-name>, <string-name><surname>Wei</surname> <given-names>B</given-names></string-name>, <string-name><surname>Dhanasekaran</surname> <given-names>A</given-names></string-name>, <string-name><surname>Subramanian</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Afroze</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Nanostructured electrocatalysts for advanced applications in fuel cells</article-title>. <source>Energies</source>. <year>2023</year>;<volume>16</volume>(<issue>4</issue>):<fpage>1876</fpage>. doi:<pub-id pub-id-type="doi">10.3390/en16041876</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ding</surname> <given-names>R</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Tan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Unlocking new insights for electrocatalyst design: a unique data science workflow leveraging Internet-sourced big data</article-title>. <source>ACS Catal</source>. <year>2023</year>;<volume>13</volume>(<issue>20</issue>):<fpage>13267</fpage>&#x2013;<lpage>81</lpage>. doi:<pub-id pub-id-type="doi">10.1021/acscatal.3c01914</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ali</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ahmad</surname> <given-names>A</given-names></string-name>, <string-name><surname>Hussain</surname> <given-names>I</given-names></string-name>, <string-name><surname>Ahmad Shah</surname> <given-names>SS</given-names></string-name>, <string-name><surname>Sufyan Javed</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ali</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Experimental and theoretical aspects of MXenes-based energy storage and energy conversion devices</article-title>. <source>J Chem Env</source>. <year>2023</year>;<volume>2</volume>(<issue>2</issue>):<fpage>54</fpage>&#x2013;<lpage>81</lpage>. doi:<pub-id pub-id-type="doi">10.56946/jce.v2i2.214</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Pan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Shan</surname> <given-names>X</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>F</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Accelerating the discovery of oxygen reduction electrocatalysts: high-throughput screening of element combinations in Pt-based high-entropy alloys</article-title>. <source>Angew Chem Int Ed</source>. <year>2024</year>;<volume>63</volume>(<issue>37</issue>):<fpage>e202407116</fpage>. doi:<pub-id pub-id-type="doi">10.1002/anie.202407116</pub-id>; <pub-id pub-id-type="pmid">38934207</pub-id></mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jijaba Kadam</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ramchandra Jadhav</surname> <given-names>M</given-names></string-name>, <string-name><surname>Abasaheb Howal</surname> <given-names>S</given-names></string-name>, <string-name><surname>Martand Kharmate</surname> <given-names>G</given-names></string-name>, <string-name><surname>Uttam Pandit</surname> <given-names>V</given-names></string-name></person-group>. <article-title>Machine learning-based optimization method for the oxygen evolution and reduction reaction of the high-entropy alloy catalysts</article-title>. <source>Int J Recent Innov Trends Comput Commun</source>. <year>2023</year>;<volume>11</volume>(<issue>9</issue>):<fpage>2123</fpage>&#x2013;<lpage>35</lpage>. doi:<pub-id pub-id-type="doi">10.17762/ijritcc.v11i9.9214</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wei</surname> <given-names>C</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Mu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Qin</surname> <given-names>R</given-names></string-name>, <string-name><surname>Yin</surname> <given-names>Y</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Large language models assisted materials development: case of predictive analytics for oxygen evolution reaction catalysts of (oxy)hydroxides</article-title>. <source>ACS Sustain Chem Eng</source>. <year>2025</year>;<volume>13</volume>(<issue>14</issue>):<fpage>5368</fpage>&#x2013;<lpage>80</lpage>. doi:<pub-id pub-id-type="doi">10.1021/acssuschemeng.5c00798</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zheng</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>O</given-names></string-name>, <string-name><surname>Borgs</surname> <given-names>C</given-names></string-name>, <string-name><surname>Chayes</surname> <given-names>JT</given-names></string-name>, <string-name><surname>Yaghi</surname> <given-names>OM</given-names></string-name></person-group>. <article-title>ChatGPT chemistry assistant for text mining and the prediction of MOF synthesis</article-title>. <source>J Am Chem Soc</source>. <year>2023</year>;<volume>145</volume>(<issue>32</issue>):<fpage>18048</fpage>&#x2013;<lpage>62</lpage>. doi:<pub-id pub-id-type="doi">10.1021/jacs.3c05819</pub-id>; <pub-id pub-id-type="pmid">37548379</pub-id></mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Akbashev</surname> <given-names>AR</given-names></string-name></person-group>. <article-title>Electrocatalysis goes nuts</article-title>. <source>ACS Catal</source>. <year>2022</year>;<volume>12</volume>(<issue>8</issue>):<fpage>4296</fpage>&#x2013;<lpage>301</lpage>. doi:<pub-id pub-id-type="doi">10.1021/acscatal.2c00123</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lin</surname> <given-names>X</given-names></string-name>, <string-name><surname>Du</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhen</surname> <given-names>S</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Pei</surname> <given-names>C</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Machine learning-assisted dual-atom sites design with interpretable descriptors unifying electrocatalytic reactions</article-title>. <source>Nat Commun</source>. <year>2024</year>;<volume>15</volume>(<issue>1</issue>):<fpage>8169</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41467-024-52519-8</pub-id>; <pub-id pub-id-type="pmid">39289388</pub-id></mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Du</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Revisiting electrocatalyst design by a knowledge graph of Cu-based catalysts for CO<sub>2</sub> reduction</article-title>. <source>ACS Catal</source>. <year>2023</year>;<volume>13</volume>(<issue>13</issue>):<fpage>8525</fpage>&#x2013;<lpage>34</lpage>. doi:<pub-id pub-id-type="doi">10.1021/acscatal.3c00759</pub-id>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Huerta</surname> <given-names>GV</given-names></string-name>, <string-name><surname>Hisama</surname> <given-names>K</given-names></string-name>, <string-name><surname>Koyama</surname> <given-names>M</given-names></string-name></person-group>. <article-title>A comprehensive computer-aided review of trends in heterogeneous catalysis and its theoretical modelling for engineering catalytic activity</article-title>. <source>J Chem Eng Jpn</source>. <year>2024</year>;<volume>57</volume>(<issue>1</issue>):<fpage>2360928</fpage>. doi:<pub-id pub-id-type="doi">10.1080/00219592.2024.2360928</pub-id>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>X</given-names></string-name>, <string-name><surname>Fu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>F</given-names></string-name>, <string-name><surname>Lv</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Cyclodextrin-based architectures for electrochemical sensing: from molecular recognition to functional hybrids</article-title>. <source>Anal Methods</source>. <year>2025</year>;<volume>17</volume>(<issue>21</issue>):<fpage>4300</fpage>&#x2013;<lpage>20</lpage>. doi:<pub-id pub-id-type="doi">10.1039/d5ay00612k</pub-id>; <pub-id pub-id-type="pmid">40392560</pub-id></mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>Large language model in electrocatalysis</article-title>. <source>Chin J Catal</source>. <year>2024</year>;<volume>59</volume>:<fpage>7</fpage>&#x2013;<lpage>14</lpage>. doi:<pub-id pub-id-type="doi">10.1016/S1872-2067(23)64612-1</pub-id>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Du</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cui</surname> <given-names>W</given-names></string-name></person-group>. <article-title>CataLM: empowering catalyst design through large language models</article-title>. <source>Int J Mach Learn Cybern</source>. <year>2025</year>;<volume>16</volume>(<issue>5&#x2013;6</issue>):<fpage>3681</fpage>&#x2013;<lpage>91</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s13042-024-02473-0</pub-id>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Park</surname> <given-names>YJ</given-names></string-name>, <string-name><surname>Jerng</surname> <given-names>SE</given-names></string-name>, <string-name><surname>Yoon</surname> <given-names>S</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name></person-group>. <article-title>1.5 million materials narratives generated by chatbots</article-title>. <source>Sci Data</source>. <year>2024</year>;<volume>11</volume>(<issue>1</issue>):<fpage>1060</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41597-024-03886-w</pub-id>; <pub-id pub-id-type="pmid">39341807</pub-id></mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>S</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>K</given-names></string-name></person-group>. <article-title>How does a generative large language model perform on domain-specific information Extraction?&#x2014;a comparison between GPT-4 and a rule-based method on band gap extraction</article-title>. <source>J Chem Inf Model</source>. <year>2024</year>;<volume>64</volume>(<issue>20</issue>):<fpage>7895</fpage>&#x2013;<lpage>904</lpage>. doi:<pub-id pub-id-type="doi">10.1021/acs.jcim.4c00882</pub-id>; <pub-id pub-id-type="pmid">39375999</pub-id></mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Eger</surname> <given-names>S</given-names></string-name>, <string-name><surname>Cao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>D&#x0027;Souza</surname> <given-names>J</given-names></string-name>, <string-name><surname>Geiger</surname> <given-names>A</given-names></string-name>, <string-name><surname>Greisinger</surname> <given-names>C</given-names></string-name>, <string-name><surname>Gross</surname> <given-names>S</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Transforming science with large language models: a survey on AI-assisted scientific discovery, experimentation, content generation, and evaluation</article-title>. <comment>arXiv:2502.05151. 2025</comment>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zimmermann</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Bazgir</surname> <given-names>A</given-names></string-name>, <string-name><surname>Afzal</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Agbere</surname> <given-names>F</given-names></string-name>, <string-name><surname>Ai</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Alampara</surname> <given-names>N</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Reflections from the 2024 large language model (LLM) hackathon for applications in materials science and chemistry</article-title>. <comment>arXiv:2411.15221. 2024</comment>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>W</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>B</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>T</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ouyang</surname> <given-names>W</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>MOOSE-CHEM: large language models for rediscovering unseen chemistry scientific hypotheses</article-title>. <comment>arXiv:2410.07076. 2024</comment>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Song</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Ju</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>C</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Li</surname> <given-names>C</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Q</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Leveraging large language models for universal scientific formula and theory discovery</article-title>. <comment>arXiv:2503.06512. 2025</comment>.</mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zheng</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Florit</surname> <given-names>F</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>B</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>H</given-names></string-name>, <string-name><surname>Li</surname> <given-names>SC</given-names></string-name>, <string-name><surname>Nandiwale</surname> <given-names>KY</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Integrating machine learning and large language models to advance exploration of electrochemical reactions</article-title>. <source>Angew Chem Int Ed</source>. <year>2025</year>;<volume>64</volume>(<issue>6</issue>):<fpage>e202418074</fpage>. doi:<pub-id pub-id-type="doi">10.1002/anie.202418074</pub-id>; <pub-id pub-id-type="pmid">39625837</pub-id></mixed-citation></ref>
<ref id="ref-52"><label>[52]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Su</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ye</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xie</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Jiang</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Automation and machine learning augmented by large language models in a catalysis study</article-title>. <source>Chem Sci</source>. <year>2024</year>;<volume>15</volume>(<issue>31</issue>):<fpage>12200</fpage>&#x2013;<lpage>33</lpage>. doi:<pub-id pub-id-type="doi">10.1039/d3sc07012c</pub-id>; <pub-id pub-id-type="pmid">39118602</pub-id></mixed-citation></ref>
<ref id="ref-53"><label>[53]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Duignan</surname> <given-names>TT</given-names></string-name></person-group>. <article-title>The potential of neural network potentials</article-title>. <source>ACS Phys Chem Au</source>. <year>2024</year>;<volume>4</volume>(<issue>3</issue>):<fpage>232</fpage>&#x2013;<lpage>41</lpage>. doi:<pub-id pub-id-type="doi">10.1021/acsphyschemau.4c00004</pub-id>; <pub-id pub-id-type="pmid">38800721</pub-id></mixed-citation></ref>
<ref id="ref-54"><label>[54]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Cui</surname> <given-names>W</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>J</given-names></string-name>, <string-name><surname>Du</surname> <given-names>Y</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Large language model enhanced corpus of CO<sub>2</sub> reduction electrocatalysts and synthesis procedures</article-title>. <source>Sci Data</source>. <year>2024</year>;<volume>11</volume>(<issue>1</issue>):<fpage>347</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41597-024-03180-9</pub-id>; <pub-id pub-id-type="pmid">38582751</pub-id></mixed-citation></ref>
<ref id="ref-55"><label>[55]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Weston</surname> <given-names>L</given-names></string-name>, <string-name><surname>Tshitoyan</surname> <given-names>V</given-names></string-name>, <string-name><surname>Dagdelen</surname> <given-names>J</given-names></string-name>, <string-name><surname>Kononova</surname> <given-names>O</given-names></string-name>, <string-name><surname>Trewartha</surname> <given-names>A</given-names></string-name>, <string-name><surname>Persson</surname> <given-names>KA</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Named entity recognition and normalization applied to large-scale information extraction from the materials science literature</article-title>. <source>J Chem Inf Model</source>. <year>2019</year>;<volume>59</volume>(<issue>9</issue>):<fpage>3692</fpage>&#x2013;<lpage>702</lpage>. doi:<pub-id pub-id-type="doi">10.1021/acs.jcim.9b00470</pub-id>; <pub-id pub-id-type="pmid">31361962</pub-id></mixed-citation></ref>
<ref id="ref-56"><label>[56]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>X</given-names></string-name>, <string-name><surname>Cui</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>X</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A corpus of CO<sub>2</sub> electrocatalytic reduction process extracted from the scientific literature</article-title>. <source>Sci Data</source>. <year>2023</year>;<volume>10</volume>(<issue>1</issue>):<fpage>175</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41597-023-02089-z</pub-id>; <pub-id pub-id-type="pmid">36991006</pub-id></mixed-citation></ref>
<ref id="ref-57"><label>[57]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kim</surname> <given-names>E</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Saunders</surname> <given-names>A</given-names></string-name>, <string-name><surname>McCallum</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ceder</surname> <given-names>G</given-names></string-name>, <string-name><surname>Olivetti</surname> <given-names>E</given-names></string-name></person-group>. <article-title>Materials synthesis insights from scientific literature via text extraction and machine learning</article-title>. <source>Chem Mater</source>. <year>2017</year>;<volume>29</volume>(<issue>21</issue>):<fpage>9436</fpage>&#x2013;<lpage>44</lpage>. doi:<pub-id pub-id-type="doi">10.1021/acs.chemmater.7b03500</pub-id>.</mixed-citation></ref>
<ref id="ref-58"><label>[58]</label><mixed-citation publication-type="other"><article-title>PEESEgroup/Awesome-materials-aware-large-language-models [Internet]</article-title>. <comment>[cited 2025 Jun 11]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://github.com/PEESEgroup/Awesome-Materials-Aware-Large-Language-Models">https://github.com/PEESEgroup/Awesome-Materials-Aware-Large-Language-Models</ext-link>.</mixed-citation></ref>
<ref id="ref-59"><label>[59]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kim</surname> <given-names>E</given-names></string-name>, <string-name><surname>Jensen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>van Grootel</surname> <given-names>A</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Staib</surname> <given-names>M</given-names></string-name>, <string-name><surname>Mysore</surname> <given-names>S</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Inorganic materials synthesis planning with literature-trained neural networks</article-title>. <source>J Chem Inf Model</source>. <year>2020</year>;<volume>60</volume>(<issue>3</issue>):<fpage>1194</fpage>&#x2013;<lpage>201</lpage>. doi:<pub-id pub-id-type="doi">10.1021/acs.jcim.9b00995</pub-id>; <pub-id pub-id-type="pmid">31909619</pub-id></mixed-citation></ref>
<ref id="ref-60"><label>[60]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zuo</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>C</given-names></string-name>, <string-name><surname>Gong</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zeng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>R</given-names></string-name>, <string-name><surname>Ong</surname> <given-names>SP</given-names></string-name></person-group>. <article-title>Accurate, interpretable predictions of materials properties within machine-learned interatomic potentials using explainable graph neural networks</article-title>. <source>npj Comput Mater</source>. <year>2023</year>;<volume>9</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>12</lpage>. doi:<pub-id pub-id-type="doi">10.1038/s41524-023-01095-4</pub-id>.</mixed-citation></ref>
<ref id="ref-61"><label>[61]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ock</surname> <given-names>J</given-names></string-name>, <string-name><surname>Guntuboina</surname> <given-names>C</given-names></string-name>, <string-name><surname>Barati Farimani</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Catalyst energy prediction with CatBERTa: unveiling feature exploration strategies through large language models</article-title>. <source>ACS Catal</source>. <year>2023</year>;<volume>13</volume>(<issue>24</issue>):<fpage>16032</fpage>&#x2013;<lpage>44</lpage>. doi:<pub-id pub-id-type="doi">10.1021/acscatal.3c04956</pub-id>.</mixed-citation></ref>
<ref id="ref-62"><label>[62]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Shoghi</surname> <given-names>N</given-names></string-name>, <string-name><surname>Kolluru</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kitchin</surname> <given-names>JR</given-names></string-name>, <string-name><surname>Ulissi</surname> <given-names>ZW</given-names></string-name>, <string-name><surname>Zitnick</surname> <given-names>CL</given-names></string-name>, <string-name><surname>Wood</surname> <given-names>BM</given-names></string-name></person-group>. <article-title>From molecules to materials: pre-training large generalizable models for atomic property prediction</article-title>. <comment>arXiv:2310.1680. 2023</comment>.</mixed-citation></ref>
<ref id="ref-63"><label>[63]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Pei</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Yin</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Language models for materials discovery and sustainability: progress, challenges, and opportunities</article-title>. <source>Prog Mater Sci</source>. <year>2025</year>;<volume>145</volume>:<fpage>14849</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.pmatsci.2025.101495</pub-id>.</mixed-citation></ref>
<ref id="ref-64"><label>[64]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Mok</surname> <given-names>DH</given-names></string-name>, <string-name><surname>Back</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Language model-based generative model for catalyst discovery Paper presented at: 2024 MRS Fall Meeting &#x0026; Exhibit</article-title>; <year>2024</year>; <publisher-loc>Boston, MA, USA</publisher-loc>: <publisher-name>Materials Research Society</publisher-name> [Intermet]. <comment>[cited 2025 Jun 1]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.mrs.org/meetings-events/annual-meetings/archive/meeting/presentations/view/2024-fall-meeting/2024-fall-meeting-4148755">https://www.mrs.org/meetings-events/annual-meetings/archive/meeting/presentations/view/2024-fall-meeting/2024-fall-meeting-4148755</ext-link>.</mixed-citation></ref>
<ref id="ref-65"><label>[65]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mok</surname> <given-names>DH</given-names></string-name>, <string-name><surname>Back</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Generative pretrained transformer for heterogeneous catalysts</article-title>. <source>J Am Chem Soc</source>. <year>2024</year>;<volume>146</volume>(<issue>49</issue>):<fpage>33712</fpage>&#x2013;<lpage>22</lpage>. doi:<pub-id pub-id-type="doi">10.1021/jacs.4c11504</pub-id>; <pub-id pub-id-type="pmid">39576215</pub-id></mixed-citation></ref>
<ref id="ref-66"><label>[66]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yu</surname> <given-names>G</given-names></string-name>, <string-name><surname>Mok</surname> <given-names>DH</given-names></string-name>, <string-name><surname>Jang</surname> <given-names>HY</given-names></string-name>, <string-name><surname>Jung</surname> <given-names>HD</given-names></string-name>, <string-name><surname>Siahrostami</surname> <given-names>S</given-names></string-name>, <string-name><surname>Back</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Leveraging Machine learning and active motifs-based catalyst design for discovery of oxygen reduction electrocatalysts for hydrogen peroxide production</article-title>. <source>J Catal</source>. <year>2025</year>;<volume>442</volume>(<issue>5</issue>):<fpage>115906</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.jcat.2024.115906</pub-id>.</mixed-citation></ref>
<ref id="ref-67"><label>[67]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>C</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>B</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Lian</surname> <given-names>X</given-names></string-name>, <string-name><surname>Deng</surname> <given-names>H</given-names></string-name></person-group>. <article-title>SynAsk: unleashing the power of large language models in organic synthesis</article-title>. <source>Chem Sci</source>. <year>2025</year>;<volume>16</volume>(<issue>1</issue>):<fpage>43</fpage>&#x2013;<lpage>56</lpage>. doi:<pub-id pub-id-type="doi">10.1039/d4sc04757e</pub-id>; <pub-id pub-id-type="pmid">39600494</pub-id></mixed-citation></ref>
<ref id="ref-68"><label>[68]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Choi</surname> <given-names>SE</given-names></string-name>, <string-name><surname>Jang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Yoon</surname> <given-names>S</given-names></string-name>, <string-name><surname>Yoo</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ahn</surname> <given-names>J</given-names></string-name>, <string-name><surname>Kim</surname> <given-names>M</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>LLM-driven synthesis planning for quantum dot materials development</article-title>. <source>J Chem Inf Model</source>. <year>2025</year>;<volume>65</volume>(<issue>6</issue>):<fpage>2748</fpage>&#x2013;<lpage>58</lpage>. doi:<pub-id pub-id-type="doi">10.1021/acs.jcim.4c01529</pub-id>; <pub-id pub-id-type="pmid">40069968</pub-id></mixed-citation></ref>
<ref id="ref-69"><label>[69]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Pan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Xiao</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>F</given-names></string-name>, <string-name><surname>Li</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ullah</surname> <given-names>S</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>Chat-microreactor: a large-language-model-based assistant for designing continuous flow systems</article-title>. <source>Chem Eng Sci</source>. <year>2025</year>;<volume>311</volume>:<fpage>121567</fpage>. doi:<pub-id pub-id-type="doi">10.26434/chemrxiv-2024-52g9l</pub-id>.</mixed-citation></ref>
<ref id="ref-70"><label>[70]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bandeira</surname> <given-names>L</given-names></string-name>, <string-name><surname>Ferreira</surname> <given-names>H</given-names></string-name>, <string-name><surname>de Almeida</surname> <given-names>JM</given-names></string-name>, <string-name><surname>Jardim de Paula</surname> <given-names>A</given-names></string-name>, <string-name><surname>Dalpian</surname> <given-names>GM</given-names></string-name></person-group>. <article-title>CO<sub>2</sub> reduction beyond copper-based catalysts: a natural language processing review from the scientific literature</article-title>. <source>ACS Sustain Chem Eng</source>. <year>2024</year>;<volume>12</volume>(<issue>11</issue>):<fpage>4411</fpage>&#x2013;<lpage>22</lpage>. doi:<pub-id pub-id-type="doi">10.1021/acssuschemeng.3c06920</pub-id>.</mixed-citation></ref>
<ref id="ref-71"><label>[71]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Zitnick</surname> <given-names>CL</given-names></string-name>, <string-name><surname>Chanussot</surname> <given-names>L</given-names></string-name>, <string-name><surname>Das</surname> <given-names>A</given-names></string-name>, <string-name><surname>Goyal</surname> <given-names>S</given-names></string-name>, <string-name><surname>Heras-Domingo</surname> <given-names>J</given-names></string-name>, <string-name><surname>Ho</surname> <given-names>C</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>An introduction to electrocatalyst design using machine learning for renewable energy storage</article-title>. <comment>arXiv:2010.09435. 2020</comment>.</mixed-citation></ref>
<ref id="ref-72"><label>[72]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>L</given-names></string-name>, <string-name><surname>Tian</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>X</given-names></string-name>, <string-name><surname>Yao</surname> <given-names>S</given-names></string-name>, <string-name><surname>Lu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A universal machine learning framework for electrocatalyst innovation: a case study of discovering alloys for hydrogen evolution reaction</article-title>. <source>Adv Funct Mater</source>. <year>2022</year>;<volume>32</volume>(<issue>47</issue>):<fpage>2208418</fpage>. doi:<pub-id pub-id-type="doi">10.1002/adfm.202208418</pub-id>.</mixed-citation></ref>
<ref id="ref-73"><label>[73]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Shan</surname> <given-names>X</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cai</surname> <given-names>F</given-names></string-name>, <string-name><surname>Gao</surname> <given-names>H</given-names></string-name>, <string-name><surname>Xu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Accelerating the discovery of efficient high-entropy alloy electrocatalysts: high-throughput experimentation and data-driven strategies</article-title>. <source>Nano Lett</source>. <year>2024</year>;<volume>24</volume>(<issue>37</issue>):<fpage>11632</fpage>&#x2013;<lpage>40</lpage>. doi:<pub-id pub-id-type="doi">10.1021/acs.nanolett.4c03208</pub-id>; <pub-id pub-id-type="pmid">39225654</pub-id></mixed-citation></ref>
<ref id="ref-74"><label>[74]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Phillips</surname> <given-names>PJ</given-names></string-name>, <string-name><surname>Phillips</surname> <given-names>PJ</given-names></string-name>, <string-name><surname>Hahn</surname> <given-names>CA</given-names></string-name>, <string-name><surname>Fontana</surname> <given-names>PC</given-names></string-name>, <string-name><surname>Yates</surname> <given-names>AN</given-names></string-name>, <string-name><surname>Greene</surname> <given-names>K</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Four principles of explainable artificial intelligence</article-title>. <source>Natl Inst Stand Technol</source>. <year>2021</year>;<volume>9</volume>:<fpage>1</fpage>&#x2013;<lpage>23</lpage>. doi:<pub-id pub-id-type="doi">10.6028/nist.ir.8312</pub-id>.</mixed-citation></ref>
<ref id="ref-75"><label>[75]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hysmith</surname> <given-names>H</given-names></string-name>, <string-name><surname>Foadian</surname> <given-names>E</given-names></string-name>, <string-name><surname>Padhy</surname> <given-names>SP</given-names></string-name>, <string-name><surname>Kalinin</surname> <given-names>SV</given-names></string-name>, <string-name><surname>Moore</surname> <given-names>RG</given-names></string-name>, <string-name><surname>Ovchinnikova</surname> <given-names>OS</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>The future of self-driving laboratories: from human in the loop interactive AI to gamification</article-title>. <source>Digit Discov</source>. <year>2024</year>;<volume>3</volume>(<issue>4</issue>):<fpage>621</fpage>&#x2013;<lpage>36</lpage>. doi:<pub-id pub-id-type="doi">10.1039/d4dd00040d</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>








