<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">63551</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2025.063551</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>The Future of Artificial Intelligence in the Face of Data Scarcity</article-title>
<alt-title alt-title-type="left-running-head">The Future of Artificial Intelligence in the Face of Data Scarcity</alt-title>
<alt-title alt-title-type="right-running-head">The Future of Artificial Intelligence in the Face of Data Scarcity</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Abdalla</surname><given-names>Hemn Barzan</given-names></name><xref ref-type="aff" rid="aff-1">1</xref><email>habdalla@kean.edu</email></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Kumar</surname><given-names>Yulia</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Marchena</surname><given-names>Jose</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Guzman</surname><given-names>Stephany</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-5" contrib-type="author">
<name name-style="western"><surname>Awlla</surname><given-names>Ardalan</given-names></name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<contrib id="author-6" contrib-type="author">
<name name-style="western"><surname>Gheisari</surname><given-names>Mehdi</given-names></name><xref ref-type="aff" rid="aff-4">4</xref></contrib>
<contrib id="author-7" contrib-type="author">
<name name-style="western"><surname>Cheraghy</surname><given-names>Maryam</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<aff id="aff-1"><label>1</label><institution>Department of Computer Science, Wenzhou-Kean University</institution>, <addr-line>Wenzhou, 325015</addr-line>, <country>China</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of Computer Science and Technology, Kean University</institution>, <addr-line>Union, NJ 07083</addr-line>, <country>USA</country></aff>
<aff id="aff-3"><label>3</label><institution>Department of Computer Science, Cihan University Sulaimaniya</institution>, <country>Sulaymaniyah</country>, <addr-line>46001</addr-line>, <country>Iraq</country></aff>
<aff id="aff-4"><label>4</label><institution>Institute of Artificial Intelligence, Shaoxing University</institution>, <addr-line>Shaoxing, 312010</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Hemn Barzan Abdalla. Email: <email>habdalla@kean.edu</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2025</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>09</day><month>06</month><year>2025</year>
</pub-date>
<volume>84</volume>
<issue>1</issue>
<fpage>1073</fpage>
<lpage>1099</lpage>
<history>
<date date-type="received">
<day>17</day>
<month>1</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>28</day>
<month>3</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 The Authors.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_63551.pdf"></self-uri>
<abstract>
<p>Dealing with data scarcity is the biggest challenge faced by Artificial Intelligence (AI), and it will be interesting to see how we overcome this obstacle in the future, but for now, &#x201C;THE SHOW MUST GO ON!!!&#x201D; As AI spreads and transforms more industries, the lack of data is a significant obstacle: the best methods for teaching machines how real-world processes work. This paper explores the considerable implications of data scarcity for the AI industry, which threatens to restrict its growth and potential, and proposes plausible solutions and perspectives. In addition, this article focuses highly on different ethical considerations: privacy, consent, and non-discrimination principles during AI model developments under limited conditions. Besides, innovative technologies are investigated through the paper in aspects that need implementation by incorporating transfer learning, few-shot learning, and data augmentation to adapt models so they could fit effective use processes in low-resource settings. This thus emphasizes the need for collaborative frameworks and sound methodologies that ensure applicability and fairness, tackling the technical and ethical challenges associated with data scarcity in AI. This article also discusses prospective approaches to dealing with data scarcity, emphasizing the blend of synthetic data and traditional models and the use of advanced machine learning techniques such as transfer learning and few-shot learning. These techniques aim to enhance the flexibility and effectiveness of AI systems across various industries while ensuring sustainable AI technology development amid ongoing data scarcity.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Data scarcity</kwd>
<kwd>artificial intelligence</kwd>
<kwd>application of artificial intelligence</kwd>
<kwd>ethical considerations</kwd>
<kwd>artificial general intelligence</kwd>
<kwd>synthetic data</kwd>
</kwd-group>
<funding-group>
<award-group id="awg1">
<funding-source>Internal Research Support Program</funding-source>
<award-id>IRSPG202202</award-id>
</award-group>
</funding-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>With the advancement of AI, many sectors have progressed to a new, innovative future, including healthcare, finance, and other sectors. However, with continued advancements in AI, a basic limiting factor has slowly come to light, which may reshape the evolution of this technology: not enough high-quality, real-world data is available. Here is the fundamental concept of how this works&#x2014;Data &#x2192; AI Engine &#x2192; Learning, Adapting, and Decision Making. However, the need for data is growing faster than authentic and varied sources can provide. The lack of naturally occurring data, i.e., unadulterated natural information based upon direct human interactions and experiences as they occur in the world, is putting AI development at serious risk [<xref ref-type="bibr" rid="ref-1">1</xref>&#x2013;<xref ref-type="bibr" rid="ref-5">5</xref>].</p>
<p>The consequences of less data are monumental, including ethical questions and the efficacy and accuracy of AI models. To truly grasp the severity of this matter, it is essential to think about how AI categorizes things not as mere numbers or letters but as patterns formed by behaviors learned from enormous data [<xref ref-type="bibr" rid="ref-6">6</xref>&#x2013;<xref ref-type="bibr" rid="ref-10">10</xref>]. However, when inputting datasets are restricted, is the AI we rely on left to function with insufficient information or underwhelming accuracy? Data scarcity is an existential problem for AI in all sectors [<xref ref-type="bibr" rid="ref-11">11</xref>&#x2013;<xref ref-type="bibr" rid="ref-14">14</xref>], though it does not exist technically. Nonetheless, it significantly affects the marginal AI systems that businesses use to enhance performance [<xref ref-type="bibr" rid="ref-15">15</xref>&#x2013;<xref ref-type="bibr" rid="ref-19">19</xref>].</p>
<p>With data learning being a dilemma, the future of AI is at a critical juncture. Given that AI systems become progressively dependent on enormous amounts of data to acquire knowledge and produce predictions, this restriction in the form of insufficient data can considerably impact their effectiveness and applicability in different domains. This issue is especially pronounced in disciplines like healthcare, where the precision and reliability of AI-powered solutions are critical. For example, we used relatively small datasets for many of the AI applications in healthcare, which have yielded accuracies below those seen in clinical settings [<xref ref-type="bibr" rid="ref-18">18</xref>]; this makes the AI models less robust and generalized, which calls for developing novel methodologies to deal with limited data situations.</p>
<p>In addition, the ethical consequences of AI creation in the setting of a lack of data are significant and must not be ignored. Data reliance and the need for extensive datasets in machine learning [<xref ref-type="bibr" rid="ref-16">16</xref>] have made questions of privacy, consent, and bias in AI algorithms to reproduce existing inequalities theoretically even more pertinent [<xref ref-type="bibr" rid="ref-1">1</xref>]. Hence, organizations are required to overcome such socio-technical issues, which will, in turn, facilitate a robust and moral AI system [<xref ref-type="bibr" rid="ref-18">18</xref>]. This joint effort was necessary with many stakeholders of different sectors to create ethical standards and regulatory frameworks to rule the responsible generation and implementation of AI technologies [<xref ref-type="bibr" rid="ref-4">4</xref>].</p>
<p>Beyond ethical considerations, the future of AI depends on advances in data engineering practices and AI model lifecycle management, given the scarcity of data. Organizations must develop the capabilities to update their AI systems as the data evolves to make their prediction relevant and effective over time [<xref ref-type="bibr" rid="ref-9">9</xref>]. Additionally, the convergence of AI with other technologies (like machine learning and intelligence augmentation) can help in better utilization of data as well as in improving AI productivity [<xref ref-type="bibr" rid="ref-2">2</xref>]. However, by encouraging the cross-fertilization of ideas between different fields of science, stakeholders can counteract the negative consequences of data shortages and facilitate the realization of AI&#x2019;s potential in a range of sectors [<xref ref-type="bibr" rid="ref-16">16</xref>].</p>
<p>As AI increasingly matters across domains, this article discusses data scarcity (See Appendix A, <xref ref-type="table" rid="table-15">Table A1</xref>) as a key obstacle to AI progress. <xref ref-type="sec" rid="s2">Section 2</xref> delves into the impact of data scarcity, emphasizing its role in limiting AI&#x2019;s ability to generalize, adapt, and perform effectively in real-world applications. <xref ref-type="sec" rid="s3">Section 3</xref> discusses some of the ethical implications of limited data, including bias, privacy, and fairness risks, and calls for responsible AI development.</p>
<p><xref ref-type="sec" rid="s4">Section 4</xref> explores technological innovations such as synthetic data generation, transfer learning, and short learning as potential paths toward overcoming the problem of data scarcity. <xref ref-type="sec" rid="s5">Section 5</xref> provides a case study that illustrates how the generation of synthetic data can be used to augment the training dataset using SMOTE (See Appendix A, <xref ref-type="table" rid="table-15">Table A1</xref>) in the case of a very imbalanced dataset, which improves model performance in terms of generalization. These experimental results confirm the essential role of synthetic data in tackling data scarcity [<xref ref-type="bibr" rid="ref-20">20</xref>,<xref ref-type="bibr" rid="ref-21">21</xref>].</p>

<p>Ultimately, the conclusion argues for a hybrid approach, whereby synthetic data is integrated into existing paradigms, underpinned by ethical frameworks and cross-indexing across all sectors, to empower AI systems to succeed in low-resource settings. This in-depth survey sheds light on addressing data challenges and scaling AI responsibly.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Literature Review</title>
<p><xref ref-type="table" rid="table-1">Table 1</xref> summarizes information about data scarcity across various fields of AI and draws from multiple references throughout the document to highlight ongoing research and methods to address this challenge, as shown in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Literature review</title>
</caption>
<table>
<colgroup>
<col/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>References</th>
<th align="center">Key points</th>
<th align="center">Findings</th>
</tr>
</thead>
<tbody>
<tr>
<td>[<xref ref-type="bibr" rid="ref-17">17</xref>,<xref ref-type="bibr" rid="ref-18">18</xref>]</td>
<td>Focused on evaluating various machine learning data augmentation techniques to address data scarcity.</td>
<td>Emphasized the need for more sophisticated methods that can simulate real-world variability effectively.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-22">22</xref>]</td>
<td>Explored strategies including collaborative filtering and content-based methods to enhance recommendation systems in the face of data scarcity.</td>
<td>Highlighted the critical impact of these strategies on improving user engagement and recommendation diversity.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-23">23</xref>]</td>
<td>Utilized transfer learning and synthetic data generation to adapt technical systems to operate under conditions of data scarcity.</td>
<td>Detailed the development of methodologies robust enough to handle diverse conditions of data scarcity.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-24">24</xref>]</td>
<td>Discussed the implementation of rejection mechanisms in ML (Machine Learning) deployments to improve decision reliability in low-resource settings.</td>
<td>Argued for the necessity of such mechanisms to ensure reliability in critical AI applications.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-25">25</xref>]</td>
<td>Reviewed data-centric approaches, contrasting them with traditional model-centric methods, to address data scarcity in AI.</td>
<td>Identified key aspects where data-centric innovations are crucial for the advancement of AI.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-26">26</xref>]</td>
<td>Proposed the creation of a low-cost mirror environment to simulate real-world conditions for AI training and development.</td>
<td>Demonstrated that such environments can significantly enhance AI system reliability and performance.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-28">28</xref>]</td>
<td>Implemented AI-based methods to integrate diverse data sources for improved pharmacovigilance.</td>
<td>Identified challenges in drug safety monitoring due to sparse data and proposed solutions.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-29">29</xref>]</td>
<td>Developed a novel algorithm to generate synthetic data for training machine learning models, specifically in medical imaging.</td>
<td>Demonstrated an increase in training data availability and improved accuracy in medical diagnostics.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-30">30</xref>]</td>
<td>Utilized deep generative models for the private data synthesis, addressing both privacy concerns and data scarcity.</td>
<td>Indicated that these models could effectively tackle privacy and scarcity issues in AI.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-31">31</xref>,<xref ref-type="bibr" rid="ref-32">32</xref>]</td>
<td>Discussed deep generative models for private data synthesis, focusing on improving classifier performance under privacy constraints.</td>
<td>Found that these models could enhance classifier performance by generating diverse and privacy-compliant data.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-33">33</xref>]</td>
<td>Employed the BERT (Bidirectional Encoder Representations from Transformers) model to generate text data based on topic relevance, aiming to enrich data availability for ML training.</td>
<td>Demonstrated an enhancement in data availability and relevance, improving overall model performance.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-34">34</xref>]</td>
<td>Explored data augmentation techniques to enhance the training and performance of large language models.</td>
<td>Found that these techniques significantly mitigate data scarcity impacts on model training.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-35">35</xref>]</td>
<td>Conducted a comprehensive review of deep learning tools designed to handle data scarcity, examining various strategies and their applications.</td>
<td>Provided an overview of tools&#x2019; capabilities and effectiveness in addressing scarce data challenges.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-36">36</xref>]</td>
<td>Addressed the issue of data scarcity in the context of political communication research, proposing methodological adaptations.</td>
<td>Detailed adaptations in research methodology but specific findings were not provided.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-37">37</xref>]</td>
<td>Proposed the creation of synthetic profiles using AI to enrich data availability for various applications.</td>
<td>Demonstrated the potential of synthetic data to enhance availability across diverse AI applications.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-38">38</xref>]</td>
<td>Discussion on the global impact of data scarcity on AI development, focusing on challenges developers face.</td>
<td>Highlighted the pervasive nature of data scarcity and its global impact on AI development.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-39">39</xref>]</td>
<td>Proposed heuristic training methods as part of Resource Constrained Training (RCT) to optimize training under resource limitations.</td>
<td>Showcased the effectiveness of RCT methods in enhancing training efficiency in resource-poor environments.</td>
</tr>
<tr>
<td>[<xref ref-type="bibr" rid="ref-8">8</xref>]</td>
<td>The importance of data sharing was emphasized, particularly in the context of medical AI development.</td>
<td>Stressed that enhancing data sharing practices is crucial for advancing AI in medicine.</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3">
<label>3</label>
<title>Navigating Data Challenges in AI Development</title>
<sec id="s3_1">
<label>3.1</label>
<title>The Critical Role of Natural Data in AI Development</title>
<p>Training AI models to function properly in unpredictable environments requires natural data, also called raw or unprocessed real-world scenario data [<xref ref-type="bibr" rid="ref-2">2</xref>,<xref ref-type="bibr" rid="ref-6">6</xref>]. The flaws, assumptions, and nuances in this data are perfect for AI to predict and adapt well to a massive range of unique scenarios. Synthetic data, on the other hand, is loosely defined as artificially generated to imitate real-world data and often does not have the authenticity or richness necessary for creating robust and reliable AI.</p>
<p>What sits behind natural data is the truth in how humans interact, speak, and behave. For example, the subtleties of language (i.e., slang in different regions, typos, and differences between syntaxes) are only learned through immersion in linguistic diversity [<xref ref-type="bibr" rid="ref-39">39</xref>]. As a result, models trained only on synthetic data may not generalize well to other demographics and contexts, which could produce biased or incorrect results. Thus, having natural data is crucial in constructing AI systems capable of interacting and servicing a wide range of user bases, addressing the fundamental challenge of responsible progress on AI.</p>
<p>Workflow of processing natural data for AI development:
<list list-type="simple">
<list-item>
<label>&#x2022;</label><p><bold>Start:</bold> Initiation of data collection.</p></list-item>
<list-item>
<label>&#x2022;</label><p><bold>Activities:</bold>
<list list-type="simple">
<list-item>
<label>-</label><p><bold>Collect Natural Data:</bold> Gathering data from various sources.</p></list-item>
<list-item>
<label>-</label><p><bold>Clean Data:</bold> Removing irrelevant or erroneous data.</p></list-item>
<list-item>
<label>-</label><p><bold>Analyze Data:</bold> Understanding patterns and anomalies.</p></list-item>
<list-item>
<label>-</label><p><bold>Train AI Model:</bold> Using the processed data to train the model.</p></list-item>
<list-item>
<label>-</label><p><bold>Evaluate Model:</bold> Testing the model against a validation dataset.</p></list-item>
</list></p></list-item>
<list-item>
<label>&#x2022;</label><p><bold>Decisions:</bold> Based on the outcome of the model evaluation, decide whether to retrain the model or proceed to deployment.</p></list-item>
<list-item>
<label>&#x2022;</label><p><bold>End:</bold> Conclude the process if the model meets the desired metrics.</p></list-item>
</list></p>
<p>In the modern age of digitalization, data and artificial intelligence (AI) are playing a central role in ushering innovation across industries, as shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. As the quantity of information keeps increasing due to digital devices and online platforms, along with rapid AI progression, our days are getting used to seeing eye-popping improvements.</p>
<p><list list-type="bullet">
<list-item>
<p><bold>Unpacking the Synergy of Data and AI:</bold></p></list-item>
</list></p>
<p><bold>The Data Deluge:</bold> In today&#x2019;s world, a vast amount of data is being produced by devices and connectivity. It comes with a volume that is difficult to manage; otherwise, AI can process and unearth insights at an unimaginable scale.</p>
<p><bold>The AI Revolution:</bold> What was once a concept is more of a reality, as AI now has become the feature that lifts our world. It covers doings/moves such as machine learning and natural language processing, which refers to the capability of machines to learn with data execution tasks requiring human-level cognition. AI and data are related to each other by:
<list list-type="bullet">
<list-item>
<p><bold>Data Feeds AI:</bold> High-quality data trains AI models to recognize patterns and make informed decisions.</p></list-item>
<list-item>
<p><bold>AI Enhances Data:</bold> AI processes data faster than humanly possible, organizing and extracting key insights.</p></list-item>
</list></p>
<p><xref ref-type="table" rid="table-2">Table 2</xref> illustrates how the use of natural data in AI development is influenced by trends towards open-source, the high computational and environmental costs of training models, the financial focus on generative AI technologies, and growing concerns about AI ethics and fairness. These elements highlight the ongoing evolution and challenges of using natural data for AI advancements.</p>


<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>The critical roles of data and AI</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_63551-fig-1.tif"/>
</fig>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>The key data points on the critical role of natural data in AI development for 2023</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col/>
<col align="center"/>
<col/>
</colgroup>
<thead>
<tr>
<th align="center">Category</th>
<th>Value</th>
<th align="center">Description</th>
<th>References</th>
</tr>
</thead>
<tbody>
<tr>
<td>Foundation models released</td>
<td>149</td>
<td>Number of foundation models released in 2023</td>
<td>[<xref ref-type="bibr" rid="ref-40">40</xref>]</td>
</tr>
<tr>
<td>Generative AI investment (Billion)</td>
<td>22.4%</td>
<td>Investment in generative AI in 2023 (in billions)</td>
<td>[<xref ref-type="bibr" rid="ref-41">41</xref>]</td>
</tr>
<tr>
<td>Environmental impact</td>
<td>41 yrs</td>
<td>Years of power for an average American home provided by the energy used to train GPT-3 (General Pre-trained Transformer-3)</td>
<td>[<xref ref-type="bibr" rid="ref-40">40</xref>]</td>
</tr>
<tr>
<td>AI ethics submissions increase</td>
<td>10</td>
<td>Tenfold increase in submissions to AI ethics conferences since 2018</td>
<td>[<xref ref-type="bibr" rid="ref-42">42</xref>]</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>The Growing Demand for Data in AI</title>
<p>Now that AI has become associated with everything from personalized product recommendations to medical diagnostics, there is a demand for enormous stores of diverse data. Large models (like, for instance, GPT-4) [<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-43">43</xref>] need the order of billions of words and images to operate at their best levels in terms of both accuracy and reliability. The burgeoning number of applications in different domains further expedites this demand. It is no secret that industries, be it finance, healthcare, or e-commerce, rely heavily on AI-powered analytics for better insights, to take control of trends, and to make strategic decisions [<xref ref-type="bibr" rid="ref-43">43</xref>].</p>
<p>On the one hand, they count on this level of demand. On the one hand, it hammers out rapid iteration as companies fight to make their AI shine brighter than others. On the other hand, it puts enormous pressure on natural data reserves, leading to a supply-demand imbalance that eventually slows down AI progress. Synthetic data, an alternative to genuine transaction data for testing and developing new analytics such as advanced Machine Learning (ML) [<xref ref-type="bibr" rid="ref-44">44</xref>] or AI, has, of late, become many organizations&#x2019; answer where the continued advancement, albeit very slowly brings with positive outcomes in part because it provides test cases where otherwise notary from real-world legitimate events were are even less use for plan these type algorithms. The absence of suitable sources for real data can reduce AI applications&#x2019; reliability and ethical integrity, making them less attractive to use in critical sectors.</p>
<p><xref ref-type="table" rid="table-3">Table 3</xref> summarizes the latest data on the growing demand for data in AI as of 2023, along with key trends and developments across various industries:</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>The growing demand for data in AI as of 2023</title>
</caption>
<table>
<colgroup>
<col/>
<col align="center"/>
<col/>
</colgroup>
<thead>
<tr>
<th>Sector</th>
<th align="center">Key data and trends</th>
<th>References</th>
</tr>
</thead>
<tbody>
<tr>
<td>Overall AI</td>
<td>72% of all new foundation models are developed by industry.</td>
<td>[<xref ref-type="bibr" rid="ref-43">43</xref>]</td>
</tr>
<tr>
<td>Machine learning</td>
<td>Logged models grew 54%, and registered models grew 411% since February 2022.</td>
<td>[<xref ref-type="bibr" rid="ref-45">45</xref>]</td>
</tr>
<tr>
<td>Generative AI</td>
<td>Investment surged to $25.2 billion in 2023.</td>
<td>[<xref ref-type="bibr" rid="ref-46">46</xref>]</td>
</tr>
<tr>
<td>Healthcare</td>
<td>Big Data enables personalized medicine and predictive healthcare.</td>
<td>[<xref ref-type="bibr" rid="ref-47">47</xref>]</td>
</tr>
<tr>
<td>Retail</td>
<td>Big Data drives insights into consumer behavior and market trends.</td>
<td>[<xref ref-type="bibr" rid="ref-48">48</xref>]</td>
</tr>
<tr>
<td>Finance</td>
<td>Big Data is used for fraud detection and risk assessment.</td>
<td>[<xref ref-type="bibr" rid="ref-47">47</xref>]</td>
</tr>
<tr>
<td>Manufacturing</td>
<td>Big Data optimizes supply chain operations and production efficiency.</td>
<td>[<xref ref-type="bibr" rid="ref-49">49</xref>]</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>These insights demonstrate the critical role of data in driving advancements in AI across various sectors, highlighting the immense growth in data demand and the evolving capabilities of AI technologies.</p>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Challenges of Data Scarcity for AI</title>
<p>AI has been training AI systems on ever-larger datasets, which is why we now have high-performing models such as ChatGPT or DALL-E 3. Simultaneously, research indicates that the growth of online data stocks is significantly slower than that of datasets utilized for AI training. In a paper published last year, researchers predicted we would run out of high-quality text data by 2026 if the current AI training trends continue. They also estimated that low-quality language data would be exhausted between 2030 and 2050 and low-quality image data between 2030 and 2060. According to the accounting and consulting firm PwC, artificial intelligence has the potential to contribute a staggering US$15.7 trillion (A$24.1 trillion) to the global economy by 2030. However, running short of usable data could slow down its development [<xref ref-type="bibr" rid="ref-50">50</xref>].</p>
<p>However, it is not just a by-product of its limitations as it has ethical and socio-economic implications. Without natural, high-quality data, AI models cannot remain perpetually privy to the dynamics of change and instead become captives in their synthetic ivory towers. For example, large language models require continuous access to diverse, up-to-date data to remain relevant and accurate. Otherwise, their outputs might become obsolete or inaccurate in time.</p>
<p>In addition, without ample natural data, AI systems may also become biased. Because these models are trained on incomplete or homogeneous datasets, they learn and replicate the biases built into that data, e.g., leading to discriminatory/unfair outcomes. One example might include an AI system used in hiring, which may lack diverse data, leading to certain demographics being preferred over others and ultimately favoring one group, thereby perpetuating societal imbalance. The consequences of these biases could be devastating as AI systems become central to different decision-making processes, from court judgments to healthcare recommendations [<xref ref-type="bibr" rid="ref-51">51</xref>].</p>
<p>Data scarcity is also very expensive from a monetary perspective. Many may struggle to innovate as companies who rely on AI for efficient operations, cost reduction, or enhanced customer experiences find it difficult when they are not sure about the origin of their data. This, in turn, could stall AI deployment in industry and, with it, the development of technologies that improve economic advancement. <xref ref-type="table" rid="table-4">Table 4</xref> lists key challenges of data scarcity in AI&#x2014;i.e., data exhaustion, risk of bias, regulatory constraints, and cost&#x2014;and their potential negative implications on model performance, fairness, and innovation.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Challenges posed by data scarcity in AI</title>
</caption>
<table>
<colgroup>
<col/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Challenge</th>
<th align="center">Description</th>
<th align="center">Potential impact</th>
</tr>
</thead>
<tbody>
<tr>
<td>Data exhaustion</td>
<td>High-quality language and image data are projected to be depleted within the next two decades.</td>
<td>Slower AI model development, reduced innovation, and performance stagnation.</td>
</tr>
<tr>
<td>Increasing bias risk</td>
<td>Limited data diversity can reinforce existing biases in AI models.</td>
<td>Biased decision-making in sectors like hiring, law, and healthcare.</td>
</tr>
<tr>
<td>Regulatory constraints</td>
<td>Strict data regulations limit access to high-quality, real-world data.</td>
<td>Reduced availability of natural data, affecting model accuracy and fairness.</td>
</tr>
<tr>
<td>Resource imbalance</td>
<td>Smaller companies may struggle more with limited data access than larger, resource-rich organizations.</td>
<td>Possible monopolization of AI advancements by well-funded companies.</td>
</tr>
<tr>
<td>Cost of data collection</td>
<td>High costs associated with sourcing, curating, and annotating high-quality data.</td>
<td>Increased operational expenses, slowing down AI project deployment.</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_4">
<label>3.4</label>
<title>Risks of Synthetic Data as a Solution</title>
<p>As natural data becomes less available, synthetic data has become a possible solution. Synthetic data is fake; it is artificially created to mirror real-world data, and the possibilities of being generated are endless. However, replacing natural data comes with some risks and imperfections. One is synthetic data because it cannot replicate the nuanced complexities that humans experience in interactions. AI models do great in a controlled environment but will struggle to generalize on chaotic, real-life situations. Many synthetic data also risk exacerbating pre-existing biases by default, as they tend to be generated based on patterns in existing data that often exclude specific user groups.</p>
<p>Furthermore, using synthetic data touches on its ethical dilemmas. This is a problem in machine learning as the AI may generate data that it assumes is correct by using another self-trained algorithm. Concepts like &#x201C;AI training AI&#x201D; are not accountable or transparent, so these models&#x2019; authenticity and accuracy cannot be easily proven. The biases contained within synthetic data will eventually dilute into the AI systems themselves, and in situations where accuracy is critical for a design, this could have adverse effects. <xref ref-type="table" rid="table-5">Table 5</xref> highlights how synthetic data can address data scarcity by increasing accessibility, affordability, control over bias, scalability, and ethics while also outlining associated risks like loss of authenticity, potential for bias, and ethical concerns.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Synthetic data as a solution to data scarcity</title>
</caption>
<table>
<colgroup>
<col/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Aspect</th>
<th align="center">Benefits</th>
<th align="center">Risks</th>
</tr>
</thead>
<tbody>
<tr>
<td>Accessibility</td>
<td>Synthetic data can be generated to fill data gaps.</td>
<td>Generated data may lack real-world authenticity, affecting model performance.</td>
</tr>
<tr>
<td>Cost-effectiveness</td>
<td>Reduces costs associated with data collection and annotation.</td>
<td>Quality concerns may require additional validation, adding costs.</td>
</tr>
<tr>
<td>Bias management</td>
<td>Synthetic data can be tailored to improve dataset diversity.</td>
<td>Potential for new biases introduced if synthetic data is derived from biased data sources.</td>
</tr>
<tr>
<td>Scalability</td>
<td>Easy to produce large volumes for training.</td>
<td>Excessive reliance on synthetic data risks a feedback loop in machine training, limiting diversity.</td>
</tr>
<tr>
<td>Ethical considerations</td>
<td>Avoids privacy concerns associated with real-world data.</td>
<td>Ethical ambiguity around training models without real-world grounding.</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_5">
<label>3.5</label>
<title>Ethical and Privacy Concerns in the Data-Driven AI Landscape</title>
<p>The lack of natural data is also hindered by privacy and ethical challenges. Human-generated data must be acquired and employed, often at the cost of sensitive private-output validation with significant end-user privacy implications. Instances like the Cambridge Analytica scandal have seeped into internet culture and public consciousness concerning data privacy, rising to quadruple scrutiny among nations. This has led to severe regulations and left little scope for vendors or site owners to snoop around without permission.</p>
<p>AI companies have to navigate data regulations that change dramatically between countries; they must also balance this with their deep wish for more user data and our demand for a private age. When the data is used to train AI systems that were not explicitly consented to or carried out with explicit permissions, ethical challenges start showing their face. Others are calling for more robust frameworks, especially concerning data collection, reiterating the importance of ethical approaches to AI regarding standards, as shown in <xref ref-type="table" rid="table-6">Table 6</xref>.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Ethical and privacy considerations in AI data usage</title>
</caption>
<table>
<colgroup>
<col/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Ethical concern</th>
<th align="center">Description</th>
<th align="center">Importance of AI development</th>
</tr>
</thead>
<tbody>
<tr>
<td>Data privacy</td>
<td>Ensuring data collection and usage comply with privacy laws and respect individual rights.</td>
<td>Builds public trust in AI systems and prevents legal repercussions.</td>
</tr>
<tr>
<td>Bias reduction</td>
<td>Avoiding biases that could lead to discrimination or unfair treatment.</td>
<td>Ensures AI applications serve all demographic groups equitably.</td>
</tr>
<tr>
<td>Transparency</td>
<td>Providing clarity on data sources and AI training methodologies.</td>
<td>Fosters trust and accountability in AI applications.</td>
</tr>
<tr>
<td>Accountability</td>
<td>Responsibility for ethical AI outcomes, especially in sensitive sectors like healthcare.</td>
<td>Minimizes risks of harm from biased or erroneous model outputs.</td>
</tr>
<tr>
<td>Public consent</td>
<td>Involving public opinion and securing consent for data usage.</td>
<td>Increases societal acceptance of AI and aligns AI development with societal values.</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Technological Innovations to Address Data Scarcity</title>
<p>Advances in technology offer promising solutions to the challenges of data scarcity in AI. This section outlines several innovative approaches:
<list list-type="bullet">
<list-item>
<p><bold>Advanced Machine Learning Algorithms:</bold> Techniques such as few-shot learning, transfer learning, and self-supervised learning enable AI models to learn effectively from limited data. For instance, few-shot learning has been successfully employed in natural language processing (NLP) tasks, achieving up to a 20% improvement in accuracy with only a handful of training examples.</p></list-item>
<list-item>
<p><bold>Data Compression and Augmentation Techniques:</bold> Data compression reduces dataset size while maintaining data integrity, enabling efficient storage and processing. Meanwhile, data augmentation artificially enhances dataset variety without requiring new data collection. For example, image augmentation techniques such as rotation, scaling, and color adjustment generate diverse training samples from a single image.</p></list-item>
<list-item>
<p><bold>Novel Data Acquisition Methods:</bold> Leveraging unconventional data sources like IoT devices and user-generated content can substantially increase data availability. Examples include using smartphones and wearable devices to collect real-time health data for personalized medicine applications.</p></list-item>
</list></p>
<p><xref ref-type="table" rid="table-7">Table 7</xref> illustrates various innovative technologies and their applications in AI, showcasing their impacts and providing key definitions, impact, and examples where applicable.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Examples of technological innovations in AI</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center">Technology</th>
<th align="center">Definition</th>
<th align="center">Impact</th>
<th align="center">Example</th>
</tr>
</thead>
<tbody>
<tr>
<td>Few-shot learning</td>
<td>Training AI models with only a few examples instead of thousands allows them to recognize patterns efficiently with minimal data [<xref ref-type="bibr" rid="ref-52">52</xref>].</td>
<td>A novel strategy that leverages (GANs, Generative Adversarial Network) and advanced optimization techniques</td>
<td>Bridges data scarcity with high-performing model adaptability and generalization</td>
</tr>
<tr>
<td>Data augmentation</td>
<td>Making small picture changes (flipping, rotating, changing brightness) to help AI learn better from limited data [<xref ref-type="bibr" rid="ref-53">53</xref>].</td>
<td>Enhanced training set diversity</td>
<td>Training autonomous driving systems with modified real-world images</td>
</tr>
<tr>
<td>IoT devices</td>
<td>Smartwatches or medical devices that track heart rate and send alerts if something is wrong [<xref ref-type="bibr" rid="ref-54">54</xref>].</td>
<td>Real-time health monitoring</td>
<td>Using wearable devices to monitor patient vitals in real-time</td>
</tr>
<tr>
<td>Synthetic data generation</td>
<td>Creating fake but realistic data so AI can learn without using real people&#x2019;s sensitive information [<xref ref-type="bibr" rid="ref-55">55</xref>].</td>
<td>Training without exposing personal data</td>
<td>Creating synthetic financial profiles for fraud detection testing</td>
</tr>
<tr>
<td>Self-supervised learning</td>
<td>AI teaches itself using raw data, like a person learning from experience instead of reading a manual [<xref ref-type="bibr" rid="ref-56">56</xref>].</td>
<td>Reduces the need for labeled datasets</td>
<td>Content moderation on social media platforms without predefined labels</td>
</tr>
<tr>
<td>Transfer learning</td>
<td>Taking what an AI learned in one area and using it elsewhere, like teaching a soccer player how to play basketball [<xref ref-type="bibr" rid="ref-57">57</xref>].</td>
<td>Adapting models to new areas without retraining</td>
<td>Applying financial market predictions to healthcare trends</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5">
<label>5</label>
<title>Strategic Solutions to the Data Crisis</title>
<p>To address data scarcity, AI developers and companies can adopt strategic approaches that optimize data efficiency, expand data availability, and maintain ethical standards.</p>
<p><list list-type="bullet">
<list-item><p><bold>Optimizing Data Efficiency:</bold> When data is limited, models should be optimized to extract maximum insight from the available information. Techniques like data augmentation, transfer learning, and reinforcement learning (See Appendix A, <xref ref-type="table" rid="table-15">Table A1</xref>) help reduce the dependency on large datasets while maintaining model accuracy.</p>
</list-item>
<list-item><p><bold>Collaborative Data Sharing:</bold> A competitive way to democratize data access is via company partnerships. In pooling resources, groups will realize a more balanced and diverse training dataset in the models they develop, leading to less bias in general. A quintessential open-source initiative is Uber, which has made available its self-driving dataset to developers around the globe, thus facilitating a co-working environment of innovation and data dissemination subject to regulated ethical narratives.</p></list-item>
<list-item><p><bold>Integrating Synthetic and Natural Data:</bold> Although synthetic data is not sufficient by itself, a hybrid of natural and artificial may overcome the weaknesses inherent in both types. If natural data are used as the base, synthetic data can fill any gaps, especially in niche or underrepresented areas. When companies take this more balanced approach, they can respond to the demands for data without degrading selection accuracy or equity.</p></list-item>
<list-item><p><bold>Exploring Alternative Data Sources:</bold> Companies are already looking to newer organic data channels, such as customer feedback, offline footprints, and proprietary datasets. By digitizing and analyzing their resources, companies can create useful training data without stepping over the privacy boundary.</p></list-item>
<list-item><p><xref ref-type="table" rid="table-8">Table 8</xref> lists the primary approaches to AI data scarcity by enhancing training efficiency, data sharing, hybrid data usage, alternative sources, and regulation support for enabling improved availability, fairness, and ethics compliance.</p>
</list-item></list></p>
<table-wrap id="table-8">
<label>Table 8</label>
<caption>
<title>Strategic solutions to address data scarcity in AI</title>
</caption>
<table>
<colgroup>
<col/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Solution</th>
<th align="center">Description</th>
<th align="center">Benefits</th>
</tr>
</thead>
<tbody>
<tr>
<td>Data efficiency techniques</td>
<td>Focus on enhancing model training through data augmentation, transfer learning, and reinforcement learning.</td>
<td>Reduces reliance on extensive datasets, enabling effective learning with limited data resources [<xref ref-type="bibr" rid="ref-11">11</xref>].</td>
</tr>
<tr>
<td>Collaborative data sharing</td>
<td>Companies partner to share anonymized datasets, expanding diverse data pools.</td>
<td>Enhances data availability, mitigates bias risks, and fosters AI innovation.</td>
</tr>
<tr>
<td>Hybrid data use</td>
<td>Combines real-world and synthetic data to expand AI training capabilities.</td>
<td>Maintains data authenticity, improves model adaptability, and enhances fairness.</td>
</tr>
<tr>
<td>Exploring new data sources</td>
<td>Alternative sources like customer feedback, sensor data, and offline repositories are used.</td>
<td>Expands available data diversity, improving real-world model applications.</td>
</tr>
<tr>
<td>Policy and regulatory support</td>
<td>Establishing responsible data-sharing frameworks in partnership with governments and policymakers.</td>
<td>Ensures ethical AI deployment while maintaining compliance with legal standards.</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s5_1">
<label>5.1</label>
<title>The Road Ahead: Building Sustainable AI with Limited Data</title>
<p>Given these challenges, the AI industry should also work to maintain energy&#x2014;and data-efficient systems. With so many seemingly successful algorithms, developing data-efficient algorithms that extract as much value from each data point will be vital to progress once natural resources are scarce. Policymakers will also have a role in making that happen, such as through regulations to facilitate responsible data use and ensure businesses maintain high ethical standards.</p>
<p>Much like oil, data is a resource that could be the key to unlocking even more potential human progress&#x2014;but as we move forward into an uncertain future where unlimited data may no longer be possible or generally allowed, people can turn their attention from the acquisition of massive amounts of low-quality convenience samples and direct our efforts on authoring and cataloging diversity in &#x201C;good&#x201D; taste if it were. Ultimately, it comes down to combining lawful data quality with trustworthy AI systems and responsible development to ensure top-notch security features. If the industry keeps these values in mind, it should be able to innovate and grow responsibly, with AI as a positive force for good.</p>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>Strategic Approaches and Partnerships</title>
<p>Collaboration and strategic partnerships between organizations can be pivotal in overcoming data scarcity:</p>
<p><bold>Data Sharing Initiatives:</bold> Companies like Google and IBM have pioneered sharing large datasets. For example, Google&#x2019;s release of the ImageNet database has revolutionized computer vision research.</p>
<p><bold>Public-Private Partnerships:</bold> Collaborations between governments and tech companies can facilitate the development of AI technologies with shared datasets. An example is the partnership between the U.S. Department of Health and AI startups to analyze health data securely.</p>
<p><xref ref-type="table" rid="table-9">Table 9</xref> highlights significant partnerships between organizations aimed at leveraging collaboration to address challenges, including data scarcity.</p>
<table-wrap id="table-9">
<label>Table 9</label>
<caption>
<title>Key strategic partnerships in AI</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col/>
</colgroup>
<thead>
<tr>
<th align="center">Partners</th>
<th align="center">Initiative</th>
<th align="center">Purpose</th>
<th align="center">Contribution</th>
<th>References</th>
</tr>
</thead>
<tbody>
<tr>
<td>Google and academic institutions</td>
<td>ImageNet database</td>
<td>Boost research in computer vision</td>
<td>Pioneered advancements in image recognition</td>
<td>[<xref ref-type="bibr" rid="ref-58">58</xref>,<xref ref-type="bibr" rid="ref-59">59</xref>]</td>
</tr>
<tr>
<td>U.S. department of health and startups</td>
<td>Health data analysis</td>
<td>Enhance predictive capabilities in healthcare</td>
<td>Improved diagnostics and treatment plans</td>
<td>[<xref ref-type="bibr" rid="ref-60">60</xref>]</td>
</tr>
<tr>
<td>IBM and weather channel</td>
<td>Weather data collaboration</td>
<td>Enhance meteorological predictions</td>
<td>Refined forecasting models in meteorology</td>
<td>[<xref ref-type="bibr" rid="ref-61">61</xref>]</td>
</tr>
<tr>
<td>Facebook and universities</td>
<td>Social data analysis</td>
<td>Study behavioral patterns</td>
<td>Provided insights into user interaction dynamics</td>
<td>[<xref ref-type="bibr" rid="ref-62">62</xref>]</td>
</tr>
<tr>
<td>Automotive companies and tech firms</td>
<td>Autonomous vehicle data sharing</td>
<td>Accelerate autonomous vehicle technology</td>
<td>Enhanced safety and navigation systems</td>
<td>[<xref ref-type="bibr" rid="ref-63">63</xref>]</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5_3">
<label>5.3</label>
<title>Policy and Regulation Considerations</title>
<p>Effective policy and regulation are crucial to ensuring data is used responsibly:</p>
<p><bold>Data Privacy Regulations:</bold> The implementation of GDPR in Europe [<xref ref-type="bibr" rid="ref-64">64</xref>] and the CCPA in California has set new benchmarks for data privacy, influencing AI data handling practices worldwide.</p>
<p><bold>Incentives for Data Sharing:</bold> Governments can incentivize data sharing with tax breaks and grants, as seen in the EU&#x2019;s Data Governance Act, which aims to foster data sharing while ensuring privacy.</p>
<p><xref ref-type="table" rid="table-10">Table 10</xref> explores how specific data privacy regulations affect AI development and application across various regions.</p>
<table-wrap id="table-10">
<label>Table 10</label>
<caption>
<title>Impact of data privacy regulations on AI</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col/>
<col align="center"/>
<col align="center"/>
<col/>
</colgroup>
<thead>
<tr>
<th align="center">Regulation</th>
<th>Region</th>
<th align="center">Impact</th>
<th align="center">Details</th>
<th>References</th>
</tr>
</thead>
<tbody>
<tr>
<td>GDPR (General Data Protection Regulation)</td>
<td>Europe</td>
<td>Tightened data protection</td>
<td>Requires stringent consent for data use in AI</td>
<td>[<xref ref-type="bibr" rid="ref-64">64</xref>]</td>
</tr>
<tr>
<td>CCPA (California Consumer Privacy Act)</td>
<td>California</td>
<td>Strengthened consumer data rights</td>
<td>Enables consumers to opt out of data selling</td>
<td>[<xref ref-type="bibr" rid="ref-65">65</xref>]</td>
</tr>
<tr>
<td>LGPD (General Data Protection Law)</td>
<td>Brazil</td>
<td>Enhanced privacy protections similar to GDPR</td>
<td>Mandates transparent data usage policies</td>
<td>[<xref ref-type="bibr" rid="ref-40">40</xref>]</td>
</tr>
<tr>
<td>PIPL (Personal Information Protection Law)</td>
<td>China</td>
<td>Strict data management and export controls</td>
<td>Imposes controls on cross-border data transfers</td>
<td>[<xref ref-type="bibr" rid="ref-66">66</xref>]</td>
</tr>
<tr>
<td>HIPAA (Health Insurance Portability and Accountability Act)</td>
<td>USA</td>
<td>Privacy Rule permits important uses of information</td>
<td>These regulations impose strict requirements on data handling and user consent, thereby influencing how AI systems are developed and implemented</td>
<td>[<xref ref-type="bibr" rid="ref-67">67</xref>]</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5_4">
<label>5.4</label>
<title>Case Study: Addressing Data Scarcity in Fraud Detection</title>
<p>This is a common challenge where rare but critical events, such as fraudulent transactions, need to be identified, like fraudulent transactions. For a hands-on understanding of synthetic data generation, we used the Credit Card Fraud Detection Dataset [<xref ref-type="bibr" rid="ref-68">68</xref>], a benchmark dataset in which only 0.17% of transactions are marked as fraudulent. This case study is the application of synthetic oversampling techniques, such as SMOTE (Synthetic Minority Oversampling Technique), to mitigate data imbalance (See Appendix A, <xref ref-type="table" rid="table-15">Table A1</xref>) and enhance AI model accuracy.</p>

<p><bold>Dataset Overview:</bold> The dataset [<xref ref-type="bibr" rid="ref-68">68</xref>] consists of anonymized transaction data with significant class imbalance, as shown in <xref ref-type="disp-formula" rid="eqn-1">Eqs. (1)</xref> and <xref ref-type="disp-formula" rid="eqn-2">(2)</xref>:
<list list-type="bullet">
<list-item>
<p><bold>Normal Transactions (Support: 85,295):</bold></p></list-item>
</list>
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mrow><mml:mtext>Percentage&#xA0;</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>85295</mml:mn><mml:mn>85443</mml:mn></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mn>100</mml:mn><mml:mo>=</mml:mo><mml:mn>99.83</mml:mn><mml:mrow><mml:mtext>%&#xA0;</mml:mtext></mml:mrow></mml:math></disp-formula>
<list list-type="bullet">
<list-item>
<p><bold>Fraudulent Transactions (Support: 148):</bold></p></list-item>
</list>
<disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mrow><mml:mtext>Percentage&#xA0;</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mn>148</mml:mn><mml:mn>85443</mml:mn></mml:mfrac><mml:mo>&#x00D7;</mml:mo><mml:mn>100</mml:mn><mml:mo>=</mml:mo><mml:mn>0.17</mml:mn><mml:mrow><mml:mtext>%&#xA0;</mml:mtext></mml:mrow></mml:math></disp-formula></p>
<p>This imbalance is a clear representation of practical data scarcity, where the availability of examples for critical classifications (fraudulent cases) is minimal.</p>
<sec id="s5_4_1">
<label>5.4.1</label>
<title>Experimental Results</title>
<p>We used SMOTE to create synthetic examples for the minority class. A Random Forest Classifier was trained on both the original and the augmented dataset, and its quality was assessed. <xref ref-type="table" rid="table-11">Table 11</xref> shows the main performance for the augmented dataset:</p>
<p><list list-type="bullet">
<list-item>
<p>0 represents the Normal (Non-Fraudulent Transactions) class.</p></list-item>
<list-item>
<p>1 represents the Fraudulent Transactions class.</p></list-item>
<list-item>
<p>Support refers to the number of true instances for each class in the dataset. It represents the count of actual samples in the test set for each class (0 for non-fraudulent transactions and 1 for fraudulent transactions).</p></list-item>
</list></p>

<table-wrap id="table-11">
<label>Table 11</label>
<caption>
<title>Model performance metrics after applying SMOTE</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th></th>
<th>Precision %</th>
<th>Recall %</th>
<th>F1-score %</th>
<th>Support</th>
</tr>
</thead>
<tbody>
<tr>
<td>0</td>
<td>100.0</td>
<td>100.0</td>
<td>100.0</td>
<td>85,295</td>
</tr>
<tr>
<td>1</td>
<td>89.2</td>
<td>78.4</td>
<td>83.5</td>
<td>148</td>
</tr>
<tr>
<td><bold>Accuracy</bold></td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>99.9</td>
<td>&#x2013;</td>
</tr>
<tr>
<td><bold>Macro avg</bold></td>
<td>94.6</td>
<td>89.2</td>
<td>91.7</td>
<td>85,443</td>
</tr>
<tr>
<td><bold>Weighted avg</bold></td>
<td>99.9</td>
<td>99.9</td>
<td>99.9</td>
<td>85,443</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s5_4_2">
<label>5.4.2</label>
<title>Metric Definitions</title>
<p><list list-type="simple">
<list-item>
<label>A.</label><p><bold>Precision:</bold></p></list-item>
</list></p>
<p>Precision measures the number of correctly identified instances of fraud divided by the total number of cases labeled as fraudulent. This particular measure focuses on how often the model is correct when it says it is positive, as shown in <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref>.
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mtext>TP</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>TP</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mtext>FP</mml:mtext></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>For example, in the Credit Card Fraud Detection Dataset, 0.17% of transactions are fraudulent. Second, precision is essential so that flagged transactions are indeed fraudulent and the false alarm rate (fraudulent transactions classified as non-fraudulent transactions) is minimized.</p>
<p><bold>Application to Experimental Results:</bold></p>
<p>The model achieved a precision of 89.00% for fraudulent transactions. That means 89% of transactions that the model flagged as fraudulent were fraudulent, showing it to be effective in reducing the false positive rate. This level of precision can be beneficial in fraud detection systems, where false positives can waste lots of time and resources.
<list list-type="simple">
<list-item>
<label>B.</label><p><bold>Recall:</bold></p></list-item>
</list></p>
<p>Recall the number of actual fraudulent transactions detected by the model. This metric is critical for grasping how well the model can capture low-probability events, as shown in <xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref>.
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mtext>TP</mml:mtext></mml:mrow><mml:mrow><mml:mrow><mml:mtext>TP</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mtext>FN</mml:mtext></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>In fraud detection, recall is essential because the cost of missing fraudulent transactions (false negatives) could lead to financial losses and undermine trust in the system.</p>
<p>It resulted in a 78.00% recall for fraudulent transactions. This means the model successfully detected 78% of all fraudulent cases. While this shows improvement, some fraudulent cases were still missed, highlighting the need for further optimization.
<list list-type="simple">
<list-item>
<label>C.</label><p><bold>F1-Score:</bold></p></list-item>
</list></p>
<p>The F1-Score, designed as a harmonic mean of precision and recall, offers a means to manage the trade-off between these measures. This is especially useful in imbalanced datasets where both of these metrics are important; the formula for the F1-score is typically given as <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref>, is:
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mi>F</mml:mi><mml:mn>1</mml:mn><mml:mrow><mml:mtext>-</mml:mtext></mml:mrow><mml:mrow><mml:mtext>score</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x2217;</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mtext>&#xA0;precision&#xA0;</mml:mtext></mml:mrow><mml:mo>&#x2217;</mml:mo><mml:mrow><mml:mtext>&#xA0;recall&#xA0;</mml:mtext></mml:mrow></mml:mrow><mml:mrow><mml:mrow><mml:mtext>&#xA0;precision&#xA0;</mml:mtext></mml:mrow><mml:mo>+</mml:mo><mml:mrow><mml:mtext>&#xA0;recall&#xA0;</mml:mtext></mml:mrow></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>A common metric used for diagnosis datasets with high-class imbalance (for example, fraud detection) is the F1-score because it provides a comprehensive evaluation of the model performance without weighing precision or recall too heavily.</p>
<p>The F1-Score of fraudulent transactions is 83.00%, indicating that the trade-off between precision and recall is balanced. This also shows that our model can detect fraud but does not raise many false positives.</p>
<p>These results confirm the efficiency of SMOTE, which improves the model&#x2019;s ability to recognize rare events while preserving high accuracy for all other cases.
<list list-type="simple">
<list-item>
<label>D.</label><p><bold>Accuracy</bold></p></list-item>
</list></p>
<p>Accuracy measures the proportion of all transactions (both normal and fraudulent) that the model correctly classified; the formula for accuracy is given in <xref ref-type="disp-formula" rid="eqn-6">Eq. (6)</xref> below.
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mrow><mml:mtext>Accuracy&#xA0;</mml:mtext></mml:mrow><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula></p>
<p>Because of the imbalance in the dataset, the accuracy can be misleading on its own, as it depends too much on the majority class.</p>
<p>The model reached 99.94% accuracy, showing that almost all transactions were classified correctly. However, such a high accuracy only represents the model&#x2019;s performance on the majority class (normal transactions) and is not indicative of its capability to detect fraudulent transactions.</p>
</sec>
</sec>
<sec id="s5_5">
<label>5.5</label>
<title>Implications of the Case Study</title>
<p><italic>Addressing Practical Data Scarcity</italic></p>
<p>This case study shows that data scarcity can be alleviated substantially through synthetic data generation techniques like SMOTE. By generating more samples of the minority class, the model was able to:
<list list-type="order">
<list-item>
<p><bold>Reduce Bias (Recall&#x2014;Fraud Transactions):</bold> Overcome the inherent bias of the model towards the majority class.</p></list-item>
<list-item>
<p><bold>Enhance Generalization</bold>: The model achieved high precision and improved recall for rare events, enhancing its ability to generalize across datasets.</p></list-item>
</list></p>
<p><bold>Applicability across Domains</bold></p>
<p>The success of SMOTE, in this case, has broader implications:
<list list-type="bullet">
<list-item>
<p>In <bold>healthcare,</bold> synthetic data can help detect rare diseases with limited availability.</p></list-item>
<list-item>
<p>In <bold>manufacturing</bold>, synthetic data can be used to enhance anomaly detection when a few failures occur in a production line.</p></list-item>
<list-item>
<p>In <bold>finance</bold>, fraud detection systems are improved through oversampling imbalanced datasets.</p></list-item>
</list></p>
</sec>
<sec id="s5_6">
<label>5.6</label>
<title>Case Study: Evaluating AI Performance in Medical Diagnosis</title>
<p>Though theoretically, the development of AI appears promising, a number of real-world applications regarding health care have shown the effectiveness of these technologies in the real world. Applying a RandomForestClassifier to the Breast Cancer Wisconsin dataset represents another case study demonstrating how AI approaches the dual challenge of data sparsity and decision ethics in diagnostics. The model yielded high precision and recall with a balanced dataset of benign and malignant cases. This could be helpful in the early detection of a disease and its accurate diagnosis, which is one of the most important aspects of any patient care and treatment. Applications of this type constitute milestones in transforming healthcare by equipping medical professionals with reliable tools for decision-making. This case study describes an application of RandomForestClassifier using the Breast Cancer Wisconsin dataset [<xref ref-type="bibr" rid="ref-69">69</xref>] to estimate a diagnosis as benign or malignant. The dataset [<xref ref-type="bibr" rid="ref-69">69</xref>] is a multivariate one, consisting of 569 samples with 30 features each, which are further divided into a training set of 70% and a test set of 30%. The model RandomForest was trained and afterward tested on the test set in order to calculate the precision, recall, f1-score, and support of both diagnostic classes. Results, represented in <xref ref-type="table" rid="table-12">Table 12</xref>, suggest very high accuracy regarding the diagnostic capability of the model: 98.33% precision and 93.65% recall for malignant cases and precision of 96.39% with a recall of 99.07% for benign cases. These metrics underline the robustness of the model, with F1-scores of 95.93% for malignant and 97.72% for benign diagnoses, therefore postulating that this model is accurate and reliable in distinguishing the two conditions quite well, as seen from <xref ref-type="table" rid="table-12">Table 12</xref>. The support values are 63 for malignant and 108 for benign, showing the number of instances evaluated in each class and the balance in the dataset used.</p>
<table-wrap id="table-12">
<label>Table 12</label>
<caption>
<title>Classifier performance metrics</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Diagnosis</th>
<th>Precision %</th>
<th>Recall %</th>
<th>F1-score %</th>
<th>Support</th>
</tr>
</thead>
<tbody>
<tr>
<td>Malignant</td>
<td>98.33</td>
<td>93.65</td>
<td>95.93</td>
<td>63</td>
</tr>
<tr>
<td>Benign</td>
<td>96.39</td>
<td>99.07</td>
<td>97.72</td>
<td>108</td>
</tr>
<tr>
<td>Overall</td>
<td>Accuracy: 97.07</td>
<td>Macro avg: 97.36</td>
<td>Weighted avg: 97.11</td>
<td></td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The above case study shows the vast potential of machine learning to improve diagnostic accuracy in the medical field. Therefore, it justifies the development of AI tools that can support clinical decision-making and hopefully improve patient outcomes.</p>
<p><italic>Ethical Considerations</italic></p>
<p>Although synthetic data resolves most scarcity-related issues, it introduces its own set of potential risks:
<list list-type="bullet">
<list-item>
<p><bold>Bias Propagation:</bold> Synthetic samples can also inherit biases from the original data.</p></list-item>
<list-item>
<p><bold>Reduced Realism:</bold> Generative data may not capture the intricacies of actual environmental interactions, leading to diminished fidelity of predictions when used for real-world applications.</p></list-item>
</list></p>
<p>This raises the need for synthetic data to be utilized along with a strong validation mechanism and an ethical framework, as discussed in <xref ref-type="sec" rid="s3_4">Section 3.4</xref>.</p>
</sec>
<sec id="s5_7">
<label>5.7</label>
<title>Proposed Solutions and Their Applications</title>
<p>The pursuit of Artificial General Intelligence (AGI) and agentic AIs entering many industries as a new type of workforce are some of the hottest topics in AI-related spaces in the mid-2020s. Such well-known AI personas as Altman, Huang, Sutskever et al. claim that AGI has already been achieved and/or they know exactly how to do it [<xref ref-type="bibr" rid="ref-70">70</xref>&#x2013;<xref ref-type="bibr" rid="ref-74">74</xref>]. While this is currently a top-secret, the possible solutions are those presented in <xref ref-type="fig" rid="fig-2">Fig. 2</xref> or a combination of them, which is even more likely. Top Large Language Models (LLMs) confirmed the possible solutions through the short survey to gather their input.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Top possible solutions to the data scarcity problem</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_63551-fig-2.tif"/>
</fig>
<p>As can be seen from <xref ref-type="fig" rid="fig-2">Fig. 2</xref>, the top option with the highest weight is Synthetic Data Generation. The authors agree with this direction [<xref ref-type="bibr" rid="ref-74">74</xref>] and believe that creating high-quality synthetic data is impossible without AI-human collaboration or, in other words, a Human-in-the-loop.</p>

<sec id="s5_7_1">
<label>5.7.1</label>
<title>Synthetic Data Creation with Human-in-the-Loop</title>
<p>While some models can be specially trained to generate synthetic data, others can verify them and then be used by top LLMs; without people involved, this process will make no sense. We will see something like software development without a client. <xref ref-type="fig" rid="fig-3">Fig. 3</xref> represents a Human-in-the-Loop Synthetic Data Workflow.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Human-in-the-loop synthetic data workflow</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_63551-fig-3.tif"/>
</fig>
<p>As apparent in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>, humanity will collaborate with AI as top content creators and algorithm developers to achieve the best possible outcome. This highly will likely include quantum computing and other ways of compressing models and learning more and better from fewer data; all approaches to distributed machine learning, such as federated learning (See Appendix A, <xref ref-type="table" rid="table-15">Table A1</xref>), are obviously on a plate.</p>

<p>The integration of AI systems should be embedded, at the development stage itself, with privacy design principles for better practicality. This could include techniques such as differential privacy (See Appendix A, <xref ref-type="table" rid="table-15">Table A1</xref>) in data processing, for which no trace of individual data points can be traced to owners. Similarly, clear policies on data governance, including consent management, data access rights, and transparency regarding data usage, would help foster trust and enforce ethics. Another benefit would be setting up an Ethics Review Committee for AI projects to ensure that the highest ethics are considered while overseeing complex issues.</p>

</sec>
<sec id="s5_7_2">
<label>5.7.2</label>
<title>Smaller Language Models</title>
<p>One of the most promising solutions for optimizing AI efficiency is Smaller Language Models (SLMs), which have gained significant attention due to their ability to operate effectively in resource-limited environments. Unlike Large Language Models (LLMs), which require extensive computational power, SLMs provide a more practical, cost-effective, and scalable approach to AI implementation. This makes them particularly suitable for deployment in low-resource settings, such as mobile devices, edge computing, and cloud-based platforms. Recently released by Microsoft, Phi-4 stands out as one of the leading SLMs demonstrating the potential of efficient AI models. It is a 40-layer, transformer-based model with a hidden size of 5120, supporting over 100k token embeddings and containing 14.1 billion trainable parameters. Unlike traditional LLMs, Phi-4 achieves high performance while significantly reducing resource consumption, making it a viable alternative for real-world applications. <xref ref-type="fig" rid="fig-4">Fig. 4</xref> illustrates Phi-4 running on Google Colab, where the model was executed on the PRO&#x002B; tier of the notebook (on an A100 runtime). The entire process was completed within 751.6 s, demonstrating high efficiency.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Phi-4 SLM running on Google Collab</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_63551-fig-4.tif"/>
</fig>
<p>The accessibility of SLMs like Phi-4 is a major advantage, allowing researchers and developers to run highly capable AI models on cloud platforms at a lower cost. For example, the Google Colab PRO&#x002B; subscription provides access to high-performance GPUs (Graphics Processing Units) for just USD 50 monthly, making SLMs an affordable research, development, and small-scale production solution. Phi-4 has demonstrated strong performance in practical applications across multiple NLP tasks, including text classification, summarization, question-answering (Q&#x0026;A), and Retrieval-Augmented Generation (RAG). Notably, Phi-4 can suggest solutions to data scarcity challenges, utilizing techniques such as Synthetic Data Generation, Few-Shot and Transfer Learning, Data Augmentation, Self-Supervised Learning, Federated Learning, Zero-Shot Learning, and Privacy-Preserving AI Techniques (See Appendix A, <xref ref-type="table" rid="table-15">Table A1</xref>). Additionally, SLMs have proven highly effective in RAG-based applications, significantly enhancing document summarization and knowledge retrieval. Researchers have successfully integrated Phi-4 into RAG, GraphRAG, and LazyGraphRAG applications [<xref ref-type="bibr" rid="ref-19">19</xref>,<xref ref-type="bibr" rid="ref-45">45</xref>,<xref ref-type="bibr" rid="ref-75">75</xref>,<xref ref-type="bibr" rid="ref-76">76</xref>], enabling AI models to interact with structured knowledge bases. These applications allow users to upload documents (e.g., PDFs, TXT files) and query them interactively, making AI more versatile and practical for real-world data processing.</p>

<p><xref ref-type="fig" rid="fig-5">Fig. 5</xref> presents a visual representation of the GraphRAG approach [<xref ref-type="bibr" rid="ref-19">19</xref>], illustrating how it structures relationships within a document to enhance AI&#x2019;s ability to retrieve, comprehend, and summarize structured knowledge. The diagram showcases interconnected nodes, representing key concepts and their relationships, which help improve contextual understanding and efficient retrieval of information.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>A visual representation of the GraphRAG approach, illustrating how a document&#x2019;s content is structured into a graph-based format. This diagram highlights key concepts and their relationships, improving AI-driven retrieval, comprehension, and summarization of structured knowledge</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_63551-fig-5.tif"/>
</fig>
<p><xref ref-type="table" rid="table-13">Table 13</xref> summarizes all three RAG-based approaches, comparing their core concepts, storage utilization, retrieval mechanisms, and efficiency trade-offs.</p>
<table-wrap id="table-13">
<label>Table 13</label>
<caption>
<title>RAG methods comparison</title>
</caption>
<table>
<colgroup>
<col/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Feature</th>
<th align="center">RAG</th>
<th align="center">Full graph RAG</th>
<th align="center">Lazy graph RAG</th>
</tr>
</thead>
<tbody>
<tr>
<td>Concept</td>
<td>It uses a retriever-generator model to fetch and process text chunks.</td>
<td>Organizes information in a graph structure, improving relational understanding.</td>
<td>&#x201C;Lazily&#x201D; explores or expands the graph at query time, retrieving only the necessary subgraph.</td>
</tr>
<tr>
<td>Storage</td>
<td>Uses dense vector indexes for direct chunk retrieval.</td>
<td>Stores entities, documents, and relationships as graph nodes &#x0026; edges.</td>
<td>Minimizes memory footprint by loading only necessary segments.</td>
</tr>
<tr>
<td>Retrieval</td>
<td>Searches for top-k text chunks and generates an answer.</td>
<td>Traverses graph relationships to extract relevant context.</td>
<td>Selects relevant nodes dynamically, reducing unnecessary retrieval overhead.</td>
</tr>
<tr>
<td>Efficiency</td>
<td>Fast, but lacks deep contextual relationships.</td>
<td>It is more resource-intensive, as graph traversal requires extra computations.</td>
<td>Optimized for efficiency, balancing context depth and computational cost.</td>
</tr>
<tr>
<td>Context quality</td>
<td>Depending on the chunk ranking, it may lose relational meaning.</td>
<td>Captures document relationships, improving contextual understanding.</td>
<td>Retains graph-based advantages while reducing computational load.</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The emergence of Smaller Language Models (SLMs) marks a significant shift in AI development, offering a balance between performance, efficiency, and accessibility. Models like Phi-4 exemplify how resource-friendly AI can power advanced applications, such as Retrieval-Augmented Generation (RAG), document summarization, and interactive knowledge retrieval. While challenges remain, ongoing research in knowledge adaptation, quantization, and efficient training methodologies will further enhance the capabilities and widespread adoption of SLMs across various domains. As AI evolves, SLMs will play an increasingly critical role in democratizing machine learning access, making AI applications more scalable, efficient, and environmentally sustainable.</p>
</sec>
<sec id="s5_7_3">
<label>5.7.3</label>
<title>Quantization and Pruning</title>
<p>Quantization and pruning are key optimization techniques that significantly reduce machine learning models&#x2019; size and computational cost, making AI more accessible and sustainable. These methods are beneficial for Small Language Models (SLMs) like Phi-4, which aim to achieve high efficiency without compromising performance. Quantization converts high-precision numerical values (32-bit floating points) into lower-precision formats (8-bit or 4-bit), reducing memory usage and power consumption. Meanwhile, pruning removes unnecessary parameters (weights, neurons, or channels) from a model, reducing complexity and computational load while maintaining accuracy.</p>
<p>Both methods align with the Green AI approach [<xref ref-type="bibr" rid="ref-18">18</xref>], ensuring lower energy consumption, faster inference speeds, and minimal hardware requirements while keeping AI models scalable for real-world applications. <xref ref-type="table" rid="table-14">Table 14</xref> presents the latest quantization and pruning approaches applied to Phi-4 (as previously tested on Phi-1.5 and Phi-2 models) and detailed explanations of their key features.</p>
<table-wrap id="table-14">
<label>Table 14</label>
<caption>
<title>Quantization and pruning methods comparison</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center">Approach</th>
<th align="center">Library/Tool</th>
<th align="center">Precision/Sparsity</th>
<th align="center">Key features</th>
</tr>
</thead>
<tbody>
<tr>
<td>4-bit Quant (NF4)</td>
<td>BitsAndBytes (bnb) &#x002B; Hugging Face Transformers</td>
<td>4-bit Weights (NormalFloat4)</td>
<td>Maximizes memory savings while maintaining good accuracy retention. Used in LLMs &#x0026; SLMs for extreme efficiency.</td>
</tr>
<tr>
<td>8-bit Quant (LLM.int8())</td>
<td>BitsAndBytes &#x002B; Accelerate/HF Transformers</td>
<td>8-bit Matrix Multiplications</td>
<td>Reduces GPU memory usage significantly with a minor accuracy drop vs. FP16. Best for general AI applications.</td>
</tr>
<tr>
<td>Dynamic Quant (8-bit/16-bit)</td>
<td>Native PyTorch Quantization</td>
<td>8-bit or 16-bit (activations/weights)</td>
<td>Applies on-the-fly quantization, requiring minimal code changes. Accuracy may vary depending on the model&#x2019;s sensitivity. Suitable for low-power devices.</td>
</tr>
<tr>
<td>Quantization-Aware Training (QAT)</td>
<td>PyTorch or TF Model Optimization</td>
<td>8-bit or 16-bit (weights &#x002B; activations)</td>
<td>Simulates quantization during forward/backward, yields higher accuracy, more complex setup.</td>
</tr>
<tr>
<td>Pruning</td>
<td>PyTorch Pruning Utilities</td>
<td>Any model/layer (set weights to 0)</td>
<td>Simulates quantization effects during training, improving accuracy in low-precision models. Used in production AI applications.</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The significance of these techniques lies in their ability to reduce model size and computational overhead, leading to faster real-time responses. Additionally, they play a crucial role in minimizing energy consumption, contributing to the sustainability goals outlined in Green AI research. As AI models continue to evolve, smaller, more efficient architectures will lower the cost of AI deployment, making advanced machine learning more widely accessible; even though quantization allows AI models to store and process numerical data efficiently while pruning eliminates redundant connections, further research is required to evaluate their energy efficiency across different AI architectures. This is especially relevant for Large Language Models (LLMs) and SLMs, where balancing performance, computational efficiency, and accuracy remains an ongoing challenge.</p>
</sec>
</sec>
<sec id="s5_8">
<label>5.8</label>
<title>The Future of AI with Synthetic Data</title>
<p>This study underscores the significance of synthetic data generation as a pivotal technological innovation. Synthetic data addresses one of the most pressing challenges in AI development by enabling AI to function effectively in domains with limited access to natural data. Focusing on enhancing recall and overall model efficacy further emphasizes the benefits of complementing conventional model training with synthetic data to surmount challenges associated with natural data availability. As AI systems continue to extend into the areas where lack of data becomes a bottleneck, hybrid schemes utilizing synthetic data, transfer learning, and few-shot learning will gain even more traction. This study illustrates that you can bring AI to reality in low-resource environments and achieve responsible and sustainable AI development goals. Developers can overcome data bottlenecks and improve model generalization by carefully generating diverse and representative synthetic samples.</p>
<p>Small language models [<xref ref-type="bibr" rid="ref-13">13</xref>&#x2013;<xref ref-type="bibr" rid="ref-15">15</xref>] are designed with fewer parameters, reducing their reliance on massive datasets compared to larger models. Their compact architecture allows for effective training with limited data, making them suitable for niche applications or domains where data collection is challenging. SLMs also provide advantages in terms of reduced computational cost and faster inference, further enhancing their practicality.</p>
<p>Quantization and pruning help mitigate data scarcity by enabling the deployment of models with reduced memory footprints and lower computational demands. Quantization achieves this by representing model weights and activations with lower precision, while pruning removes less important connections in the network. Consequently, these techniques allow for efficient training and deployment of models even with limited data, as the reduced model complexity lessens the risk of overfitting and improves generalization on smaller datasets.</p>
</sec>
<sec id="s5_9">
<label>5.9</label>
<title>Future Research &#x0026; Considerations</title>
<p>Future research could explore how AI integrates with emerging technologies such as quantum computing, which might enhance model training efficiency and solve complex optimization problems faster than classical methods. This integration could be particularly beneficial in data-intensive domains like drug discovery [<xref ref-type="bibr" rid="ref-29">29</xref>] and financial modeling [<xref ref-type="bibr" rid="ref-74">74</xref>]. However, practical implementations remain a challenge and require further investigation. Additionally, research should focus on dynamic AI governance frameworks (See Appendix A, <xref ref-type="table" rid="table-15">Table A1</xref>), such as the EU&#x2019;s Data Governance Act, which can adapt to the rapid pace of technological change [<xref ref-type="bibr" rid="ref-16">16</xref>]. Future work could explore how regulatory sandbox environments (See Appendix A, <xref ref-type="table" rid="table-15">Table A1</xref>) allow AI systems to be tested while ensuring ethical compliance and risk mitigation. Another promising area is advancing synthetic data generation techniques. While current methods like Generative Adversarial Networks (GANs) [<xref ref-type="bibr" rid="ref-77">77</xref>] and diffusion models are effective, they still struggle with maintaining diversity, realism, and fairness. Future research could focus on hybrid approaches that combine synthetic data with human feedback (Human-in-the-Loop AI) to improve data quality [<xref ref-type="bibr" rid="ref-78">78</xref>]. This collaborative approach leverages human expertise to guide AI systems, enhancing the realism and applicability of synthetic data [<xref ref-type="bibr" rid="ref-79">79</xref>]. Moreover, Green AI, as discussed in this paper, is gaining importance in developing computationally efficient models [<xref ref-type="bibr" rid="ref-76">76</xref>]. Investigating techniques like model pruning, quantization, and smaller AI architectures like Small Language Models (SLMs) like Phi-4 and Mistral could provide sustainable AI solutions for low-resource environments. These approaches aim to reduce the environmental impact of AI development while maintaining performance and adaptability. Addressing these areas will ensure that AI remains ethical, efficient, and adaptable in overcoming data scarcity challenges.</p>

</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusion</title>
<p>As artificial intelligence continues to evolve, the industry must proactively develop sustainable, data-efficient systems to address the challenges posed by data scarcity. Developing efficient algorithms that maximize the value of each data point will be essential as access to high-quality natural data becomes increasingly limited. This requires shifting the focus of AI development from simply accumulating large datasets to curating diverse, ethically sourced, high-quality data that ensures fairness, reliability, and adaptability.</p>
<p>To achieve this, AI systems must be designed with dynamic governance frameworks that can adapt to evolving ethical and regulatory landscapes, ensuring responsible development and deployment. Regulatory sandbox environments, as discussed, offer structured mechanisms to test AI models under controlled conditions, fostering both compliance and innovation.</p>
<p>Moreover, the AI industry must embrace computationally efficient techniques to minimize environmental impact while maintaining scalability. Strategies such as model pruning, quantization, and Small Language Models (SLMs) provide sustainable solutions, ensuring AI remains accessible even in low-resource environments. Simultaneously, quantum computing is expected to revolutionize AI, particularly in data-intensive fields. While these advancements hold promise, they also introduce new technical and ethical challenges that necessitate ongoing research and interdisciplinary collaboration.</p>
<p>Ultimately, the future of AI depends on a balanced integration of technological innovation, ethical supervision (See Appendix A, <xref ref-type="table" rid="table-15">Table A1</xref>), and sustainable data strategies. By prioritizing responsible AI governance, optimized data utilization, and energy-efficient models, the AI industry can ensure equitable, ethical, and adaptive progress in overcoming data scarcity challenges and fostering AI&#x2019;s long-term sustainability.</p>

</sec>
</body>
<back>
<ack>
<p>The authors gratefully acknowledge the financial support from Wenzhou-Kean University.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>This work was supported by Internal Research Support Program (IRSPG202202). It is important to note that the USA team involved in this project was not funded by any China-based grants.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>Conceptualization: Hemn Barzan Abdalla, Yulia Kumar. Investigation: Hemn Barzan Abdalla, Ardalan Awlla, Maryam Cheraghy, Jose Marchena, Mehdi Gheisari. Methodology: Hemn Barzan Abdalla, Jose Marchena, Stephany Guzman. Formal analysis: Hemn Barzan Abdalla, Yulia Kumar, Jose Marchena, Stephany Guzman. Writing&#x2014;original draft: Hemn Barzan Abdalla, Yulia Kumar, Jose Marchena, Stephany Guzman, Maryam Cheraghy. Writing&#x2014;review &#x0026; editing: Hemn Barzan Abdalla, Yulia Kumar, Ardalan Awlla. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>All data based on references.</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<app-group id="appg-1">
<app id="app-1">
<title>Appendix A</title>

<table-wrap id="table-15">
<label>Table A1</label>
<caption>
<title>Key vocabulary definitions</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center">Term</th>
<th align="center">Definition</th>
</tr>
</thead>
<tbody>
<tr>
<td><bold>AI governance frameworks</bold></td>
<td>Policies, regulations, and ethical guidelines that ensure AI systems are developed and deployed responsibly.</td>
</tr>
<tr>
<td><bold>Data imbalance</bold></td>
<td>A common issue in AI datasets where one class of data (e.g., fraudulent transactions) is significantly underrepresented compared to another, leading to biased model predictions.</td>
</tr>
<tr>
<td><bold>Data scarcity</bold></td>
<td>The lack of sufficient labeled training data challenges the AI model performance and requires alternative strategies like transfer learning, synthetic data, and self-supervised learning.</td>
</tr>
<tr>
<td><bold>Differential privacy</bold></td>
<td>A data protection technique that ensures individual user data remains anonymous while still enabling AI model training.</td>
</tr>
<tr>
<td><bold>Ethical supervision</bold></td>
<td>Rules and guidelines ensure that AI systems are fair, unbiased, and used responsibly to avoid harm.</td>
</tr>
<tr>
<td><bold>Federated learning</bold></td>
<td>A method that allows AI to learn from data on different devices without collecting or storing it in one place, keeping user information private.</td>
</tr>
<tr>
<td><bold>Green AI</bold></td>
<td>An approach to AI that prioritizes energy efficiency, sustainability, and minimizing environmental impact.</td>
</tr>
<tr>
<td><bold>Privacy-preserving AI techniques</bold></td>
<td>AI systems are designed to process and analyze data while maintaining user privacy, using techniques like federated learning and differential privacy to minimize risks.</td>
</tr>
<tr>
<td><bold>Reinforcement learning</bold></td>
<td>A machine learning paradigm is one in which an AI agent learns by interacting with its environment and receiving feedback in the form of rewards or penalties.</td>
</tr>
<tr>
<td><bold>Regulatory sandbox</bold></td>
<td>A controlled testing environment that allows AI developers to experiment with models while ensuring good performance with ethical standards.</td>
</tr>
<tr>
<td><bold>SMOTE</bold></td>
<td>SMOTE (Synthetic Minority Over-sampling Technique) is a method for balancing imbalanced datasets by generating synthetic samples for underrepresented classes.</td>
</tr>
<tr>
<td><bold>Zero-shot learning</bold></td>
<td>An AI capability that allows models to make predictions on data they have never seen before by transferring knowledge from related tasks.</td>
</tr>
</tbody>
</table>
</table-wrap>
</app>
</app-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Singh</surname> <given-names>J</given-names></string-name></person-group>. <article-title>The rise of synthetic data: enhancing AI and machine learning model training to address data scarcity and mitigate privacy risks</article-title>. <source>J Artif Intell Res Appl</source>. <year>2021</year>;<volume>1</volume>(<issue>2</issue>):<fpage>292</fpage>&#x2013;<lpage>332</lpage>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Panagiotou</surname> <given-names>E</given-names></string-name>, <string-name><surname>Qian</surname> <given-names>H</given-names></string-name>, <string-name><surname>Marx</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ntoutsi</surname> <given-names>E</given-names></string-name></person-group>. <article-title>Generative AI based augmentation for offshore jacket design: an integrated approach for mixed tabular data generation under data scarcity and imbalance. [cited 2025 Jan 1]</article-title>. Available from: <pub-id pub-id-type="doi">10.2139/ssrn.4703856</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>CW</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>R</given-names></string-name>, <string-name><surname>Goh</surname> <given-names>TN</given-names></string-name></person-group>. <article-title>Reliability assessment of high-quality new products with data scarcity</article-title>. <source>Int J Prod Res</source>. <year>2021</year>;<volume>59</volume>(<issue>14</issue>):<fpage>4175</fpage>&#x2013;<lpage>87</lpage>. doi:<pub-id pub-id-type="doi">10.1080/00207543.2020.1758355</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Danaher</surname> <given-names>J</given-names></string-name>, <string-name><surname>Nyholm</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Digital duplicates and the scarcity problem: might AI make us less scarce and therefore less valuable?</article-title> <source>Philos Technol</source>. <year>2024</year>;<volume>37</volume>(<issue>3</issue>):<fpage>106</fpage>. doi:<pub-id pub-id-type="doi">10.1007/s13347-024-00795-z</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Shen</surname> <given-names>JT</given-names></string-name></person-group>. <source>AI in education: effective machine learning methods to improve data scarcity and knowledge generalization</source>. <publisher-loc>University Park, PA, USA</publisher-loc>: <publisher-name>The Pennsylvania State University</publisher-name>; <year>2023</year>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Bai</surname> <given-names>J</given-names></string-name>, <string-name><surname>Alzubaidi</surname> <given-names>L</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Kuhl</surname> <given-names>E</given-names></string-name>, <string-name><surname>Bennamoun</surname> <given-names>M</given-names></string-name>, <string-name><surname>Gu</surname> <given-names>Y</given-names></string-name></person-group>. <article-title>Utilising physics-guided deep learning to overcome data scarcity</article-title>. <comment>arXiv:2211.15664. 2022</comment>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Kaai</surname> <given-names>K</given-names></string-name></person-group>. <source>Addressing data scarcity in domain generalization for computer vision applications in image classification</source>. <publisher-loc>Waterloo, ON, Canada</publisher-loc>: <publisher-name>University of Waterloo</publisher-name>; <year>2024</year>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sufi</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Addressing data scarcity in the medical domain: a GPT-based approach for synthetic data generation and feature extraction</article-title>. <source>Information</source>. <year>2024</year>;<volume>15</volume>(<issue>5</issue>):<fpage>264</fpage>. doi:<pub-id pub-id-type="doi">10.3390/info15050264</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ramakrishnan</surname> <given-names>R</given-names></string-name></person-group>. <article-title>How to build good AI solutions when data is scarce</article-title>. <source>MIT Sloan Manag Rev</source>. <year>2022</year>;<volume>64</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Aldoseri</surname> <given-names>A</given-names></string-name>, <string-name><surname>Al-Khalifa</surname> <given-names>KN</given-names></string-name>, <string-name><surname>Hamouda</surname> <given-names>AM</given-names></string-name></person-group>. <article-title>Re-thinking data strategy and integration for artificial intelligence: concepts, opportunities, and challenges</article-title>. <source>Appl Sci</source>. <year>2023</year>;<volume>13</volume>(<issue>12</issue>):<fpage>7082</fpage>. doi:<pub-id pub-id-type="doi">10.3390/app13127082</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nandy</surname> <given-names>A</given-names></string-name>, <string-name><surname>Duan</surname> <given-names>C</given-names></string-name>, <string-name><surname>Kulik</surname> <given-names>HJ</given-names></string-name></person-group>. <article-title>Audacity of huge: overcoming challenges of data scarcity and data quality for machine learning in computational materials discovery</article-title>. <source>Curr Opin Chem Eng</source>. <year>2022</year>;<volume>36</volume>:<fpage>100778</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.coche.2021.100778</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>OpenAI</collab>, <string-name><surname>Achiam</surname> <given-names>J</given-names></string-name>, <string-name><surname>Adler</surname> <given-names>S</given-names></string-name>, <string-name><surname>Agarwal</surname> <given-names>S</given-names></string-name>, <string-name><surname>Ahmad</surname> <given-names>L</given-names></string-name>, <string-name><surname>Akkaya</surname> <given-names>I</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>GPT-4 technical report. arXiv:2303.08774. 2023</article-title>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Aghajanyan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>L</given-names></string-name>, <string-name><surname>Conneau</surname> <given-names>A</given-names></string-name>, <string-name><surname>Hsu</surname> <given-names>WN</given-names></string-name>, <string-name><surname>Hambardzumyan</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Scaling laws for generative mixed-modal language models</article-title>. In: <conf-name>International Conference on Machine Learning</conf-name>; <year>2023 Jul 23&#x2013;29</year>; <publisher-loc>Honolulu, HI, USA</publisher-loc>. p. <fpage>265</fpage>&#x2013;<lpage>79</lpage>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Villalobos</surname> <given-names>P</given-names></string-name>, <string-name><surname>Ho</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sevilla</surname> <given-names>J</given-names></string-name>, <string-name><surname>Besiroglu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Heim</surname> <given-names>L</given-names></string-name>, <string-name><surname>Hobbhahn</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Position: will we run out of data? Limits of LLM scaling based on human-generated data</article-title>. In: <conf-name>Forty-First International Conference on Machine Learning</conf-name>; <year>2024 Jul 21&#x2013;27</year>; <publisher-loc>Vienna, Austria</publisher-loc>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Villalobos</surname> <given-names>P</given-names></string-name>, <string-name><surname>Ho</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sevilla</surname> <given-names>J</given-names></string-name>, <string-name><surname>Besiroglu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Heim</surname> <given-names>L</given-names></string-name>, <string-name><surname>Hobbhahn</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Will we run out of data? Limits of LLM scaling based on human-generated data. arXiv:2211.04325. 2022</article-title>. doi:<pub-id pub-id-type="doi">10.48550/arXiv.2211.04325</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chowdhery</surname> <given-names>A</given-names></string-name>, <string-name><surname>Narang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Devlin</surname> <given-names>J</given-names></string-name>, <string-name><surname>Bosma</surname> <given-names>M</given-names></string-name>, <string-name><surname>Mishra</surname> <given-names>G</given-names></string-name>, <string-name><surname>Roberts</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>PaLM: scaling language modeling with pathways</article-title>. <source>J Mach Learn Res</source>. <year>2023</year>;<volume>24</volume>(<issue>240</issue>):<fpage>1</fpage>&#x2013;<lpage>13</lpage>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hausen</surname> <given-names>R</given-names></string-name>, <string-name><surname>Azarbonyad</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Discovering data sets through machine learning: an ensemble approach to uncovering the prevalence of government-funded data sets</article-title>. <source>Harv Data Sci Rev</source>. <year>2024</year>;<volume>4</volume>:<fpage>1</fpage>&#x2013;<lpage>18</lpage>. doi:<pub-id pub-id-type="doi">10.1162/99608f92</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Babbar</surname> <given-names>R</given-names></string-name>, <string-name><surname>Sch&#x00F6;lkopf</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Data scarcity, robustness and extreme multi-label classification</article-title>. <source>Mach Learn</source>. <year>2019</year>;<volume>108</volume>(<issue>8</issue>):<fpage>1329</fpage>&#x2013;<lpage>51</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10994-019-05791-5</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Larson</surname> <given-names>J</given-names></string-name>, <string-name><surname>Steven Truitt</surname> <given-names>S</given-names></string-name></person-group>. <article-title>GraphRAG: unlocking LLM discovery on narrative private data. [cited 2025 Jan 15]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.microsoft.com/en-us/research/blog/graphrag-unlocking-llm-discovery-on-narrative-private-data/">https://www.microsoft.com/en-us/research/blog/graphrag-unlocking-llm-discovery-on-narrative-private-data/</ext-link>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Pons</surname> <given-names>G</given-names></string-name>, <string-name><surname>Bilalli</surname> <given-names>B</given-names></string-name>, <string-name><surname>Abell&#x00F3;</surname> <given-names>A</given-names></string-name>, <string-name><surname>Blanco S&#x00E1;nchez</surname> <given-names>S</given-names></string-name></person-group>. <article-title>On the use of trajectory data for tackling data scarcity</article-title>. <source>Inf Syst</source>. <year>2025</year>;<volume>130</volume>(<issue>5</issue>):<fpage>102523</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.is.2025.102523</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Karst</surname> <given-names>F</given-names></string-name>, <string-name><surname>Li</surname> <given-names>M</given-names></string-name>, <string-name><surname>Reinhard</surname> <given-names>P</given-names></string-name>, <string-name><surname>Leimeister</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Not enough data to be fair? Evaluating fairness implications of data scarcity solutions</article-title>. In: <conf-name>Proceedings of the 58th Hawaii International Conference on System Sciences</conf-name>; <year>2025 Jan 7&#x2013;10</year>; <publisher-loc>Big Island, HI, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Gan</surname> <given-names>W</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>K</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>H</given-names></string-name></person-group>. <article-title>Data scarcity in recommendation systems: a survey</article-title>. <source>ACM Trans Recomm Syst</source>. <year>2025</year>;<volume>3</volume>(<issue>3</issue>):<fpage>1</fpage>&#x2013;<lpage>31</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3700890</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Khalil</surname> <given-names>M</given-names></string-name>, <string-name><surname>Vadiee</surname> <given-names>F</given-names></string-name>, <string-name><surname>Shakya</surname> <given-names>R</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Q</given-names></string-name></person-group>. <article-title>Creating artificial students that never existed: leveraging large language models and CTGANs for synthetic data generation</article-title>. <comment>arXiv:2501.01793. 2025</comment>. doi:<pub-id pub-id-type="doi">10.1145/3706468</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>White</surname> <given-names>J</given-names></string-name>, <string-name><surname>Madaan</surname> <given-names>P</given-names></string-name>, <string-name><surname>Shenoy</surname> <given-names>N</given-names></string-name>, <string-name><surname>Agnihotri</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sharma</surname> <given-names>M</given-names></string-name>, <string-name><surname>Doshi</surname> <given-names>J</given-names></string-name></person-group>. <article-title>A case for rejection in low resource ML deployment</article-title>. <comment>arXiv:2208.06359. 2022</comment>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Singh</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Systematic review of data-centric approaches in artificial intelligence and machine learning</article-title>. <source>Data Sci Manag</source>. <year>2023</year>;<volume>6</volume>(<issue>3</issue>):<fpage>144</fpage>&#x2013;<lpage>57</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.dsm.2023.06.001</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>D</given-names></string-name>, <string-name><surname>Akashi</surname> <given-names>K</given-names></string-name>, <string-name><surname>Nozue</surname> <given-names>H</given-names></string-name>, <string-name><surname>Tayama</surname> <given-names>K</given-names></string-name></person-group>. <article-title>A mirror environment to produce artificial intelligence training data</article-title>. <source>IEEE Access</source>. <year>2022</year>;<volume>10</volume>:<fpage>24578</fpage>&#x2013;<lpage>86</lpage>. doi:<pub-id pub-id-type="doi">10.1109/ACCESS.2022.3154825</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lautrup</surname> <given-names>AD</given-names></string-name>, <string-name><surname>Hyrup</surname> <given-names>T</given-names></string-name>, <string-name><surname>Zimek</surname> <given-names>A</given-names></string-name>, <string-name><surname>Schneider-Kamp</surname> <given-names>P</given-names></string-name></person-group>. <article-title>SynthEval: a framework for detailed utility and privacy evaluation of tabular synthetic data</article-title>. <source>Data Min Knowl Discov</source>. <year>2025</year>;<volume>39</volume>(<issue>1</issue>):<fpage>1</fpage>&#x2013;<lpage>25</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10618-024-01081-4</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gangwal</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ansari</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ahmad</surname> <given-names>I</given-names></string-name>, <string-name><surname>Azad</surname> <given-names>AK</given-names></string-name>, <string-name><surname>Wan Sulaiman</surname> <given-names>WMA</given-names></string-name></person-group>. <article-title>Current strategies to address data scarcity in artificial intelligence-based drug discovery: a comprehensive review</article-title>. <source>Comput Biol Med</source>. <year>2024</year>;<volume>179</volume>:<fpage>108734</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.compbiomed.2024.108734</pub-id>; <pub-id pub-id-type="pmid">38964243</pub-id></mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Niel</surname> <given-names>O</given-names></string-name></person-group>. <article-title>A novel algorithm can generate data to train machine learning models in conditions of extreme scarcity of real world data</article-title>. <comment>arXiv:2305.00987. 2023</comment>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chen</surname> <given-names>YT</given-names></string-name>, <string-name><surname>Hsu</surname> <given-names>CY</given-names></string-name>, <string-name><surname>Yu</surname> <given-names>CM</given-names></string-name>, <string-name><surname>Barhamgi</surname> <given-names>M</given-names></string-name>, <string-name><surname>Perera</surname> <given-names>C</given-names></string-name></person-group>. <article-title>On the private data synthesis through deep generative models for data scarcity of industrial Internet of Things</article-title>. <source>IEEE Trans Ind Inform</source>. <year>2023</year>;<volume>19</volume>(<issue>1</issue>):<fpage>551</fpage>&#x2013;<lpage>60</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TII.2021.3133625</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bansal</surname> <given-names>MA</given-names></string-name>, <string-name><surname>Sharma</surname> <given-names>DR</given-names></string-name>, <string-name><surname>Kathuria</surname> <given-names>DM</given-names></string-name></person-group>. <article-title>A systematic review on data scarcity problem in deep learning: solution and applications</article-title>. <source>ACM Comput Surv</source>. <year>2022</year>;<volume>54</volume>(<issue>10s</issue>):<fpage>1</fpage>&#x2013;<lpage>29</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3502287</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Nahid</surname> <given-names>MMH</given-names></string-name>, <string-name><surname>Bin Hasan</surname> <given-names>S</given-names></string-name></person-group>. <article-title>SafeSynthDP: leveraging large language models for privacy-preserving synthetic data generation using differential privacy</article-title>. <comment>arXiv:2412.20641. 2024</comment>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Zou</surname> <given-names>MX</given-names></string-name>, <string-name><surname>Lo</surname> <given-names>CC</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>CH</given-names></string-name>, <string-name><surname>Shieh</surname> <given-names>CS</given-names></string-name>, <string-name><surname>Horng</surname> <given-names>MF</given-names></string-name></person-group>. <article-title>Data augmentation based on topic relevance to enhance text classification in scarcity of training data</article-title>. In: <conf-name>International Conference on Intelligent Information Hiding and Multimedia Signal Processing</conf-name>; <year>2022 Dec 16&#x2013;18</year>; <publisher-loc>Kitakyushu, Japan. Singapore</publisher-loc>: <publisher-name>Springer Nature Singapore</publisher-name>; <comment>2022</comment>. p. <fpage>347</fpage>&#x2013;<lpage>57</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-981-99-0105-0_31</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Hoang</surname> <given-names>V</given-names></string-name></person-group>. <article-title>Mitigating data scarcity for large language models</article-title>. <comment>arXiv:2302.01806. 2023</comment>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Alzubaidi</surname> <given-names>L</given-names></string-name>, <string-name><surname>Bai</surname> <given-names>J</given-names></string-name>, <string-name><surname>Al-Sabaawi</surname> <given-names>A</given-names></string-name>, <string-name><surname>Santamar&#x00ED;a</surname> <given-names>J</given-names></string-name>, <string-name><surname>Albahri</surname> <given-names>AS</given-names></string-name>, <string-name><surname>Al-dabbagh</surname> <given-names>BSN</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A survey on deep learning tools dealing with data scarcity: definitions, challenges, solutions, tips, and applications</article-title>. <source>J Big Data</source>. <year>2023</year>;<volume>10</volume>(<issue>1</issue>):<fpage>46</fpage>. doi:<pub-id pub-id-type="doi">10.1186/s40537-023-00727-2</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Laurer</surname> <given-names>M</given-names></string-name>, <string-name><surname>van Atteveldt</surname> <given-names>W</given-names></string-name>, <string-name><surname>Casas</surname> <given-names>A</given-names></string-name>, <string-name><surname>Welbers</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Less annotating, more classifying: addressing the data scarcity issue of supervised machine learning with deep transfer learning and BERT-NLI</article-title>. <source>Polit Anal</source>. <year>2024</year>;<volume>32</volume>(<issue>1</issue>):<fpage>84</fpage>&#x2013;<lpage>100</lpage>. doi:<pub-id pub-id-type="doi">10.1017/pan.2023.20</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Quito</surname> <given-names>JA</given-names></string-name>, <string-name><surname>Andrade</surname> <given-names>LJ</given-names></string-name></person-group>. <article-title>Proposal for the generation of profiles using a synthetic database</article-title>. In: <conf-name>AHFE, 2022 Conference on Applied Human Factors and Ergonomics International</conf-name>; <year>2022 Jul 24&#x2013;28</year>; <publisher-loc>New York City, NY, USA</publisher-loc>. doi:<pub-id pub-id-type="doi">10.54941/ahfe1001462</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Oxford</surname> <given-names>Analytica</given-names></string-name></person-group>. <source>Data scarcity challenges AI developers globally</source>. <publisher-loc>Bingley, UK</publisher-loc>: <publisher-name>Emerald Expert Briefings</publisher-name>; <year>2024</year>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Luo</surname> <given-names>T</given-names></string-name>, <string-name><surname>Yan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>JT</given-names></string-name>, <string-name><surname>Goh</surname> <given-names>R</given-names></string-name></person-group>. <article-title>RCT: resource constrained training for edge AI</article-title>. <source>IEEE Trans Neural Netw Learn Syst</source>. <year>2022</year>;<volume>35</volume>(<issue>2</issue>):<fpage>2575</fpage>&#x2013;<lpage>87</lpage>. doi:<pub-id pub-id-type="doi">10.1109/TNNLS.2022.3190451</pub-id>; <pub-id pub-id-type="pmid">35998171</pub-id></mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="other"><comment>[cited 2025 Mar 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://brusselsprivacyhub.com/wp-content/uploads/2024/02/Personal-Data-Protection-in-Brazil.pdf">https://brusselsprivacyhub.com/wp-content/uploads/2024/02/Personal-Data-Protection-in-Brazil.pdf</ext-link>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="other"><comment>[cited 2025 Mar 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.novaoneadvisor.com/report/artificial-intelligence-market">https://www.novaoneadvisor.com/report/artificial-intelligence-market</ext-link>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="other"><comment>[cited 2025 Mar 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://ourworldindata.org/data-insights/investment-in-generative-ai-has-surged-recently">https://ourworldindata.org/data-insights/investment-in-generative-ai-has-surged-recently</ext-link>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="other"><comment>[cited 2025 Mar 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://hai.stanford.edu/news/ai-index-state-ai-13-charts">https://hai.stanford.edu/news/ai-index-state-ai-13-charts</ext-link>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Healy</surname> <given-names>M</given-names></string-name>, <string-name><surname>Baum</surname> <given-names>A</given-names></string-name>, <string-name><surname>Musumeci</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Addressing data scarcity in ML-based failure-cause identification in optical networks through generative models</article-title>. <source>Opt Fiber Technol</source>. <year>2025</year>;<volume>90</volume>(<issue>12</issue>):<fpage>104137</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.yofte.2025.104137</pub-id>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="other"><comment>[cited 2025 Mar 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.fbcinc.com/source/virtualhall_images/2024_Virtual_Events/USDA_Innovation/Databricks/State_of_Data___AI_Resource.pdf">https://www.fbcinc.com/source/virtualhall_images/2024_Virtual_Events/USDA_Innovation/Databricks/State_of_Data___AI_Resource.pdf</ext-link>.</mixed-citation></ref>
<ref id="ref-46"><label>[46]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Belkada</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Dettmers</surname> <given-names>T</given-names></string-name>, <string-name><surname>Pagnoni</surname> <given-names>A</given-names></string-name>, <string-name><surname>Gugger</surname> <given-names>S</given-names></string-name>, <string-name><surname>Mangrulkar</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Making LLMs even more accessible with bitsandbytes, 4-bit quantization and QLoRA. [cited 2025 Jan 15]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://huggingface.co/blog/4bit-transformers-bitsandbytes">https://huggingface.co/blog/4bit-transformers-bitsandbytes</ext-link>.</mixed-citation></ref>
<ref id="ref-47"><label>[47]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Abdalla</surname> <given-names>HB</given-names></string-name></person-group>. <article-title>A brief survey on big data: technologies, terminologies and data-intensive applications</article-title>. <source>J Big Data</source>. <year>2022</year>;<volume>9</volume>(<issue>1</issue>):<fpage>107</fpage>. doi:<pub-id pub-id-type="doi">10.1186/s40537-022-00659-3</pub-id>.</mixed-citation></ref>
<ref id="ref-48"><label>[48]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Abdalla</surname> <given-names>HB</given-names></string-name>, <string-name><surname>Abuhaija</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Comprehensive analysis of various big data classification techniques: a challenging overview</article-title>. <source>J Info Know Mgmt</source>. <year>2023</year>;<volume>22</volume>(<issue>1</issue>):<fpage>2250083</fpage>. doi:<pub-id pub-id-type="doi">10.1142/S0219649222500836</pub-id>.</mixed-citation></ref>
<ref id="ref-49"><label>[49]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Abdalla</surname> <given-names>HB</given-names></string-name>, <string-name><surname>Awlla</surname> <given-names>AH</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Cheraghy</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Big data: past, present, and future insights</article-title>. In: <conf-name>Proceedings of the 2024 Asia Pacific Conference on Computing Technologies, Communications and Networking</conf-name>; <year>2024 Jul 26&#x2013;27</year>; <publisher-loc>Chengdu, China</publisher-loc>. p. <fpage>60</fpage>&#x2013;<lpage>70</lpage>. doi:<pub-id pub-id-type="doi">10.1145/3685767.3685777</pub-id>.</mixed-citation></ref>
<ref id="ref-50"><label>[50]</label><mixed-citation publication-type="other"><comment>[cited 2025 Mar 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.sciencealert.com/the-world-is-running-out-of-data-to-feed-ai-experts-warn">https://www.sciencealert.com/the-world-is-running-out-of-data-to-feed-ai-experts-warn</ext-link>.</mixed-citation></ref>
<ref id="ref-51"><label>[51]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Giuffr&#x00E8;</surname> <given-names>M</given-names></string-name>, <string-name><surname>Shung</surname> <given-names>DL</given-names></string-name></person-group>. <article-title>Harnessing the power of synthetic data in healthcare: innovation, application, and privacy</article-title>. <source>NPJ Digit Med</source>. <year>2023</year>;<volume>6</volume>(<issue>1</issue>):<fpage>186</fpage>. doi:<pub-id pub-id-type="doi">10.1038/s41746-023-00927-3</pub-id>; <pub-id pub-id-type="pmid">37813960</pub-id></mixed-citation></ref>
<ref id="ref-52"><label>[52]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Dang</surname> <given-names>H</given-names></string-name>, <string-name><surname>Mecke</surname> <given-names>L</given-names></string-name>, <string-name><surname>Lehmann</surname> <given-names>F</given-names></string-name>, <string-name><surname>Goller</surname> <given-names>S</given-names></string-name>, <string-name><surname>Buschek</surname> <given-names>D</given-names></string-name></person-group>. <article-title>How to prompt? Opportunities and challenges of zero- and few-shot learning for human-AI interaction in creative applications of generative models</article-title>. <comment>arXiv:2209.01390. 2022</comment>.</mixed-citation></ref>
<ref id="ref-53"><label>[53]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Yang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Xiao</surname> <given-names>W</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>M</given-names></string-name>, <string-name><surname>Guo</surname> <given-names>S</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>J</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>F</given-names></string-name></person-group>. <article-title>Image data augmentation for deep learning: A survey</article-title>. <comment>arXiv:2204.08610. 2022</comment>.</mixed-citation></ref>
<ref id="ref-54"><label>[54]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Abdulmalek</surname> <given-names>S</given-names></string-name>, <string-name><surname>Nasir</surname> <given-names>A</given-names></string-name>, <string-name><surname>Jabbar</surname> <given-names>WA</given-names></string-name>, <string-name><surname>Almuhaya</surname> <given-names>MAM</given-names></string-name>, <string-name><surname>Bairagi</surname> <given-names>AK</given-names></string-name>, <string-name><surname>Khan</surname> <given-names>MA</given-names></string-name>, <etal>et al.</etal></person-group> <article-title>IoT-based healthcare-monitoring system towards improving quality of life: a review</article-title>. <source>Healthcare</source>. <year>2022</year>;<volume>10</volume>(<issue>10</issue>):<fpage>1993</fpage>. doi:<pub-id pub-id-type="doi">10.3390/healthcare10101993</pub-id>; <pub-id pub-id-type="pmid">36292441</pub-id></mixed-citation></ref>
<ref id="ref-55"><label>[55]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Goyal</surname> <given-names>M</given-names></string-name>, <string-name><surname>Mahmoud</surname> <given-names>QH</given-names></string-name></person-group>. <article-title>A systematic review of synthetic data generation techniques using generative AI</article-title>. <source>Electronics</source>. <year>2024</year>;<volume>13</volume>(<issue>17</issue>):<fpage>3509</fpage>. doi:<pub-id pub-id-type="doi">10.3390/electronics13173509</pub-id>.</mixed-citation></ref>
<ref id="ref-56"><label>[56]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rani</surname> <given-names>V</given-names></string-name>, <string-name><surname>Nabi</surname> <given-names>ST</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>M</given-names></string-name>, <string-name><surname>Mittal</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kumar</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Self-supervised learning: a succinct review</article-title>. <source>Arch Comput Methods Eng</source>. <year>2023</year>;<volume>30</volume>(<issue>4</issue>):<fpage>2761</fpage>&#x2013;<lpage>75</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11831-023-09884-2</pub-id>; <pub-id pub-id-type="pmid">36713767</pub-id></mixed-citation></ref>
<ref id="ref-57"><label>[57]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>W</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>R</given-names></string-name>, <string-name><surname>Li</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liao</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>Z</given-names></string-name>, <string-name><surname>He</surname> <given-names>G</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A perspective survey on deep transfer learning for fault diagnosis in industrial scenarios: theories, applications and challenges</article-title>. <source>Mech Syst Signal Process</source>. <year>2022</year>;<volume>167</volume>(<issue>12</issue>):<fpage>108487</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.ymssp.2021.108487</pub-id>.</mixed-citation></ref>
<ref id="ref-58"><label>[58]</label><mixed-citation publication-type="other"><comment>[cited 2025 Mar 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://technologymagazine.com/articles/how-googles-ai-breakthroughs-earned-nobel-recognition">https://technologymagazine.com/articles/how-googles-ai-breakthroughs-earned-nobel-recognition</ext-link>.</mixed-citation></ref>
<ref id="ref-59"><label>[59]</label><mixed-citation publication-type="other"><comment>[cited 2025 Mar 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://research.google/blog/automl-for-large-scale-image-classification-and-object-detection">https://research.google/blog/automl-for-large-scale-image-classification-and-object-detection</ext-link>.</mixed-citation></ref>
<ref id="ref-60"><label>[60]</label><mixed-citation publication-type="other"><comment>[cited 2025 Mar 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://innovationexchange.mayoclinic.org/an-introduction-to-federal-funding-for-innovation/">https://innovationexchange.mayoclinic.org/an-introduction-to-federal-funding-for-innovation/</ext-link>.</mixed-citation></ref>
<ref id="ref-61"><label>[61]</label><mixed-citation publication-type="other"><comment>[cited 2025 Mar 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.techmonitor.ai/risks/extreme-weather-events/ibm-weather-and-climate-model?cf-view">https://www.techmonitor.ai/risks/extreme-weather-events/ibm-weather-and-climate-model?cf-view</ext-link>.</mixed-citation></ref>
<ref id="ref-62"><label>[62]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lund</surname> <given-names>B</given-names></string-name></person-group>. <article-title>Universities engaging social media users: an investigation of quantitative relationships between universities&#x2019; Facebook followers/interactions and university attributes</article-title>. <source>J Mark High Educ</source>. <year>2019</year>;<volume>29</volume>(<issue>2</issue>):<fpage>251</fpage>&#x2013;<lpage>67</lpage>. doi:<pub-id pub-id-type="doi">10.1080/08841241.2019.1641875</pub-id>.</mixed-citation></ref>
<ref id="ref-63"><label>[63]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Miller</surname> <given-names>T</given-names></string-name>, <string-name><surname>Durlik</surname> <given-names>I</given-names></string-name>, <string-name><surname>Kostecka</surname> <given-names>E</given-names></string-name>, <string-name><surname>Borkowski</surname> <given-names>P</given-names></string-name>, <string-name><surname>&#x0141;obodzi&#x0144;ska</surname> <given-names>A</given-names></string-name></person-group>. <article-title>A critical AI view on autonomous vehicle navigation: the growing danger</article-title>. <source>Electronics</source>. <year>2024</year>;<volume>13</volume>(<issue>18</issue>):<fpage>3660</fpage>. doi:<pub-id pub-id-type="doi">10.3390/electronics13183660</pub-id>.</mixed-citation></ref>
<ref id="ref-64"><label>[64]</label><mixed-citation publication-type="other"><comment>[cited 2025 Mar 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.europarl.europa.eu/RegData/etudes/STUD/2020/641530/EPRS_STU(2020)641530_EN.pdf">https://www.europarl.europa.eu/RegData/etudes/STUD/2020/641530/EPRS_STU(2020)641530_EN.pdf</ext-link>.</mixed-citation></ref>
<ref id="ref-65"><label>[65]</label><mixed-citation publication-type="other"><comment>[cited 2025 Mar 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://oag.ca.gov/privacy/ccpa">https://oag.ca.gov/privacy/ccpa</ext-link>.</mixed-citation></ref>
<ref id="ref-66"><label>[66]</label><mixed-citation publication-type="other"><article-title>[cited 2025 Mar 12]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.twobirds.com/en/insights/2024/china/ai-governance-in-china-strategies-initiatives-and-key">https://www.twobirds.com/en/insights/2024/china/ai-governance-in-china-strategies-initiatives-and-key</ext-link>.</mixed-citation></ref>
<ref id="ref-67"><label>[67]</label><mixed-citation publication-type="other"><comment>[cited 2025 Mar 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.hhs.gov/hipaa/for-professionals/privacy/laws">https://www.hhs.gov/hipaa/for-professionals/privacy/laws</ext-link>.</mixed-citation></ref>
<ref id="ref-68"><label>[68]</label><mixed-citation publication-type="other"><comment>[cited 2025 Mar 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.kaggle.com/datasets/mlg-ulb/creditcardfraud/data">https://www.kaggle.com/datasets/mlg-ulb/creditcardfraud/data</ext-link>.</mixed-citation></ref>
<ref id="ref-69"><label>[69]</label><mixed-citation publication-type="other"><comment>[cited 2025 Mar 12]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://archive.ics.uci.edu/dataset/17/breast&#x002B;cancer&#x002B;wisconsin&#x002B;diagnostic">https://archive.ics.uci.edu/dataset/17/breast&#x002B;cancer&#x002B;wisconsin&#x002B;diagnostic</ext-link>.</mixed-citation></ref>
<ref id="ref-70"><label>[70]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kumar</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Lin</surname> <given-names>M</given-names></string-name>, <string-name><surname>Paredes</surname> <given-names>C</given-names></string-name>, <string-name><surname>Li</surname> <given-names>D</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Kruger</surname> <given-names>D</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>A comprehensive review of AI advancement using testFAILS and testFAILS-2 for the pursuit of AGI</article-title>. <source>Electronics</source>. <year>2024</year>;<volume>13</volume>(<issue>24</issue>):<fpage>4991</fpage>. doi:<pub-id pub-id-type="doi">10.3390/electronics13244991</pub-id>.</mixed-citation></ref>
<ref id="ref-71"><label>[71]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Haruni</surname> <given-names>R</given-names></string-name></person-group>. <article-title>Nvidia CEO maps out bold vision: AGI and robotics set to merge. [cited 2025 Jan 15]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://wallstreetpit.com/120150-nvidia-ceo-maps-out-bold-vision-agi-and-robotics-set-to-merge/">https://wallstreetpit.com/120150-nvidia-ceo-maps-out-bold-vision-agi-and-robotics-set-to-merge/</ext-link>.</mixed-citation></ref>
<ref id="ref-72"><label>[72]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Barlow</surname> <given-names>G</given-names></string-name></person-group>. <article-title>Altman predicts artificial superintelligence (AGI) will happen this year</article-title>. <comment>[cited 2025 Jan 15]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.techradar.com/computing/artificial-intelligence/sam-altman-predicts-artificial-superintelligence-agi-will-happen-this-year">https://www.techradar.com/computing/artificial-intelligence/sam-altman-predicts-artificial-superintelligence-agi-will-happen-this-year</ext-link>.</mixed-citation></ref>
<ref id="ref-73"><label>[73]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Robison</surname> <given-names>K</given-names></string-name></person-group>. <article-title>OpenAI cofounder Ilya Sutskever says the way AI is built is about to change. [cited 2025 Jan 15]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.theverge.com/2024/12/13/24320811/what-ilya-sutskever-sees-openai-model-data-training">https://www.theverge.com/2024/12/13/24320811/what-ilya-sutskever-sees-openai-model-data-training</ext-link>.</mixed-citation></ref>
<ref id="ref-74"><label>[74]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kumar</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Huang</surname> <given-names>K</given-names></string-name>, <string-name><surname>Perez</surname> <given-names>A</given-names></string-name>, <string-name><surname>Yang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Li</surname> <given-names>JJ</given-names></string-name>, <string-name><surname>Morreale</surname> <given-names>P</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Bias and cyberbullying detection and data generation using transformer artificial intelligence models and top large language models</article-title>. <source>Electronics</source>. <year>2024</year>;<volume>13</volume>(<issue>17</issue>):<fpage>3431</fpage>. doi:<pub-id pub-id-type="doi">10.3390/electronics13173431</pub-id>.</mixed-citation></ref>
<ref id="ref-75"><label>[75]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kumar</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Marchena</surname> <given-names>J</given-names></string-name>, <string-name><surname>Awlla</surname> <given-names>AH</given-names></string-name>, <string-name><surname>Li</surname> <given-names>JJ</given-names></string-name>, <string-name><surname>Abdalla</surname> <given-names>HB</given-names></string-name></person-group>. <article-title>The AI-powered evolution of big data</article-title>. <source>Appl Sci</source>. <year>2024</year>;<volume>14</volume>(<issue>22</issue>):<fpage>10176</fpage>. doi:<pub-id pub-id-type="doi">10.3390/app142210176</pub-id>.</mixed-citation></ref>
<ref id="ref-76"><label>[76]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Edge</surname> <given-names>D</given-names></string-name>, <string-name><surname>Larson</surname> <given-names>J</given-names></string-name></person-group>. <article-title>LazyGraphRAG: setting a new standard for quality and cost. [cited 2025 Jan 15]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.microsoft.com/en-us/research/blog/lazygraphrag-setting-a-new-standard-for-quality-and-cost/">https://www.microsoft.com/en-us/research/blog/lazygraphrag-setting-a-new-standard-for-quality-and-cost/</ext-link>.</mixed-citation></ref>
<ref id="ref-77"><label>[77]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Feng</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Shen</surname> <given-names>A</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Liang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>S</given-names></string-name>, <string-name><surname>Du</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Enhancing few-shot learning with integrated data and GAN model approaches</article-title>. <comment>arXiv:2411.16567. 2024</comment>.</mixed-citation></ref>
<ref id="ref-78"><label>[78]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Wiethof</surname> <given-names>C</given-names></string-name>, <string-name><surname>Bittner</surname> <given-names>EA</given-names></string-name></person-group>. <article-title>Hybrid intelligence&#x2014;combining the human in the loop with the computer in the loop: a systematic literature review. [cited 2025 Jan 1]</article-title>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.researchgate.net/publication/356209722">https://www.researchgate.net/publication/356209722</ext-link>.</mixed-citation></ref>
<ref id="ref-79"><label>[79]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Verdecchia</surname> <given-names>R</given-names></string-name>, <string-name><surname>Sallou</surname> <given-names>J</given-names></string-name>, <string-name><surname>Cruz</surname> <given-names>L</given-names></string-name></person-group>. <article-title>A systematic review of green AI</article-title>. <source>Wiley Interdiscip Rev Data Min Knowl Discov</source>. <year>2023</year>;<volume>13</volume>(<issue>4</issue>):<fpage>e1507</fpage>. doi:<pub-id pub-id-type="doi">10.1002/widm.1507</pub-id>.</mixed-citation></ref>
</ref-list>
</back></article>