<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xml:lang="en" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">67840</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2025.067840</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Robust Multi-Label Cartoon Character Classification on the Novel Kral Sakir Dataset Using Deep Learning Techniques</article-title>
<alt-title alt-title-type="left-running-head">Robust Multi-Label Cartoon Character Classification on the Novel Kral Sakir Dataset Using Deep Learning Techniques</alt-title>
<alt-title alt-title-type="right-running-head">Robust Multi-Label Cartoon Character Classification on the Novel Kral Sakir Dataset Using Deep Learning Techniques</alt-title>
</title-group>
<contrib-group>
<contrib id="author-1" contrib-type="author">
<name name-style="western"><surname>Tumer</surname><given-names>Candan</given-names></name><xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Guvenoglu</surname><given-names>Erdal</given-names></name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Tunali</surname><given-names>Volkan</given-names></name><xref ref-type="aff" rid="aff-3">3</xref><email>volkan.tunali@uws.ac.uk</email></contrib>
<aff id="aff-1"><label>1</label><institution>Graduate School, Maltepe University</institution>, <addr-line>Istanbul, 34857</addr-line>, <country>Turkiye</country></aff>
<aff id="aff-2"><label>2</label><institution>Department of Computer Programming, Vocational School, Maltepe University</institution>, <addr-line>Istanbul, 34857</addr-line>, <country>Turkiye</country></aff>
<aff id="aff-3"><label>3</label><institution>Division of Computing, School of Computing, Engineering and Physical Sciences, University of the West of Scotland, London Campus</institution>, <addr-line>London, E14 2BE</addr-line>, <country>UK</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Volkan Tunali. Email: <email>volkan.tunali@uws.ac.uk</email></corresp>
</author-notes>
<pub-date date-type="collection" publication-format="electronic">
<year>2025</year>
</pub-date>
<pub-date date-type="pub" publication-format="electronic">
<day>23</day><month>10</month><year>2025</year>
</pub-date>
<volume>85</volume>
<issue>3</issue>
<fpage>5135</fpage>
<lpage>5158</lpage>
<history>
<date date-type="received">
<day>14</day>
<month>5</month>
<year>2025</year>
</date>
<date date-type="accepted">
<day>19</day>
<month>8</month>
<year>2025</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2025 The Authors.</copyright-statement>
<copyright-year>2025</copyright-year>
<copyright-holder>Published by Tech Science Press.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_67840.pdf"></self-uri>
<abstract>
<p>Automated cartoon character recognition is crucial for applications in content indexing, filtering, and copyright protection, yet it faces a significant challenge in animated media due to high intra-class visual variability, where characters frequently alter their appearance. To address this problem, we introduce the novel Kral Sakir dataset, a public benchmark of 16,725 images specifically curated for the task of multi-label cartoon character classification under these varied conditions. This paper conducts a comprehensive benchmark study, evaluating the performance of state-of-the-art pretrained Convolutional Neural Networks (CNNs), including DenseNet, ResNet, and VGG, against a custom baseline model trained from scratch. Our experiments, evaluated using metrics of F1-Score, accuracy, and Area Under the ROC Curve (AUC), demonstrate that fine-tuning pretrained models is a highly effective strategy. The best-performing model, DenseNet121, achieved an F1-Score of 0.9890 and an accuracy of 0.9898, significantly outperforming our baseline CNN (F1-Score of 0.9545). The findings validate the power of transfer learning for this domain and establish a strong performance benchmark. The introduced dataset provides a valuable resource for future research into developing robust and accurate character recognition systems.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Cartoon character recognition</kwd>
<kwd>multi-label classification</kwd>
<kwd>deep learning</kwd>
<kwd>transfer learning</kwd>
<kwd>predictive modelling</kwd>
<kwd>artificial intelligence-enhanced (AI-Enhanced) systems</kwd>
<kwd>Kral Sakir dataset</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Animated content, including cartoons, constitutes a significant and globally popular form of entertainment and communication. The ability to automatically analyse this content, particularly to identify and track characters, offers substantial benefits across various applications, from content indexing and retrieval to copyright protection and interactive media experiences [<xref ref-type="bibr" rid="ref-1">1</xref>]. However, automated cartoon character recognition presents unique and significant challenges that differentiate it from general object recognition tasks.</p>
<p>A key challenge in the application of AI to image recognition, and a central focus of our research, stems from the substantial visual variability inherent in domains like cartoon character portrayal. Unlike the more constrained variations typically observed in real-world object classes, cartoon characters often undergo transformations driven by artistic license and narrative progression. These can manifest as considerable changes in attire, colour schemes, accessories, depicted age, and even fundamental drawing style or perspective [<xref ref-type="bibr" rid="ref-2">2</xref>]. This high degree of intra-class variance presents a notable difficulty for current AI models, particularly deep learning architectures that learn features directly from pixel data, as consistent low-level visual patterns may become less prevalent or frequently altered [<xref ref-type="bibr" rid="ref-3">3</xref>]. While human perception demonstrates remarkable robustness in identifying characters across these disparate renditions, improving the ability of AI systems to handle such artistic freedom is important for enhancing their performance in specialised visual contexts. Our work addresses this by investigating how established AI techniques can achieve more reliable recognition in these dynamic visual environments, thereby contributing to the understanding of how AI can manage high intra-class variance and informing efforts to develop more robust AI-driven image recognition systems for such domains.</p>
<p>Progress in overcoming these complexities is intrinsically linked to the availability of diverse and well-curated datasets that adequately represent these variations. The development of new, specialised datasets is crucial, as they provide the necessary foundation for training and comprehensively evaluating more sophisticated AI models capable of handling such complexities. Such datasets enable researchers to benchmark different approaches and push the boundaries of what is achievable in automated cartoon analysis. To this end, this paper introduces a novel dataset, the Kral Sakir dataset, specifically designed to capture the multi-label nature and visual diversity of characters from a popular animated series. Furthermore, we investigate the effectiveness of deep learning techniques, particularly transfer learning from pretrained Convolutional Neural Networks (CNNs), for robust multi-label cartoon character classification on this new dataset, aiming to provide insights into effective strategies for this challenging task.</p>
<p>The main contributions and novelty of our work can be summarised as follows:
<list list-type="simple">
<list-item><label>1.</label><p>We introduce the novel Kral Sakir dataset, derived from a popular Turkish animated series, for multi-label cartoon character recognition. This dataset features significant visual diversity from a specific cultural context and serves as a benchmark for future research in automated cartoon visual analysis.</p></list-item>
<list-item><label>2.</label><p>We systematically evaluate state-of-the-art pretrained CNNs against a custom baseline, demonstrating transfer learning&#x2019;s robustness and efficiency for multi-label cartoon character classification and offering insights into effective model choices.</p></list-item>
<list-item><label>3.</label><p>We conduct a targeted investigation into classifying cartoon characters under high intra-class visual variance, demonstrating effective deep learning strategies on our diverse dataset for robust, multi-label recognition and thereby contributing to the development of more resilient automated systems.</p></list-item>
</list></p>
<p>The remainder of this paper is structured as follows. <xref ref-type="sec" rid="s2">Section 2</xref> reviews the relevant literature on cartoon character recognition, outlining key developments and established techniques. <xref ref-type="sec" rid="s3">Section 3</xref> details the materials and methods employed in this study, including our proposed dataset and experimental setup. The experimental results are presented and discussed in <xref ref-type="sec" rid="s4">Section 4</xref>. Finally, <xref ref-type="sec" rid="s5">Section 5</xref> concludes the paper, summarising our key findings and contributions.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Related Work</title>
<p>The automated detection and recognition of characters within cartoons, comics, and animated media has been a significant area of research, driven by applications ranging from copyright protection to content analysis and retrieval. Early investigations often centered on extracting handcrafted visual features, such as colour, edges, and shape contexts, to identify characters in static images and line drawings. More recently, the field has witnessed a pronounced shift towards deep learning methodologies, with various CNN architectures, including Region-based Convolutional Neural Networks (R-CNN) variants and You Only Look Once (YOLO), demonstrating robust performance in identifying faces and full characters across diverse comic and manga datasets. While much of this work has focused on static images, the challenge of recognising characters with high visual diversity in dynamic video content, which is the focus of our study, builds upon these foundational efforts. This review will explore these key developments and techniques employed in the broader domain of cartoon character recognition, also summarising them in <xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Summary of the previous studies in chronological order</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center">Reference</th>
<th align="center">Proposed technique</th>
<th align="center">Detection type</th>
<th align="center">Dataset</th>
</tr>
</thead>
<tbody>
<tr>
<td>Sun and Kise (2011) [<xref ref-type="bibr" rid="ref-4">4</xref>]</td>
<td>ROI detection</td>
<td>Face detection</td>
<td>Dataset created</td>
</tr>
<tr>
<td>Yu and Seah (2011) [<xref ref-type="bibr" rid="ref-9">9</xref>]</td>
<td>FDM</td>
<td>Character edge detection</td>
<td>Cartoon dataset (Tom and Jerry)</td>
</tr>
<tr>
<td>Takayama et al. (2012) [<xref ref-type="bibr" rid="ref-10">10</xref>]</td>
<td>Object detection</td>
<td>Face detection</td>
<td>Lucky Paradise, TV TOKYO, MBS</td>
</tr>
<tr>
<td>Khan et al. (2012) [<xref ref-type="bibr" rid="ref-7">7</xref>]</td>
<td>HOG</td>
<td>Character detection</td>
<td>Dataset created</td>
</tr>
<tr>
<td>Zhang et al. (2013) [<xref ref-type="bibr" rid="ref-5">5</xref>]</td>
<td>SSC</td>
<td>Character detection</td>
<td>Dataset created</td>
</tr>
<tr>
<td>Ho et al. (2013) [<xref ref-type="bibr" rid="ref-11">11</xref>]</td>
<td>Adjacency graphs</td>
<td>Recurring character detection</td>
<td>Cosmozone</td>
</tr>
<tr>
<td>Sun et al. (2013) [<xref ref-type="bibr" rid="ref-6">6</xref>]</td>
<td>SIFT</td>
<td>Character detection</td>
<td>eBDtheque dataset</td>
</tr>
<tr>
<td>Iwata et al. 2014) [<xref ref-type="bibr" rid="ref-13">13</xref>]</td>
<td>HOG</td>
<td>Character detection</td>
<td>Manga dataset</td>
</tr>
<tr>
<td>Rigaud et al. (2014) [<xref ref-type="bibr" rid="ref-8">8</xref>]</td>
<td>Colour-based object detection</td>
<td>Character detection</td>
<td>eBDtheque dataset</td>
</tr>
<tr>
<td>Le et al. (2015) [<xref ref-type="bibr" rid="ref-12">12</xref>]</td>
<td>RAGs, FSM</td>
<td>Character detection</td>
<td>Cosmozone</td>
</tr>
<tr>
<td>Matsui et al. (2017) [<xref ref-type="bibr" rid="ref-21">21</xref>]</td>
<td>ANN</td>
<td>Image identification</td>
<td>Manga109</td>
</tr>
<tr>
<td>Qin et al. (2017) [<xref ref-type="bibr" rid="ref-14">14</xref>]</td>
<td>R-CNN</td>
<td>Face detection</td>
<td>JC2463, AEC912</td>
</tr>
<tr>
<td>Nguyen et al. (2017) [<xref ref-type="bibr" rid="ref-15">15</xref>]</td>
<td>YOLOv2</td>
<td>Character detection</td>
<td>Sequencity612, Fahad18, Ho42, Sun60</td>
</tr>
<tr>
<td>Yanagisawa et al. (2018) [<xref ref-type="bibr" rid="ref-16">16</xref>]</td>
<td>Fast R-CNN, Faster R-CNN</td>
<td>Panel layout, speech bubble, character face and text detection</td>
<td>Manga109</td>
</tr>
<tr>
<td>Nguyen et al. (2018) [<xref ref-type="bibr" rid="ref-24">24</xref>]</td>
<td>SSD, Faster R-CNN, YOLOv2</td>
<td>Panels, balloons, text character face and character detection</td>
<td>DCM772, eBDtheque, Fahad18</td>
</tr>
<tr>
<td>Nguyen et al. (2019) [<xref ref-type="bibr" rid="ref-25">25</xref>]</td>
<td>CNN, Mask R-CNN ComicMTL</td>
<td>Panels, balloons and identification of characters</td>
<td>DCM772, eBDtheque</td>
</tr>
<tr>
<td>Chu and Li (2019) [<xref ref-type="bibr" rid="ref-17">17</xref>]</td>
<td>CNN</td>
<td>Face detection</td>
<td>Manga FaceNet</td>
</tr>
<tr>
<td>Dutta and Biswas (2019) [<xref ref-type="bibr" rid="ref-18">18</xref>]</td>
<td>YOLO, CNN</td>
<td>Detection of panels and characters</td>
<td>BCBId</td>
</tr>
<tr>
<td>Aizawa et al. (2020) [<xref ref-type="bibr" rid="ref-22">22</xref>]</td>
<td>GAN, EOH, SSD</td>
<td>Panels, texts, character faces and character bodies detection</td>
<td>Manga109, COMICS</td>
</tr>
<tr>
<td>Mashood Nasir et al. (2020) [<xref ref-type="bibr" rid="ref-19">19</xref>]</td>
<td>CNN</td>
<td>Character detection</td>
<td>Product images dataset</td>
</tr>
<tr>
<td>Kim et al. (2020) [<xref ref-type="bibr" rid="ref-2">2</xref>]</td>
<td>AA-R-CNN</td>
<td>Character detection</td>
<td>Images Disney Enterprises</td>
</tr>
<tr>
<td>Kim et al. (2021) [<xref ref-type="bibr" rid="ref-20">20</xref>]</td>
<td>Pose mapping method</td>
<td>Face detection</td>
<td>AnimeCeleb</td>
</tr>
<tr>
<td>Nir et al. (2022) [<xref ref-type="bibr" rid="ref-26">26</xref>]</td>
<td>CNN, CAST</td>
<td>Character detection</td>
<td>Manga</td>
</tr>
<tr>
<td>Jain et al. (2022) [<xref ref-type="bibr" rid="ref-27">27</xref>]</td>
<td>Mask R-CNN, CNN</td>
<td>Character detection and emotion classification</td>
<td>Cartoon dataset (Tom and Jerry)</td>
</tr>
<tr>
<td>Rios et al. (2022) [<xref ref-type="bibr" rid="ref-30">30</xref>]</td>
<td>CNN, Self-attention based ViT</td>
<td>Anime character recognition</td>
<td>DanbooruAnimeFaces</td>
</tr>
<tr>
<td>Agrawal and Tripathi (2023) [<xref ref-type="bibr" rid="ref-28">28</xref>]</td>
<td>Mask R-CNN, CNN</td>
<td>Character detection and emotion classification</td>
<td>Cartoon dataset (Tom and Jerry)</td>
</tr>
<tr>
<td>Yi et al. (2023) [<xref ref-type="bibr" rid="ref-29">29</xref>]</td>
<td>ViT, BERT</td>
<td>Anime character recognition</td>
<td>Dataset created</td>
</tr>
<tr>
<td>Qi et al. (2024) [<xref ref-type="bibr" rid="ref-23">23</xref>]</td>
<td>MP-YOLO</td>
<td>Character detection</td>
<td>Dataset created</td>
</tr>
<tr>
<td>Li et al. (2025) [<xref ref-type="bibr" rid="ref-1">1</xref>]</td>
<td>Efficient channel attention</td>
<td>Character recognition</td>
<td>Dataset created from iCartoonFace</td>
</tr>
</tbody>
</table>
</table-wrap>
<sec id="s2_1">
<label>2.1</label>
<title>Traditional and Feature-Based Approaches</title>
<p>Early research in character recognition relied heavily on handcrafted visual features, focusing on shape, colour, texture, and structural relationships to identify characters. These methods often addressed specific tasks like copy detection, similarity estimation, or retrieval in static images. For instance, Sun and Kise (2011) addressed copyright protection by proposing a method for detecting partial copies in monochrome line drawings, using a cascade classifier for region-of-interest detection followed by feature matching [<xref ref-type="bibr" rid="ref-4">4</xref>]. Similarly, Zhang et al. (2013) focused on detecting 2D cartoon characters to prevent pirate uploads, designing a local shape feature named Scalable-Shape Context (SSC) combined with a Hough-voting scheme [<xref ref-type="bibr" rid="ref-5">5</xref>]. Another feature-based approach for copyright protection in comics was presented by Sun et al. (2013), who utilised the Scale-invariant Feature Transform (SIFT) identifier to match characters despite various transformations [<xref ref-type="bibr" rid="ref-6">6</xref>].</p>
<p>Other works focused on colour and similarity. Khan et al. (2012) demonstrated that using explicit colour attributes, in combination with traditional shape features, could significantly improve object detection performance, which they validated on a new cartoon character dataset where colour was a pivotal feature [<xref ref-type="bibr" rid="ref-7">7</xref>]. Rigaud et al. (2014) also presented a colour-based approach for retrieving comic characters, using a colour palette and content-based drawing retrieval [<xref ref-type="bibr" rid="ref-8">8</xref>]. In a different vein, Yu and Seah (2011) proposed a novel method called fuzzy diffusion maps (FDM) to estimate cartoon similarity by learning a diffusion distance in a low-dimensional manifold, showing robustness in recognition and clustering tasks [<xref ref-type="bibr" rid="ref-9">9</xref>]. Face detection was another key area, with Takayama et al. (2012) proposing methods for both face detection and recognition of cartoon characters by specifically considering their unique facial features compared to real human faces [<xref ref-type="bibr" rid="ref-10">10</xref>].</p>
<p>Structural and graph-based methods were also explored to capture relationships within comic book panels. Ho et al. (2013) introduced a method to detect recurrent characters in comics by representing panels as attributed adjacency graphs and using inexact graph matching to find similar subgraph structures [<xref ref-type="bibr" rid="ref-11">11</xref>]. Building on this, Le et al. (2015) proposed a retrieval system based on representing comic pages as multilayer graphs and using Frequent Subgraph Mining (FSM) to identify and rank images [<xref ref-type="bibr" rid="ref-12">12</xref>]. Finally, Iwata et al. (2014) investigated manga character retrieval by modifying and applying existing methods to the specific domain of Japanese comics [<xref ref-type="bibr" rid="ref-13">13</xref>].</p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Deep Learning for Character Recognition</title>
<p>With the resurgence of deep learning, the field saw a paradigm shift towards using Convolutional Neural Networks (CNNs) to automatically learn features, significantly outperforming traditional methods. Researchers began applying and adapting state-of-the-art object detection models for cartoon and comic analysis. Qin et al. (2017) proposed a Faster R-CNN-based method for face detection in comics, empirically finding that a sigmoid classifier outperformed softmax for this task [<xref ref-type="bibr" rid="ref-14">14</xref>]. Nguyen et al. (2017) also leveraged deep learning for comic character detection, training modern object detection networks on a proposed dataset and establishing a strong baseline for future work [<xref ref-type="bibr" rid="ref-15">15</xref>]. The comparative effectiveness of different deep learning models was explored by Yanagisawa et al. (2018), who studied object detection in manga and found that Faster R-CNN was more effective for character faces, while Fast R-CNN was better suited for panel layouts [<xref ref-type="bibr" rid="ref-16">16</xref>].</p>
<p>This trend continued with more specialised applications. Chu and Li (2019) proposed a deep neural network that fused global and local information to detect manga faces of various appearances [<xref ref-type="bibr" rid="ref-17">17</xref>]. Dutta and Biswas (2019) developed a YOLO and CNN-based architecture to extract both panels and characters from Bengali comic book images, demonstrating its effectiveness on their custom BCBId dataset as well as other public datasets [<xref ref-type="bibr" rid="ref-18">18</xref>]. Mashood Nasir et al. (2020) proposed a hybrid 17-layer deep CNN for classifying superhero fashion products, achieving a high accuracy of 97.9% and showcasing the power of custom deep learning models [<xref ref-type="bibr" rid="ref-19">19</xref>]. More recently, Kim et al. (2021) addressed the challenge of detecting characters in animated movies by introducing an animation adaptive region-based CNN that adds a hierarchical adaptation module to the Faster R-CNN model to handle diverse animation styles [<xref ref-type="bibr" rid="ref-20">20</xref>].</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Dataset Creation and Benchmarking</title>
<p>The advancement of deep learning is intrinsically linked to the availability of large, high-quality datasets. Recognising a gap in the animation and comics domain, several research efforts have focused on creating and annotating new public benchmarks. Matsui et al. (2017) made a significant contribution by building and releasing Manga109, a large dataset of 109 manga books with over 21,000 pages, to facilitate research in manga retrieval and analysis [<xref ref-type="bibr" rid="ref-21">21</xref>]. Aizawa et al. (2020) further detailed the Manga109 dataset, providing extensive annotations for frames, text, faces, and bodies to support a wide range of multimedia applications [<xref ref-type="bibr" rid="ref-22">22</xref>]. Addressing the lack of data for animated faces, Kim et al. (2020) introduced AnimeCeleb, a large-scale animation face dataset generated using controllable 3D synthetic models, designed to boost research in tasks like head reenactment [<xref ref-type="bibr" rid="ref-2">2</xref>]. More recently, Qi et al. (2024) introduced CCDaS, a large benchmark dataset for cartoon character detection in practical scenarios like merchandise and advertising, and proposed an accompanying multi-path YOLO model (MP-YOLO) to handle multi-scale and facially similar objects [<xref ref-type="bibr" rid="ref-23">23</xref>].</p>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Advanced and Specialised Deep Learning Applications</title>
<p>Building on foundational deep learning models, recent research has explored more sophisticated techniques to tackle nuanced challenges in character analysis. Multi-task learning has emerged as a way to create more efficient and holistic models. Nguyen et al. (2018, 2019) extended their work into digital comics indexing and proposed Comic MTL, a multi-task learning model capable of handling panel detection, character detection, and even the relationship between speech balloons and speakers simultaneously [<xref ref-type="bibr" rid="ref-24">24</xref>,<xref ref-type="bibr" rid="ref-25">25</xref>].</p>
<p>Other advanced paradigms include self-supervised learning and style adaptation. Nir et al. (2022) presented CAST, a method using self-supervision and multi-object tracking to learn refined semantic representations of characters within a specific animation style, allowing for better clustering of characters with diverse appearances [<xref ref-type="bibr" rid="ref-26">26</xref>]. Li et al. (2025) proposed a framework for character recognition that fuses portrait style, using an ECA-based residual attention module to improve feature learning and a style-transfer mechanism to create a simulated dataset for copyright protection [<xref ref-type="bibr" rid="ref-1">1</xref>].</p>
<p>The field has also expanded to include emotion and multimodal analysis. Jain et al. (2022) and Agrawal et al. (2023) both tackled the problem of cartoon emotion recognition, developing integrated deep neural network approaches to identify characters and classify their emotions from images, using datasets based on the &#x201C;Tom and Jerry&#x201D; cartoons [<xref ref-type="bibr" rid="ref-27">27</xref>,<xref ref-type="bibr" rid="ref-28">28</xref>]. Looking beyond visual data, Yi et al. (2023) proposed a multimodal deep learning network for anime character identification and tag prediction, exploiting both image and text data and introducing curriculum learning to handle missing text annotations [<xref ref-type="bibr" rid="ref-29">29</xref>]. Finally, looking towards next-generation architectures, Rios et al. (2022) studied the problem of anime character recognition using Vision Transformers (ViTs), proposing a novel Intermediate Features Aggregation head to improve classification accuracy [<xref ref-type="bibr" rid="ref-30">30</xref>].</p>
</sec>
<sec id="s2_5">
<label>2.5</label>
<title>Limitations of Existing Work and Motivation</title>
<p>Despite these significant contributions, a review of the existing literature reveals several gaps that motivate our current study. First, a notable focus exists on specific media formats, particularly Japanese manga (e.g., Manga109) and static comic books. These formats possess distinct visual aesthetics, such as monochrome art and structured panel layouts, which differ considerably from the dynamic, full-colour animation style addressed in our work. This focus limits the direct applicability and generalisability of findings to the diverse world of animated series. Second, much of the prior research has concentrated on tasks such as face detection or single-character object detection. Less emphasis has been placed on the more complex scenario of multi-label character classification from unstructured video frames, where multiple characters frequently interact within a single scene. Finally, while intra-class variance is an implicit challenge in all character recognition, few studies have been built around datasets specifically designed to test model robustness against the wide range of artistic and narrative-driven transformations (e.g., costume and style changes) found in a long-running animated series. These gaps&#x2014;the focus on specific artistic styles, the prevalence of static image analysis, and the limited attention to multi-label scenarios&#x2014;underscore the need for new, diverse datasets and benchmark studies. Our work directly addresses these limitations by introducing the Kral Sakir dataset, derived from video content, and providing a comprehensive analysis focused specifically on the multi-label character classification problem under high visual variance.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Materials and Methods</title>
<sec id="s3_1">
<label>3.1</label>
<title>Proposed Dataset: Kral Sakir</title>
<p>We constructed the novel Kral Sakir dataset for this study. This dataset features four principal characters from the animated series Kral Sakir streaming on Turkish channels: namely, Sakir, Remzi, Necati, and Canan. For the data collection procedure, we extracted frames from official series content available via YouTube [<xref ref-type="bibr" rid="ref-31">31</xref>]. We executed frame capture using VLC Player (VideoLAN Client). We then applied a manual curation and labelling process to each image to guarantee precise character representation, accommodating both single-character and multi-character scenes. The final dataset contains 16,725 colour images, with varying distributions across the four character classes and multi-label combinations. <xref ref-type="fig" rid="fig-1">Fig. 1</xref> displays the four main characters and <xref ref-type="table" rid="table-2">Table 2</xref> outlines the statistical distributions per class.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Four principal characters from Kral Sakir animated series</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67840-fig-1.tif"/>
</fig><table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Distribution of images across the four character classes and multi-label combinations</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Class/Combination</th>
<th>Number of images</th>
</tr>
</thead>
<tbody>
<tr>
<td>Sakir</td>
<td>1945</td>
</tr>
<tr>
<td>Remzi</td>
<td>1505</td>
</tr>
<tr>
<td>Necati</td>
<td>1686</td>
</tr>
<tr>
<td>Canan</td>
<td>1512</td>
</tr>
<tr>
<td>Sakir-Remzi</td>
<td>1557</td>
</tr>
<tr>
<td>Sakir-Necati</td>
<td>1129</td>
</tr>
<tr>
<td>Sakir-Canan</td>
<td>1173</td>
</tr>
<tr>
<td>Remzi-Necati</td>
<td>1550</td>
</tr>
<tr>
<td>Remzi-Canan</td>
<td>362</td>
</tr>
<tr>
<td>Necati-Canan</td>
<td>558</td>
</tr>
<tr>
<td>Sakir-Remzi-Necati</td>
<td>1958</td>
</tr>
<tr>
<td>Sakir-Remzi-Canan</td>
<td>470</td>
</tr>
<tr>
<td>Sakir-Necati-Canan</td>
<td>347</td>
</tr>
<tr>
<td>Remzi-Necati-Canan</td>
<td>222</td>
</tr>
<tr>
<td>Sakir-Remzi-Necati-Canan</td>
<td>751</td>
</tr>
<tr>
<td>Total</td>
<td>16,725</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>A key challenge in automated cartoon character recognition stems from the significant visual variability inherent in character portrayal, often driven by artistic licence and narrative progression, unlike the more constrained variations typically seen in real-world object classes. These transformations can manifest as substantial changes in attire, colour schemes, accessories, depicted age, and even fundamental drawing style or perspective between appearances. While the human visual system demonstrates remarkable robustness, readily identifying characters across these disparate renditions by focusing on core visual cues and abstract concepts, replicating this capability with machine learning presents a considerable difficulty. Algorithms learning features directly from pixel data must contend with this high degree of intra-class variance, where consistent low-level visual patterns may be less prevalent or frequently altered, making robust recognition a non-trivial task. <xref ref-type="fig" rid="fig-2">Figs. 2</xref>&#x2013;<xref ref-type="fig" rid="fig-5">5</xref> illustrate this challenge by presenting multiple, visually distinct examples of the main characters as found within our dataset.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Selected examples demonstrating the visual variability of the main character Sakir</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67840-fig-2.tif"/>
</fig><fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Selected examples demonstrating the visual variability of the character Remzi</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67840-fig-3.tif"/>
</fig><fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Selected examples demonstrating the visual variability of the character Necati</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67840-fig-4.tif"/>
</fig><fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Selected examples demonstrating the visual variability of the character Canan</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67840-fig-5.tif"/>
</fig>
<p>In summary, the novelty and relevance of the Kral Sakir dataset lie in several key aspects as shown in <xref ref-type="table" rid="table-3">Table 3</xref>. Its focus on multi-label classification from dynamic video content addresses a more complex and realistic scenario than many existing benchmarks. The dataset is specifically curated to feature high diversity in character portrayal, making it an ideal resource for evaluating model robustness against intra-class variance. Furthermore, the high annotation quality, ensured through a manual curation and labelling process, provides a reliable foundation for training and testing. These properties make the dataset not only a challenging benchmark for academic research but also a valuable asset for developing and validating systems with potential for real-world applications like automated content tagging and analysis. To facilitate reproducibility and encourage further research in this area, the novel Kral Sakir dataset is publicly accessible in our repository at <ext-link ext-link-type="uri" xlink:href="https://github.com/candantumer/kral_sakir">https://github.com/candantumer/kral_sakir</ext-link> (accessed on 18 August 2025).</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Key properties and novelty of the Kral Sakir dataset</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center">Feature</th>
<th align="center">Description</th>
<th align="center">Relevance and novelty</th>
</tr>
</thead>
<tbody>
<tr>
<td>Task formulation</td>
<td>Multi-label classification.</td>
<td>Addresses the realistic scenario of multiple characters appearing in a single scene, a less-explored area compared to single-character detection.</td>
</tr>
<tr>
<td>Data source</td>
<td>Video frames from a long-running animated TV series.</td>
<td>Captures dynamic scene compositions and a wide range of character interactions, unlike datasets derived from static media like manga or comics.</td>
</tr>
<tr>
<td>Visual diversity</td>
<td>High intra-class variance due to changes in attire, perspective, expression, and artistic style driven by narrative.</td>
<td>Specifically designed to test model robustness and generalisation, a key challenge in character recognition that is a primary focus of this work.</td>
</tr>
<tr>
<td>Annotation quality</td>
<td>Each of the 16,725 images was manually curated and labelled by human annotators to ensure precise character representation.</td>
<td>A sequence-based split was implemented for training/testing. Guarantees high-quality ground truth for reliable model training and evaluation. The splitting protocol ensures methodologically sound performance assessment by preventing data leakage.</td>
</tr>
<tr>
<td>Potential applications</td>
<td>Benchmarking classification models, developing content-based video retrieval systems, training models for automated content analysis and tagging. The dataset&#x2019;s size is manageable for academic research.</td>
<td>Provides a foundational resource for a range of tasks beyond simple classification, fostering further research in automated media analysis. The model performance suggests potential for near-real-time applications.</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Deep Learning Architectures</title>
<sec id="s3_2_1">
<label>3.2.1</label>
<title>Deep Learning</title>
<p>Deep Learning (DL) constitutes a specialised subfield within machine learning, drawing inspiration for its algorithms from the complex network of connections observed in the human brain, known as Artificial Neural Networks (ANNs) [<xref ref-type="bibr" rid="ref-32">32</xref>,<xref ref-type="bibr" rid="ref-33">33</xref>]. ANNs employ multiple layers of interconnected processing units (neurons) designed to learn representations of data through hierarchical feature extraction. This capacity for learning intricate patterns has enabled the development of diverse DL architectures which have yielded significant breakthroughs across numerous domains, including computer vision, automated speech recognition, and natural language processing. Among these architectures, Convolutional Neural Networks (CNNs) are particularly prevalent and effective for tasks involving grid-like data, such as images [<xref ref-type="bibr" rid="ref-34">34</xref>]. CNNs utilise specific operations, like convolution and pooling, to automatically learn spatial hierarchies of features directly from pixel data, making them exceptionally well-suited for image feature extraction necessary for classification, detection, and other visual understanding tasks.</p>
</sec>
<sec id="s3_2_2">
<label>3.2.2</label>
<title>Transfer Learning</title>
<p>Transfer learning is a technique in machine learning designed to expedite the development and improve the performance of models on a target task by reusing knowledge from a related source task where a model has already been trained [<xref ref-type="bibr" rid="ref-35">35</xref>]. Operationally, this involves adopting the network architecture and initialising the weights of the target model using the parameters learned by the &#x201C;pretrained&#x201D; source model. This initialisation leverages the hierarchical features extracted by the source model, assuming their relevance to the target domain. The utility of transfer learning is most pronounced when faced with data limitations for the target problem, as training deep models from scratch demands substantial data to learn meaningful representations effectively. By inheriting a robust set of initial parameters, the learning process on the target task is essentially &#x201C;bootstrapped,&#x201D; often requiring fewer training epochs to converge and thereby reducing the associated computational burden and time investment compared to random weight initialisation. Common strategies include using the pretrained model as a fixed feature extractor or fine-tuning some or all of its layers on the target dataset [<xref ref-type="bibr" rid="ref-36">36</xref>].</p>
</sec>
<sec id="s3_2_3">
<label>3.2.3</label>
<title>Baseline CNN Model</title>
<p>For the task of multi-label image classification, we designed and implemented a CNN architecture as a baseline. This model was intentionally designed with a relatively simple architecture to serve two purposes: it provides a computationally efficient benchmark and allows for a clear, direct assessment of the performance gains attributable to the pretrained features of the transfer learning models. The core of the model consists of a feature extraction module comprising four sequential convolutional layers with filter depths of 16, 32, 64, and 64, respectively. Standard Rectified Linear Unit (ReLU) activation functions were applied after each convolution to introduce non-linearity. To progressively reduce spatial dimensions and enhance feature robustness, three Max Pooling layers were strategically interspersed within the convolutional blocks. Following the final feature extraction layer, a Global Average Pooling (GAP) layer was utilised. This GAP layer aggregates the spatial information across each feature map, producing a fixed-length feature vector irrespective of input dimensions. This feature vector subsequently feeds into the final dense output layer, configured with four units and employing a sigmoid activation function. The use of sigmoid activation allows each output unit to predict the independent probability of the corresponding label being present, appropriately addressing the multi-label nature of the classification problem.</p>
</sec>
<sec id="s3_2_4">
<label>3.2.4</label>
<title>Pretrained Models Used</title>
<p>The remarkable success of deep learning in computer vision owes much to the development and widespread adoption of powerful CNN architectures. Many of the most impactful models are made available in a &#x201C;pretrained&#x201D; state, having been extensively trained on large-scale image datasets like ImageNet [<xref ref-type="bibr" rid="ref-37">37</xref>]. This pretraining equips them with robust feature extraction capabilities, capturing hierarchical visual patterns that are often transferable to new tasks. Generic architecture of the proposed baseline CNN model and the pretrained models is presented in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>. This research utilised the following pretrained models, all employing weights derived from training on the ImageNet dataset:</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Generic architecture of the proposed baseline CNN model and the pretrained models</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67840-fig-6.tif"/>
</fig>
<p><bold>ResNet (Residual Network):</bold> Introduced residual (skip) connections, enabling the effective training of extremely deep networks by allowing gradients to bypass layers and mitigate the vanishing gradient problem [<xref ref-type="bibr" rid="ref-38">38</xref>]. We used ResNet50V2, ResNet101V2, and ResNet152V2 models from ResNet family in this research.</p>
<p><bold>DenseNet (Densely Connected Network):</bold> Features dense connectivity where each layer receives inputs from all preceding layers within a block, promoting feature reuse, stronger gradient flow, and often greater parameter efficiency [<xref ref-type="bibr" rid="ref-39">39</xref>]. DenseNet121, DenseNet169, and DenseNet201 were the pretrained models we utilised from DenseNet family in our research.</p>
<p><bold>EfficientNet:</bold> Developed using a compound scaling method that systematically and optimally balances network depth, width, and image resolution to achieve high accuracy with significantly fewer parameters and computations compared to previous models [<xref ref-type="bibr" rid="ref-40">40</xref>]. EfficientNetB7 model from EfficientNet family was chosen for the research.</p>
<p><bold>MobileNet:</bold> Optimised specifically for mobile and embedded vision applications, employing depthwise separable convolutions to drastically reduce computational cost (FLOPs) and model size, enabling efficient on-device inference [<xref ref-type="bibr" rid="ref-41">41</xref>]. We employed MobileNet and MobileNetV2 models from MobileNet family in our research.</p>
<p><bold>VGG (Visual Geometry Group):</bold> Known for its simple and uniform architecture using stacked small (<inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mn>3</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>3</mml:mn></mml:math></inline-formula>) convolutional filters, demonstrating the effectiveness of increased network depth for performance, though generally less parameter-efficient than more modern architectures [<xref ref-type="bibr" rid="ref-42">42</xref>]. VGG16 and VGG19 were the models chosen from the VGG family in this research.</p>
</sec>
</sec>
<sec id="s3_3">
<label>3.3</label>
<title>Experimental Settings</title>
<sec id="s3_3_1">
<label>3.3.1</label>
<title>Dataset Loading and Preprocessing</title>
<p>We partitioned our dataset using an 80% allocation for training and 20% for testing. Recognising the potential for data leakage due to the sequential origin of the images from animations, we employed a sequence-based splitting protocol. This protocol ensured that frames derived from the same animation sequence were never present in both the training and testing sets simultaneously; each sequence was assigned entirely to one partition. This approach prevents overly optimistic results arising from testing on frames highly similar to training data and allows for a more valid assessment of the model&#x2019;s ability to generalise to unseen sequences, thereby ensuring the integrity of the experimental findings. Distribution of images across the four character classes and multi-label combinations after the split is shown in <xref ref-type="table" rid="table-4">Table 4</xref>.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Distribution of images across the four character classes and multi-label combinations</title>
</caption>
<table>
<colgroup>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th align="center">Class</th>
<th align="center">Number of images</th>
<th align="center">Number of sequences</th>
<th align="center">Train images (80%)</th>
<th align="center">Train sequences</th>
<th align="center">Test images (20%)</th>
<th align="center">Test sequences</th>
</tr>
</thead>
<tbody>
<tr>
<td>Sakir</td>
<td>1945</td>
<td>185</td>
<td>1556</td>
<td>148</td>
<td>389</td>
<td>37</td>
</tr>
<tr>
<td>Remzi</td>
<td>1505</td>
<td>140</td>
<td>1204</td>
<td>106</td>
<td>301</td>
<td>34</td>
</tr>
<tr>
<td>Necati</td>
<td>1686</td>
<td>128</td>
<td>1349</td>
<td>96</td>
<td>337</td>
<td>32</td>
</tr>
<tr>
<td>Canan</td>
<td>1512</td>
<td>117</td>
<td>1210</td>
<td>83</td>
<td>302</td>
<td>34</td>
</tr>
<tr>
<td>Sakir-Remzi</td>
<td>1557</td>
<td>127</td>
<td>1246</td>
<td>96</td>
<td>311</td>
<td>31</td>
</tr>
<tr>
<td>Sakir-Necati</td>
<td>1129</td>
<td>73</td>
<td>903</td>
<td>53</td>
<td>226</td>
<td>20</td>
</tr>
<tr>
<td>Sakir-Canan</td>
<td>1173</td>
<td>103</td>
<td>938</td>
<td>79</td>
<td>235</td>
<td>24</td>
</tr>
<tr>
<td>Remzi-Necati</td>
<td>1550</td>
<td>113</td>
<td>1240</td>
<td>76</td>
<td>310</td>
<td>37</td>
</tr>
<tr>
<td>Remzi-Canan</td>
<td>362</td>
<td>37</td>
<td>290</td>
<td>32</td>
<td>72</td>
<td>5</td>
</tr>
<tr>
<td>Necati-Canan</td>
<td>558</td>
<td>30</td>
<td>446</td>
<td>15</td>
<td>112</td>
<td>15</td>
</tr>
<tr>
<td>Sakir-Remzi-Necati</td>
<td>1958</td>
<td>136</td>
<td>1566</td>
<td>99</td>
<td>392</td>
<td>37</td>
</tr>
<tr>
<td>Sakir-Remzi-Canan</td>
<td>470</td>
<td>44</td>
<td>376</td>
<td>31</td>
<td>94</td>
<td>13</td>
</tr>
<tr>
<td>Sakir-Necati-Canan</td>
<td>347</td>
<td>32</td>
<td>278</td>
<td>26</td>
<td>69</td>
<td>6</td>
</tr>
<tr>
<td>Remzi-Necati-Canan</td>
<td>222</td>
<td>18</td>
<td>178</td>
<td>12</td>
<td>44</td>
<td>6</td>
</tr>
<tr>
<td>Sakir-Remzi-Necati-Canan</td>
<td>751</td>
<td>69</td>
<td>601</td>
<td>53</td>
<td>150</td>
<td>16</td>
</tr>
<tr>
<td>Total</td>
<td>16,725</td>
<td>1352</td>
<td>13,381</td>
<td>1005</td>
<td>3344</td>
<td>347</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The images within our original dataset possess native dimensions of <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mn>320</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>180</mml:mn></mml:math></inline-formula> pixels. However, to ensure compatibility with the selected pretrained architectures for transfer learning, a resizing preprocessing step was necessary during both model training and evaluation phases. Each image was programmatically resized to match the standard input dimensions required by the specific model being utilised&#x2014;either <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mn>224</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>224</mml:mn></mml:math></inline-formula> or <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mn>299</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>299</mml:mn></mml:math></inline-formula> pixels. This resizing, along with other requisite normalization procedures tailored to each architecture, was performed using the respective preprocessing functions integrated within the Keras library.</p>
</sec>
<sec id="s3_3_2">
<label>3.3.2</label>
<title>Model Training</title>
<p>All experiments presented in this study were conducted within the Kaggle Notebook environment. The computational tasks, primarily model training and evaluation, were accelerated using an NVIDIA P100 GPU. Our software stack was built upon Python version 3.11, with Keras version 3.8 and TensorFlow version 2.18 (with CUDA support) serving as the core deep learning framework.</p>
<p>Outlined below is the training methodology for each evaluated model, detailing the hyperparameters used for both the network architectures and the optimisation process.</p>
<p>For optimisation, we employed the Adam optimiser [<xref ref-type="bibr" rid="ref-43">43</xref>], recognised for its adaptive learning rate capabilities and effectiveness in training deep neural networks. We initialised Adam with a relatively conservative learning rate of <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>4</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula>, a value often suitable for fine-tuning tasks to prevent disruption of the learned ImageNet features.</p>
<p>Binary Cross-Entropy was selected as the loss function, directly aligning with the multi-label classification objective where each output neuron, activated by a sigmoid function, predicts the independent probability of a specific class label. To regularise the custom classification head added after the base model&#x2019;s feature extraction layers (which consisted of GAP, a 128-unit ReLU Dense layer, and the final output layer), we incorporated a Dropout layer with a rate of 0.1 between the intermediate Dense layer and the final output layer. This helps mitigate overfitting by randomly setting a fraction of neuron activations to zero during training.</p>
<p>To further control the training dynamics and prevent overfitting, we integrated two key Keras callbacks. <monospace>ReduceLROnPlateau</monospace> was used to adaptively adjust the learning rate during training. This callback monitored the validation loss; if no improvement was observed for a patience period of 3 epochs, the learning rate was reduced by a factor of 0.015. This allows the model to take smaller steps towards the minimum of the loss function when progress stalls, with a defined minimum learning rate floor of <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:msup><mml:mn>10</mml:mn><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>8</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> to prevent excessively small updates [<xref ref-type="bibr" rid="ref-44">44</xref>].</p>
<p>Additionally, <monospace>EarlyStopping</monospace> served as an important mechanism against overfitting, a common challenge especially when adapting models to potentially smaller, specialised datasets. This callback also monitored the validation loss. If the validation loss failed to improve for 7 consecutive epochs, the training process was automatically terminated. Importantly, we configured <monospace>EarlyStopping</monospace> with <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>b</mml:mi><mml:mi>e</mml:mi><mml:mi>s</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">&#x005F;</mml:mi><mml:mi>w</mml:mi><mml:mi>e</mml:mi><mml:mi>i</mml:mi><mml:mi>g</mml:mi><mml:mi>h</mml:mi><mml:mi>t</mml:mi><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mi>T</mml:mi><mml:mi>r</mml:mi><mml:mi>u</mml:mi><mml:mi>e</mml:mi></mml:math></inline-formula>. This ensures that upon termination (or normal completion), the model retains the weights from the epoch that yielded the lowest validation loss, effectively capturing the model state with the best generalisation performance observed during training [<xref ref-type="bibr" rid="ref-45">45</xref>].</p>
<p>Each model variant was trained for a maximum of 50 epochs under this configuration, using the designated training and validation splits. The best performing weights for each architecture, as determined by the Early Stopping callback, were saved for subsequent analysis and comparison.</p>
</sec>
<sec id="s3_3_3">
<label>3.3.3</label>
<title>Model Evaluation Metrics</title>
<p>The performance of our models on the multi-label cartoon character classification task was evaluated using a suite of standard metrics. Specifically, we utilised Accuracy, Precision, Recall, F1-Score, and Area Under the ROC Curve (AUC), appropriately adapted for the multi-label scenario, to comprehensively assess model effectiveness. This section describes the computation details of these metrics for multi-label classification.</p>
<p>Let <italic>N</italic> be the total number of samples and <italic>L</italic> be the total number of labels. Let <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2208;</mml:mo><mml:mo fence="false" stretchy="false">{</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo fence="false" stretchy="false">}</mml:mo></mml:math></inline-formula> be the true binary label for sample <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>i</mml:mi></mml:math></inline-formula> and label <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>l</mml:mi></mml:math></inline-formula>. Let <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:msub><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> be the predicted probability for sample <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi>i</mml:mi></mml:math></inline-formula> and label <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>l</mml:mi></mml:math></inline-formula>. The binarised prediction <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is obtained by thresholding as in <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>.
<disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mtable columnalign="left left" rowspacing=".2em" columnspacing="1em" displaystyle="false"><mml:mtr><mml:mtd><mml:mn>1</mml:mn></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if</mml:mtext></mml:mrow><mml:mtext>&#xA0;</mml:mtext><mml:msub><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>&#x003E;</mml:mo><mml:mn>0.5</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn>0</mml:mn></mml:mtd><mml:mtd><mml:mrow><mml:mtext>if&#xA0;</mml:mtext></mml:mrow><mml:mtext>&#xA0;</mml:mtext><mml:msub><mml:mrow><mml:mover><mml:mi>p</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2264;</mml:mo><mml:mn>0.5</mml:mn></mml:mtd></mml:mtr></mml:mtable><mml:mo fence="true" stretchy="true" symmetric="true"></mml:mo></mml:mrow></mml:math></disp-formula></p>
<p>From testing a classifier, we derive counts of True Positives (TP), False Positives (FP), True Negatives (TN), and False Negatives (FN), calculated as in <xref ref-type="disp-formula" rid="eqn-2">Eqs. (2)</xref>&#x2013;<xref ref-type="disp-formula" rid="eqn-5">(5)</xref>, respectively. These fundamental counts enable the computation of multi-label performance metrics including Accuracy, Precision, Recall, F1-Score, and AUC.</p>
<p><disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mi mathvariant="double-struck">I</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2227;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mi mathvariant="double-struck">I</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>&#x2227;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mtable columnalign="right left right left right left right left right left right left" rowspacing="3pt" columnspacing="0em 2em 0em 2em 0em 2em 0em 2em 0em 2em 0em" displaystyle="true"><mml:mtr><mml:mtd /><mml:mtd><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mi mathvariant="double-struck">I</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo>&#x2227;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo>=</mml:mo><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mi mathvariant="double-struck">I</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2227;</mml:mo><mml:msub><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>0</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:math></disp-formula>where <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mrow><mml:mi mathvariant="double-struck">I</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mo>&#x22C5;</mml:mo><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> is the indicator function. From these counts, multi-label Accuracy, Precision, Recall, F1-Score, and AUC are calculated as in <xref ref-type="disp-formula" rid="eqn-6">Eqs. (6)</xref>&#x2013;<xref ref-type="disp-formula" rid="eqn-10">(10)</xref>, respectively.
<disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mi>A</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mi>u</mml:mi><mml:mi>r</mml:mi><mml:mi>a</mml:mi><mml:mi>c</mml:mi><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mi>F</mml:mi><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mi>S</mml:mi><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mo>=</mml:mo><mml:mn>2</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mfrac><mml:mrow><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>i</mml:mi><mml:mi>s</mml:mi><mml:mi>i</mml:mi><mml:mi>o</mml:mi><mml:mi>n</mml:mi><mml:mo>+</mml:mo><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>c</mml:mi><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>l</mml:mi></mml:mrow></mml:mfrac></mml:math></disp-formula>
<disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mi>A</mml:mi><mml:mi>U</mml:mi><mml:mi>C</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>L</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mo>&#x222B;</mml:mo><mml:mn>0</mml:mn><mml:mn>1</mml:mn></mml:msubsup><mml:mi>T</mml:mi><mml:msub><mml:mi>P</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mo stretchy="false">(</mml:mo><mml:mi>F</mml:mi><mml:msubsup><mml:mi>P</mml:mi><mml:mi>l</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msubsup><mml:mo stretchy="false">(</mml:mo><mml:mi>x</mml:mi><mml:mo stretchy="false">)</mml:mo><mml:mo stretchy="false">)</mml:mo><mml:mi>d</mml:mi><mml:mi>x</mml:mi></mml:math></disp-formula></p>
<p>To assess the statistical stability of our results, we calculated 95% confidence intervals for the primary performance metrics using the percentile bootstrap method. For each model, this involved generating 1000 bootstrap samples from the test set predictions and calculating the metric for each sample to create an empirical distribution from which the confidence intervals were derived.</p>
</sec>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Experimental Results and Discussion</title>
<p>The performance of the custom baseline CNN and various pretrained deep learning models, fine-tuned for the multi-label cartoon character classification task on our Kral Sakir dataset, is summarised in <xref ref-type="table" rid="table-5">Table 5</xref>. The metrics reported include loss, accuracy, Area Under the ROC Curve (AUC), precision, recall, and F1-Score, all evaluated on the hold-out test set. The performance scores reported here were obtained by following the comprehensive training methodology, data pipeline, and annotation process detailed in <xref ref-type="sec" rid="s3">Section 3</xref>.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Performance scores of all models, with 95% confidence intervals for key metrics</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col align="center"/>
<col/>
<col/>
<col/>
<col align="center"/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>Epochs</th>
<th>Loss</th>
<th align="center">Accuracy</th>
<th>AUC</th>
<th>Precision</th>
<th>Recall</th>
<th align="center">F1-Score</th>
</tr>
</thead>
<tbody>
<tr>
<td>Baseline CNN</td>
<td>66</td>
<td>0.1701</td>
<td>0.9579<break/> [0.9540, 0.9615]</td>
<td>0.9861</td>
<td>0.9661</td>
<td>0.9431</td>
<td>0.9545<break/> [0.9511, 0.9578]</td>
</tr>
<tr>
<td>MobileNet</td>
<td>9</td>
<td>0.0774</td>
<td>0.9830<break/> [0.9801, 0.9856]</td>
<td>0.9940</td>
<td>0.9879</td>
<td>0.9756</td>
<td>0.9817<break/> [0.9789, 0.9841]</td>
</tr>
<tr>
<td>MobileNetV2</td>
<td>21</td>
<td>0.0729</td>
<td>0.9862<break/> [0.9835, 0.9884]</td>
<td>0.9938</td>
<td>0.9900</td>
<td>0.9805</td>
<td>0.9852<break/> [0.9826, 0.9875]</td>
</tr>
<tr>
<td>VGG16</td>
<td>15</td>
<td>0.0610</td>
<td>0.9889<break/> [0.9867, 0.9908]</td>
<td>0.9951</td>
<td>0.9902</td>
<td><bold>0.9859</bold></td>
<td>0.9881<break/> [0.9859, 0.9900]</td>
</tr>
<tr>
<td>VGG19</td>
<td>18</td>
<td>0.0623</td>
<td>0.9865<break/> [0.9841, 0.9886]</td>
<td>0.9953</td>
<td>0.9883</td>
<td>0.9829</td>
<td>0.9856<break/> [0.9832, 0.9877]</td>
</tr>
<tr>
<td>ResNet50V2</td>
<td>9</td>
<td>0.0790</td>
<td>0.9831<break/> [0.9805, 0.9854]</td>
<td>0.9934</td>
<td>0.9861</td>
<td>0.9776</td>
<td>0.9819<break/> [0.9794, 0.9842]</td>
</tr>
<tr>
<td>ResNet101V2</td>
<td>9</td>
<td>0.0707</td>
<td>0.9839<break/> [0.9812, 0.9861]</td>
<td>0.9944</td>
<td>0.9904</td>
<td>0.9751</td>
<td>0.9827<break/> [0.9801, 0.9850]</td>
</tr>
<tr>
<td>ResNet152V2</td>
<td>11</td>
<td>0.0615</td>
<td>0.9858<break/> [0.9834, 0.9879]</td>
<td>0.9953</td>
<td>0.9890</td>
<td>0.9805</td>
<td>0.9848<break/> [0.9824, 0.9869]</td>
</tr>
<tr>
<td>EfficientNetB7</td>
<td>9</td>
<td>0.0599</td>
<td>0.9884<break/> [0.9861, 0.9903]</td>
<td>0.9954</td>
<td><bold>0.9915</bold></td>
<td>0.9837</td>
<td>0.9876<break/> [0.9854, 0.9895]</td>
</tr>
<tr>
<td>DenseNet121</td>
<td>9</td>
<td><bold>0.0484</bold></td>
<td><bold>0.9898 [0.9879, 0.9914]</bold></td>
<td><bold>0.9963</bold></td>
<td>0.9912</td>
<td><bold>0.9869</bold></td>
<td><bold>0.9890 [0.9871, 0.9907]</bold></td>
</tr>
<tr>
<td>DenseNet169</td>
<td>8</td>
<td><bold>0.0450</bold></td>
<td><bold>0.9894 [0.9873, 0.9911]</bold></td>
<td><bold>0.9970</bold></td>
<td><bold>0.9920</bold></td>
<td>0.9853</td>
<td><bold>0.9886 [0.9865, 0.9904]</bold></td>
</tr>
<tr>
<td>DenseNet201</td>
<td>8</td>
<td><bold>0.0487</bold></td>
<td><bold>0.9895 [0.9875, 0.9912]</bold></td>
<td><bold>0.9960</bold></td>
<td><bold>0.9913</bold></td>
<td><bold>0.9863</bold></td>
<td><bold>0.9888 [0.9868, 0.9906]</bold></td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-5fn1" fn-type="other">
<p>Note: Top three values in each column are marked in boldface.</p>
</fn>
</table-wrap-foot>
</table-wrap>
 
<sec id="s4_1">
<label>4.1</label>
<title>Overall Performance and Impact of Transfer Learning</title>
<p>It is important to interpret the following results in the context of this study&#x2019;s primary objective. As outlined in the introduction, our goal is not to propose a novel model architecture, but rather to conduct a rigorous benchmark of existing state-of-the-art algorithms on our new dataset. Therefore, the discussion focuses on the comparative performance and the insights gained from applying these established techniques to the unique challenges of multi-label cartoon character recognition</p>
<p>A prominent observation is the significant performance advantage demonstrated by all transfer learning models when compared to the custom baseline CNN. The baseline CNN, while achieving a respectable F1-Score of 0.9545 and an AUC of 0.9861, was consistently outperformed across all metrics by the pretrained architectures. For instance, the loss for the baseline model was 0.1701, whereas all transfer learning models achieved substantially lower loss values, generally below 0.08. This underscores the efficacy of leveraging features learned from large-scale datasets like ImageNet for specialised tasks such as cartoon character recognition, even when the domain (cartoons vs. natural images) differs. Furthermore, the pretrained models typically converged to optimal performance in significantly fewer epochs (most between 8&#x2013;18 epochs) compared to the baseline CNN which required 66 epochs, highlighting the efficiency gains from transfer learning.</p>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Top-Performing Models</title>
<p>The DenseNet family of models consistently emerged as the top performers. Specifically, DenseNet121 achieved the highest overall accuracy (0.9898) and F1-Score (0.9890), with a very low loss of 0.0484 and an excellent AUC of 0.9963. Closely following, DenseNet169 yielded the lowest loss (0.0450) and the highest AUC (0.9970) among all models, along with an F1-Score of 0.9886 and accuracy of 0.9894. DenseNet201 also demonstrated strong performance, with metrics nearly identical to its family members (e.g., F1-Score of 0.9888). These results suggest that the dense connectivity pattern, promoting feature reuse and gradient flow, is particularly well-suited for this classification task.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>Competitive Architectures</title>
<p>EfficientNetB7 and VGG16 also proved to be highly competitive. EfficientNetB7 achieved an F1-Score of 0.9876, an AUC of 0.9954, and an accuracy of 0.9884, with a low loss of 0.0599, placing it amongst the top-tier models. Notably, VGG16, despite being an older architecture, delivered a remarkable F1-Score of 0.9881 and an accuracy of 0.9889. Its performance was comparable to EfficientNetB7 and some DenseNet variants, though with a slightly higher loss (0.0610). VGG19 also performed well, with an F1-Score of 0.9856 and an AUC of 0.9953, slightly trailing VGG16.</p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>Performance of ResNet and MobileNet Families</title>
<p>The ResNet variants (ResNet50V2, ResNet101V2, and ResNet152V2) showed robust improvements over the baseline, with ResNet152V2 being the best performer in this family, achieving an F1-Score of 0.9848 and an AUC of 0.9953. While effective, they were generally slightly outperformed by the DenseNet, EfficientNet, and VGG models in terms of F1-Score and loss.</p>
<p>The MobileNet architectures (MobileNet and MobileNetV2), designed for efficiency, also surpassed the baseline. MobileNetV2, with an F1-Score of 0.9852, performed better than MobileNet (F1-Score 0.9817) and was competitive with some of the larger ResNet models. This indicates that even lightweight models can achieve high accuracy on this dataset, though the very top performance levels were reached by the more complex architectures.</p>
</sec>
<sec id="s4_5">
<label>4.5</label>
<title>Summary of Results</title>
<p>The experimental results clearly indicate that transfer learning with pretrained models offers substantial benefits for multi-label cartoon character recognition on the Kral Sakir dataset. The DenseNet architecture, particularly DenseNet121 and DenseNet169, exhibited superior performance, achieving the best balance of low loss, high accuracy, AUC, and F1-Score. EfficientNetB7 and VGG16 also demonstrated excellent capabilities. The significant improvement over the custom baseline CNN, coupled with faster convergence, validates the choice of transfer learning as an effective strategy for this domain. All models achieved high precision and recall values, generally above 0.975, suggesting a strong ability to correctly identify relevant characters and minimise false positives and false negatives. The high F1-scores further support the models&#x2019; balanced performance in terms of precision and recall.</p>
</sec>
<sec id="s4_6">
<label>4.6</label>
<title>Contextual Comparison with State-of-the-Art</title>
<p>While a direct numerical benchmark against models trained on other cartoon datasets is challenging due to fundamental differences in tasks, datasets, and metrics, it is crucial to situate our findings within the broader landscape of cartoon character analysis.</p>
<p>For instance, the work by Mashood Nasir et al. (2020) [<xref ref-type="bibr" rid="ref-19">19</xref>] on comic book character identification achieved a classification accuracy of 97%. Our best models, with accuracies exceeding 98.9%, are highly competitive, though we acknowledge our task is multi-label classification on video frames, which presents different challenges than single-label classification on curated comic images. Similarly, recent methods applied to the popular Manga109 dataset [<xref ref-type="bibr" rid="ref-22">22</xref>] focus heavily on object detection and retrieval within monochrome, structured panels, a different domain from our full-colour, dynamic animation scenes. More closely related is the work of Kim et al. (2021) [<xref ref-type="bibr" rid="ref-20">20</xref>], who developed an animation-adaptive model for character detection in animated movies. While their task was detection (finding bounding boxes) rather than multi-label classification, their success in adapting models to varied animation styles echoes our finding that transfer learning is a powerful strategy.</p>
<p>Considering these different contexts, the high performance of our models (e.g., an F1-Score of 0.9890 and an accuracy of 0.9898) demonstrates that our approach is not only effective but achieves a standard of performance consistent with or exceeding state-of-the-art results in related but distinct tasks within cartoon analysis. This validates the use of our methodology for the specific challenge of multi-label classification in animated series.</p>
</sec>
<sec id="s4_7">
<label>4.7</label>
<title>Ablation Study of the Baseline Model Architecture</title>
<p>In addition to comparing against pretrained models, we conducted an ablation study to justify the architectural choices of our custom baseline CNN. To assess the contribution of key components, we trained two additional variants: (1) the baseline model with the Dropout layer removed, and (2) a simpler version of the baseline with only two convolutional layers instead of four. The performance of these variants on the test set is summarised in <xref ref-type="table" rid="table-6">Table 6</xref>.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Ablation study results for the baseline CNN model</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>Epochs</th>
<th>Loss</th>
<th>Accuracy</th>
<th>AUC</th>
<th>Precision</th>
<th>Recall</th>
<th>F1-Score</th>
</tr>
</thead>
<tbody>
<tr>
<td>4 layers, With Dropout</td>
<td>66</td>
<td><bold>0.1701</bold></td>
<td><bold>0.9579</bold></td>
<td><bold>0.9861</bold></td>
<td><bold>0.9661</bold></td>
<td><bold>0.9431</bold></td>
<td><bold>0.9545</bold></td>
</tr>
<tr>
<td>4 layers, No Dropout</td>
<td>26</td>
<td>0.1761</td>
<td>0.9384</td>
<td>0.9775</td>
<td>0.9466</td>
<td>0.9203</td>
<td>0.9349</td>
</tr>
<tr>
<td>2 layers, No Dropout</td>
<td>23</td>
<td>0.3671</td>
<td>0.8449</td>
<td>0.9136</td>
<td>0.8430</td>
<td>0.8217</td>
<td>0.8322</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-6fn1" fn-type="other">
<p>Note: Best value in each column is marked in boldface.</p>
</fn>
</table-wrap-foot>
</table-wrap>
<p>The results clearly validate our design choices. Removing the Dropout layer led to a slight decrease in F1-score and accuracy, indicating its positive regularising effect. More significantly, reducing the network&#x2019;s depth from four to two layers resulted in a substantial performance drop, confirming that the deeper feature hierarchy of our proposed baseline is necessary to effectively learn the character features from scratch. This ablation study provides context for the baseline&#x2019;s performance and reinforces the subsequent findings on the superiority of transfer learning.</p>
</sec>
<sec id="s4_8">
<label>4.8</label>
<title>Ablation Study on the Impact of Pretrained Weights (Transfer Learning vs. Training from Scratch)</title>
<p>To precisely isolate and quantify the benefit of transfer learning, we conducted a second, more controlled ablation study. This experiment compares the performance of a single, powerful architecture, DenseNet121, under two distinct training conditions: (1) using weights pretrained on ImageNet, as done in our main experiments, and (2) training the exact same architecture from scratch with random weight initialisation. Both models were trained using the identical hyperparameter configuration (optimiser, learning rate, and callbacks) to ensure a fair comparison.</p>
<p>The results of this ablation study, presented in <xref ref-type="table" rid="table-7">Table 7</xref>, confirm that transfer learning provides a distinct advantage, even on our high-quality, domain-specific dataset. The DenseNet121 model initialised with pretrained ImageNet weights achieved a slightly higher F1-Score (0.9890) and converged faster (9 epochs) compared to the same architecture trained from scratch, which reached an F1-Score of 0.9808 in 11 epochs. While the model trained from scratch still achieved a remarkably strong performance&#x2014;a testament to the quality and size of the Kral Sakir dataset&#x2014;the pretrained model consistently held a small but significant edge in both final accuracy and training efficiency. This suggests that even when a dataset is sufficient to train a deep network effectively, the foundational visual features learned from a diverse source like ImageNet provide a more optimal starting point, enabling the model to fine-tune to a superior solution with less computational effort. This finding robustly validates the choice of transfer learning as the preferred strategy for this task.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<caption>
<title>Ablation study results on the impact of pretrained weights using the DenseNet121 architecture</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Model</th>
<th>Epochs</th>
<th>Loss</th>
<th>Accuracy</th>
<th>AUC</th>
<th>Precision</th>
<th>Recall</th>
<th>F1-Score</th>
</tr>
</thead>
<tbody>
<tr>
<td>Transfer Learning</td>
<td>9</td>
<td><bold>0.0484</bold></td>
<td><bold>0.9898</bold></td>
<td><bold>0.9963</bold></td>
<td><bold>0.9912</bold></td>
<td><bold>0.9869</bold></td>
<td><bold>0.9890</bold></td>
</tr>
<tr>
<td>From Scratch</td>
<td>11</td>
<td>0.0658</td>
<td>0.9821</td>
<td>0.9950</td>
<td>0.9874</td>
<td>0.9743</td>
<td>0.9808</td>
</tr>
</tbody>
</table>
<table-wrap-foot>
<fn id="table-7fn1" fn-type="other">
<p>Note: Best value in each column is marked in boldface.</p>
</fn>
</table-wrap-foot>
</table-wrap>
</sec>
<sec id="s4_9">
<label>4.9</label>
<title>Error Analysis and Qualitative Insights</title>
<p>To complement the quantitative metrics, we conducted a qualitative analysis of the predictions made by our best-performing model (DenseNet121). This analysis focuses on common failure modes to provide deeper insights into the model&#x2019;s behaviour and the inherent challenges of the dataset. <xref ref-type="fig" rid="fig-7">Fig. 7</xref> illustrates several examples of these incorrect predictions.</p>
<fig id="fig-7">
<label>Figure 7</label>
<caption>
<title>Examples of these incorrect predictions (T &#x003D; True Class, P &#x003D; Predicted Class, FP &#x003D; False Positive, FN &#x003D; False Negative)</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67840-fig-7.tif"/>
</fig>
<p>Our analysis identified three primary categories of errors:
<list list-type="simple">
<list-item><label>1.</label><p><bold>Inter-Class Confusion Due to Character Similarity:</bold> A notable source of error was the confusion between the characters Sakir and Necati. As shown in <xref ref-type="fig" rid="fig-7">Fig. 7a</xref>, these two characters share similar body shapes and certain facial features. In some frames, distinguishing between them is challenging even for a human observer without the broader scene context, leading the model to produce false positive or false negative predictions for these characters.</p>
</list-item>
<list-item><label>2.</label><p><bold>Failure to Detect Occluded or Partially Visible Characters:</bold> The model often struggled to correctly identify characters that were heavily occluded or only partially visible in the frame. As seen in <xref ref-type="fig" rid="fig-7">Fig. 7b</xref>&#x2013;<xref ref-type="fig" rid="fig-7">e</xref>, when a character is obscured by another character or object, or is positioned at the edge of the screen, the model may fail to capture enough distinctive features to make a positive identification. This typically results in a false negative, where the model misses a character that is present.</p>
</list-item>
<list-item><label>3.</label><p><bold>Spurious Feature Activation Leading to False Positives:</bold> In some instances, the model incorrectly identified a character when they were not present in the scene at all. This type of false positive is likely caused by the model activating on background textures, objects, or colour patterns that coincidentally resemble features learned to be associated with that character. For example, we observed cases where the model falsely predicted the presence of the purple elephant character, Necati, when a non-character object with a purple colour and circular patterns appeared in the scene, as illustrated in <xref ref-type="fig" rid="fig-7">Fig. 7f</xref>. Similarly, spurious detections for the character Canan were also noted, as seen in <xref ref-type="fig" rid="fig-7">Fig. 7g</xref>,<xref ref-type="fig" rid="fig-7">h</xref>. This demonstrates that the model can sometimes over-rely on simple features like colour and basic shapes rather than more complex structural character features.</p>
</list-item>
</list></p>
<p>This qualitative analysis underscores that the model&#x2019;s primary limitations are rooted in classic and difficult computer vision challenges: high inter-class similarity, object occlusion, and discriminating between genuine character features and spurious background patterns. These insights are valuable for guiding future research, which could focus on developing models with better contextual understanding or attention mechanisms to mitigate these specific types of errors.</p>
</sec>
<sec id="s4_10">
<label>4.10</label>
<title>Model Explainability with Grad-CAM</title>
<p>To understand which visual features our model uses for character identification, we employed Gradient-weighted Class Activation Mapping (Grad-CAM) on our best-performing model, DenseNet121. This technique produces heatmaps that highlight the image regions most influential to the model&#x2019;s predictions. The analysis, illustrated in <xref ref-type="fig" rid="fig-8">Fig. 8</xref>, reveals distinct attention patterns for each character.</p>
<fig id="fig-8">
<label>Figure 8</label>
<caption>
<title>Sample Grad-CAM heatmaps from DenseNet121</title>
</caption>
<graphic mimetype="image" mime-subtype="tif" xlink:href="CMC_67840-fig-8.tif"/>
</fig>
<p>For the characters Canan, Necati, and Remzi, the model consistently demonstrates robust and accurate localisation of attention. The heatmaps show strong activation on highly discriminative features, such as Canan&#x2019;s unique hair and facial structure, Necati&#x2019;s elephantine form, and Remzi&#x2019;s distinctive mane and face. This indicates that the model has successfully learned to identify these characters based on their core, salient features.</p>
<p>The analysis also provides a clear explanation for the observed confusion between Sakir and Remzi. As shown in the heatmaps, when identifying both Sakir and Remzi, the model&#x2019;s attention is strongly focused on their facial regions, which share significant structural similarities. This reliance on similar facial cues appears to be the primary source of confusion between the two characters. For the model to successfully differentiate them, it must learn to weigh other, more subtle distinguishing features present in the image.</p>
<p>Overall, this explainability analysis confirms that our model learns meaningful and interpretable features for each character and provides a clear visual reason for the specific inter-class confusion noted in our quantitative results.</p>
</sec>
<sec id="s4_11">
<label>4.11</label>
<title>Scientific Contributions and Practical Implications</title>
<p>While our specific findings are detailed throughout this paper, it is useful to consolidate our core scientific contributions and discuss how they can be utilised in practical systems.</p>
<sec id="s4_11_1">
<label>4.11.1</label>
<title>Scientific Contributions</title>
<p>The primary scientific contribution of this work is the establishment of a rigorous benchmark for the challenging, under-addressed problem of multi-label cartoon character recognition under high visual variance. This is achieved through three key elements:
<list list-type="simple">
<list-item><label>1.</label><p><bold>A Novel Public Dataset:</bold> The Kral Sakir dataset provides the community with a new, publicly accessible resource specifically designed to test model generalisation against artistic style changes, a key hurdle in non-photorealistic computer vision. It serves as a standard testbed for future algorithmic comparisons.</p></list-item>
<list-item><label>2.</label><p><bold>A Foundational Performance Baseline:</bold> Our systematic evaluation provides a crucial insight: transfer learning from models pretrained on natural images is highly effective, with architectures featuring dense connectivity (e.g., DenseNet) showing a particular advantage. This knowledge guides future researchers in selecting appropriate models for this domain, saving significant experimental effort.</p></list-item>
<list-item><label>3.</label><p><bold>A Methodological Blueprint:</bold> Through our use of sequence-aware data splitting and detailed ablation studies, we provide a methodologically sound template for future research in character recognition from video media, ensuring more reliable and comparable results across the field.</p></list-item>
</list></p>
</sec>
<sec id="s4_11_2">
<label>4.11.2</label>
<title>Utilisation and Adaptation in Practical Systems</title>
<p>Our work can be directly utilised or adapted to build real-world systems in several ways:
<list list-type="simple">
<list-item><label>1.</label><p><bold>Direct Integration for Content Analysis:</bold> The trained models, particularly our high-performing DenseNet121 variant, can be directly integrated as a core component in larger systems. For example, a media company could use it to automatically generate metadata tags for their video library, identifying which characters appear in every scene. This enables powerful content search and retrieval (e.g., &#x201C;find all scenes featuring Remzi and Necati&#x201D;).</p></list-item>
<list-item><label>2.</label><p><bold>Adaptation for Copyright and Brand Protection:</bold> The entire fine-tuning pipeline serves as a blueprint. A company wishing to protect its intellectual property could follow our methodology to train a model on their own characters. This adapted system could then scan user-generated content platforms (like YouTube or TikTok) to automatically detect and flag unauthorised use of their animated characters.</p></list-item>
<list-item><label>3.</label><p><bold>Foundation for Interactive Media and Parental Controls:</bold> The ability to reliably identify on-screen characters is a foundational step for more advanced applications. For instance, a streaming service could build a feature allowing parents to filter content based on specific characters, or interactive educational apps could trigger events based on which character is speaking or present.</p></list-item>
</list></p>
<p>In essence, this research provides not just a set of results, but a validated dataset, a foundational performance benchmark, and an adaptable methodology that together lower the barrier for creating practical, high-performing AI systems for cartoon media analysis.</p>
</sec>
</sec>
<sec id="s4_12">
<label>4.12</label>
<title>Limitations</title>
<p>A potential limitation of this study is that the Kral Sakir dataset is sourced from a single animated series. While this allows for a deep and focused analysis of character recognition under high visual variance within a consistent artistic style, the findings may not directly generalise to other cartoons with different visual aesthetics. The models trained on our dataset have become specialised to the specific drawing styles, colour palettes, and character designs of Kral Sakir. Consequently, their performance on animated series from different creators or cultural contexts would likely be diminished without further training or domain adaptation techniques.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Conclusion and Future Directions</title>
<p>This study successfully addressed the challenge of multi-label cartoon character classification by introducing the novel Kral Sakir dataset and thoroughly evaluating a range of deep learning techniques. Our experimental investigation, which compared a custom baseline CNN against several state-of-the-art pretrained architectures (ResNet, DenseNet, EfficientNet, MobileNet, and VGG) fine-tuned using transfer learning, underscored the significant advantages of utilising pre-existing knowledge.</p>
<p>The experimental results clearly demonstrated that employing transfer learning with pretrained CNNs substantially improved performance over a custom baseline model across all key metrics, alongside achieving faster convergence. Architectures such as DenseNet proved particularly effective, though other models like EfficientNetB7 and VGG16 also delivered highly competitive results, underscoring the general applicability of transfer learning for this task. Notably, all evaluated models, including the baseline, exhibited a strong capacity for accurate character identification, consistently achieving high precision and recall.</p>
<p>In conclusion, a primary contribution of this research is the establishment of the novel Kral Sakir dataset. This resource, specifically designed for multi-label cartoon character recognition, offers a significant asset for the research community. It can facilitate future investigations, serve as a benchmark for developing and comparing new methodologies, and ultimately contribute to advancing the capabilities of automated systems in this specialised area of computer vision. Furthermore, our study provides strong validation for the robustness and efficiency of transfer learning with advanced CNN architectures for this task. The findings highlight the potential of these methods for developing accurate and reliable automated systems for cartoon content analysis, opening new avenues for exploration in the field.</p>
<p>For future work, we plan to address the single-source limitation of our current dataset. A key priority will be to expand our research by incorporating characters from a diverse range of animated series, encompassing various artistic styles and cultural origins. This will allow for the development and testing of more generalised models. Furthermore, we intend to investigate domain adaptation techniques that could allow a model trained on one series to be effectively fine-tuned for another with minimal additional data, thereby improving the overall robustness and applicability of automated cartoon character recognition systems.</p>
</sec>
</body>
<back>
<ack>
<p>This research was made possible through the use of the Kral Sakir cartoon characters, created by cartoonist and writer Varol Yasaroglu. We are deeply thankful to Varol Yasaroglu for his generous permission and his support of scientific progress through this work. We also thank the Grafi2000 production team for their role. Lastly, we extend special thanks to Kaan Bicakci, whose invaluable suggestions were instrumental to this research.</p>
</ack>
<sec>
<title>Funding Statement</title>
<p>The authors received no specific funding for this study.</p>
</sec>
<sec>
<title>Author Contributions</title>
<p>The authors confirm contribution to the paper as follows: Conceptualisation, Candan Tumer; methodology, Candan Tumer, Erdal Guvenoglu and Volkan Tunali; software, Candan Tumer; validation, Candan Tumer, Erdal Guvenoglu and Volkan Tunali; formal analysis, Candan Tumer; investigation, Erdal Guvenoglu and Volkan Tunali; resources, Candan Tumer; data curation, Candan Tumer; writing&#x2014;original draft preparation, Candan Tumer; writing&#x2014;review and editing, Candan Tumer, Erdal Guvenoglu and Volkan Tunali; visualisation, Candan Tumer; supervision, Erdal Guvenoglu and Volkan Tunali; project administration, Candan Tumer; funding acquisition, Candan Tumer. All authors reviewed the results and approved the final version of the manuscript.</p>
</sec>
<sec sec-type="data-availability">
<title>Availability of Data and Materials</title>
<p>The data that support the findings of this study are openly available on GitHub at <ext-link ext-link-type="uri" xlink:href="https://github.com/candantumer/kral_sakir">https://github.com/candantumer/kral_sakir</ext-link> (accessed on 18 August 2025).</p>
</sec>
<sec>
<title>Ethics Approval</title>
<p>Not applicable.</p>
</sec>
<sec sec-type="COI-statement">
<title>Conflicts of Interest</title>
<p>The authors declare no conflicts of interest to report regarding the present study.</p>
</sec>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>D</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Cartoon character recognition based on portrait style fusion</article-title>. <source>Comput Vis Image Underst</source>. <year>2025</year>;<volume>253</volume>(<issue>6</issue>):<fpage>104316</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.cviu.2025.104316</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kim</surname> <given-names>H</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>EC</given-names></string-name>, <string-name><surname>Seo</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Im</surname> <given-names>DH</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>IK</given-names></string-name></person-group>. <article-title>Character detection in animated movies using multi-style adaptation and visual attention</article-title>. <source>IEEE Transact Multim</source>. <year>2020</year>;<volume>23</volume>:<fpage>1990</fpage>&#x2013;<lpage>2004</lpage>. doi:<pub-id pub-id-type="doi">10.1109/tmm.2020.3006372</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Li</surname> <given-names>D</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>L</given-names></string-name>, <string-name><surname>Jin</surname> <given-names>X</given-names></string-name></person-group>. <article-title>Cartoon copyright recognition method based on character personality action</article-title>. <source>EURASIP J Image Video Process</source>. <year>2024</year>;<volume>2024</volume>(<issue>1</issue>):<fpage>11</fpage>. doi:<pub-id pub-id-type="doi">10.1186/s13640-024-00627-2</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Sun</surname> <given-names>W</given-names></string-name>, <string-name><surname>Kise</surname> <given-names>K</given-names></string-name></person-group>. <chapter-title>Similar partial copy detection of line drawings using a cascade classifier and feature matching</chapter-title>. In: <person-group person-group-type="editor"><string-name><surname>Sako</surname> <given-names>H</given-names></string-name>, <string-name><surname>Franke</surname> <given-names>KY</given-names></string-name>, <string-name><surname>Saitoh</surname> <given-names>S</given-names></string-name></person-group>, editors. <source>Computational forensics</source>. <publisher-loc>Berlin/Heidelberg, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2011</year>. p. <fpage>126</fpage>&#x2013;<lpage>37</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-642-19376-7_11</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Zhang</surname> <given-names>T</given-names></string-name>, <string-name><surname>Han</surname> <given-names>Q</given-names></string-name>, <string-name><surname>Abd El-Latif</surname> <given-names>AA</given-names></string-name>, <string-name><surname>Bai</surname> <given-names>X</given-names></string-name>, <string-name><surname>Niu</surname> <given-names>X</given-names></string-name></person-group>. <article-title>2-D cartoon character detection based on scalable-shape context and hough voting</article-title>. <source>Inf Technol J</source>. <year>2013</year>;<volume>12</volume>(<issue>12</issue>):<fpage>2342</fpage>. doi:<pub-id pub-id-type="doi">10.3923/itj.2013.2342.2349</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Sun</surname> <given-names>W</given-names></string-name>, <string-name><surname>Burie</surname> <given-names>JC</given-names></string-name>, <string-name><surname>Ogier</surname> <given-names>JM</given-names></string-name>, <string-name><surname>Kise</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Specific comic character detection using local feature matching</article-title>. In: <conf-name>2013 12th International Conference on Document Analysis and Recognition; 2013 Aug 25&#x2013;28</conf-name>; <publisher-loc>Washington, DC, USA</publisher-loc>. p. <fpage>275</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Khan</surname> <given-names>FS</given-names></string-name>, <string-name><surname>Anwer</surname> <given-names>RM</given-names></string-name>, <string-name><surname>van de Weijer</surname> <given-names>J</given-names></string-name>, <string-name><surname>Bagdanov</surname> <given-names>AD</given-names></string-name>, <string-name><surname>Vanrell</surname> <given-names>M</given-names></string-name>, <string-name><surname>Lopez</surname> <given-names>AM</given-names></string-name></person-group>. <article-title>Color attributes for object detection</article-title>. In: <conf-name>2012 IEEE Conference on Computer Vision and Pattern Recognition; 2012 Jun 16&#x2013;21</conf-name>; <publisher-loc>Providence, RI, USA</publisher-loc>. p. <fpage>3306</fpage>&#x2013;<lpage>13</lpage>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Rigaud</surname> <given-names>C</given-names></string-name>, <string-name><surname>Karatzas</surname> <given-names>D</given-names></string-name>, <string-name><surname>Burie</surname> <given-names>JC</given-names></string-name>, <string-name><surname>Ogier</surname> <given-names>JM</given-names></string-name></person-group>. <article-title>Color descriptor for content-based drawing retrieval</article-title>. In: <conf-name>2014 11th IAPR International Workshop on Document Analysis Systems; 2014 Apr 7&#x2013;10</conf-name>; <publisher-loc>Tours, France</publisher-loc>. p. <fpage>267</fpage>&#x2013;<lpage>71</lpage>. doi:<pub-id pub-id-type="doi">10.1109/das.2014.70</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Seah</surname> <given-names>HS</given-names></string-name></person-group>. <article-title>Fuzzy diffusion distance learning for cartoon similarity estimation</article-title>. <source>J Comput Sci Technol</source>. <year>2011</year>;<volume>26</volume>(<issue>2</issue>):<fpage>203</fpage>&#x2013;<lpage>16</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11390-011-9427-4</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Takayama</surname> <given-names>K</given-names></string-name>, <string-name><surname>Johan</surname> <given-names>H</given-names></string-name>, <string-name><surname>Nishita</surname> <given-names>T</given-names></string-name></person-group>. <article-title>Face detection and face recognition of cartoon characters using feature extraction</article-title>. In: <conf-name>Proceedings of the 2012 IIEEJ Image Electronics and Visual Computing Workshop</conf-name>. <publisher-loc>Kuching, Malaysia</publisher-loc>: <publisher-name>The Institute of Image Electronics Engineers of Japan</publisher-name>; <year>2012</year>. p. <fpage>1</fpage>&#x2013;<lpage>5</lpage>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Ho</surname> <given-names>HN</given-names></string-name>, <string-name><surname>Rigaud</surname> <given-names>C</given-names></string-name>, <string-name><surname>Burie</surname> <given-names>JC</given-names></string-name>, <string-name><surname>Ogier</surname> <given-names>JM</given-names></string-name></person-group>. <article-title>Redundant structure detection in attributed adjacency graphs for character detection in comics books</article-title>. In: <conf-name>10th IAPR International Workshop on Graphics Recognition; 2013 Aug 20&#x2013;21</conf-name>; <publisher-loc>Bethlehem, PA, USA</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>6</lpage>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Le</surname> <given-names>TN</given-names></string-name>, <string-name><surname>Luqman</surname> <given-names>MM</given-names></string-name>, <string-name><surname>Burie</surname> <given-names>JC</given-names></string-name>, <string-name><surname>Ogier</surname> <given-names>JM</given-names></string-name></person-group>. <article-title>A comic retrieval system based on multilayer graph representation and graph mining</article-title>. In: <conf-name>Graph-Based Representations in Pattern Recognition: 10th IAPR-TC-15 International Workshop, GbRPR 2015</conf-name>. <publisher-loc>Cham, Switzerland</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>2015</year>. p. <fpage>355</fpage>&#x2013;<lpage>64</lpage>. doi:<pub-id pub-id-type="doi">10.1007/978-3-319-18224-7_35</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Iwata</surname> <given-names>M</given-names></string-name>, <string-name><surname>Ito</surname> <given-names>A</given-names></string-name>, <string-name><surname>Kise</surname> <given-names>K</given-names></string-name></person-group>. <article-title>A study to achieve manga character retrieval method for manga images</article-title>. In: <conf-name>2014 11th IAPR International Workshop on Document Analysis Systems; 2014 Apr 7&#x2013;10</conf-name>; <publisher-loc>Tours, France</publisher-loc>. p. <fpage>309</fpage>&#x2013;<lpage>13</lpage>. doi:<pub-id pub-id-type="doi">10.1109/das.2014.60</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Qin</surname> <given-names>X</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>Y</given-names></string-name>, <string-name><surname>He</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Tang</surname> <given-names>Z</given-names></string-name></person-group>. <article-title>A faster R-CNN based method for comic characters face detection</article-title>. In: <conf-name>2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR); 2017 Nov 9&#x2013;15</conf-name>; <publisher-loc>Kyoto, Japan</publisher-loc>. p. <fpage>1074</fpage>&#x2013;<lpage>80</lpage>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Nguyen</surname> <given-names>NV</given-names></string-name>, <string-name><surname>Rigaud</surname> <given-names>C</given-names></string-name>, <string-name><surname>Burie</surname> <given-names>JC</given-names></string-name></person-group>. <article-title>Comic characters detection using deep learning</article-title>. In: <conf-name>2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR); 2017 Nov 9&#x2013;15</conf-name>; <publisher-loc>Kyoto, Japan</publisher-loc>. p. <fpage>41</fpage>&#x2013;<lpage>6</lpage>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yanagisawa</surname> <given-names>H</given-names></string-name>, <string-name><surname>Yamashita</surname> <given-names>T</given-names></string-name>, <string-name><surname>Watanabe</surname> <given-names>H</given-names></string-name></person-group>. <article-title>A study on object detection method from manga images using CNN</article-title>. In: <conf-name>2018 International Workshop on Advanced Image Technology (IWAIT); 2018 Jan 7&#x2013;9</conf-name>; <publisher-loc>Chiang Mai, Thailand</publisher-loc>; <year>2018</year>. p. <fpage>1</fpage>&#x2013;<lpage>4</lpage>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Chu</surname> <given-names>WT</given-names></string-name>, <string-name><surname>Li</surname> <given-names>WW</given-names></string-name></person-group>. <article-title>Manga face detection based on deep neural networks fusing global and local information</article-title>. <source>Pattern Recognition</source>. <year>2019</year>;<volume>86</volume>(<issue>20</issue>):<fpage>62</fpage>&#x2013;<lpage>72</lpage>. doi:<pub-id pub-id-type="doi">10.1016/j.patcog.2018.08.008</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Dutta</surname> <given-names>A</given-names></string-name>, <string-name><surname>Biswas</surname> <given-names>S</given-names></string-name></person-group>. <article-title>CNN based extraction of panels/characters from Bengali comic book page images</article-title>. In: <conf-name>2019 International Conference on Document Analysis and Recognition Workshops (ICDARW); 2019 Sep 22&#x2013;25</conf-name>; <publisher-loc>Sydney, NSW, Australia</publisher-loc>. p. <fpage>38</fpage>&#x2013;<lpage>43</lpage>.</mixed-citation></ref>
<ref id="ref-19"><label>[19]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mashood Nasir</surname> <given-names>I</given-names></string-name>, <string-name><surname>Attique Khan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Alhaisoni</surname> <given-names>M</given-names></string-name>, <string-name><surname>Saba</surname> <given-names>T</given-names></string-name>, <string-name><surname>Rehman</surname> <given-names>A</given-names></string-name>, <string-name><surname>Iqbal</surname> <given-names>T</given-names></string-name></person-group>. <article-title>A hybrid deep learning architecture for the classification of superhero fashion products: an application for medical-tech classification</article-title>. <source>Comput Model Eng Sci</source>. <year>2020</year>;<volume>124</volume>(<issue>3</issue>):<fpage>1017</fpage>&#x2013;<lpage>33</lpage>. doi:<pub-id pub-id-type="doi">10.32604/cmes.2020.010943</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>[20]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Kim</surname> <given-names>K</given-names></string-name>, <string-name><surname>Park</surname> <given-names>S</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>J</given-names></string-name>, <string-name><surname>Chung</surname> <given-names>S</given-names></string-name>, <string-name><surname>Lee</surname> <given-names>J</given-names></string-name>, <string-name><surname>Choo</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Animeceleb: large-scale animation celebfaces dataset via controllable 3D synthetic models</article-title>. <comment>arXiv:2111.07640. 2021</comment>.</mixed-citation></ref>
<ref id="ref-21"><label>[21]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Matsui</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Ito</surname> <given-names>K</given-names></string-name>, <string-name><surname>Aramaki</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Fujimoto</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ogawa</surname> <given-names>T</given-names></string-name>, <string-name><surname>Yamasaki</surname> <given-names>T</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Sketch-based manga retrieval using manga109 dataset</article-title>. <source>Multimed Tools Appl</source>. <year>2017</year>;<volume>76</volume>(<issue>20</issue>):<fpage>21811</fpage>&#x2013;<lpage>38</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s11042-016-4020-z</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>[22]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Aizawa</surname> <given-names>K</given-names></string-name>, <string-name><surname>Fujimoto</surname> <given-names>A</given-names></string-name>, <string-name><surname>Otsubo</surname> <given-names>A</given-names></string-name>, <string-name><surname>Ogawa</surname> <given-names>T</given-names></string-name>, <string-name><surname>Matsui</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Tsubota</surname> <given-names>K</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Building a manga dataset &#x201C;manga109&#x201D; with annotations for multimedia applications</article-title>. <source>IEEE Multimedia</source>. <year>2020</year>;<volume>27</volume>(<issue>2</issue>):<fpage>8</fpage>&#x2013;<lpage>18</lpage>. doi:<pub-id pub-id-type="doi">10.1109/mmul.2020.2987895</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>[23]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Qi</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Pan</surname> <given-names>D</given-names></string-name>, <string-name><surname>Niu</surname> <given-names>T</given-names></string-name>, <string-name><surname>Ying</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Shi</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Bridge the gap between practical application scenarios and cartoon character detection: a benchmark dataset and deep learning model</article-title>. <source>Displays</source>. <year>2024</year>;<volume>84</volume>(<issue>20</issue>):<fpage>102793</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.displa.2024.102793</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>[24]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nguyen</surname> <given-names>NV</given-names></string-name>, <string-name><surname>Rigaud</surname> <given-names>C</given-names></string-name>, <string-name><surname>Burie</surname> <given-names>JC</given-names></string-name></person-group>. <article-title>Digital comics image indexing based on deep learning</article-title>. <source>J Imaging</source>. <year>2018</year>;<volume>4</volume>(<issue>7</issue>):<fpage>89</fpage>. doi:<pub-id pub-id-type="doi">10.3390/jimaging4070089</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>[25]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Nguyen</surname> <given-names>NV</given-names></string-name>, <string-name><surname>Rigaud</surname> <given-names>C</given-names></string-name>, <string-name><surname>Burie</surname> <given-names>JC</given-names></string-name></person-group>. <article-title>Comic MTL: optimized multi-task learning for comic book image analysis</article-title>. <source>Int J Docum Anal Recognit (IJDAR)</source>. <year>2019</year>;<volume>22</volume>(<issue>3</issue>):<fpage>265</fpage>&#x2013;<lpage>84</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10032-019-00330-3</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>[26]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Nir</surname> <given-names>O</given-names></string-name>, <string-name><surname>Rapoport</surname> <given-names>G</given-names></string-name>, <string-name><surname>Shamir</surname> <given-names>A</given-names></string-name></person-group>. <chapter-title>CAST: character labeling in animation using self-supervision by tracking</chapter-title>. In: <source>Computer graphics forum</source>. Vol. <volume>41</volume>. <publisher-loc>Hoboken, NJ, USA</publisher-loc>: <publisher-name>John Wiley &#x0026; Sons, Inc.</publisher-name>; <year>2022</year>. p. <fpage>135</fpage>&#x2013;<lpage>45</lpage>.</mixed-citation></ref>
<ref id="ref-27"><label>[27]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jain</surname> <given-names>N</given-names></string-name>, <string-name><surname>Gupta</surname> <given-names>V</given-names></string-name>, <string-name><surname>Shubham</surname> <given-names>S</given-names></string-name>, <string-name><surname>Madan</surname> <given-names>A</given-names></string-name>, <string-name><surname>Chaudhary</surname> <given-names>A</given-names></string-name>, <string-name><surname>Santosh</surname> <given-names>K</given-names></string-name></person-group>. <article-title>Understanding cartoon emotion using integrated deep neural network on large dataset</article-title>. <source>Neural Comput Appl</source>. <year>2022</year>;<volume>34</volume>(<issue>24</issue>):<fpage>21481</fpage>&#x2013;<lpage>21501</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s00521-021-06003-9</pub-id>; <pub-id pub-id-type="pmid">33903785</pub-id></mixed-citation></ref>
<ref id="ref-28"><label>[28]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Agrawal</surname> <given-names>SC</given-names></string-name>, <string-name><surname>Tripathi</surname> <given-names>RK</given-names></string-name></person-group>. <article-title>Cartoon face detection and recognition with emotion recognition</article-title>. In: <conf-name>2023 International Conference on Data Science, Agents and Artificial Intelligence (ICDSAAI); 2023 Dec 21&#x2013;23</conf-name>; <publisher-loc>Chennai, India</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>4</lpage>.</mixed-citation></ref>
<ref id="ref-29"><label>[29]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Yi</surname> <given-names>F</given-names></string-name>, <string-name><surname>Wu</surname> <given-names>J</given-names></string-name>, <string-name><surname>Zhao</surname> <given-names>M</given-names></string-name>, <string-name><surname>Zhou</surname> <given-names>S</given-names></string-name></person-group>. <article-title>Anime character identification and tag prediction by multimodality modeling: dataset and model</article-title>. In: <conf-name>2023 International Joint Conference on Neural Networks (IJCNN); 2023 Jul 18&#x2013;22</conf-name>; <publisher-loc>Gold Coast, QLD, Australia</publisher-loc>. p. <fpage>1</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-30"><label>[30]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Rios</surname> <given-names>EA</given-names></string-name>, <string-name><surname>Hu</surname> <given-names>MC</given-names></string-name>, <string-name><surname>Lai</surname> <given-names>BC</given-names></string-name></person-group>. <article-title>Anime character recognition using intermediate features aggregation</article-title>. In: <conf-name>2022 IEEE International Symposium on Circuits and Systems (ISCAS); 2022 May 28&#x2013;Jun 1</conf-name>; <publisher-loc>Austin, TX, USA</publisher-loc>. p. <fpage>424</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-31"><label>[31]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><collab>Grafi2000</collab></person-group>. <article-title>Kral Sakir&#x2014;YouTube [Internet]</article-title>; <year>2023</year> <comment>[cited 2025 Aug 18]</comment>. Available from: <ext-link ext-link-type="uri" xlink:href="https://www.youtube.com/@KralSakirResmi">https://www.youtube.com/@KralSakirResmi</ext-link>.</mixed-citation></ref>
<ref id="ref-32"><label>[32]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Razzaq</surname> <given-names>K</given-names></string-name>, <string-name><surname>Shah</surname> <given-names>M</given-names></string-name></person-group>. <article-title>Machine learning and deep learning paradigms: from techniques to practical applications and research frontiers</article-title>. <source>Computers</source>. <year>2025</year>;<volume>14</volume>(<issue>3</issue>):<fpage>93</fpage>. doi:<pub-id pub-id-type="doi">10.3390/computers14030093</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>[33]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ahmed</surname> <given-names>SF</given-names></string-name>, <string-name><surname>Alam</surname> <given-names>MSB</given-names></string-name>, <string-name><surname>Kabir</surname> <given-names>M</given-names></string-name>, <string-name><surname>Afrin</surname> <given-names>S</given-names></string-name>, <string-name><surname>Rafa</surname> <given-names>SJ</given-names></string-name>, <string-name><surname>Mehjabin</surname> <given-names>A</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Unveiling the frontiers of deep learning: innovations shaping diverse domains</article-title>. <source>Appl Intell</source>. <year>2025</year>;<volume>55</volume>(<issue>7</issue>):<fpage>1</fpage>&#x2013;<lpage>55</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s10489-025-06259-x</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>[34]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lecun</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Bottou</surname> <given-names>L</given-names></string-name>, <string-name><surname>Bengio</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Haffner</surname> <given-names>P</given-names></string-name></person-group>. <article-title>Gradient-based learning applied to document recognition</article-title>. <source>Proc IEEE</source>. <year>1998</year>;<volume>86</volume>(<issue>11</issue>):<fpage>2278</fpage>&#x2013;<lpage>324</lpage>. doi:<pub-id pub-id-type="doi">10.1109/5.726791</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>[35]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gholizade</surname> <given-names>M</given-names></string-name>, <string-name><surname>Soltanizadeh</surname> <given-names>H</given-names></string-name>, <string-name><surname>Rahmanimanesh</surname> <given-names>M</given-names></string-name>, <string-name><surname>Sana</surname> <given-names>SS</given-names></string-name></person-group>. <article-title>A review of recent advances and strategies in transfer learning</article-title>. <source>Int J Syst Assur Eng Manag</source>. <year>2025</year>;<volume>16</volume>(<issue>3</issue>):<fpage>1123</fpage>&#x2013;<lpage>62</lpage>. doi:<pub-id pub-id-type="doi">10.1007/s13198-024-02684-2</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>[36]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Sohail</surname> <given-names>SS</given-names></string-name>, <string-name><surname>Himeur</surname> <given-names>Y</given-names></string-name>, <string-name><surname>Kheddar</surname> <given-names>H</given-names></string-name>, <string-name><surname>Amira</surname> <given-names>A</given-names></string-name>, <string-name><surname>Fadli</surname> <given-names>F</given-names></string-name>, <string-name><surname>Atalla</surname> <given-names>S</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Advancing 3D point cloud understanding through deep transfer learning: a comprehensive survey</article-title>. <source>Inf Fusion</source>. <year>2025</year>;<volume>113</volume>(<issue>10</issue>):<fpage>102601</fpage>. doi:<pub-id pub-id-type="doi">10.1016/j.inffus.2024.102601</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>[37]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Deng</surname> <given-names>J</given-names></string-name>, <string-name><surname>Dong</surname> <given-names>W</given-names></string-name>, <string-name><surname>Socher</surname> <given-names>R</given-names></string-name>, <string-name><surname>Li</surname> <given-names>LJ</given-names></string-name>, <string-name><surname>Li</surname> <given-names>K</given-names></string-name>, <string-name><surname>Fei-Fei</surname> <given-names>L</given-names></string-name></person-group>. <article-title>ImageNet: a large-scale hierarchical image database</article-title>. In: <conf-name>2009 IEEE Conference on Computer Vision and Pattern Recognition; 2009 Jun 20&#x2013;25</conf-name>; <publisher-loc>Miami, FL, USA</publisher-loc>. p. <fpage>248</fpage>&#x2013;<lpage>55</lpage>.</mixed-citation></ref>
<ref id="ref-38"><label>[38]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>He</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zhang</surname> <given-names>X</given-names></string-name>, <string-name><surname>Ren</surname> <given-names>S</given-names></string-name>, <string-name><surname>Sun</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Deep residual learning for image recognition</article-title>. In: <conf-name>2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2016 Jun 27&#x2013;30</conf-name>; <publisher-loc>Las Vegas, NV, USA</publisher-loc>. p. <fpage>770</fpage>&#x2013;<lpage>8</lpage>.</mixed-citation></ref>
<ref id="ref-39"><label>[39]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Huang</surname> <given-names>G</given-names></string-name>, <string-name><surname>Liu</surname> <given-names>Z</given-names></string-name>, <string-name><surname>Van Der Maaten</surname> <given-names>L</given-names></string-name>, <string-name><surname>Weinberger</surname> <given-names>KQ</given-names></string-name></person-group>. <article-title>Densely connected convolutional networks</article-title>. In: <conf-name>2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2017 Jul 21&#x2013;26</conf-name>; <publisher-loc>Honolulu, HI, USA</publisher-loc>. p. <fpage>2261</fpage>&#x2013;<lpage>9</lpage>.</mixed-citation></ref>
<ref id="ref-40"><label>[40]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Tan</surname> <given-names>M</given-names></string-name>, <string-name><surname>Le</surname> <given-names>QV</given-names></string-name></person-group>. <article-title>Efficientnet: rethinking model scaling for convolutional neural networks</article-title>. In: <conf-name>36th International Conference on Machine Learning (ICML); 2019 Jun 9&#x2013;15</conf-name>; <publisher-loc>Long Beach, CA, USA</publisher-loc>. p. <fpage>6105</fpage>&#x2013;<lpage>14</lpage>.</mixed-citation></ref>
<ref id="ref-41"><label>[41]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><surname>Howard</surname> <given-names>AG</given-names></string-name>, <string-name><surname>Zhu</surname> <given-names>M</given-names></string-name>, <string-name><surname>Chen</surname> <given-names>B</given-names></string-name>, <string-name><surname>Kalenichenko</surname> <given-names>D</given-names></string-name>, <string-name><surname>Wang</surname> <given-names>W</given-names></string-name>, <string-name><surname>Weyand</surname> <given-names>T</given-names></string-name>, <etal>et al</etal></person-group>. <article-title>Mobilenets: efficient convolutional neural networks for mobile vision applications</article-title>. <comment>arXiv:1704.04861. 2017</comment>.</mixed-citation></ref>
<ref id="ref-42"><label>[42]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Simonyan</surname> <given-names>K</given-names></string-name>, <string-name><surname>Zisserman</surname> <given-names>A</given-names></string-name></person-group>. <article-title>Very deep convolutional networks for large-scale image recognition</article-title>. In: <conf-name>3rd International Conference on Learning Representations (ICLR); 2015 May 7&#x2013;9</conf-name>; <publisher-loc>San Diego, CA, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-43"><label>[43]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Kingma</surname> <given-names>DP</given-names></string-name>, <string-name><surname>Ba</surname> <given-names>J</given-names></string-name></person-group>. <article-title>Adam: a method for stochastic optimization</article-title>. In: <conf-name>3rd International Conference on Learning Representations (ICLR); 2015 May 7&#x2013;9</conf-name>; <publisher-loc>San Diego, CA, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-44"><label>[44]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>Krizhevsky</surname> <given-names>A</given-names></string-name>, <string-name><surname>Sutskever</surname> <given-names>I</given-names></string-name>, <string-name><surname>Hinton</surname> <given-names>GE</given-names></string-name></person-group>. <article-title>ImageNet classification with deep convolutional neural networks</article-title>. In: <conf-name>Proceedings of the 26th International Conference on Neural Information Processing Systems; 2012 Dec 3&#x2013;6</conf-name>; <publisher-loc>Lake Tahoe, NV, USA</publisher-loc>. p. <fpage>1097</fpage>&#x2013;<lpage>105</lpage>.</mixed-citation></ref>
<ref id="ref-45"><label>[45]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Prechelt</surname> <given-names>L</given-names></string-name></person-group>. <article-title>Early stopping&#x2014;but when?</article-title> In: <person-group person-group-type="editor"><string-name><surname>Orr</surname> <given-names>GB</given-names></string-name>, <string-name><surname>M&#x00FC;ller</surname> <given-names>KR</given-names></string-name></person-group>, editors. <source>Neural networks: Tricks of the trade, lecture notes in computer science</source>. <publisher-loc>Berlin/Heidelberg, Germany</publisher-loc>: <publisher-name>Springer</publisher-name>; <year>1998</year>. Vol. <volume>7700</volume>, p. <fpage>55</fpage>&#x2013;<lpage>69</lpage>.</mixed-citation></ref>
</ref-list>
</back></article>