<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMES</journal-id>
<journal-id journal-id-type="nlm-ta">CMES</journal-id>
<journal-id journal-id-type="publisher-id">CMES</journal-id>
<journal-title-group>
<journal-title>Computer Modeling in Engineering &#x0026; Sciences</journal-title>
</journal-title-group>
<issn pub-type="epub">1526-1506</issn>
<issn pub-type="ppub">1526-1492</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">19097</article-id>
<article-id pub-id-type="doi">10.32604/cmes.2022.019097</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>Prediction of Intrinsically Disordered Proteins Based on Deep Neural Network-ResNet18</article-title>
<alt-title alt-title-type="left-running-head">Prediction of Intrinsically Disordered Proteins Based on Deep Neural Network-ResNet18</alt-title>
<alt-title alt-title-type="right-running-head">Prediction of Intrinsically Disordered Proteins Based on Deep Neural Network-ResNet18</alt-title>
</title-group>
<contrib-group content-type="authors">
<contrib id="author-1" contrib-type="author">
<name name-style="western">
<surname>Zhang</surname>
<given-names>Jie</given-names>
</name>
</contrib>
<contrib id="author-2" contrib-type="author" corresp="yes">
<name name-style="western">
<surname>Zhao</surname>
<given-names>Jiaxiang</given-names>
</name><email>zhaojx@nankai.edu.cn</email>
</contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western">
<surname>Xu</surname>
<given-names>Pengchang</given-names>
</name>
</contrib>
<aff><institution>School of Electronic Information and Optical Engineering, Nankai University, Tianjin Key Laboratory of Optoelectronic Sensor and Sensing Network Technology</institution>, <addr-line>Tianjin, 300350</addr-line>, <country>China</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Jiaxiang Zhao. Email: <email>zhaojx@nankai.edu.cn</email></corresp>
</author-notes>
<pub-date pub-type="epub" date-type="pub" iso-8601-date="2022-03-11">
<day>11</day>
<month>03</month>
<year>2022</year>
</pub-date>
<volume>131</volume>
<issue>2</issue>
<fpage>905</fpage>
<lpage>917</lpage>
<history>
<date date-type="received">
<day>02</day>
<month>9</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>01</day>
<month>11</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2022 Zhang, Zhao and Xu</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Zhang, Zhao and Xu</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMES_19097.pdf"></self-uri>
<abstract>
<p>Accurately, reliably and rapidly identifying intrinsically disordered (IDPs) proteins is essential as they often play important roles in various human diseases; moreover, they are related to numerous important biological activities. However, current computational methods have yet to develop a network that is sufficiently deep to make predictions about IDPs and demonstrate an improvement in performance. During this study, we constructed a deep neural network that consisted of five identical variant models, ResNet18, combined with an MLP network, for classification. Resnet18 was applied for the first time as a deep model for predicting IDPs, which allowed the extraction of information from IDP residues in greater detail and depth, and this information was then passed through the MLP network for the final identification process. Two well-known datasets, MXD494 and R80, were used as the blind independent datasets to compare their performance with that of our method. The simulation results showed that Matthew&#x2019;s correlation coefficient obtained using our deep network model was 0.517 on the blind R80 dataset and 0.450 on the MXD494 dataset; thus, our method outperformed existing methods.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>ResNet18</kwd>
<kwd>MLP</kwd>
<kwd>intrinsically disordered protein</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Intrinsically disordered proteins (IDPs) (<italic>i.e</italic>., proteins that contain disordered regions [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-2">2</xref>]) have been confirmed to be related to many important biological activities and involved in several important cell functions, such as nucleic acid folding [<xref ref-type="bibr" rid="ref-3">3</xref>] and cell signaling and regulation [<xref ref-type="bibr" rid="ref-4">4</xref>,<xref ref-type="bibr" rid="ref-5">5</xref>]. Moreover, various human diseases, such as certain types of cancers [<xref ref-type="bibr" rid="ref-6">6</xref>], genetic diseases, and Alzheimer&#x2019;s disease [<xref ref-type="bibr" rid="ref-7">7</xref>,<xref ref-type="bibr" rid="ref-8">8</xref>] are associated with IDPs. Furthermore, IDPs are more easily blocked by small molecules, and are therefore potential targets for drug design, providing a good basis for drug treatment [<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-10">10</xref>]. However, accurately, reliably, and rapidly identifying IDPs remains a challenging problem.</p>
<p>Numerous schemes for detecting IDPs have been proffered over the past several decades and can be categorized into two types: those based on physical and chemical properties of amino acids and those based on computational methods. (i) examples based on physicochemical properties include FoldIndex [<xref ref-type="bibr" rid="ref-11">11</xref>], GlobPlot [<xref ref-type="bibr" rid="ref-12">12</xref>], FoldUnfold [<xref ref-type="bibr" rid="ref-13">13</xref>], and IsUnstruct [<xref ref-type="bibr" rid="ref-14">14</xref>]. FoldIndex [<xref ref-type="bibr" rid="ref-11">11</xref>] predicts disordered proteins by calculating the ratio of the average hydrophobicity to the average net charge of the residues. GlobPlot [<xref ref-type="bibr" rid="ref-12">12</xref>] predicts the disordered region by analyzing the tendency of all amino acids in the protein sequence to disordered residues and ordered residues. FoldUnfold [<xref ref-type="bibr" rid="ref-13">13</xref>] treats areas with a weaker density as unnecessary areas by predicting the average packing density of residues. IsUnstruct [<xref ref-type="bibr" rid="ref-14">14</xref>] is a prediction method based on statistical physics and uses the Ising model for disorder-order transformations in protein sequences and replaces the interactions of adjacent terms with penalties for changes between boundary energy states. These methods promote the field of IDPs, yet ignore the overall structure of the protein, which leads to inaccurate prediction results. (ii) More recently, the use of machine learning has increased in the field of bioinformatics, especially to solve problems that are closely related to human life and health [<xref ref-type="bibr" rid="ref-15">15</xref>,<xref ref-type="bibr" rid="ref-16">16</xref>]. Methods to identify IDPs through machine learning, especially deep learning techniques, such as PONDR [<xref ref-type="bibr" rid="ref-17">17</xref>], DISOPRED2 [<xref ref-type="bibr" rid="ref-18">18</xref>], RONN [<xref ref-type="bibr" rid="ref-19">19</xref>], DISKNN [<xref ref-type="bibr" rid="ref-20">20</xref>], IDP-Seq2Seq [<xref ref-type="bibr" rid="ref-21">21</xref>], NetSurfP-2.0 [<xref ref-type="bibr" rid="ref-22">22</xref>], SPOT-Disorder2 [<xref ref-type="bibr" rid="ref-23">23</xref>], and RFPR-IDP [<xref ref-type="bibr" rid="ref-24">24</xref>], have also been developed. The PONDR [<xref ref-type="bibr" rid="ref-17">17</xref>] series is the first publicly available and established method for predicting IDPs internationally, which distinguishes disordered proteins from ordered proteins based primarily on differences in their amino acid composition. DISOPRED2 [<xref ref-type="bibr" rid="ref-18">18</xref>] is a dynamic prediction method for disordered proteins, and the output of the support vector machine is used as the prediction results. RONN [<xref ref-type="bibr" rid="ref-19">19</xref>] is a function-based array and neural network prediction algorithm. Its main idea is that if two proteins have similar biological functions and similar tendencies toward being disordered or ordered, then their sequences are also similar. DISKNN [<xref ref-type="bibr" rid="ref-20">20</xref>] applies KNN with several protein features to predict the disordered regions of proteins. However, the performance comparison in this article was unconvincing because it was only regarding one protein. IDP-Seq2Seq [<xref ref-type="bibr" rid="ref-21">21</xref>] draws on sequence-to-sequence learning in natural language processing to map protein sequences and use associations between residues as features for prediction. NetSurfP-2.0 [<xref ref-type="bibr" rid="ref-22">22</xref>] applies convergence strategies with the convolutional neural network (CNN) and long and short-term memory networks (LSTM) based on protein structural features. SPOT-Disorder2 [<xref ref-type="bibr" rid="ref-23">23</xref>] is an improvement of the SPOT-Disorder by Hanson et al. [<xref ref-type="bibr" rid="ref-25">25</xref>] (a profile-based method), which mainly uses the LSTM model to predict intrinsically disordered proteins. RFPR-IDP [<xref ref-type="bibr" rid="ref-24">24</xref>] is based on a combination of the CNN and the bidirectional LSTM. In addition, there are also various meta-methods to predict IDPs, such as IDP-FSP [<xref ref-type="bibr" rid="ref-26">26</xref>], MFDp [<xref ref-type="bibr" rid="ref-27">27</xref>], Spark-IDPP [<xref ref-type="bibr" rid="ref-28">28</xref>], and Meta-Disorder [<xref ref-type="bibr" rid="ref-29">29</xref>], which run multiple independent prediction schemes and merge their results to obtain final prediction results. Mishra et al. used a deep learning-integrated method combining 9890 features for protein function prediction; the large number of features used was novel [<xref ref-type="bibr" rid="ref-30">30</xref>]. These schemes based on computational methods do not sufficiently capture details between protein residues and only capture those at the sequence level, which leads to inaccurate predictions.</p>
<p>Thus, in the field of predicting IDPs, several problems remain: (i) Predictions using physical and chemical methods is not only a complex and time-consuming process but also have poor prediction performance; (ii) Previous studies based on computational methods did not employ a sufficiently deep network to capture more accurate characteristics among residues, and consequently, did not demonstrate a significant improvement in performance; (iii) Although some previous methods have advantages in predicting IDPs, Matthew&#x0027;s correlation coefficient (MCC) [<xref ref-type="bibr" rid="ref-31">31</xref>], which directly indicates the quality of the prediction result, remain relatively low on blind independent test datasets.</p>
<p><italic>Motivation:</italic> Because IDPs are essential given their important roles in various human diseases and their association with many important biological activities, addressing the above three problems in an accurate, reliable, and rapid manner is of great significance and research value. Therefore, using simple computational methods, rather than complex physicochemical methods, is crucial, and constructing a neural network with deep layers for predicting IDPs and demonstrating an improvement in prediction performance is essential. ResNet18 [<xref ref-type="bibr" rid="ref-32">32</xref>] has sufficiently deep layers and has been used in myriad domains owing to its excellent performance; however, it is the first application of ResNet18 for predicting IDPs. It would be constructive to apply the ResNet18 deep neural network and achieve good results, with a high as possible MCC value, for predicting IDPs.</p>
<p><italic>Contribution:</italic> During this study, we proposed a novel method for predicting IDPs using deep neural networks, which outperformed existing methods. The innovative value of our contributions are as follows:</p>
<p>&#x2022; We constructed a sufficiently deep structure, which consisted of five identical ResNet18 networks, and combined it with a constructed multilayer perceptron (MLP) network. This constructure differed from all previous methods and was a completely new architecture. <?A3B2 "fig1",5,"anchor"?><xref ref-type="fig" rid="fig-1">Fig. 1</xref> depicts the paradigm of our deep neural structure.</p>
<p>&#x2022; Resnet18 was applied for the first time as a deep network model for predicting IDPs. The study did not simply use the ResNet18 network directly but replaced the output layer with a dense layer, and the output was then input into the MLP network.</p>
<p>&#x2022; The simulation results showed that the MCC value obtained using our deep network model was a significant improvement on other existing methods, which has important implications for more precise investigations of IDPs.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Paradigm of our deep neural structure</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19097-fig-1.png"/>
</fig>
<p>The following steps illustrate the specific performing process of our proposed method in detail:</p>
<p>Step 1. Data preparation: A total of 1616 pieces of the latest protein sequences were downloaded from the DisProt database as a training dataset for our deep network model.</p>
<p>Step 2. Feature selection: For the features of each amino acid of the IDP sequence, we calculated five structural properties, seven physicochemical properties, and 20 protein evolution information of the sequence with a total of 32 features; the five structural properties included Shannon entropy [<xref ref-type="bibr" rid="ref-33">33</xref>], topological entropy [<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>], and three amino-acid propensity scales; the seven physicochemical properties were obtained from Meiler et al. [<xref ref-type="bibr" rid="ref-35">35</xref>], and the position-specific scoring matrices (PSSMs) were generated from the PSI-BLAST software using the latest NCBI non-redundant database [<xref ref-type="bibr" rid="ref-36">36</xref>] (updated in June 2020).</p>
<p>Step 3. Feature pre-processing: An important step after selecting features is preprocessing features using the sliding window approach. With the sliding calculation of the window over each protein sequence, we obtained the feature matrix <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mrow><mml:mi mathvariant="bold">X</mml:mi></mml:mrow></mml:math></inline-formula>, which was fed into our constructed deep neural network model.</p>
<p>Step 4. Model processing: We constructed a sufficiently deep structure, which consisted of five identical ResNet18 networks, and combined it with a constructed MLP network. For the original ResNet18 model, we replaced the fully connected (FC) layer with a dense layer, and the output was then input into the MLP network for the final identification process.</p>
<p>Step 5. Analyze the blind dataset performance: In contrast to other well-known methods, the R80 and MXD494 datasets were used as our blind datasets to analyze the performance of our deep network model.</p>
<p>The different sections of this article describe the contents of the proposed method in a step-by-step manner and are organized according to the above steps as follows: <xref ref-type="sec" rid="s2">Section 2</xref> describes above Steps 1 to 4 and the materials and methods proposed in this article, which include the preparation of our method, the architecture of the network, and how the method works. <xref ref-type="sec" rid="s3">Section 3</xref> compares other well-known methods that predict IDPs using five recognized performance measures. <xref ref-type="sec" rid="s4">Section 4</xref> provides our conclusions and the future scope of our research.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Materials and Methods</title>
<p>We have listed all datasets used for training and blind testing, and we introduce the architecture of our deep network model depicted in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. The model is comprised of a pre-processing block, five copies of the ResNet18, and a self-constructed MLP network.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Datasets</title>
<p>We employed DIS1616 from the latest version of the DisProt database as our training dataset. The DIS1616 dataset consists of 1616 protein sequences, which comprise 888678 residues. Of these 888678 residues, 182316 are disordered residues and 706362 are ordered residues. Here, we randomly shuffled and divided the dataset into 10 separate subsets, with the test dataset containing 166 sequences. Then, 10-fold cross-validation was performed. To analyze the performance of our deep network, the R80 [<xref ref-type="bibr" rid="ref-19">19</xref>] and MXD494 datasets [<xref ref-type="bibr" rid="ref-21">21</xref>,<xref ref-type="bibr" rid="ref-37">37</xref>] were used as our blind datasets. The R80 dataset consisted of 80 protein sequences with 3566 disordered residues and 29243 ordered residues. The MXD494 dataset contained 494 protein sequences with 44087 disordered residues and 152474 ordered residues. <?A3B2 "fig2",5,"anchor"?><xref ref-type="fig" rid="fig-2">Fig. 2</xref> illustrates two completely ordered proteins (2FG1 and 3BBB) with stable three-dimensional structures from the MXD494 dataset.</p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Two protein pictures from the MXD494 dataset</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19097-fig-2.png"/>
</fig>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Protein Feature Selection and Pre-Processing Procedure</title>
<p>Three types of features were selected: five features associated with the structural properties, seven features corresponding to the physicochemical properties, and the remaining features related to the evolutionary information. Structural features were Shannon entropy [<xref ref-type="bibr" rid="ref-33">33</xref>], topological entropy [<xref ref-type="bibr" rid="ref-33">33</xref>,<xref ref-type="bibr" rid="ref-34">34</xref>], and three amino acid propensity scales provided in the GlobPlot NAR paper [<xref ref-type="bibr" rid="ref-12">12</xref>]. The physicochemical features were obtained from Meiler et al. [<xref ref-type="bibr" rid="ref-35">35</xref>]. The PSSMs were used to describe the protein evolution information generated from the PSI-BLAST software using the latest NCBI non-redundant database [<xref ref-type="bibr" rid="ref-36">36</xref>] (updated in June 2020).</p>
<p>As shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>, our deep neural network contained a pre-processing block that computed the input feature matrix <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mrow><mml:mi mathvariant="bold">X</mml:mi></mml:mrow></mml:math></inline-formula> as below:</p>
<p>1) For each residue in a given protein sequence of length L, a window of size M centered around this residue was chosen. We added <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mrow><mml:mo>&#x230A;</mml:mo><mml:mi>M</mml:mi><mml:mrow><mml:mo>/</mml:mo></mml:mrow><mml:mn>2</mml:mn><mml:mo>&#x230B;</mml:mo></mml:mrow></mml:math></inline-formula> zeros to each end of the protein sequence. We then computed the features associated with the structural properties, the physicochemical properties, and the evolutionary information described above for each residue within this window. The characteristic values of these calculated residues were averaged over this specific window and assigned to residues in the center of the window as their characteristic values. Therefore, each sequence was associated with a <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mn>32</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mi>L</mml:mi></mml:math></inline-formula> characteristic matrix</p>
<p><disp-formula id="eqn-1">
<label>(1)</label>
<mml:math id="mml-eqn-1" display="block"><mml:mrow><mml:mi mathvariant="bold">F</mml:mi></mml:mrow><mml:mrow><mml:mo mathvariant="bold">=</mml:mo></mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mn mathvariant="bold">1</mml:mn></mml:mrow></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mn mathvariant="bold">2</mml:mn></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x22EF;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mrow><mml:mi mathvariant="bold-italic">i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22EF;</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="bold">L</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math>
</disp-formula></p>
<p>where <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:msub><mml:mrow><mml:mtext>x</mml:mtext></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> with <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mn>1</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>i</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>L</mml:mi></mml:math></inline-formula> denotes the 32 characteristic values (five structural properties, seven physicochemical properties, and 20 evolutionary information) associated with the <italic>i</italic>-th residue. The entry <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> with <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mn>1</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>i</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>L</mml:mi></mml:math></inline-formula> in <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref> is a <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mn>32</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> vector that can be expressed as</p>
<p><disp-formula id="eqn-2">
<label>(2)</label>
<mml:math id="mml-eqn-2" display="block"><mml:msub><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mspace width="thinmathspace" /><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>&#x22EF;</mml:mo><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>&#x22EF;</mml:mo><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>32</mml:mn></mml:mrow></mml:msub><mml:mo>]</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo>,</mml:mo></mml:math>
</disp-formula></p>
<p>where <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> denotes the <italic>k</italic>-th characteristic value with <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mn>1</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>k</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mn>32</mml:mn></mml:math></inline-formula> of the <italic>i</italic>-th residue (<inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mn>1</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>i</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>L</mml:mi></mml:math></inline-formula>) assigned over the associated window. In <xref ref-type="disp-formula" rid="eqn-2">Eq. (2)</xref>, <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msub></mml:math></inline-formula> represent Shannon entropy and topological entropy associated with the <italic>i</italic>-th residue, respectively. Their computations follow the process presented in Eqs. (1) and (14) of He et al. [<xref ref-type="bibr" rid="ref-33">33</xref>].</p>
<p>2) For the <italic>i</italic>-th residue (<inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mn>1</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>i</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>L</mml:mi></mml:math></inline-formula>) in the protein sequence of length L, we varied the size of the sliding window centered around this residue to yield a feature matrix:</p>
<p><disp-formula id="eqn-3">
<label>(3)</label>
<mml:math id="mml-eqn-3" display="block"><mml:msub><mml:mrow><mml:mi mathvariant="bold">X</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x22EF;</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mtable columnalign="left" rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msubsup><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mtd><mml:mtd><mml:msubsup><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mtd><mml:mtd><mml:mo>&#x2026;</mml:mo></mml:mtd><mml:mtd><mml:msubsup><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msubsup><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mtd><mml:mtd><mml:msubsup><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mtd><mml:mtd><mml:mo>&#x22EF;</mml:mo></mml:mtd><mml:mtd><mml:msubsup><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x22EE;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x22EE;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x22F1;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x22EE;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msubsup><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mtd><mml:mtd><mml:msubsup><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mtd><mml:mtd><mml:mo>&#x2026;</mml:mo></mml:mtd><mml:mtd><mml:msubsup><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x22EE;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x22EE;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x22F1;</mml:mo></mml:mtd><mml:mtd><mml:mo>&#x22EE;</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msubsup><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>32</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mtd><mml:mtd><mml:msubsup><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>32</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mtd><mml:mtd><mml:mo>&#x2026;</mml:mo></mml:mtd><mml:mtd><mml:msubsup><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>32</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:mtd></mml:mtr></mml:mtable><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math>
</disp-formula></p>
<p>where <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msubsup><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup></mml:math></inline-formula> defined in <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref> with <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mn>1</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>l</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>t</mml:mi></mml:math></inline-formula> represents the <italic>k</italic>-th characteristic value when the sliding window of size <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> centered around <italic>i</italic>-th residue (<inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mn>1</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>i</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>L</mml:mi></mml:math></inline-formula>) is employed.</p>
<p><xref ref-type="fig" rid="fig-1">Fig. 1</xref> shows that the output of the pre-processing block yielded the feature matrix in <xref ref-type="disp-formula" rid="eqn-3">Eq. (3)</xref> of the <italic>i</italic>-th (<inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mn>1</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>i</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mi>L</mml:mi></mml:math></inline-formula>) residue. If we use <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:msub><mml:mrow><mml:mi mathvariant="bold">M</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msubsup><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mspace width="thinmathspace" /><mml:msubsup><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x2026;</mml:mo><mml:msubsup><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>n</mml:mi><mml:mrow><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> with <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mn>1</mml:mn><mml:mo>&#x2264;</mml:mo><mml:mi>k</mml:mi><mml:mo>&#x2264;</mml:mo><mml:mn>32</mml:mn></mml:math></inline-formula> to represent the <italic>k</italic>-th row of matrix <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:msub><mml:mrow><mml:mi mathvariant="bold">X</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, then we can rewrite <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:msub><mml:mrow><mml:mi mathvariant="bold">X</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">M</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x22EF;</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">M</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x22EF;</mml:mo><mml:msubsup><mml:mrow><mml:mi mathvariant="bold">M</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>32</mml:mn></mml:mrow><mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:mrow></mml:msubsup><mml:mo>]</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:mtext>T</mml:mtext></mml:mrow></mml:mrow></mml:msup><mml:mo>.</mml:mo></mml:math></inline-formula> Therefore, a set of <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:msub><mml:mrow><mml:mi mathvariant="bold">M</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">M</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>l</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">M</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>l</mml:mi><mml:mo>+</mml:mo><mml:mn>2</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mtext>&#xA0;with&#xA0;</mml:mtext></mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>4</mml:mn><mml:mo>,</mml:mo><mml:mn>7</mml:mn><mml:mo>,</mml:mo><mml:mn>10</mml:mn></mml:math></inline-formula> (<italic>i.e</italic>., the output of multi-feature1 to multi-feature4 in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>) was chosen as the input dataset, which was fed into ResNet1 to ResNet4. In addition, a set of <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msub><mml:mrow><mml:mi mathvariant="bold">M</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi mathvariant="bold">M</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi><mml:mo>,</mml:mo><mml:mn>32</mml:mn></mml:mrow></mml:msub><mml:mrow><mml:mtext>&#xA0;with&#xA0;</mml:mtext></mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>13</mml:mn></mml:math></inline-formula> was chosen as the input dataset, which was fed into ResNet5.</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Designing and Training the ResNet18 and MLP Models</title>
<p>We constructed a sufficiently deep structure, which consisted of five variant ResNet18 networks, and combined it with a constructed MLP network. <?A3B2 "fig3",5,"anchor"?><xref ref-type="fig" rid="fig-3">Fig. 3</xref> shows each of the five identical variant deep neural structures that replaced the FC layer with a dense layer containing 16 perceptrons. The outputs of the five variant networks were then concatenated and input into the MLP network that we constructed for the prediction.</p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Frame diagram of the variant ResNet18 model</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19097-fig-3.png"/>
</fig>
<p>The MLP network we constructed had two hidden layers, where the binary cross-entropy cost function was employed:</p>
<p><disp-formula id="eqn-4">
<label>(4)</label>
<mml:math id="mml-eqn-4" display="block"><mml:mi>L</mml:mi><mml:mi>o</mml:mi><mml:mi>s</mml:mi><mml:mi>s</mml:mi><mml:mo>=</mml:mo><mml:mo>&#x2212;</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mi>n</mml:mi></mml:mfrac><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi></mml:mrow></mml:munderover><mml:mrow><mml:mo>[</mml:mo><mml:msup><mml:mi>y</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:msup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>+</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mi>y</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>)</mml:mo></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>.</mml:mo></mml:math>
</disp-formula></p>
<p>In <xref ref-type="disp-formula" rid="eqn-4">Eq. (4)</xref>, <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msup><mml:mi>y</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mn>1</mml:mn><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> is the label; 0 indicates that the <italic>i</italic>-th residue is ordered, and 1 indicates that the residue is disordered. <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:msup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> is the obtained probability of the <italic>i</italic>-th residue using our constructed MLP network.</p>
<p>The output of each of our variant ResNet18 models in <xref ref-type="fig" rid="fig-1">Fig. 1</xref> was a <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>16</mml:mn></mml:math></inline-formula> matrix, and therefore, the outputs of these five matrices were concatenated to yield a <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mn>1</mml:mn><mml:mo>&#x00D7;</mml:mo><mml:mn>80</mml:mn></mml:math></inline-formula> matrix. This was then used as the input for our MLP network. The MLP network was comprised of two hidden layers, which included 75 and 15 perceptrons, respectively, and a rectified linear unit was used as the activation function. Moreover, in the two hidden layers, we adopted the dropout mechanism, which randomly drops 60% perceptrons during each iteration. One perceptron was contained in the output layer, which utilizes a sigmoid as the activation function. The structure of the MLP network is depicted in <?A3B2 "fig4",5,"anchor"?><xref ref-type="fig" rid="fig-4">Fig. 4</xref>, where the sigmoid function is used in the perceptron of the output layer.</p>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Structure of the MLP network</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19097-fig-4.png"/>
</fig>
<p>We randomly initialized the parameters of the MLP model and employed an stochastic gradient descent (SGD) optimizer to perform back-propagation to update the parameters during the training process. The output of the sigmoid function was thus used as the predicted probability <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula>, defined in <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref> of the <italic>i</italic>-th residue.</p>
<p><disp-formula id="eqn-5">
<label>(5)</label>
<mml:math id="mml-eqn-5" display="block"><mml:msup><mml:mrow><mml:mover><mml:mi>y</mml:mi><mml:mo stretchy="false">&#x005E;</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mi mathvariant="italic">S</mml:mi><mml:mi mathvariant="italic">i</mml:mi><mml:mi mathvariant="italic">g</mml:mi><mml:mi mathvariant="italic">m</mml:mi><mml:mi mathvariant="italic">o</mml:mi><mml:mi mathvariant="italic">i</mml:mi><mml:mi mathvariant="italic">d</mml:mi></mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:msup><mml:mi>a</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo stretchy="false">)</mml:mo><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mn>1</mml:mn><mml:mo>+</mml:mo><mml:msup><mml:mi>e</mml:mi><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:msup><mml:mi>a</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:mrow></mml:msup></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:math>
</disp-formula></p>
<p>where <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:msup><mml:mi>a</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>i</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula>denotes the output through the network of the <italic>i</italic>-th residue.</p>
<p>The training process with our constructed model was as follows: first, we randomly shuffled and split the training dataset into multiple batches and calculated the probabilities of all residues in each batch through the constructed network using <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref>. Subsequently, the binary cross-entropy loss function was employed to compute the loss of the given batch, to optimize the network parameters through the SGD optimization mechanism for the back-propagation process, where the learning rate was set to 0.001. The above procedure was repeated for every batch of residues until one epoch was completed (<italic>i.e</italic>., all residues were trained by the network, and probabilities were calculated). After one epoch was completed, we then randomly shuffled and split the training dataset into multiple batches again and repeated the whole procedure until the loss stopped converging or the training epoch reached the setting number. The main parameters of our algorithm are summarized in <?A3B2 "tbl1",5,"anchor"?><xref ref-type="table" rid="table-1">Table 1</xref>.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Experimental parameters used in our algorithm</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Experimental parameters</th>
<th>Values</th>
</tr>
</thead>
<tbody>
<tr>
<td>Language</td>
<td>Python</td>
</tr>
<tr>
<td>Environment</td>
<td>Google Colaboratory</td>
</tr>
<tr>
<td>ResNet18 FC layer</td>
<td>16 perceptrons</td>
</tr>
<tr>
<td>MLP network hidden layer</td>
<td>2 layers</td>
</tr>
<tr>
<td>Hidden layer in MLP</td>
<td>75 and 15 perceptrons</td>
</tr>
<tr>
<td>Dropout rate</td>
<td>60%</td>
</tr>
<tr>
<td>Initial learning rate</td>
<td>0.001</td>
</tr>
<tr>
<td>Activation function in output layer</td>
<td>Sigmoid</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>After the training process was complete and our network parameters were determined, we used the test dataset to test our network. The predicted probabilities obtained using <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref> were used to determine whether the residue was disordered or ordered. Finally, we conducted the performance evaluation.</p>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Performance Measures</title>
<p>To evaluate the accuracy of our model for predicting IDPs, we mainly used five authoritative and universal evaluation criteria in the IDP prediction field: sensitivity (Sen), specificity (Spe), binary accuracy (BAcc), weight score (Sw), and MCC. TP, FP, TN, and FN were also used to respectively denote the numbers of true disordered, false disordered, true ordered, and false ordered samples. The following formulae were used to calculate these five evaluation criteria:</p>
<p><disp-formula id="eqn-6">
<label>(6)</label>
<mml:math id="mml-eqn-6" display="block"><mml:mi>S</mml:mi><mml:mi>e</mml:mi><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-7">
<label>(7)</label>
<mml:math id="mml-eqn-7" display="block"><mml:mi>S</mml:mi><mml:mi>p</mml:mi><mml:mi>e</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac><mml:mo>,</mml:mo></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-8">
<label>(8)</label>
<mml:math id="mml-eqn-8" display="block"><mml:mi>B</mml:mi><mml:mi>A</mml:mi><mml:mi>c</mml:mi><mml:mi>c</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mn>1</mml:mn><mml:mn>2</mml:mn></mml:mfrac><mml:mrow><mml:mo>(</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac><mml:mo>+</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-9">
<label>(9)</label>
<mml:math id="mml-eqn-9" display="block"><mml:mi>S</mml:mi><mml:mi>w</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:mrow></mml:mfrac><mml:mo>+</mml:mo><mml:mfrac><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi></mml:mrow></mml:mfrac><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo></mml:math>
</disp-formula></p>
<p><disp-formula id="eqn-10">
<label>(10)</label>
<mml:math id="mml-eqn-10" display="block"><mml:mi>M</mml:mi><mml:mi>C</mml:mi><mml:mi>C</mml:mi><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>&#x00D7;</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:msqrt><mml:mrow><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>N</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>T</mml:mi><mml:mi>N</mml:mi><mml:mo>+</mml:mo><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:msqrt></mml:mfrac><mml:mo>.</mml:mo></mml:math>
</disp-formula></p>
<p>Among these, the MCC value was the most important and effective criterion to measure the performance of the prediction of IDPs, which varies from &#x2212;1 to 1. A high MCC value indicates outstanding classification performance.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Results and Discussion</title>
<sec id="s3_1">
<label>3.1</label>
<title>Comparison of Performance with Other Well-known Methods</title>
<p>For ease of description, DISRES was employed as the acronym of our deep network model. To analyze the performance of our deep network model and highlight the advantages of our method, DISRES was compared with seven other state-of-the-art methods used for predicting IDPs, using the five performance measures described above using two blind datasets: MXD494 and R80. The MCC value obtained by our trained model was 0.450 for the blind MXD494 dataset and 0.517 for the R80 dataset. Therefore, the MCC values obtained using our network model showed that our model outperformed the existing methods of DISpre, SPOT-Disorder2, RFPR-IDP, DISOPRED2, DISpro, RONN, PONDR, and FoldIndex. All relevant evaluation criteria for comparing the performance of the models using the two blind datasets, MXD494 and R80, are shown in <?A3B2 "tbl2",5,"anchor"?><xref ref-type="table" rid="table-2">Tables 2</xref> and <?A3B2 "tbl3",5,"anchor"?><xref ref-type="table" rid="table-3">3</xref>. To visualize the comparison results, <?A3B2 "fig5",5,"anchor"?><xref ref-type="fig" rid="fig-5">Figs. 5</xref> and <?A3B2 "fig6",5,"anchor"?><xref ref-type="fig" rid="fig-6">6</xref> show the comparisons of performance between different methods using the two blind datasets, respectively, where the red line represents MCC, the most important evaluation criteria, for prediction performance.</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Performance comparison of the various methods using the blind R80 dataset</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Methods</th>
<th>Sen</th>
<th>Spe</th>
<th>BAcc</th>
<th>Sw</th>
<th>MCC</th>
<th>Rank</th>
</tr>
<tr>
<th></th>
<th></th>
<th></th>
<th></th>
<th></th>
<th></th>
<th>MCC</th>
</tr>
</thead>
<tbody>
<tr>
<td>DISRES</td>
<td>0.605</td>
<td>0.937</td>
<td>0.771</td>
<td>0.542</td>
<td>0.517</td>
<td>1</td>
</tr>
<tr>
<td>RFPR-IDP</td>
<td>0.546</td>
<td>0.954</td>
<td>0.750</td>
<td>0.501</td>
<td>0.513</td>
<td>2</td>
</tr>
<tr>
<td>DISpre</td>
<td>0.748</td>
<td>0.862</td>
<td>0.805</td>
<td>0.610</td>
<td>0.471</td>
<td>3</td>
</tr>
<tr>
<td>DISOPRED2</td>
<td>0.405</td>
<td>0.972</td>
<td>0.688</td>
<td>0.377</td>
<td>0.470</td>
<td>4</td>
</tr>
<tr>
<td>SpotDisorder2</td>
<td>0.494</td>
<td>0.944</td>
<td>0.719</td>
<td>0.438</td>
<td>0.449</td>
<td>5</td>
</tr>
<tr>
<td>RONN</td>
<td>0.603</td>
<td>0.878</td>
<td>0.740</td>
<td>0.481</td>
<td>0.395</td>
<td>6</td>
</tr>
<tr>
<td>PONDR</td>
<td>0.557</td>
<td>0.816</td>
<td>0.686</td>
<td>0.373</td>
<td>0.278</td>
<td>7</td>
</tr>
<tr>
<td>FoldIndex</td>
<td>0.488</td>
<td>0.811</td>
<td>0.649</td>
<td>0.299</td>
<td>0.224</td>
<td>8</td>
</tr>
</tbody>
</table>
</table-wrap>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Performance comparison of the various methods using the blind MXD494 dataset</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th>Methods</th>
<th>Sen</th>
<th>Spe</th>
<th>BAcc</th>
<th>Sw</th>
<th>MCC</th>
<th>Rank</th>
</tr>
<tr>
<th></th>
<th></th>
<th></th>
<th></th>
<th></th>
<th></th>
<th>MCC</th>
</tr>
</thead>
<tbody>
<tr>
<td>DISRES</td>
<td>0.683</td>
<td>0.811</td>
<td>0.747</td>
<td>0.494</td>
<td>0.450</td>
<td>1</td>
</tr>
<tr>
<td>SpotDisorder2</td>
<td>0.637</td>
<td>0.819</td>
<td>0.728</td>
<td>0.457</td>
<td>0.448</td>
<td>2</td>
</tr>
<tr>
<td>RFPR-IDP</td>
<td>0.749</td>
<td>0.758</td>
<td>0.753</td>
<td>0.507</td>
<td>0.442</td>
<td>3</td>
</tr>
<tr>
<td>DISOPRED2</td>
<td>0.647</td>
<td>0.800</td>
<td>0.723</td>
<td>0.447</td>
<td>0.406</td>
<td>4</td>
</tr>
<tr>
<td>PONDR</td>
<td>0.744</td>
<td>0.698</td>
<td>0.721</td>
<td>0.442</td>
<td>0.401</td>
<td>5</td>
</tr>
<tr>
<td>RONN</td>
<td>0.664</td>
<td>0.754</td>
<td>0.709</td>
<td>0.418</td>
<td>0.368</td>
<td>6</td>
</tr>
<tr>
<td>DISpro</td>
<td>0.303</td>
<td>0.970</td>
<td>0.637</td>
<td>0.273</td>
<td>0.318</td>
<td>7</td>
</tr>
<tr>
<td>FoldIndex</td>
<td>0.602</td>
<td>0.717</td>
<td>0.659</td>
<td>0.319</td>
<td>0.278</td>
<td>8</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>As shown in <xref ref-type="table" rid="table-2">Table 2</xref>, for the blind R80 test dataset, DISRES showed superiority over the other seven well-known methods for predicting IDPs, with an MCC value of 0.517, and ranking first among all methods. DISpre is a method we developed previously in relation to MCC using the MLP network alone. <xref ref-type="fig" rid="fig-5">Fig. 5</xref> shows that DISpre achieved the highest sensitivity value but obtained lower specificity and MCC value, which resulted in poor predictive performance. The MCC value significantly improved with the addition of the ResNet18 deep neural network, and the value obtained by DISRES was on the highest point of the red line, which demonstrated that our method yielded the best performance in predicting IDPs.</p>
<fig id="fig-5">
<label>Figure 5</label>
<caption>
<title>Comparisons of the different methods using the blind R80 dataset</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19097-fig-5.png"/>
</fig>

<p>To make the results more convincing, we used another widely used blind test dataset, the MXD494, which contains more data than R80. As shown in <xref ref-type="table" rid="table-3">Table 3</xref>, DISRES achieved an MCC of 0.450, which was higher than that of several other prediction methods; moreover, it still ranked first among all methods, which can largely be attributed to the deep neural network ResNet18 applied in DISRES. Although the PONDR obtained the best specificity, as shown in <xref ref-type="fig" rid="fig-6">Fig. 6</xref>, its sensitivity and MCC value were low, which indicated that the PONDR did not perform well in predicting IDPs. DISRES still achieved the highest point on the red line of MCC, which showed that it outperformed the other methods.</p>
<fig id="fig-6">
<label>Figure 6</label>
<caption>
<title>Comparisons of the different methods using the blind MXD494 dataset</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19097-fig-6.png"/>
</fig>
</sec>
<sec id="s3_2">
<label>3.2</label>
<title>Limitations of the Current Study</title>
<p>For the first time, we used the deep neural network model, ResNet18, and combined it with an MLP network to predict disordered regions of IDPs with good performance. However, it is well known that the size of the training dataset in a neural network affects the final prediction results. Although the dataset we obtained from the authoritative DisProt database is the most recent data available, there were still only 1616 protein sequences, which limits the potential of improving the performance of our model.</p>
</sec>
</sec>
<sec id="s4">
<label>4</label>
<title>Conclusions</title>
<p>IDPs are of great importance as they play essential roles in numerous human diseases, such as certain types of cancers, genetic diseases, and Alzheimer&#x2019;s disease; moreover, they are related to many important biological activities, such as nucleic acid folding and cell signaling and regulation. Thus, developing a method that can accurately, reliably, and rapidly identify IDPs is important to understand the mechanisms underlying biological activities and study the role of IDPs in major diseases. Therefore, our proposed method for efficiently detecting IDPs has important practical implications for research on biological activities.</p>
<p>In contrast to other previously proposed methods, our model has the following advantages: (i) fewer features were selected to achieve better results than other methods; (ii) a sufficiently deep neural network, ResNet18, was introduced for the prediction and achieved the most accurate predictions; (iii) most convincingly, the MCC values obtained using our method were the highest.</p>
<p>The outperforming of our constructed model over other methods was mainly attributed to the following points: (i) the construction of a sufficiently deep structure, which consisted of five identical ResNet18 networks, and its combination with a constructed MLP network. This constructure differed from all previous methods and was a completely new architecture; (ii) Resnet18 was applied for the first time as a deep network model for predicting IDPs, which enabled the extraction of information from IDP residues in greater detail and depth than those of other methods; (iii) using two well-known datasets, MXD494 and R80, as blind test datasets, simulation results showed that the MCC values obtained using our method were 0.517 for the blind R80 dataset and 0.450 from the MXD494 dataset, which demonstrated that our method outperformed existing methods.</p>
<p>In the future, we will approach subsequent research from two perspectives: (i) the extraction of protein features; exploring additional properties of amino acids may improve prediction performance; (ii) models developed based on other deep learning methods to further improve prediction performance.</p>
</sec>
</body>
<back>
<fn-group>
<fn fn-type="other">
<p><bold>Funding Statement:</bold> The authors received no specific funding for this study.</p>
</fn>
<fn fn-type="conflict">
<p><bold>Conflicts of Interest:</bold> The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</fn>
</fn-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>1.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Deng</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Eickholt</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Cheng</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2011</year>). <article-title>A comprehensive overview of computational protein disorder prediction methods</article-title>. <source>Molecular Biosystems</source><italic>,</italic> <volume>8</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>114</fpage>&#x2013;<lpage>121</lpage>. DOI <pub-id pub-id-type="doi">10.1039/C1MB05207A</pub-id>.</mixed-citation></ref>
<ref id="ref-2"><label>2.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>B.</given-names></string-name></person-group> (<year>2017</year>). <article-title>A comprehensive review and comparison of existing computational methods for intrinsically disordered protein and region prediction</article-title>. <source>Briefings in Bioinformatics</source><italic>,</italic> <italic>(</italic><issue>1</issue><italic>),</italic> <fpage>330</fpage>&#x2013;<lpage>346</lpage>. DOI <pub-id pub-id-type="doi">10.1093/bib/bbx126</pub-id>.</mixed-citation></ref>
<ref id="ref-3"><label>3.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Holmstrom</surname>, <given-names>E. D.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Nettels</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Best</surname>, <given-names>R. B.</given-names></string-name>, <string-name><surname>Schule</surname>, <given-names>B.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Disordered RNA chaperones can enhance nucleic acid folding via local charge screening</article-title>. <source>Nature Communications</source><italic>,</italic> <volume>10</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>2453</fpage>. DOI <pub-id pub-id-type="doi">10.1038/s41467-019-10356-0</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>4.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Wright</surname>, <given-names>P. E.</given-names></string-name>, <string-name><surname>Dyson</surname>, <given-names>H. J.</given-names></string-name></person-group> (<year>2015</year>). <article-title>Intrinsically disordered proteins in cellular signaling and regulation</article-title>. <source>Nature Reviews Molecular Cell Biology</source><italic>,</italic> <volume>16</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>18</fpage>&#x2013;<lpage>29</lpage>. DOI <pub-id pub-id-type="doi">10.1038/nrm3920</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>5.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Iakoucheva</surname>, <given-names>L. M.</given-names></string-name>, <string-name><surname>Brown</surname>, <given-names>C. J.</given-names></string-name>, <string-name><surname>Lawson</surname>, <given-names>J. D.</given-names></string-name>, <string-name><surname>Obradovi</surname>, <given-names>Z.</given-names></string-name>, <string-name><surname>Dunker</surname>, <given-names>A. K.</given-names></string-name></person-group> (<year>2002</year>). <article-title>Intrinsic disorder in cell-signaling and cancer-associated proteins</article-title>. <source>Journal of Molecular Biology</source><italic>,</italic> <volume>323</volume><italic>(</italic><issue>3</issue><italic>),</italic> <fpage>573</fpage>&#x2013;<lpage>584</lpage>. DOI <pub-id pub-id-type="doi">10.1016/S0022-2836(02)00969-5</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>6.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kulkarni</surname>, <given-names>V.</given-names></string-name>, <string-name><surname>Kulkarni</surname>, <given-names>P.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Intrinsically disordered proteins and phenotypic switching: Implications in cancer</article-title>. <source>Progress in Molecular Biology and Translational Science</source><italic>,</italic> <volume>166</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>63</fpage>&#x2013;<lpage>84</lpage>. DOI <pub-id pub-id-type="doi">10.1016/bs.pmbts.2019.03.013</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>7.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Pankratz</surname>, <given-names>N.</given-names></string-name>, <string-name><surname>Nichols</surname>, <given-names>W. C.</given-names></string-name>, <string-name><surname>Elsaesser</surname>, <given-names>V. E.</given-names></string-name>, <string-name><surname>Pauciulo</surname>, <given-names>M. W.</given-names></string-name>, <string-name><surname>Foroud</surname>, <given-names>T.</given-names></string-name></person-group> (<year>2010</year>). <article-title>Alpha-synuclein and familial Parkinson&#x2019;s disease</article-title>. <source>Movement Disorders</source><italic>,</italic> <volume>24</volume><italic>(</italic><issue>8</issue><italic>),</italic> <fpage>1125</fpage>&#x2013;<lpage>1131</lpage>. DOI <pub-id pub-id-type="doi">10.1002/mds.22524</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>8.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Uversky</surname>, <given-names>V. N.</given-names></string-name>, <string-name><surname>Oldfield</surname>, <given-names>C. J.</given-names></string-name>, <string-name><surname>Midic</surname>, <given-names>U.</given-names></string-name>, <string-name><surname>Xie</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Xue</surname>, <given-names>B.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2009</year>). <article-title>Unfoldomics of human diseases: Linking protein intrinsic disorder with diseases</article-title>. <source>BMC Genomics</source><italic>,</italic> <volume>10</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>1</fpage>&#x2013;<lpage>17</lpage>. DOI <pub-id pub-id-type="doi">10.1186/1471-2164-10-S1-S7</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>9.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Uversky</surname>, <given-names>V. N.</given-names></string-name></person-group> (<year>2014</year>). <article-title>Introduction to intrinsically disordered proteins (IDPS)</article-title>. <source>Chemical Reviews</source><italic>,</italic> <volume>114</volume><italic>(</italic><issue>13</issue><italic>),</italic> <fpage>6557</fpage>&#x2013;<lpage>6560</lpage>. DOI <pub-id pub-id-type="doi">10.1021/cr500288y</pub-id>.</mixed-citation></ref>
<ref id="ref-10"><label>10.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>He</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Zhao</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Sun</surname>, <given-names>G.</given-names></string-name></person-group> (<year>2019</year>). <article-title>The prediction of intrinsically disordered proteins based on feature selection</article-title>. <source>Algorithms</source><italic>,</italic> <volume>12</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>46</fpage>. DOI <pub-id pub-id-type="doi">10.3390/a12020046</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>11.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Prilusky</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Felder</surname>, <given-names>C. E.</given-names></string-name>, <string-name><surname>Zeev-Ben-Mordehai</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Rydberg</surname>, <given-names>E. H.</given-names></string-name>, <string-name><surname>Sussman</surname>, <given-names>J. L.</given-names></string-name></person-group> (<year>2005</year>). <article-title>FoldIndex: A simple tool to predict whether a given protein sequence is intrinsically unfolded</article-title>. <source>Bioinformatics</source><italic>,</italic> <volume>21</volume><italic>(</italic><issue>16</issue><italic>),</italic> <fpage>3435</fpage>&#x2013;<lpage>3438</lpage>. DOI <pub-id pub-id-type="doi">10.1093/bioinformatics/bti537</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>12.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Rune</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Russell</surname>, <given-names>R. B.</given-names></string-name>, <string-name><surname>Victor</surname>, <given-names>N.</given-names></string-name>, <string-name><surname>Gibson</surname>, <given-names>T. J.</given-names></string-name></person-group> (<year>2003</year>). <article-title>Globplot: Exploring protein sequences for globularity and disorder</article-title>. <source>Nucleic Acids Research</source><italic>,</italic> <volume>31</volume><italic>(</italic><issue>13</issue><italic>),</italic> <fpage>3701</fpage>&#x2013;<lpage>3708</lpage>. DOI <pub-id pub-id-type="doi">10.1093/nar/gkg519</pub-id>.</mixed-citation></ref>
<ref id="ref-13"><label>13.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Galzitskaya</surname>, <given-names>O. V.</given-names></string-name>, <string-name><surname>Lobanov</surname>, <given-names>G. M. Y.</given-names></string-name></person-group> (<year>2006</year>). <article-title>Foldunfold: Web server for the prediction of disordered regions in protein chain</article-title>. <source>Bioinformatics</source><italic>,</italic> <volume>22</volume><italic>(</italic><issue>23</issue><italic>),</italic> <fpage>2948</fpage>&#x2013;<lpage>2949</lpage>. DOI <pub-id pub-id-type="doi">10.1093/bioinformatics/btl504</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>14.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lobanov</surname>, <given-names>M. Y.</given-names></string-name>, <string-name><surname>Galzitskaya</surname>, <given-names>O. V.</given-names></string-name></person-group> (<year>2011</year>). <article-title>The Ising model for prediction of disordered residues from protein sequence alone</article-title>. <source>Physical Biology</source><italic>,</italic> <volume>8</volume><italic>(</italic><issue>3</issue><italic>),</italic> <fpage>035004</fpage>. DOI <pub-id pub-id-type="doi">10.1088/1478-3975/8/3/035004</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>15.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Alyasseri</surname>, <given-names>Z. A.</given-names></string-name>, <string-name><surname>Al-Betar</surname>, <given-names>M. A.</given-names></string-name>, <string-name><surname>Doush</surname>, <given-names>I. A.</given-names></string-name>, <string-name><surname>Awadallah</surname>, <given-names>M. A.</given-names></string-name>, <string-name><surname>Abasi</surname>, <given-names>A. K.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2021</year>). <article-title>Review on COVID-19 diagnosis models based on machine learning and deep learning approaches</article-title>. <source>Expert Systems</source><italic>,</italic> <volume>80</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>1</fpage>. DOI <pub-id pub-id-type="doi">10.1111/exsy.12759</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>16.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lakhan</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Mohammed</surname>, <given-names>M. A.</given-names></string-name>, <string-name><surname>Kozlov</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Rodrigues</surname>, <given-names>J. J. P. C.</given-names></string-name></person-group> (<year>2021</year>). <article-title>Mobile-fog-cloud assisted deep reinforcement learning and blockchain-enable IoMT system for healthcare workflows</article-title>. <source>Transactions on Emerging Telecommunications Technologies</source><italic>,</italic> <volume>19</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>1</fpage>. DOI <pub-id pub-id-type="doi">10.1002/ett.4363</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>17.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Peng</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Vucetic</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Radivojac</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Brown</surname>, <given-names>C. J.</given-names></string-name>, <string-name><surname>Dunker</surname>, <given-names>A. K.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2011</year>). <article-title>Optimizing long intrinsic disorder predictors with protein evolutionary information</article-title>. <source>Journal of Bioinformatics and Computational Biology</source><italic>,</italic> <volume>3</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>35</fpage>&#x2013;<lpage>60</lpage>. DOI <pub-id pub-id-type="doi">10.1142/S0219720005000886</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>18.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ward</surname>, <given-names>J. J.</given-names></string-name>, <string-name><surname>Sodhi</surname>, <given-names>J. S.</given-names></string-name>, <string-name><surname>McGuffin</surname>, <given-names>L. J.</given-names></string-name>, <string-name><surname>Buxton</surname>, <given-names>B. F.</given-names></string-name>, <string-name><surname>Jones</surname>, <given-names>D. T.</given-names></string-name></person-group> (<year>2004</year>). <article-title>Prediction and functional analysis of native disorder in proteins from the three kingdoms of life</article-title>. <source>Journal of Molecular Biology</source><italic>,</italic> <volume>337</volume><italic>(</italic><issue>3</issue><italic>),</italic> <fpage>635</fpage>&#x2013;<lpage>645</lpage>. DOI <pub-id pub-id-type="doi">10.1016/j.jmb.2004.02.002</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>19.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname>, <given-names>Z. R.</given-names></string-name>, <string-name><surname>Thomson</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>McNeil</surname>, <given-names>P.</given-names></string-name>, <string-name><surname>Esnouf</surname>, <given-names>R. M.</given-names></string-name></person-group> (<year>2005</year>). <article-title>RONN: The bio-basis function neural network technique applied to the detection of natively disordered regions in proteins</article-title>. <source>Bioinformatics</source><italic>,</italic> <volume>21</volume><italic>(</italic><issue>16</issue><italic>),</italic> <fpage>3369</fpage>&#x2013;<lpage>3376</lpage>. DOI <pub-id pub-id-type="doi">10.1093/bioinformatics/bti534</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>20.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yang</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>He</surname>, <given-names>H.</given-names></string-name></person-group> (<year>2020</year>). <article-title>Prediction of intrinsically disordered proteins with a low computational complexity method</article-title>. <source>Computer Modeling in Engineering &#x0026; Sciences</source><italic>,</italic> <volume>125</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>111</fpage>&#x2013;<lpage>123</lpage>. DOI <pub-id pub-id-type="doi">10.32604/cmes.2020.010347</pub-id>.</mixed-citation></ref>
<ref id="ref-21"><label>21.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Tang</surname>, <given-names>Y. J.</given-names></string-name>, <string-name><surname>Pang</surname>, <given-names>Y. H.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>B.</given-names></string-name></person-group> (<year>2020</year>). <article-title>Idp-seq2seq: Identification of intrinsically disordered regions based on sequence-to-sequence learning</article-title>. <source>Bioinformatics</source><italic>,</italic> <volume>17</volume><italic>(</italic><issue>21</issue><italic>),</italic> <fpage>396</fpage>&#x2013;<lpage>404</lpage>. DOI <pub-id pub-id-type="doi">10.1093/bioinformatics/btaa667</pub-id>.</mixed-citation></ref>
<ref id="ref-22"><label>22.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Klausen</surname>, <given-names>M. S.</given-names></string-name>, <string-name><surname>Jespersen</surname>, <given-names>M. C.</given-names></string-name>, <string-name><surname>Nielsen</surname>, <given-names>H.</given-names></string-name></person-group> (<year>2019</year>). <article-title>NetSurfP-2.0: Improved prediction of protein structural features by integrated deep learning</article-title>. <source>Proteins: Structure, Function, and Bioinformatics</source><italic>,</italic> <volume>87</volume><italic>(</italic><issue>6</issue><italic>),</italic> <fpage>520</fpage>&#x2013;<lpage>527</lpage>. DOI <pub-id pub-id-type="doi">10.1002/prot.25674</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>23.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hanson</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Paliwal</surname>, <given-names>K. K.</given-names></string-name>, <string-name><surname>Litfin</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Zhou</surname>, <given-names>Y.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Spot-disorder2: Improved protein intrinsic disorder prediction by ensembled deep learning</article-title>. <source>Genomics, Proteomics and Bioinformatics</source><italic>,</italic> <volume>17</volume><italic>(</italic><issue>6</issue><italic>),</italic> <fpage>645</fpage>&#x2013;<lpage>656</lpage>. DOI <pub-id pub-id-type="doi">10.1016/j.gpb.2019.01.004</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>24.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>B.</given-names></string-name></person-group> (<year>2020</year>). <article-title>RFPR-IDP: Reduce the false positive rates for intrinsically disordered protein and region prediction by incorporating both fully ordered proteins and disordered proteins</article-title>. <source>Briefings in Bioinformatics</source><italic>,</italic> <volume>22</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>2000</fpage>&#x2013;<lpage>2011</lpage>. DOI <pub-id pub-id-type="doi">10.1093/bib/bbaa018</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>25.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hanson</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Yang</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Paliwal</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Zhou</surname>, <given-names>Y.</given-names></string-name></person-group> (<year>2017</year>). <article-title>Improving protein disorder prediction by deep bidirectional long short-term memory recurrent neural networks</article-title>. <source>Bioinformatics</source><italic>,</italic> <volume>33</volume><italic>(</italic><issue>5</issue><italic>),</italic> <fpage>685</fpage>&#x2013;<lpage>692</lpage>. DOI <pub-id pub-id-type="doi">10.1093/bioinformatics/btw678</pub-id>.</mixed-citation></ref>
<ref id="ref-26"><label>26.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Liu</surname>, <given-names>B.</given-names></string-name></person-group> (<year>2019</year>). <article-title>IDP-FSP: Identification of intrinsically disordered proteins/regions by length-dependent predictors based on conditional random fields</article-title>. <source>Molecular Therapy-Nucleic Acids</source><italic>,</italic> <volume>17</volume><italic>(</italic><issue>D1</issue><italic>),</italic> <fpage>396</fpage>&#x2013;<lpage>404</lpage>. DOI <pub-id pub-id-type="doi">10.1016/j.omtn.2019.06.004</pub-id>.</mixed-citation></ref>
<ref id="ref-27"><label>27.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mizianty</surname>, <given-names>M. J.</given-names></string-name>, <string-name><surname>Stach</surname>, <given-names>W.</given-names></string-name>, <string-name><surname>Chen</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Kedarisetti</surname>, <given-names>K. D.</given-names></string-name>, <string-name><surname>Disfani</surname>, <given-names>F. M.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2010</year>). <article-title>Improved sequence-based prediction of disordered regions with multilayer fusion of multiple information sources</article-title>. <source>Bioinformatics</source><italic>,</italic> <volume>26</volume><italic>(</italic><issue>18</issue><italic>),</italic> <fpage>i489</fpage>&#x2013;<lpage>i496</lpage>. DOI <pub-id pub-id-type="doi">10.1093/bioinformatics/btq373</pub-id>.</mixed-citation></ref>
<ref id="ref-28"><label>28.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Maysiak-Mrozek</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Baron</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Mrozek</surname>, <given-names>D.</given-names></string-name></person-group> (<year>2019</year>). <article-title>Spark-IDPP: High-throughput and scalable prediction of intrinsically disordered protein regions with spark clusters on the cloud</article-title>. <source>Cluster Computing</source><italic>,</italic> <volume>22</volume><italic>(</italic><issue>17</issue><italic>),</italic> <fpage>487</fpage>&#x2013;<lpage>508</lpage>. DOI <pub-id pub-id-type="doi">10.1007/s10586-018-2857-9</pub-id>.</mixed-citation></ref>
<ref id="ref-29"><label>29.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Kozlowski</surname>, <given-names>L. P.</given-names></string-name>, <string-name><surname>Bujnicki</surname>, <given-names>J. M.</given-names></string-name></person-group> (<year>2012</year>). <article-title>Meta-disorder: A meta-server for the prediction of intrinsic disorder in proteins</article-title>. <source>BMC Bioinformatics</source><italic>,</italic> <volume>13</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>111</fpage>. DOI <pub-id pub-id-type="doi">10.1186/1471-2105-13-111</pub-id>.</mixed-citation></ref>
<ref id="ref-30"><label>30.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Mishra</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Rastogi</surname>, <given-names>Y. P.</given-names></string-name>, <string-name><surname>Jabin</surname>, <given-names>S.</given-names></string-name></person-group> (<year>2019</year>). <article-title>A deep learning ensemble for function prediction of hypothetical proteins from pathogenic bacterial species</article-title>. <source>Computational Biology and Chemistry</source><italic>,</italic> <volume>83</volume><italic>(</italic><issue>3</issue><italic>),</italic> <fpage>107</fpage>&#x2013;<lpage>147</lpage>. DOI <pub-id pub-id-type="doi">10.1016/j.compbiolchem.2019.107147</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>31.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Liu</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Xu</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Fan</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Xu</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Zhou</surname>, <given-names>J.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2015</year>). <article-title>PSEDNA-PRO: DNA-binding protein identification by combining Chou&#x2019;s PseAAC and physicochemical distance transformation</article-title>. <source>Molecular Informatics</source><italic>,</italic> <volume>34</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>8</fpage>&#x2013;<lpage>17</lpage>. DOI <pub-id pub-id-type="doi">10.1002/minf.201400025</pub-id>.</mixed-citation></ref>
<ref id="ref-32"><label>32.</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><surname>He</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Zhang</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Ren</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Sun</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2016</year>). <article-title>Deep residual learning for image recognition</article-title>. <conf-name>IEEE Conference on Computer Vision and Pattern Recognition</conf-name>, pp. <fpage>770</fpage>&#x2013;<lpage>778</lpage>. <publisher-loc>Las Vegas, USA</publisher-loc>.</mixed-citation></ref>
<ref id="ref-33"><label>33.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>He</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Zhao</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2018</year>). <article-title>A low computational complexity scheme for the prediction of intrinsically disordered protein regions</article-title>. <source>Mathematical Problems in Engineering</source><italic>,</italic> <volume>2018</volume><italic>,</italic> <fpage>1</fpage>&#x2013;<lpage>7</lpage>. DOI <pub-id pub-id-type="doi">10.1155/2018/8087391</pub-id>.</mixed-citation></ref>
<ref id="ref-34"><label>34.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Jin</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Tan</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Jiang</surname>, <given-names>Q.</given-names></string-name>, <string-name><surname>Li</surname>, <given-names>X.</given-names></string-name>, <string-name><surname>Peng</surname>, <given-names>J.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2014</year>). <article-title>A generalized topological entropy for analyzing the complexity of DNA sequences</article-title>. <source>PLoS One</source><italic>,</italic> <volume>9</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>e88519</fpage>. DOI <pub-id pub-id-type="doi">10.1371/journal.pone.0088519</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>35.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Meiler</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>M&#x00FC;ller</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Zeidler</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Schm&#x00E4;schke</surname>, <given-names>F.</given-names></string-name></person-group> (<year>2001</year>). <article-title>Generation and evaluation of dimension-reduced amino acid parameter representations by artificial neural networks</article-title>. <source>Journal of Molecular Modeling</source><italic>,</italic> <volume>7</volume><italic>(</italic><issue>9</issue><italic>),</italic> <fpage>360</fpage>&#x2013;<lpage>369</lpage>. DOI <pub-id pub-id-type="doi">10.1007/s008940100038</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>36.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Pruitt</surname>, <given-names>K. D.</given-names></string-name>, <string-name><surname>Tatiana</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>William</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Maglott</surname>, <given-names>D. R.</given-names></string-name></person-group> (<year>2009</year>). <article-title>NCBI reference sequences: Current status, policy and new initiatives</article-title>. <source>Nucleic Acids Research</source><italic>,</italic> <volume>37</volume><italic>,</italic> <fpage>D32</fpage>&#x2013;<lpage>36</lpage>. DOI <pub-id pub-id-type="doi">10.1093/nar/gkn721</pub-id>.</mixed-citation></ref>
<ref id="ref-37"><label>37.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Peng</surname>, <given-names>Z. L.</given-names></string-name>, <string-name><surname>Kurgan</surname>, <given-names>L.</given-names></string-name></person-group> (<year>2012</year>). <article-title>Comprehensive comparative assessment of in-silico predictors of disordered regions</article-title>. <source>Current Protein &#x0026; Peptide Science</source><italic>,</italic> <volume>13</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>6</fpage>&#x2013;<lpage>18</lpage>. DOI <pub-id pub-id-type="doi">10.2174/138920312799277938</pub-id>.</mixed-citation></ref>
</ref-list>
</back>
</article>
