<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.0 20120330//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="1.0">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMC</journal-id>
<journal-id journal-id-type="nlm-ta">CMC</journal-id>
<journal-id journal-id-type="publisher-id">CMC</journal-id>
<journal-title-group>
<journal-title>Computers, Materials &#x0026; Continua</journal-title>
</journal-title-group>
<issn pub-type="epub">1546-2226</issn>
<issn pub-type="ppub">1546-2218</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">17795</article-id>
<article-id pub-id-type="doi">10.32604/cmc.2021.017795</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>An Intelligent Graph Edit Distance-Based Approach for Finding Business Process Similarities</article-title>
<alt-title alt-title-type="left-running-head">An Intelligent Graph Edit Distance-Based Approach for Finding Business Process Similarities</alt-title>
<alt-title alt-title-type="right-running-head">An Intelligent Graph Edit Distance-Based Approach for Finding Business Process Similarities</alt-title>
</title-group>
<contrib-group content-type="authors">
<contrib id="author-1" contrib-type="author">
<name name-style="western">
<surname>Sohail</surname>
<given-names>Abid</given-names>
</name>
<xref ref-type="aff" rid="aff-1">1</xref>
</contrib>
<contrib id="author-2" contrib-type="author">
<name name-style="western">
<surname>Haseeb</surname>
<given-names>Ammar</given-names>
</name>
<xref ref-type="aff" rid="aff-1">1</xref></contrib>
<contrib id="author-3" contrib-type="author" corresp="yes">
<name name-style="western">
<surname>Rehman</surname>
<given-names>Mobashar</given-names>
</name>
<xref ref-type="aff" rid="aff-2">2</xref><email>mobashar@utar.edu.my</email></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western">
<surname>Dominic</surname>
<given-names>Dhanapal Durai</given-names>
</name>
<xref ref-type="aff" rid="aff-3">3</xref></contrib>  <contrib id="author-5" contrib-type="author">
<name name-style="western">
<surname>Butt</surname>
<given-names>Muhammad Arif</given-names>
</name>
<xref ref-type="aff" rid="aff-4">4</xref></contrib>
<aff id="aff-1"><label>1</label><institution>Department of Computer Science, COMSATS University Islamabad</institution>, <addr-line>Lahore</addr-line>, <country>Pakistan</country></aff>
<aff id="aff-2"><label>2</label><institution>Faculty of Information and Communication Technology, Universiti Tunku Abdul Rahman</institution>, <addr-line>Kampar, Perak</addr-line>, <country>Malaysia</country></aff>
<aff id="aff-3"><label>3</label><institution>Universiti Teknologi PETRONAS</institution>, <addr-line>Bandar Seri Iskandar, Tronoh Perak</addr-line>, <country>Malaysia</country></aff>
<aff id="aff-4"><label>4</label><institution>Punjab University College of Information Technology, University of the Punjab</institution>, <addr-line>Lahore</addr-line>, <country>Pakistan</country></aff>
</contrib-group>
<author-notes><corresp id="cor1">&#x002A;Corresponding Author: Mobashar Rehman. Email: <email>mobashar@utar.edu.my</email></corresp></author-notes>
<pub-date pub-type="epub" date-type="pub" iso-8601-date="2021-08-23">
<day>23</day>
<month>8</month>
<year>2021</year>
</pub-date>
<volume>69</volume>
<issue>3</issue>
<fpage>3603</fpage>
<lpage>3618</lpage>
<history>
<date date-type="received">
<day>11</day>
<month>2</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>20</day>
<month>4</month>
<year>2021</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2021 Sohail et al.</copyright-statement>
<copyright-year>2021</copyright-year>
<copyright-holder>Sohail et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMC_17795.pdf"></self-uri>
<abstract>
<p>There are numerous application areas of computing similarity between process models. It includes finding similar models from a repository, controlling redundancy of process models, and finding corresponding activities between a pair of process models. The similarity between two process models is computed based on their similarity between labels, structures, and execution behaviors. Several attempts have been made to develop similarity techniques between activity labels, as well as their execution behavior. However, a notable problem with the process model similarity is that two process models can also be similar if there is a structural variation between them. However, neither a benchmark dataset exists for the structural similarity between process models nor there exist an effective technique to compute structural similarity. To that end, we have developed a large collection of process models in which structural changes are handcrafted while preserving the semantics of the models. Furthermore, we have used a machine learning-based approach to compute the similarity between a pair of process models having structural and label differences. Finally, we have evaluated the proposed approach using our generated collection of process models.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Machine learning</kwd>
<kwd>intelligent data management</kwd>
<kwd>similarities of process models</kwd>
<kwd>structural metrics</kwd>
<kwd>dataset</kwd>
<kwd>graph edit distance</kwd>
<kwd>process matching</kwd>
<kwd>artificial intelligence</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<p>Business Process Management (BPM) is composed of different techniques and methods for analysis, enactment, design, and management of business processes [<xref ref-type="bibr" rid="ref-1">1</xref>]. To achieve a competitive edge, business processes and their process models are analyzed to identify the deficiencies. Subsequently, these processes are redesigned to develop effective techniques. The benefits of business process models are manifolds [<xref ref-type="bibr" rid="ref-2">2</xref>]. They are used to support communication, documenting projects, and train employees. Typically, large organizations have hundreds or thousands of process models [<xref ref-type="bibr" rid="ref-3">3</xref>]. To effectively these process models, these organizations maintain repositories of process models. The key features of process model repositories including controlling redundant models and querying process models. The accuracy of these features depends upon the effectiveness of the underlying techniques that compute the similarity between a pair of process models [<xref ref-type="bibr" rid="ref-4">4</xref>].</p>
<p>An established study has proposed that the similarity between a pair of process models can be computed in terms of <italic>label</italic>, <italic>behavioral</italic>, and <italic>structural</italic> similarity [<xref ref-type="bibr" rid="ref-5">5</xref>]. Where <italic>label</italic> similarity relies on the textual labels of process model elements. That is, two process models are considered equivalent if a certain percentage of their activities have similar labels. In contrast, behavioral similarity techniques also take into account the execution behavior of business process models, in addition to the label similarity. As opposed to the first two types, <italic>structural</italic> similarity techniques take into consideration the topology of process models, as well as their activity labels. However, there is no benchmark collection of process models that can be used to evaluate the effectiveness of <italic>structural</italic> similarity techniques. Furthermore, it is desirable to develop techniques for the effective evaluation of structural process similarity.</p>
<p>To that end, in this study, we have developed a large collection of business process models which is composed of a substantial amount of structurally different process model variants. Furthermore, we have developed a graph edit distance-based technique to compute the similarity between a pair of process models. Lastly, we have evaluated the effectiveness of the proposed technique for its ability to detect structural changes. The rest of the paper is organized as follows: Section 2 demonstrates the details of generating our process model collection having a sufficiently large number of process variations. The structural metrics measurements and results are presented in Section 3. Section 4 presents our proposed approach for computing similarity between a pair of process models. Section 5 presents the details of our evaluation, including evaluation measures, and experimental settings. The paper concludes in Section 6.</p>
</sec>
<sec id="s2">
<label>2</label>
<title>Generating Process Model Collection</title>
<p>The first contribution of this paper is to develop the first-ever and the largest-ever collection of business process models in which the structural changes are manually induced in such a way that the semantics of the process models are not changed significantly. The benefits of the collection are the following: a) the process model collection is freely available for the research community which will be useful for fostering BPM research. And, b) the collection contains multiple variations which stem from the established literature. Hence, we contend that the developed resources will be useful for the evaluation of process similarity techniques.</p>
<p>Below, we discuss the theoretical grounds for structural changes that exist in literature and thereafter use these operations for generating process model variations. In particular, we start with the process model changes from a notable study [<xref ref-type="bibr" rid="ref-6">6</xref>] and the process flexibility patterns presented in [<xref ref-type="bibr" rid="ref-7">7</xref>]. Both the model change operations were synthesized to elicit a key set of change patterns to be used in this study. A key finding of the synthesis is that we shortlisted only those patterns that do not change the semantics of the process. For instance, from the first study [<xref ref-type="bibr" rid="ref-8">8</xref>], if the C and F variants are applied to a process model, it is likely to change the semantics of the model. Similarly, the adaptive patterns referred to as AP1, AP2, and AP4 [<xref ref-type="bibr" rid="ref-7">7</xref>], are likely to transform the semantics of a process model. Therefore, both these patterns were not used for generating process model variants.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Generating Multi-Variant Dataset</title>
<p>In order to generate a large collection of business process models having structural variations we identified five types of changes based on a comprehensive literature review. The changes are the following: Type-1: Addition of gateways, Type-2: Adjusting trivial activities, Type-3: Inserting control edges, Type-4: Reordering activities, and Type-5: Changing labels of the activities.</p>
<sec id="s2_1_1">
<label>2.1.1</label>
<title>Type 1 Change: Addition of Gateways</title>
<p>The &#x2018;<italic>Type-1</italic>&#x2019; change was proposed by multiple notable studies [<xref ref-type="bibr" rid="ref-6">6</xref>&#x2013;<xref ref-type="bibr" rid="ref-11">11</xref>]. For inducing this type of change, the original process model the largest sequence of activities from the model is changed to parallel activities by adding an &#x2018;AND&#x2019; gateway between them. The sequences which were not bounded in sequentially strict order were then converted to parallel. This change pattern is adopted from the Adaptive Pattern 9 <italic>(AP9)</italic> proposed in an existing study [<xref ref-type="bibr" rid="ref-11">11</xref>]. A similar change is also proposed in another study [<xref ref-type="bibr" rid="ref-8">8</xref>]. The details of this type and its implementation rules are explained in <xref ref-type="table" rid="table-1">Tab. 1</xref>.</p>
<table-wrap id="table-1">
<label>Table 1</label>
<caption>
<title>Creating structural variation by adding a gateway</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
</colgroup>
<tbody>
<tr>
<td>Explanation:</td>
<td>Sequential activities are converted into parallel. That is, parallelism is achieved by adding additional gateway(s).</td>
<td/>
</tr>
<tr>
<td>Illustration:</td>
<td>A sample process model is shown below to demonstrate the application of this variation. In the process model the activities labeled as T5, T6, T7, and T8 were changed to parallel activities by adding an &#x2018;AND&#x2019; gateway. Subsequently, the version is labeled as &#x2018;<italic>ACME Inc. V1&#x2019;</italic>.</td>
<td/>
</tr>
<tr>
<td>Rooted from:</td>
<td>AP9 [<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-11">11</xref>], variant B [<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-10">10</xref>].</td>
<td/>
</tr>
<tr>
<td>Rules for implementation:</td>
<td>1) Identify the core sequences in the models.2) Observe the selected sequences and choose at least one sequence.3) Convert the selected sequence into parallel by adding either XOR or AND gateway.4) Observe the semantics of the generated model. If the significant change in semantics observed the generated should not added in repository.</td>
<td/>
</tr>
<tr>
<td>Graphical representation:</td>
<td><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_17795-inline-1.png"/><italic>T1 fill request form T6 receive application formT2 send form to manager for approval    T7 review formT3 evaluate form  T8 send request to vice principal for approvalT4 reject form T9 receive approvalT5 approve form T10 receive form</italic></td>
<td/>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s2_1_2">
<label>2.1.2</label>
<title>Type 2 Change: Adjusting Trivial Activities</title>
<p>The second type of variation is referred to as <italic>&#x201C;Type-2&#x201D;</italic> structural variation. This type of change stems from several notable studies [<xref ref-type="bibr" rid="ref-6">6</xref>&#x2013;<xref ref-type="bibr" rid="ref-11">11</xref>]. A key feature of this change is that it involves inserting or deleting activities from process models. The complete representation of Type-2 change is shown in <xref ref-type="table" rid="table-2">Tab. 2</xref>. At least one activity was deleted or inserted from the model given that the semantics of the original model remained unchanged. Trivial or unimportant activities were first selected and then only one activity was deleted. The Type-2 change pattern was adopted after combing the adaptive patterns AP1, AP2, and variation D proposed by [<xref ref-type="bibr" rid="ref-8">8</xref>,<xref ref-type="bibr" rid="ref-9">9</xref>].</p>
<table-wrap id="table-2">
<label>Table 2</label>
<caption>
<title>Structure variation by adjusting trivial activities</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
</colgroup>
<tbody>
<tr>
<td>Explanation:</td>
<td>Deletion or insertion of an trivial activity.</td>
<td/>
</tr>
<tr>
<td>Illustration:</td>
<td>In the process model shown below a new activity labeled T10 was added between activity T2 and T3. Whereas, the remaining semantics of process models remain the same. The activity label of T10 &#x2018;<italic>receive form</italic>&#x2019; was added before activity T3 &#x2018;<italic>evaluate form</italic>.&#x2019;</td>
<td/>
</tr>
<tr>
<td>Rooted from:</td>
<td>AP1, AP2 [<xref ref-type="bibr" rid="ref-7">7</xref>] variant D [<xref ref-type="bibr" rid="ref-10">10</xref>].</td>
<td/>
</tr>
<tr>
<td>Rules for implementation:</td>
<td>1) In the 1<inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:msup><mml:mrow></mml:mrow><mml:mrow><mml:mstyle mathvariant="normal"><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mstyle></mml:mrow></mml:msup></mml:math></inline-formula> step, two potential activities were selected for insertion or deletion. Based on the analysis of the model, new activities were added.2) The deletion of trivial activities was also done in the case where the insertion of new activity was not possible.</td>
<td/>
</tr>
<tr>
<td>Graphical representation:</td>
<td><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_17795-inline-2.png"/><italic>T1 fill request form T6 receive application formT2 send form to manager for approval    T7 review formT3 evaluate form  T8 send request to vice principal for approvalT4 reject form T9 receive approvalT5 approve form T10 receive form</italic></td>
<td/>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s2_1_3">
<label>2.1.3</label>
<title>Type 3 Change: Inserting Control Edges</title>
<p>In Type-3 change, a control edge or control flow was inserted into the process model under consideration. This change is rooted in several existing studies [<xref ref-type="bibr" rid="ref-6">6</xref>&#x2013;<xref ref-type="bibr" rid="ref-11">11</xref>]. More specifically, the AP11, AP12, variant C, variant E, and variant F, proposed in the studies were considered for the Type-3 change. The details of the Type-3 change are shown in <xref ref-type="table" rid="table-3">Tab. 3</xref>.</p>
<table-wrap id="table-3">
<label>Table 3</label>
<caption>
<title>Structure variation by inserting control edge</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
</colgroup>
<tbody>
<tr>
<td>Explanation:</td>
<td>Insertion of new control edges in the process model.</td>
<td/>
</tr>
<tr>
<td>Illustration:</td>
<td>Inserting a control edge to the process model without changing its semantics was a challenging task. This type of change was performed by adding a control flow edge.</td>
<td/>
</tr>
<tr>
<td>Rooted from:</td>
<td>AP11, AP12 [<xref ref-type="bibr" rid="ref-9">9</xref>], variant C, E, F [<xref ref-type="bibr" rid="ref-10">10</xref>].</td>
<td/>
</tr>
<tr>
<td>Rules for implementation:</td>
<td>1) The edge or flow line exists between two nodes which could be activities or gateways.2) The potential gateway nodes were selected as a candidate for adding an edge.3) The inserted edges were also labeled.</td>
<td/>
</tr>
<tr>
<td>Graphical representation:</td>
<td><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_17795-inline-3.png"/><italic>T1 fill request form T6 receive application formT2 send form to manager for approval    T7 review formT3 evaluate form  T8 send request to vice principal for approvalT4 reject form T9 receive approvalT5 approve form T10 receive form</italic></td>
<td/>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s2_1_4">
<label>2.1.4</label>
<title>Type-4 Change: Reordering Activities</title>
<p>The Type-4 change is also rooted in studies [<xref ref-type="bibr" rid="ref-6">6</xref>&#x2013;<xref ref-type="bibr" rid="ref-11">11</xref>] by mapping the Weber et. al. [<xref ref-type="bibr" rid="ref-7">7</xref>], adaptive patterns AP5, as well as in the two variants (G and H) proposed in [<xref ref-type="bibr" rid="ref-9">9</xref>]. The overall formation of Type-4 change is shown in <xref ref-type="table" rid="table-4">Tab. 4</xref>. This type of change was achieved by reordering the activities. The Type-4 change was applied to all 150 models.</p>
<table-wrap id="table-4">
<label>Table 4</label>
<caption>
<title>Structure variation via reordering of activities</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
</colgroup>
<tbody>
<tr>
<td>Explanation:</td>
<td>Reordering of activity/activities.</td>
<td/>
</tr>
<tr>
<td>Illustration:</td>
<td>In the process model shown below, two activities T7 &#x2018;<italic>review form</italic>&#x2019; and T5 &#x2018;<italic>evaluate form</italic>,&#x2019; were swapped.</td>
<td/>
</tr>
<tr>
<td>Rooted from:</td>
<td>AP5 [<xref ref-type="bibr" rid="ref-9">9</xref>], variant G, H [<xref ref-type="bibr" rid="ref-10">10</xref>].</td>
<td/>
</tr>
<tr>
<td>Rules for implementation:</td>
<td>1) A set of activities that can be swapped, was selected. The order of the process model activities and semantics of the model, were observed. The change was made in such a way that it would not affect the overall semantics.2) In some cases, where swapping was not possible in the original version, Type-2 variation was performed.</td>
<td/>
</tr>
<tr>
<td>Graphical Representation:</td>
<td><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_17795-inline-4.png"/><italic>T1 fill request form T6 receive application formT2 send form to manager for approval    T7 review formT3 evaluate form  T8 send request to vice principal for approvalT4 reject form T9 receive approvalT5 approve form T10 receive form</italic></td>
<td/>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s2_1_5">
<label>2.1.5</label>
<title>Type-5 Change: Changing Labels of the Activities</title>
<p>Type-5 involves changing labels for the process activities by replacing them with suitable synonyms in such a way that the meaning of the label remains the same. Change Type-5 which stems from adaptive pattern AP4 [<xref ref-type="bibr" rid="ref-7">7</xref>]. According to [<xref ref-type="bibr" rid="ref-7">7</xref>], the model could be different if its elements are labeled differently. Labels were changed by following some defined rules shown below in <xref ref-type="table" rid="table-5">Tab. 5</xref>. The change was made in the original models while keeping in mind their semantics.</p>
<table-wrap id="table-5">
<label>Table 5</label>
<caption>
<title>Structure variation changing labels of the activities</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
</colgroup>
<tbody>
<tr>
<td>Explanation:</td>
<td>This variation adopts the changes in the selected labels of activities.</td>
<td/>
</tr>
<tr>
<td>Illustration:</td>
<td>In the example model labels of three activities are changed.</td>
<td/>
</tr>
<tr>
<td>Rooted from:</td>
<td>AP4 [<xref ref-type="bibr" rid="ref-9">9</xref>]</td>
<td/>
</tr>
<tr>
<td>Rules for implementation:</td>
<td>1) The labels of activities were changed, whereas, the trace of the process model remained unchanged.2) In case there are less than 6 activities in a line, the label of one activity is changed.3) If there are more than 6 activities in a lane, labels of two activities were changed.4) The change in the label should be in such a way that the semantics would not change.</td>
<td/>
</tr>
<tr>
<td>Graphical representation:</td>
<td><graphic mimetype="image" mime-subtype="png" xlink:href="CMC_17795-inline-5.png"/><italic>T1 fill request form T6 receive application formT2 send form to manager for approval    T7 review formT3 evaluate form  T8 send request to vice principal for approvalT4 reject form T9 receive approvalT5 approve form T10 receive form</italic></td>
<td/>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Creation of Repository for Process Model Similarities Multi-Variants</title>
<p>As a starting point for the development of the corpus, we used an existing collection of 150 process models [<xref ref-type="bibr" rid="ref-12">12</xref>]. The choice of the collection stems from the following reasons: a) the collection has limited propriety issues as it is freely and publicly available, b) the label-based variants have already been generated and used for the process model matching tasks. Therefore, extending the dataset with structural variants will be a valuable addition to the already usable collection, c) the collection includes process models from diverse genres, hence, providing enough diversity of process models, and d) the models in the collection comply with the widely used process modeling guidelines, which states that the process models in the collection do not contain any errors. For instance, there are no connector mismatches which resulted in no cyclic complexity, and all these models have a single start and an end node, which makes the model more structured.</p>
<p>In the second step of the development, a random sample of 10% of the models was refined by a team of three experts. Specifically, the process model collection was divided into two parts and two researchers were asked to generate variants of models, in such a way that each researcher generate at least two variants of each model. Note, generating these variations was a challenging and resource-intensive task due to several reasons. For instance, generating a process model variant based on Type-4 change was a challenging task as it involves changing the position of the activity which is likely to affect the semantics of the process. All five types of changes were made to create a repository of 900 process models. Where, Type-0 represents an original model, and Type-1, Type-2, Type-3, Type-4, and Type-5 represent the five variants of the original process models. All the models were designed in a widely used process modeling tool and stored in XML format. Also, PNG files of all the models were generated for visualization.</p>
<p><xref ref-type="table" rid="table-6">Tab. 6</xref> provides a comparison of the newly developed process model collection with the existing collections that are publicly available. It includes three process model collections from the Process Model Matching Contest (PMMC&#x2019;15) [<xref ref-type="bibr" rid="ref-13">13</xref>], a state-of-the-art collection of process models, and our newly developed collection. It can be observed from the table that the number of process models in our collection is significantly more than the number of models in any of the existing collections. Furthermore, similar to the existing collections, our collection is also publicly available, and the models are designed in BPMN, which is the de jure for process modeling. Also, our collection includes process models from multiple genres, meaning that the collection contains process models from different domains making it a representative sample of several genres. A notable observation is that most of the existing collections do not include variants of process models. The only exception is a recently developed collection of process models [<xref ref-type="bibr" rid="ref-13">13</xref>]. However, the variations of the models in those collections are limited to the paraphrasing of labels. In contrast to the existing collections, our collection contains variants of process models, including structural and label-based variants, making it the most comprehensive collection that is publicly available. Hence, we contend that the developed collection is a comprehensive resource for the evaluation of the process similarity techniques.</p>
<table-wrap id="table-6">
<label>Table 6</label>
<caption>
<title>Comparison of our process collection with other collections</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th></th>
<th colspan="5">Process model collection</th>
</tr>
<tr>
<th>Criteria</th>
<th>UA</th>
<th>BR</th>
<th>AM</th>
<th>PMC</th>
<th>Our</th>
</tr>
</thead><tbody>
<tr>
<td>No of models</td>
<td>32</td>
<td>8</td>
<td>8</td>
<td>600</td>
<td>900</td>
</tr>
<tr>
<td>Public availability</td>
<td>Yes</td>
<td>Yes</td>
<td>Yes</td>
<td>Yes</td>
<td>Yes</td>
</tr>
<tr>
<td>Format of models</td>
<td>JSON</td>
<td>JSON</td>
<td>JSON</td>
<td>JSON</td>
<td>JSON</td>
</tr>
<tr>
<td>Modelling Language</td>
<td>BPMN</td>
<td>&#x2013;</td>
<td>Petri net</td>
<td>BPMN</td>
<td>BPMN</td>
</tr>
<tr>
<td>Genre wise diversity</td>
<td>No</td>
<td>No</td>
<td>No</td>
<td>Yes</td>
<td>Yes</td>
</tr>
<tr>
<td>Compliance of guidelines</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>&#x2013;</td>
<td>Yes</td>
<td>Yes</td>
</tr>
<tr>
<td>Variation of models</td>
<td>No</td>
<td>No</td>
<td>No</td>
<td>Yes</td>
<td>Yes</td>
</tr>
<tr>
<td>Structural variation</td>
<td>No</td>
<td>No</td>
<td>No</td>
<td>No</td>
<td>Yes</td>
</tr>
<tr>
<td>Label variation</td>
<td>No</td>
<td>No</td>
<td>No</td>
<td>Yes</td>
<td>Yes</td>
</tr>
</tbody>
</table>
</table-wrap>
<p>The models were saved in XML format along with the Portable Network Graphics (PNG) files. The compatible XML format was generated using Camunda, an open-source and established tool for designing process models. The 900 process models were stored after a comprehensive audit of XML codes and graphical models. The models in XML were passed as an input to the developed prototype. The whole collection of process models was passed to our developed prototype and the scores of 26 similarity metrics were computed.</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Structural Metrics Measurements and Results</title>
<p>To provide an overview of the structural properties of the collection of our process model collection, we use structural metrics. These metrics have been widely used in literature for analyzing the structural properties of the process model collection. The metrics we used to evaluate our process models were extracted from studies [<xref ref-type="bibr" rid="ref-14">14</xref>&#x2013;<xref ref-type="bibr" rid="ref-18">18</xref>]. The developed prototype is shown in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>. The 1<inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:msup><mml:mrow></mml:mrow><mml:mrow><mml:mstyle mathvariant="normal"><mml:mi>s</mml:mi><mml:mi>t</mml:mi></mml:mstyle></mml:mrow></mml:msup></mml:math></inline-formula> module of the developed tool was used to calculate the structural similarities using 26 structural similarity metrics [<xref ref-type="bibr" rid="ref-8">8</xref>].</p>
<p>To provide an overview of the process model collection, we computed the values of structural metrics of each process model and exported them to a CSV file. The standard deviation of all the structural metrics is presented in <xref ref-type="fig" rid="fig-2">Fig. 2</xref>. In the figure, the blue bars represent the mean score of structural metrics, whereas, the standard deviation is represented by red bars. Furthermore, the 26 metrics are along the x-axis, whereas, the values of these metrics are plotted along the y-axis. It can be observed from <xref ref-type="fig" rid="fig-2">Fig. 2</xref> that the mean values of metrics, such as Size, Diameter, S(N), and S(F), are comparatively higher than the other metrics. On the other hand, the mean and standard deviation values of metrics like Density, Connector Mismatch, Cyclicity, S(C)OR, S(J)OR, and S(S)OR is zero, from these values one cannot predict how one model is congruous to the other.</p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>Structural metrics calculation tool</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_17795-fig-1.png"/>
</fig>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Mean and standard deviation of the results from structural metrics</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_17795-fig-2.png"/>
</fig>
</sec>
<sec id="s4">
<label>4</label>
<title>The Proposed Approach for Structural Similarities Measurement</title>
<p>The scores of the structural metrics presented in the preceding section provide an overview of the process collection that we have generated. However, the values of the structural metrics cannot be used to compute the similarity between a pair of process models due to the following reasons: a) these metrics provide a macro view of the process structure and does not provide detailed insights about process models, b) these metrics do not provide a holistic view about the structure of two different process models. For instance, two process models having identical values of size, diameter, and other metrics, but at the same time their semantics can be different, whereas, the two process models having substantially different structural properties can have significantly similar semantics.</p>
<p>In this study, we have proposed a novel approach that relies on structural variation topology, as well as label-content of process models. The proposed approach stems from the notation of Graph Edit Distance (GED) [<xref ref-type="bibr" rid="ref-10">10</xref>]. To elaborate on the approach, firstly, we define a Business Process Graph (BPG) to omit the language-specific details. A BPG is composed of three elements, a finite set of nodes, a finite set of edges between these nodes, and labels associated with them.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Proposed Algorithm</title>
<p>The pseudocode of the proposed approach is presented below. The algorithm is subdivided into two functions. The first function &#x2018;FUNCTION-1 Maplanes()&#x2019; was developed to retrieve the activities, flow line, gateways, and pool lanes, of the model. Whereas, the second function &#x2018;FUNCTION-2 CalculateSimilarity()&#x2019; computes the similarities of a model with a single model and with all 899 models from the dataset. The proposed approach is implemented in Java.</p>
<table-wrap id="table-7">
<label>Table 7</label>
<table>
<colgroup>
<col/>
</colgroup>
<tbody>
<tr>
<td>1: Start</td>
</tr>
<tr>
<td>2: List: mappednodes [250][4];</td>
</tr>
<tr>
<td>3: Variables: Model1 <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> InputModel1 Model2 <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> InputModel2</td>
</tr>
<tr>
<td>3: Variables: Lane11 <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> Model1.lane Lane2 <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> Model2.lane</td>
</tr>
<tr>
<td>3: Maplanes() <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:msup><mml:mrow></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula>FUNCTION-1</td>
</tr>
<tr>
<td>4: CalculateSimilarity() <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:msup><mml:mrow></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula>FUNCTION-2</td>
</tr>
<tr>
<td>5: End</td>
</tr>
</tbody>
</table>
</table-wrap>
 
<table-wrap id="table-8">
<label>Table 8</label>
<table>
<colgroup>
<col/>
</colgroup>
<tbody>
<tr>
<td><bold>FUNCTION-1 Maplanes()</bold></td>
</tr>
<tr>
<td>1: Start</td>
</tr>
<tr>
<td>2: Declare variable mappedlanes [50] [4] ml <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> 0</td>
</tr>
<tr>
<td>3: Repeat1 until (End of Lane1)</td>
</tr>
<tr>
<td>4: Variable: key1 <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> Lane1.key Value1 <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> Lane1.value</td>
</tr>
<tr>
<td>5: Repeat2 until (End of Lane2)</td>
</tr>
<tr>
<td>6: Variable: key2 <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> Lane2.key Value2 <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> Lane2.value</td>
</tr>
<tr>
<td>7: if (Value1 = = Value2)</td>
</tr>
<tr>
<td>8: Mappedlanes[ml][0] <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> Key1</td>
</tr>
<tr>
<td>9: Mappedlanes[ml][1] <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> Value1</td>
</tr>
<tr>
<td>10: Mappedlanes[ml][2] <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> Key2</td>
</tr>
<tr>
<td>11: Mappedlanes[ml][3] <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> Value2</td>
</tr>
<tr>
<td>12: ml <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> ml + 1</td>
</tr>
<tr>
<td>13 End Repeat2</td>
</tr>
<tr>
<td>14 End Repeat1</td>
</tr>
<tr>
<td>13: End</td>
</tr>
</tbody>
</table>
</table-wrap>
 
<table-wrap id="table-9">
<label>Table 9</label>
<table>
<colgroup>
<col/>
</colgroup>
<tbody>
<tr>
<td><bold>FUNCTION-2 CalculateSimilarity()</bold></td>
</tr>
<tr>
<td>1: Start</td>
</tr>
<tr>
<td>2: Variables: SN <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> 0 SE <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> 0 SB <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> 0 GED <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> 0 GEDSim <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> 0</td>
</tr>
<tr>
<td>3: Variables: mappednodes [250][4] SNV <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> 0 SEV <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> 0 SBV <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> 0</td>
</tr>
<tr>
<td>4: If (No mapped Lanes)</td>
</tr>
<tr>
<td>5: SN <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> Model1.Size + Model2.Size</td>
</tr>
<tr>
<td>6: SE <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> Model1.Edges + Model2.Edges</td>
</tr>
<tr>
<td> SN <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> 0</td>
</tr>
<tr>
<td>7: Else</td>
</tr>
<tr>
<td>8: Check Tasks that have same labels in mapped lanes</td>
</tr>
<tr>
<td>9: Check Gateways that have same types (Exclusive or Parallel which includes</td>
</tr>
<tr>
<td> both split and join)</td>
</tr>
<tr>
<td>10: Check Events that have same types (Start or End Event) in mapped lanes</td>
</tr>
<tr>
<td>11: Check Tasks that have different label in mapped lanes based on previous</td>
</tr>
<tr>
<td> and next nodes</td>
</tr>
<tr>
<td>12: Check Tasks that have different label in mapped lanes based on already</td>
</tr>
<tr>
<td> mapped nodes</td>
</tr>
<tr>
<td>13: Check Remaining Tasks in mapped lanes and map them</td>
</tr>
<tr>
<td>14: Check Remaining Gateways in mapped lanes and map them</td>
</tr>
<tr>
<td>15: Check Remaining Events in mapped lanes and map them</td>
</tr>
<tr>
<td>16: Remove all the mapped Tasks from the Tasks list of both Models</td>
</tr>
<tr>
<td>17: Remove all the mapped Gateways (Parallel and Exclusive) from the Gateways</td>
</tr>
<tr>
<td> lists.</td>
</tr>
<tr>
<td>18: Remove all the mapped Events (Start and End Event) from the Events list of</td>
</tr>
<tr>
<td> both Models</td>
</tr>
<tr>
<td>19: SN <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> Model1.Tasks + Model1.Events + Model1.Gateways + Model2.Tasks</td>
</tr>
<tr>
<td> + Model2.Events + Model2.Gateways</td>
</tr>
<tr>
<td>20: Variable: i <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> 0 j <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> 0 k <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> 0 mn1 mn2 mn3 mn4</td>
</tr>
<tr>
<td>21: SE <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> Model1.Edges+Model2.Edges</td>
</tr>
<tr>
<td>22: Repeat1 until (i &#x003C; MappedNodes.Size)</td>
</tr>
<tr>
<td>23: mn1 <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> MappedNodes[i][0] mn2 <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> MappedNodes[i] [2]</td>
</tr>
<tr>
<td>24: Repeat2 until (j &#x003C; MappedNodes.Size)</td>
</tr>
<tr>
<td>25: mn3 <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> MappedNodes[j][0] mn4 <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> MappedNodes[j] [2]</td>
</tr>
<tr>
<td>26: Repeat3 until (k &#x003C; Edges)</td>
</tr>
<tr>
<td>27: If (Model1.Edge.Start = = mn1 AND Model1.Edge.End = =</td>
</tr>
<tr>
<td> mn3 AND</td>
</tr>
<tr>
<td>28: Model2.Edge.Start = = mn2 AND Model2.Edge.END = = mn4)</td>
</tr>
<tr>
<td>29: SE <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> SE-2</td>
</tr>
<tr>
<td>30: End Repeat3</td>
</tr>
<tr>
<td>31: End Repeat2</td>
</tr>
<tr>
<td>32: End Repeat1</td>
</tr>
<tr>
<td>33: Variable: i <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> 0</td>
</tr>
<tr>
<td>34: Repeat4 until (i &#x003C; MappedNodes.Size)</td>
</tr>
<tr>
<td>35: if (MappedNodes[i][1]!=MappedNodes[i][3])</td>
</tr>
<tr>
<td>36: Variable: Label1 <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> MappedNodes[i][1]</td>
</tr>
<tr>
<td>37: Variable: Label2 <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> MappedNodes[i][3]</td>
</tr>
<tr>
<td>38: Split Labels into array of words</td>
</tr>
<tr>
<td>39: Variable: max <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> Count(Label with max alphabets)</td>
</tr>
<tr>
<td>40: Remove the words that are same in both labels</td>
</tr>
<tr>
<td>41: Variable: tempsb <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> Total letters in both labels</td>
</tr>
<tr>
<td>42: Map words on their positions with same indexes</td>
</tr>
<tr>
<td>43: If (same letters are found at some index)</td>
</tr>
<tr>
<td>44: Tempsb <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> tempsb-2</td>
</tr>
<tr>
<td>45: SB <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> SB + (tempsb/max)</td>
</tr>
<tr>
<td>46: End Repeat4</td>
</tr>
<tr>
<td>47: GED <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> SN + SE + (2 <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mo>*</mml:mo></mml:math></inline-formula> SB)</td>
</tr>
<tr>
<td>48: SNV <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> SN/(Model1.Size + Model2.Size) Step 8: SEV <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> SE/(Model1.Edges +</td>
</tr>
<tr>
<td> Model2.Edges)</td>
</tr>
<tr>
<td>49: SBV <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> SB/(Model1.Size + Model2.Size-SN)</td>
</tr>
<tr>
<td>50: GEDSim <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mo>&#x2190;</mml:mo></mml:math></inline-formula> 1 &#x2212; [(SNV + SEV + SBV)/3]</td>
</tr>
<tr>
<td>51: End</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>Implementation</title>
<p>A screenshot of the implemented prototype is presented in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>. In essence, the lanes in the extracted process model are mapped based on the labels of the two models. Subsequently, the corresponding nodes of the mapped lanes are also be mapped. Finally, the overall similarity between the two process models is computed based on the mapped elements. Consider, while computing the similarity, the unmapped nodes are represented by SN, whereas, unmapped edges are represented SE. Furthermore, the edit distance SB value was computed. Using the three values, SN, SE, and SB, we calculate the SNV, SEV, and SBV. Where, SNV is the ratio between the unmapped nodes SN and the total number of nodes between the two models, and SEV is the ratio between the unmapped edges SE and the total number of edges in the two models. Furthermore, SBV calculates the betweenness of two mapped nodes in the models. A separate module contains the list of models to be used as a query, it provides a preview of the model that is selected. The lower part of the module contains the models that are similar to the query model, as well as the intermediate computation. The Precision and Recall of the selected model show how precise the results are for that specific model. In the following module, the similarity between a pair of process models is provided. The screenshot contains a preview of the pair of process models and the labels of the corresponding elements of the process model. Also, it contains the different scores, SN, SB, SE, SNV, SEV, and SBV, used to compute the similarity between a pair of models.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Experimental Results</title>
<p><xref ref-type="fig" rid="fig-4">Fig. 4</xref> shows the module to compute Precision and Recall scores of a proposed technique. The two measures, Precision and Recall, have been widely used for information retrieval, information matching, and similarity computing tasks. Precision is defined as the ratio between the number of process models that are correctly declared similar and the process models that are declared similar by the technique. The Recall is defined as the ratio between the numbers of process models that are correctly declared similar. Formally, the two measures are defined as follows.</p>
<p><italic>Precision = &#x7c;(similar process models) <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mo>&#x2229;</mml:mo></mml:math></inline-formula> (declared similar process models)&#x7c;/&#x7c;(declared similar process models)&#x7c;</italic></p>
<p><italic>Recall = &#x7c;(similar process models) <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mo>&#x2229;</mml:mo></mml:math></inline-formula> (declared similar process models)&#x7c;/&#x7c;(similar process models)&#x7c;</italic></p>
<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Individual matcher screen</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_17795-fig-3.png"/>
</fig>
<fig id="fig-4">
<label>Figure 4</label>
<caption>
<title>Results screen 
 
</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMC_17795-fig-4.png"/>
</fig>
<p>Experiments were performed using all the 900 process models. Based on the results we observed that as the precision increases the value of recall decreases in both cases (GED normalized and GED similarity). The optimum values for the thresholds are shown in <xref ref-type="table" rid="table-7">Tab. 7</xref>. The higher precision and recall results were observed at the threshold of 0.7. Furthermore, to obtain the optimum value execution of the program was readjusted with different weights for <italic>SNV, SEV, and SEB</italic>, and values were recalculated. The precision results were observed high but the recall values were observed on the lower side in comparison of GED similarity to GED normalized. The weight value 1-1-1 means the variables <italic>SNV, SEV and SEB</italic> were treated equally. The optimal results were found at 2-1-1 where the precision and recall results were observed high for both cases.</p>
<table-wrap id="table-10">
<label>Table 10</label>
<caption>
<title>Results of the implemented algorithms</title>
</caption>
<table>
<colgroup>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
<col/>
</colgroup>
<thead>
<tr>
<th colspan="4">Optimum threshold for GED similarity</th>
<th colspan="4">Optimum threshold for GED normalized</th>
</tr>
<tr>
<th>Weights</th>
<th>Threshold</th>
<th>Precision</th>
<th>Recall</th>
<th>Weights</th>
<th>Threshold</th>
<th>Precision</th>
<th>Recall</th>
</tr>
</thead><tbody>
<tr>
<td>1-1-1</td>
<td>0.7</td>
<td>0.97</td>
<td>0.88</td>
<td>1-1-1</td>
<td>0.7</td>
<td>0.97</td>
<td>0.41</td>
</tr>
<tr>
<td>2-1-1</td>
<td>0.8</td>
<td>0.99</td>
<td>0.85</td>
<td>2-1-1</td>
<td>0.7</td>
<td>1.0</td>
<td>0.41</td>
</tr>
<tr>
<td>3-2-1</td>
<td>0.7</td>
<td>0.95</td>
<td>0.91</td>
<td>3-2-1</td>
<td>0.7</td>
<td>0.88</td>
<td>0.41</td>
</tr>
</tbody>
</table>
</table-wrap>
</sec>
<sec id="s6">
<label>6</label>
<title>Conclusion</title>
<p>Several attempts have been made to develop techniques for computing similarity between process models. However, a key challenge is that there is a scarcity of process model collections that can be used to evaluate the effectiveness of process similarity techniques. Furthermore, to the best of our knowledge, there is no publicly available collection of process models that can be used to evaluate the effectiveness of process similarity techniques between structurally different process models. To that end, in this study, we have developed a large collection of 900 process models having substantially different structures but at the same time having similar semantics. To demonstrate that we have contributed a valuable resource, we have compared the specification of our developed corpus with the existing collections. The results show that our newly developed collection includes diverse processes, and the specifications of our collection are superior than the existing ones. We have also developed a technique for computing similarity between a pair of process models. The technique relies on the use of graph edit distance and similarity between labels. To demonstrate the applicability of the proposed approach, we have implemented a prototype that is composed of several modules. It includes a module that can parse XML format, import a process model, and compute structural properties of the input model. Another module computes the similarity score between a query process model with all the models in the collection. Also, it computes Precision and Recall scores. Furthermore, a third module provides details of each process model pair i.e., it identifies the corresponding lanes, as well as their corresponding activities. Finally, we evaluated the effectiveness of the proposed approach using the developed collection. The results show that the proposed technique achieved a very high effectiveness score of 0.95.</p>
</sec>
</body>
<back>
<fn-group><fn fn-type="other"><p><bold>Funding Statement:</bold> The authors received no specific funding for this study.</p></fn>
<fn fn-type="conflict"><p><bold>Conflicts of Interest:</bold> The authors declare that they have no conflicts of interest to report regarding the present study.</p></fn></fn-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>[1]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>W. M. P.</given-names> <surname>van der Aalst</surname></string-name></person-group>, &#x201C;<article-title>Business process management: A comprehensive survey</article-title>,&#x201D; <source>ISRN Software Engineering</source>, vol. <volume>2013</volume>, no. <issue>1</issue>, pp. <fpage>1</fpage>&#x2013;<lpage>37</lpage>.</mixed-citation></ref>
<ref id="ref-2"><label>[2]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Indulska</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Green</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Recker</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Rosemann</surname></string-name></person-group>, &#x201C;<article-title>Business process modeling: Perceived benefits</article-title>,&#x201D; in <conf-name>Int. Conf. on Conceptual Modeling</conf-name>, pp. <fpage>458</fpage>&#x2013;<lpage>471</lpage>, <year>2009</year>.</mixed-citation></ref>
<ref id="ref-3"><label>[3]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Polpinij</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Ghose</surname></string-name> and <string-name><given-names>H. K.</given-names> <surname>Dam</surname></string-name></person-group>, &#x201C;<article-title>Mining business rules from business process model repositories</article-title>,&#x201D; <source>Business Process Management Journal</source>, vol. <volume>21</volume>, no. <issue>4</issue>, pp. <fpage>820</fpage>&#x2013;<lpage>836</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-4"><label>[4]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>E.</given-names> <surname>Kuss</surname></string-name>, <string-name><given-names>H.</given-names> <surname>Leopold</surname></string-name>, <string-name><given-names>H.</given-names> <surname>van der Aa</surname> </string-name>, <string-name><given-names>H.</given-names> <surname>Stuckenschmidt</surname></string-name> and <string-name><given-names>H. A.</given-names> <surname>Reijers</surname></string-name></person-group>, &#x201C;<article-title>A probabilistic evaluation procedure for process model matching techniques</article-title>,&#x201D; <source>Data Knowledge Engineering</source>, vol. <volume>117</volume>, no. <issue>4</issue>, pp. <fpage>393</fpage>&#x2013;<lpage>406</lpage>, <year>2018</year>.</mixed-citation></ref>
<ref id="ref-5"><label>[5]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>P.</given-names> <surname>Fettke</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Loos</surname></string-name> and <string-name><given-names>J.</given-names> <surname>Zwicker</surname></string-name></person-group>, &#x201C;<article-title>Business process reference models: Survey and classification</article-title>,&#x201D; in <conf-name>Proc. of Business Process Management Workshops</conf-name>, <publisher-loc>Berlin, Heidelberg</publisher-loc>, <publisher-name>Springer</publisher-name>, pp. <fpage>469</fpage>&#x2013;<lpage>483</lpage>, <year>2006</year>.</mixed-citation></ref>
<ref id="ref-6"><label>[6]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Becker</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Klingner</surname></string-name></person-group>, &#x201C;<chapter-title>A criteria catalogue for evaluating business process pattern approaches</chapter-title>,&#x201D; in <source>Enterprise, Business-Process and Information Systems Modeling</source>, <publisher-loc>Berlin, Heidelberg</publisher-loc>, <publisher-name>Springer</publisher-name>, pp. <fpage>257</fpage>&#x2013;<lpage>271</lpage>, <year>2014</year>.</mixed-citation></ref>
<ref id="ref-7"><label>[7]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>B.</given-names> <surname>Weber</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Rinderle</surname></string-name> and <string-name><given-names>M.</given-names> <surname>Reichert</surname></string-name></person-group>, &#x201C;<article-title>Change patterns and change support features in process-aware information systems</article-title>,&#x201D; <source>Advanced Information Systems Engineering</source>, vol. <volume>4495</volume>, no. <issue>3</issue>, pp. <fpage>574</fpage>&#x2013;<lpage>588</lpage>, <year>2007</year>.</mixed-citation></ref>
<ref id="ref-8"><label>[8]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Armas-Cervantes</surname></string-name>, <string-name><given-names>P.</given-names> <surname>Baldan</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Dumas</surname></string-name> and <string-name><given-names>L.</given-names> <surname>Garcia-Ba&#x00F1;uelos</surname></string-name></person-group>, &#x201C;<article-title>Diagnosing behavioral differences between business process models: An approach based on event structures</article-title>,&#x201D; <source>Information Systems</source>, vol. <volume>56</volume>, no. <issue>2</issue>, pp. <fpage>304</fpage>&#x2013;<lpage>325</lpage>, <year>2016</year>.</mixed-citation></ref>
<ref id="ref-9"><label>[9]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Becker</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Klingner</surname></string-name>, <string-name><given-names>B.</given-names> <surname>Weber</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Reichert</surname></string-name> and <string-name><given-names>S.</given-names> <surname>Rinderle-Ma</surname></string-name></person-group>, &#x201C;<article-title>Change patterns and change support features-enhancing flexibility in process-aware information systems</article-title>,&#x201D; in <conf-name>Proc. of Enterprise, Business-Process and Information Systems Modeling</conf-name>, <publisher-loc>Berlin, Heidelberg</publisher-loc>, <publisher-name>Springer</publisher-name>, vol. <volume>66</volume>, pp. <fpage>438</fpage>&#x2013;<lpage>466</lpage>, <year>2008</year>.</mixed-citation></ref>
<ref id="ref-10"><label>[10]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Becker</surname></string-name> and <string-name><given-names>R.</given-names> <surname>Laue</surname></string-name></person-group>, &#x201C;<article-title>A comparative survey of business process similarity measures</article-title>,&#x201D; <source>Computers in Industry</source>, vol. <volume>63</volume>, no. <issue>2</issue>, pp. <fpage>148</fpage>&#x2013;<lpage>167</lpage>, <year>2012</year>.</mixed-citation></ref>
<ref id="ref-11"><label>[11]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>M.</given-names> <surname>Dumas</surname></string-name>, <string-name><given-names>L.</given-names> <surname>Garc&#x00ED;a-Ba&#x00F1;uelos</surname></string-name> and <string-name><given-names>R. M.</given-names> <surname>Dijkman</surname></string-name></person-group>, &#x201C;<article-title>Similarity search of business process models</article-title>,&#x201D; <source>Bulletin of the IEEE Computer Society Technical Committee on Data Engineering</source>, vol. <volume>32</volume>, no. <issue>3</issue>, pp. <fpage>23</fpage>&#x2013;<lpage>28</lpage>, <year>2009</year>.</mixed-citation></ref>
<ref id="ref-12"><label>[12]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>K.</given-names> <surname>Shahzad</surname></string-name>, <string-name><given-names>R. M. A.</given-names> <surname>Nawab</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Abid</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Sharif</surname></string-name>, <string-name><given-names>F.</given-names> <surname>Ali</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>A process model collection and gold standard correspondences for process model matching</article-title>,&#x201D; <source>IEEE Access</source>, vol. <volume>7</volume>, pp. <fpage>30708</fpage>&#x2013;<lpage>30723</lpage>, <year>2019</year>.</mixed-citation></ref>
<ref id="ref-13"><label>[13]</label><mixed-citation publication-type="conf-proc"><person-group person-group-type="author"><string-name><given-names>G.</given-names> <surname>Antunes</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Bakhshandeh</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Borbinha</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Cardoso</surname></string-name>, <string-name><given-names>S.</given-names> <surname>Dadashnia</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>The process model matching contest 2015</article-title>,&#x201D; in <conf-name>Proc. of the 6th Int. Workshop on Enterprise Modelling and Information Systems Architectures, Gesellschaft f&#x00FC;r Informatik Lecture Notes Informatics</conf-name>, <publisher-loc>Innsbruck, Austria</publisher-loc>, vol. <volume>248</volume>, pp. <fpage>127</fpage>&#x2013;<lpage>155</lpage>, <year>2015</year>.</mixed-citation></ref>
<ref id="ref-14"><label>[14]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>I.</given-names> <surname>Vanderfeesten</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Cardoso</surname></string-name>, <string-name><given-names>J.</given-names> <surname>Mendling</surname></string-name>, <string-name><given-names>H. A.</given-names> <surname>Reijers</surname></string-name> and <string-name><given-names>W. M. P.</given-names> <surname>van der Aalst</surname></string-name></person-group>, &#x201C;<chapter-title>Quality metrics for business process models</chapter-title>,&#x201D; in <source>BPM Workflow Handbook, Lecture Notes in Business Information Processing Book Series</source>, vol. <volume>144</volume>. <publisher-loc>Berlin, Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>, pp. <fpage>179</fpage>&#x2013;<lpage>190</lpage>, <year>2007</year>.</mixed-citation></ref>
<ref id="ref-15"><label>[15]</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Sohail</surname></string-name>, <string-name><given-names>K.</given-names> <surname>Shahzad</surname></string-name>, <string-name><given-names>P. D.</given-names> <surname>Dominic</surname></string-name>, <string-name><given-names>M. A.</given-names> <surname>Butt</surname></string-name>, <string-name><given-names>M.</given-names> <surname>Arif</surname></string-name> <etal>et al.</etal></person-group><italic>,</italic> &#x201C;<article-title>On computing the suitability of non-human resources for business process analysis</article-title>,&#x201D; <source>Computers, Materials &#x0026; Continua</source>, vol. <volume>67</volume>, no. <issue>1</issue>, pp. <fpage>303</fpage>&#x2013;<lpage>319</lpage>, <year>2021</year>.</mixed-citation></ref>
<ref id="ref-16"><label>[16]</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><given-names>J.</given-names> <surname>Mendling</surname></string-name></person-group>, &#x201C;<chapter-title>Metrics for business process models</chapter-title>,&#x201D; in <source>Lecture Notes in Business Information Processing Book Series on Metrics for Process Models</source>. vol. <volume>6</volume>. <publisher-loc>Berlin, Heidelberg</publisher-loc>: <publisher-name>Springer</publisher-name>, pp. <fpage>103</fpage>&#x2013;<lpage>133</lpage>, <year>2008</year>.</mixed-citation></ref>
<ref id="ref-17"><label>[17]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>Sohail</surname></string-name> and <string-name><given-names>R.</given-names> <surname>Anum</surname></string-name></person-group>, &#x201C;<article-title>A structural variants process models collection for process similarities evaluations</article-title>,&#x201D; <comment>MS dissertation</comment>. <publisher-name>COMSATS University Islamabad</publisher-name>, <publisher-loc>Lahore Campus, Lahore, Pakistan</publisher-loc>, <year>2017</year>.</mixed-citation></ref>
<ref id="ref-18"><label>[18]</label><mixed-citation publication-type="other"><person-group person-group-type="author"><string-name><given-names>A.</given-names> <surname>sohail</surname></string-name>, <string-name><given-names>A.</given-names> <surname>Haseen</surname></string-name>, <string-name><given-names>M. H.</given-names> <surname>Arshad</surname></string-name> and <string-name><given-names>M. H.</given-names> <surname>Mansoor</surname></string-name></person-group>, &#x201C;<article-title>Style Guidelines for Final Year Project ReportsSimilarity of Business Process Models: A Structural Changes based Technique</article-title>,&#x201D; <comment>FYP dissertation</comment>. <publisher-name>COMSATS University Islamabad</publisher-name>, <publisher-loc>Lahore Campus, Lahore, Pakistan</publisher-loc>, <year>2018</year>.</mixed-citation></ref>
</ref-list>
</back>
</article>