<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Publishing DTD v1.1 20151215//EN" "http://jats.nlm.nih.gov/publishing/1.1/JATS-journalpublishing1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" article-type="research-article" dtd-version="1.1">
<front>
<journal-meta>
<journal-id journal-id-type="pmc">CMES</journal-id>
<journal-id journal-id-type="nlm-ta">CMES</journal-id>
<journal-id journal-id-type="publisher-id">CMES</journal-id>
<journal-title-group>
<journal-title>Computer Modeling in Engineering &#x0026; Sciences</journal-title>
</journal-title-group>
<issn pub-type="epub">1526-1506</issn>
<issn pub-type="ppub">1526-1492</issn>
<publisher>
<publisher-name>Tech Science Press</publisher-name>
<publisher-loc>USA</publisher-loc>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="publisher-id">19708</article-id>
<article-id pub-id-type="doi">10.32604/cmes.2022.019708</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Article</subject>
</subj-group>
</article-categories>
<title-group>
<article-title>gscaLCA in R: Fitting Fuzzy Clustering Analysis Incorporated with Generalized Structured Component Analysis</article-title>
<alt-title alt-title-type="left-running-head">gscaLCA in R: Fitting Fuzzy Clustering Analysis Incorporated with Generalized Structured Component Analysis</alt-title>
<alt-title alt-title-type="right-running-head">gscaLCA in R: Fitting Fuzzy Clustering Analysis Incorporated with Generalized Structured Component Analysis</alt-title>
</title-group>
<contrib-group content-type="authors">
<contrib id="author-1" contrib-type="author" corresp="yes">
<name name-style="western"><surname>Ryoo</surname><given-names>Ji Hoon</given-names>
</name><xref ref-type="aff" rid="aff-1">1</xref><email>ryoox001@yonsei.ac.kr</email>
</contrib> <contrib id="author-2" contrib-type="author">
<name name-style="western"><surname>Park</surname><given-names>Seohee</given-names>
</name><xref ref-type="aff" rid="aff-2">2</xref></contrib>
<contrib id="author-3" contrib-type="author">
<name name-style="western"><surname>Kim</surname><given-names>Seongeun</given-names>
</name><xref ref-type="aff" rid="aff-3">3</xref></contrib>
<contrib id="author-4" contrib-type="author">
<name name-style="western"><surname>Hwang</surname><given-names>Heungsun</given-names>
</name><xref ref-type="aff" rid="aff-4">4</xref></contrib>
<aff id="aff-1"><label>1</label><institution>Department of Education, College of Educational Sciences, Yonsei University</institution>, <addr-line>Seoul, 03722</addr-line>, <country>South Korea</country></aff>
<aff id="aff-2"><label>2</label><institution>Psychometrics Department, American Board of Internal Medicine</institution>, <addr-line>Philadelphia, 19016</addr-line>, <country>USA</country></aff>
<aff id="aff-3"><label>3</label><institution>Department of Educational Research Methodology, School of Education, University of North Carolina at Greensboro</institution>, <addr-line>Greensboro, 27412</addr-line>, <country>USA</country></aff>
<aff id="aff-4"><label>4</label><institution>Department of Psychology, McGill University</institution>, <addr-line>Montreal, H3A 0G4</addr-line>, <country>Canada</country></aff>
</contrib-group>
<author-notes>
<corresp id="cor1"><label>&#x002A;</label>Corresponding Author: Ji Hoon Ryoo. Email: <email>ryoox001@yonsei.ac.kr</email></corresp>
</author-notes>
<pub-date pub-type="epub" date-type="pub" iso-8601-date="2022-06-23">
<day>23</day>
<month>06</month>
<year>2022</year>
</pub-date>
<volume>132</volume>
<issue>3</issue>
<fpage>801</fpage>
<lpage>822</lpage>
<history>
<date date-type="received">
<day>10</day>
<month>10</month>
<year>2021</year>
</date>
<date date-type="accepted">
<day>25</day>
<month>1</month>
<year>2022</year>
</date>
</history>
<permissions>
<copyright-statement>&#x00A9; 2022 Ryoo et al.</copyright-statement>
<copyright-year>2022</copyright-year>
<copyright-holder>Ryoo et al.</copyright-holder>
<license xlink:href="https://creativecommons.org/licenses/by/4.0/">
<license-p>This work is licensed under a <ext-link ext-link-type="uri" xlink:type="simple" xlink:href="https://creativecommons.org/licenses/by/4.0/">Creative Commons Attribution 4.0 International License</ext-link>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.</license-p>
</license>
</permissions>
<self-uri content-type="pdf" xlink:href="TSP_CMES_19708.pdf"></self-uri>
<abstract>
<p>Clustering analysis identifying unknown heterogenous subgroups of a population (or a sample) has become increasingly popular along with the popularity of machine learning techniques. Although there are many software packages running clustering analysis, there is a lack of packages conducting clustering analysis within a structural equation modeling framework. The package, <bold>gscaLCA</bold> which is implemented in the <bold>R</bold> statistical computing environment, was developed for conducting clustering analysis and has been extended to a latent variable modeling. More specifically, by applying both fuzzy clustering (FC) algorithm and generalized structured component analysis (GSCA), the package <bold>gscaLCA</bold> computes membership prevalence and item response probabilities as posterior probabilities, which is applicable in mixture modeling such as latent class analysis in statistics. As a hybrid model between data clustering in classifications and model-based mixture modeling approach, fuzzy clusterwise GSCA, denoted as gscaLCA, encompasses many advantages from both methods: (1) soft partitioning from FC and (2) efficiency in estimating model parameters with bootstrap method via resolution of global optimization problem from GSCA. The main function, gscaLCA, works for both binary and ordered categorical variables. In addition, gscaLCA can be used for latent class regression as well. Visualization of profiles of latent classes based on the posterior probabilities is also available in the package <bold>gscaLCA</bold>. This paper contributes to providing a methodological tool, <bold>gscaLCA</bold> that applied researchers such as social scientists and medical researchers can apply clustering analysis in their research.</p>
</abstract>
<kwd-group kwd-group-type="author">
<kwd>Fuzzy clustering</kwd>
<kwd>generalized structured component analysis</kwd>
<kwd>gscaLCA</kwd>
<kwd>latent class analysis</kwd>
</kwd-group>
</article-meta>
</front>
<body>
<sec id="s1">
<label>1</label>
<title>Introduction</title>
<sec id="s1_1">
<label>1.1</label>
<title>Motivation</title>
<p>Latent class analysis (LCA) [<xref ref-type="bibr" rid="ref-1">1</xref>,<xref ref-type="bibr" rid="ref-2">2</xref>], as a mixture modeling, has been widely used to identify homogeneous subpopulations from observed categorical variables under the assumption that the population is heterogeneous. One of the reasons for its popularity is its ability to reveal the characteristics of each homogeneous population identified via statistical modeling. Consideration of the heterogeneity also informs the characteristics of subpopulations unveiled in research areas including social, behavioral, and health sciences. Conceptually, identification of unobserved group characteristics via LCA corresponds to unsupervised learning or cluster analysis in data mining such as <inline-formula id="ieqn-1"><mml:math id="mml-ieqn-1"><mml:mi>K</mml:mi></mml:math></inline-formula>-means or <inline-formula id="ieqn-2"><mml:math id="mml-ieqn-2"><mml:mi>K</mml:mi></mml:math></inline-formula>-median algorithm. However, the popularity of LCA as a statistical model is somewhat different from unsupervised learning due to its strict property implemented in the most common estimation method, maximum likelihood estimation based on the expectation-maximization (EM) algorithm [<xref ref-type="bibr" rid="ref-3">3</xref>]. More precisely, the EM algorithm requires multivariate normality for variance-covariance matrix used in the LCA to estimate parameters. As another way of saying, the multivariate normality property prevents researchers from utilizing the concept of big data in LCA, which often produces estimation issues such as non-positive definiteness in computing a Hessian matrix [<xref ref-type="bibr" rid="ref-4">4</xref>].</p>
<p>As described in the comparison between the mixture-modeling approach and cluster analysis procedure such as <inline-formula id="ieqn-3"><mml:math id="mml-ieqn-3"><mml:mi>K</mml:mi></mml:math></inline-formula>-means [<xref ref-type="bibr" rid="ref-5">5</xref>], the mixture-modeling approach does not always perform better than <inline-formula id="ieqn-4"><mml:math id="mml-ieqn-4"><mml:mi>K</mml:mi></mml:math></inline-formula>-means cluster analysis. Rather, Steinley et al. [<xref ref-type="bibr" rid="ref-5">5</xref>] showed an equivalence in terms of statistical modeling under a certain condition. That is, for two conceptually equivalent clustering methods, it cannot be said that the one is superior to the other, which is consistent with the results from Brusco et al. [<xref ref-type="bibr" rid="ref-6">6</xref>] comparing latent class, <inline-formula id="ieqn-5"><mml:math id="mml-ieqn-5"><mml:mi>K</mml:mi></mml:math></inline-formula>-means, and <inline-formula id="ieqn-6"><mml:math id="mml-ieqn-6"><mml:mi>K</mml:mi></mml:math></inline-formula>-median methods. Although Lubke et al. [<xref ref-type="bibr" rid="ref-7">7</xref>] noted that &#x201C;Model-based methods have the advantage that more rigorous methods can be applied for the comparison of alternative models&#x201D; (p. 23), it would not be applicable when we consider big data and/or data mining due to the model complexity and strict assumption of multivariate normality. On the other hand, cluster analysis often takes advantage in estimation due to the simple estimation algorithm using an alternative least square estimation in the distance function from centroids. Steinley et al. [<xref ref-type="bibr" rid="ref-5">5</xref>] also noted that &#x201C;it is important to realize that increased complexity and flexibility do not necessarily imply that a better solution will be found if the goal of the analysis is to uncover the unknown cluster membership.&#x201D; (p. 76).</p>
<p>Although Steinley et al. [<xref ref-type="bibr" rid="ref-5">5</xref>] highlighted the advantages of <inline-formula id="ieqn-7"><mml:math id="mml-ieqn-7"><mml:mi>K</mml:mi></mml:math></inline-formula>-means, <inline-formula id="ieqn-8"><mml:math id="mml-ieqn-8"><mml:mi>K</mml:mi></mml:math></inline-formula>-means cluster analysis is not a perfect alternative to a mixture-modeling approach. It often suffers from a poor local optimum and has a limitation that its algorithm works with only convex clustering. Such limitation of <inline-formula id="ieqn-9"><mml:math id="mml-ieqn-9"><mml:mi>K</mml:mi></mml:math></inline-formula>-means cluster analysis can be got rid of by applying fuzzy clustering analysis [<xref ref-type="bibr" rid="ref-8">8</xref>]. When fuzzy clustering methods was applied to identification of homogeneous subgroups, the estimation did show less analytic issues such as non-positive definitie matrix. Furthermore, by utilizing the method of how fuzzy clustering analysis can be accompanied by generalized structured component analysis (GSCA; [<xref ref-type="bibr" rid="ref-9">9</xref>]), Ryoo et al. [<xref ref-type="bibr" rid="ref-8">8</xref>] encompassed the LCA in the GSCA framework, which is called fuzzy clusterwise GSCA. In other words, the merger of fuzzy clustering and GSCA allows researchers not only to effectively classify the homogeneous cluster via fuzzy clustering but also to move forward in the model-based classification via GSCA. One of the advantages of LCA over cluster analysis procedures is the capacity to examine the effects of other variables on the LCA parameters, which is no longer the exclusive property of the fuzzy clusterwise GSCA. Such a role of examining the effects of other variables on the LCA was replaced with GSCA in the fuzzy clusterwise GSCA that is a statistical tool of fitting various component-based structural equation models into data [<xref ref-type="bibr" rid="ref-9">9</xref>,<xref ref-type="bibr" rid="ref-10">10</xref>]. In addition, Ryoo et al. [<xref ref-type="bibr" rid="ref-11">11</xref>] provided more indexes that can be utilized in identifying homogeneous subgroups and the procedure of enumerating the number of clusters, which is out of the scope of this manuscript. The method of fuzzy clusterwise GSCA will be described in <xref ref-type="sec" rid="s2">Section 2</xref>.</p>
</sec>
<sec id="s1_2">
<label>1.2</label>
<title>Existing Methods and Tools</title>
<p>Although there are many statistical packages including R for latent class analysis and cluster analysis, the method utilizing both the fuzzy clustering analysis and GSCA for LCA is a new approach and thus, there is no comparable R package for fuzzy clusterwise GSCA. Instead, in this section, we introduce three well-known packages for LCA as a mixture-modeling approach: Mplus, poLCA in R, and SAS procedure LCA as competitors. There are many other software packages available but exhaustive search of packages are out of our scope. We focus on their key features of the three packages, which validates the necessity and coverage of our new package, gscaLCA, in R. Note that most statistical programs for LCA include the capabilities of fitting multi-groups LCA, imposing measurement invariance across groups, and implementing latent class regression (LCR). In addition, binary and multinomial logistic regression options for predicting latent class membership and the ability to take into account sampling weights and clusters are also possible.</p>
<sec id="s1_2_1">
<label>1.2.1</label>
<title>Mplus</title>
<p>Mplus is the most common software package for fitting structural equation models and provides a variety of tools for modeling. For example, within a mixture modeling framework, both latent class analysis and latent profile analysis are available using the option of TYPE&#x003D;MIXTURE in Analysis part of Mplus. Latent class analysis is for categorical observed variables, whereas latent profile analysis is used for continuous observed variables. Mplus also provides various estimation methods utilizing the maximum likelihood (ML) method [<xref ref-type="bibr" rid="ref-12">12</xref>] such as MLM (ML parameter estimates with standard errors and a mean-adjusted chi-square test statistic) and MLMV (ML parameter estimates with standard errors and a mean- and variance-adjusted chi-square test statistic). However, Mplus does not provide the stablest and most robust solution of fitting model among statistical software packages in the case of the deviation of data from normality assumption or any other little violation of assumptions such as multicollinearity. Thus, it is often required for researchers to investigate other options to fit in their studies when they are faced with such violations. Nevertheless, Mplus is still versatile in the LCA. Here, listed are a couple of LCA examples using Mplus: van Horn et al. [<xref ref-type="bibr" rid="ref-13">13</xref>] including syntax, and O&#x2019;Neill et al. [<xref ref-type="bibr" rid="ref-14">14</xref>].</p>
</sec>
<sec id="s1_2_2">
<label>1.2.2</label>
<title>poLCA</title>
<p>Among several R packages fitting LCA, &#x201C;poLCA&#x201D; is one of the most common R packages. By using expectation-maximization and Newton-Raphson algorithm, poLCA finds maximum likelihood estimates of the LCA model parameters [<xref ref-type="bibr" rid="ref-15">15</xref>]. Latent class regression (LCR; LCA with covariates) in poLCA estimates how covariates affect latent class membership probabilities. For example, Schreiber [<xref ref-type="bibr" rid="ref-16">16</xref>] used poLCA with a syntax, and Miranda et al. [<xref ref-type="bibr" rid="ref-17">17</xref>] used poLCA to evaluated female young adults&#x2019; lifestyle from the behavioral variable measurement in public health, There are two more examples of using poLCA package, van Rijnsoever et al. [<xref ref-type="bibr" rid="ref-18">18</xref>] in information science, and Xia et al. [<xref ref-type="bibr" rid="ref-19">19</xref>] in tourism management. New package, gscaLCA, also deals with the LCR in addition to most of functionalities in the poLCA package.</p>
</sec>
<sec id="s1_2_3">
<label>1.2.3</label>
<title>Proc LCA</title>
<p>Proc LCA was developed for SAS for Windows. Lanza et al. [<xref ref-type="bibr" rid="ref-20">20</xref>] listed key features including multi-groups LCA, measurement invariance across groups, LCR, binary and multinomial logistic regression options. The regression options predicts latent class membership and holds the capability to take into account sampling weights and clusters. Collins et al. [<xref ref-type="bibr" rid="ref-21">21</xref>] described the whole process of fitting LCA including the key features, although most of those key features are also available in other packages, nowadays. Listed are a couple of examples using Proc LCA: Reynolds et al. [<xref ref-type="bibr" rid="ref-22">22</xref>], and Ryoo et al. [<xref ref-type="bibr" rid="ref-23">23</xref>], they explored gifted children&#x2019;s victimization and bullying by using Proc LCA. As of Jan, 2022, Proc LCA is still requiring additional installation within SAS.</p>
</sec>
<sec id="s1_2_4">
<label>1.2.4</label>
<title>Our Contribution</title>
<p>As mentioned, there is no dominating method to identify heterogeneity of a population between mixture-modeling approach and cluster analysis because each approach has advantages or disadvantages aforementioned and estimation procedures are different. Rather, the choice of statistical model would be related to researcher&#x0027;s discretion [<xref ref-type="bibr" rid="ref-24">24</xref>,<xref ref-type="bibr" rid="ref-25">25</xref>]. Our goal and contribution is to provide an analytic tool for researchers who want to run LCA using a heuristic cluster analysis procedure with fuzzy clustering algorithm and GSCA as a hybrid method. In addition, the utilization of GSCA in gscaLCA allows researchers to analyze data within the full range of the structural equation modeling perspective. We dscribe the package gscaLCA in the four following sections: Framework of fuzzy clusterwise GSCA (FC-GSCA), FC-GSCA with covariates, description of main functions, gscaLCA and gscaLCR, and demonstration of fitting fuzzy clusterwise GSCA with and without covariates by using two empirical examples.</p>
</sec>
</sec>
</sec>
<sec id="s2">
<label>2</label>
<title>Framework of Fuzzy Clusterwise GSCA</title>
<p>Fuzziness is well understood as a soft clustering where each object belongs to every cluster with a certain degree of membership probability, whereas <inline-formula id="ieqn-10"><mml:math id="mml-ieqn-10"><mml:mi>K</mml:mi></mml:math></inline-formula>-means is a hard clustering that every object belongs only one cluster. As a centroid-based clustering approaches, fuzzy clustering overcomes one disadvantage of <inline-formula id="ieqn-11"><mml:math id="mml-ieqn-11"><mml:mi>K</mml:mi></mml:math></inline-formula>-means algorithm in that it does not work well for non-convex data. In addition to the clustering point of view, the fuzzy clustering together with GSCA in estimation process provides tools such as statistical modeling and model evaluation, which is more informative compared with <inline-formula id="ieqn-12"><mml:math id="mml-ieqn-12"><mml:mi>K</mml:mi></mml:math></inline-formula>-means. Model evaluation will be discussed later in this section, which provides cluster validity measures in fuzzy clustering.</p>
<sec id="s2_1">
<label>2.1</label>
<title>Fuzzy c Means</title>
<p>Based on the description of fuzzy clustering [<xref ref-type="bibr" rid="ref-26">26</xref>] and terminologies [<xref ref-type="bibr" rid="ref-27">27</xref>], we briefly describe fuzzy <inline-formula id="ieqn-13"><mml:math id="mml-ieqn-13"><mml:mi>c</mml:mi></mml:math></inline-formula> means (FCM) algorithm as follows: FCM minimizes the following objective (distance) function, <inline-formula id="ieqn-14"><mml:math id="mml-ieqn-14"><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, measured by the sum of squares:</p>
<p><disp-formula id="eqn-1"><label>(1)</label><mml:math id="mml-eqn-1" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mn>39</mml:mn><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msubsup><mml:msup><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-15"><mml:math id="mml-ieqn-15"><mml:mi>m</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mo stretchy="false">(</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">&#x221E;</mml:mi><mml:mo stretchy="false">]</mml:mo></mml:math></inline-formula> is a classification index, <inline-formula id="ieqn-16"><mml:math id="mml-ieqn-16"><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> indicates the membership probability of <inline-formula id="ieqn-17"><mml:math id="mml-ieqn-17"><mml:mi>i</mml:mi></mml:math></inline-formula><sup>th</sup> object in class <inline-formula id="ieqn-18"><mml:math id="mml-ieqn-18"><mml:mi>k</mml:mi></mml:math></inline-formula>, and <inline-formula id="ieqn-19"><mml:math id="mml-ieqn-19"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> indicates the centroid for class <inline-formula id="ieqn-20"><mml:math id="mml-ieqn-20"><mml:mi>k</mml:mi><mml:mo>&#x2208;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>,</mml:mo><mml:mi>K</mml:mi><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>. The <inline-formula id="ieqn-21"><mml:math id="mml-ieqn-21"><mml:mi>m</mml:mi></mml:math></inline-formula> is also known as a fuzzifier, and adjusts the probability of belonging to a class, i.e., <inline-formula id="ieqn-22"><mml:math id="mml-ieqn-22"><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> indicates that the membership probability will converge either 0 or 1 such as in <inline-formula id="ieqn-23"><mml:math id="mml-ieqn-23"><mml:mi>K</mml:mi></mml:math></inline-formula>-means, which is excluded in this FCM algorithm. On the other hand, <inline-formula id="ieqn-24"><mml:math id="mml-ieqn-24"><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="normal">&#x221E;</mml:mi></mml:math></inline-formula> indicates that all membership probabilities equal probability in belonging to any of classes, i.e., <inline-formula id="ieqn-25"><mml:math id="mml-ieqn-25"><mml:mfrac><mml:mn>1</mml:mn><mml:mi>K</mml:mi></mml:mfrac></mml:math></inline-formula>. Both <inline-formula id="ieqn-26"><mml:math id="mml-ieqn-26"><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-27"><mml:math id="mml-ieqn-27"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> are defined by</p>
<p><disp-formula id="ueqn-20"><mml:math id="mml-ueqn-20" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover><mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:msup><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup><mml:msup><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>l</mml:mi></mml:mrow></mml:msub><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:mfrac></mml:mstyle><mml:mo>]</mml:mo></mml:mrow><mml:mrow><mml:mfrac><mml:mn>1</mml:mn><mml:mrow><mml:mi>m</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mstyle></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>and</p>
<p><disp-formula id="ueqn-21"><mml:math id="mml-ueqn-21" display="block"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mfrac><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac><mml:mo>.</mml:mo></mml:math></disp-formula></p>
<p>With a termination criterion such that <inline-formula id="ieqn-28"><mml:math id="mml-ieqn-28"><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mo stretchy="false">(</mml:mo><mml:mi>L</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo stretchy="false">)</mml:mo></mml:mrow></mml:msup><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msubsup><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>L</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>|</mml:mo></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mo>&#x003C;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03F5;</mml:mi></mml:mrow></mml:math></inline-formula> for a small value <inline-formula id="ieqn-29"><mml:math id="mml-ieqn-29"><mml:mrow><mml:mi mathvariant="normal">&#x03F5;</mml:mi></mml:mrow></mml:math></inline-formula> at <inline-formula id="ieqn-30"><mml:math id="mml-ieqn-30"><mml:mi>L</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula> respects, the FCM algorithm is done as follows:
<list list-type="order">
<list-item><label>1.</label><p>(Step 1) Initialize <inline-formula id="ieqn-31"><mml:math id="mml-ieqn-31"><mml:msup><mml:mi>U</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mn>0</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup></mml:math></inline-formula> for <inline-formula id="ieqn-32"><mml:math id="mml-ieqn-32"><mml:mi>U</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msubsup><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>,</p></list-item>
<list-item><label>2.</label><p>(Step 2) Compute <inline-formula id="ieqn-33"><mml:math id="mml-ieqn-33"><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> and <inline-formula id="ieqn-34"><mml:math id="mml-ieqn-34"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> by minimizing <inline-formula id="ieqn-35"><mml:math id="mml-ieqn-35"><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, which also update <inline-formula id="ieqn-36"><mml:math id="mml-ieqn-36"><mml:mi>U</mml:mi></mml:math></inline-formula>, and</p></list-item>
<list-item><label>3.</label><p>(Step 3) If <inline-formula id="ieqn-37"><mml:math id="mml-ieqn-37"><mml:mrow><mml:mo symmetric="true">&#x2016;</mml:mo><mml:msup><mml:mi>U</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>L</mml:mi><mml:mo>+</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>&#x2212;</mml:mo><mml:msup><mml:mi>U</mml:mi><mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>L</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo symmetric="true">&#x2016;</mml:mo></mml:mrow><mml:mo>&#x2264;</mml:mo><mml:mrow><mml:mi mathvariant="normal">&#x03F5;</mml:mi></mml:mrow></mml:math></inline-formula> then &#x201C;STOP&#x201D;. Otherwise, the algorithm repeats from Step 2.</p></list-item>
</list></p>
</sec>
<sec id="s2_2">
<label>2.2</label>
<title>Generalized Structured Component Analysis (GSCA)</title>
<p>Generalized structured component analysis (GSCA) [<xref ref-type="bibr" rid="ref-9">9</xref>] is a component-based approach of structural equation modeling (SEM). Different from a maximum likelihood-based SEM (ML-SEM), the component-based SEM (CB-SEM) utilizes the underlying construct as a composite of weighted observed variables, and applies an alternating least square method to estimate model parameters. Here, we briefly describe GSCA as a CB-SEM (see Hwang et al. [<xref ref-type="bibr" rid="ref-9">9</xref>] for more detail). The alternating least square method in the component-based approach is relatively simple and straightforward procedure compared to the ML-based approach because it does not assume the multivariate normality of model parameters but minimizes a sum of squares of residuals computed from sample data directly. Along with regularization such as Ridge and Lasso, the least square method produces a more interpretable and predictive model that has possibly lower prediction error [<xref ref-type="bibr" rid="ref-28">28</xref>]. Such a great property of the least square methods is inherited into GSCA [<xref ref-type="bibr" rid="ref-9">9</xref>].</p>
<p>GSCA consists of three sub models: a measurement model describing observed indicators from each latent construct, a structural model defining the associations among latent constructs, and an weighted relation model defining latent constructs. While the typical SEM models include a measurement model and a structural model assuming the normality of the latent constructs [<xref ref-type="bibr" rid="ref-29">29</xref>], GSCA additionally includes the weighted relation model that represents a formative relation between a component and its indicators. That is, the weighted relation model defines each underlying construct as a weighted composite or component of indicators. In this paper, we used both component and latent variable, interchangeably. Such a function in the weighted relation model plays a key role in parameter estimation without a multivariate normality assumption, which eases the estimation. The three sub models of GSCA can be expressed as follows:</p>
<p><disp-formula id="eqn-2"><label>(2)</label><mml:math id="mml-eqn-2" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">M</mml:mi><mml:mi mathvariant="italic">e</mml:mi><mml:mi mathvariant="italic">a</mml:mi><mml:mi mathvariant="italic">s</mml:mi><mml:mi mathvariant="italic">u</mml:mi><mml:mi mathvariant="italic">r</mml:mi><mml:mi mathvariant="italic">e</mml:mi><mml:mi mathvariant="italic">m</mml:mi><mml:mi mathvariant="italic">e</mml:mi><mml:mi mathvariant="italic">n</mml:mi><mml:mi mathvariant="italic">t</mml:mi></mml:mrow><mml:mtext>&#xA0;</mml:mtext><mml:mrow><mml:mi mathvariant="italic">M</mml:mi><mml:mi mathvariant="italic">o</mml:mi><mml:mi mathvariant="italic">d</mml:mi><mml:mi mathvariant="italic">e</mml:mi><mml:mi mathvariant="italic">l</mml:mi></mml:mrow><mml:mo>&#x003A;</mml:mo><mml:mspace width="1em" /><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">A</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mi>&#x03B3;</mml:mi><mml:mo>+</mml:mo><mml:mi>&#x03F5;</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p><disp-formula id="eqn-3"><label>(3)</label><mml:math id="mml-eqn-3" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">S</mml:mi><mml:mi mathvariant="italic">t</mml:mi><mml:mi mathvariant="italic">r</mml:mi><mml:mi mathvariant="italic">u</mml:mi><mml:mi mathvariant="italic">c</mml:mi><mml:mi mathvariant="italic">t</mml:mi><mml:mi mathvariant="italic">u</mml:mi><mml:mi mathvariant="italic">r</mml:mi><mml:mi mathvariant="italic">a</mml:mi><mml:mi mathvariant="italic">l</mml:mi></mml:mrow><mml:mtext>&#xA0;</mml:mtext><mml:mrow><mml:mi mathvariant="italic">M</mml:mi><mml:mi mathvariant="italic">o</mml:mi><mml:mi mathvariant="italic">d</mml:mi><mml:mi mathvariant="italic">e</mml:mi><mml:mi mathvariant="italic">l</mml:mi></mml:mrow><mml:mo>&#x003A;</mml:mo><mml:mspace width="1em" /><mml:mi>&#x03B3;</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">B</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mi>&#x03B3;</mml:mi><mml:mo>+</mml:mo><mml:mi>&#x03B6;</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p><disp-formula id="eqn-4"><label>(4)</label><mml:math id="mml-eqn-4" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="italic">W</mml:mi><mml:mi mathvariant="italic">e</mml:mi><mml:mi mathvariant="italic">i</mml:mi><mml:mi mathvariant="italic">g</mml:mi><mml:mi mathvariant="italic">h</mml:mi><mml:mi mathvariant="italic">t</mml:mi><mml:mi mathvariant="italic">e</mml:mi><mml:mi mathvariant="italic">d</mml:mi></mml:mrow><mml:mtext>&#xA0;</mml:mtext><mml:mrow><mml:mi mathvariant="italic">R</mml:mi><mml:mi mathvariant="italic">e</mml:mi><mml:mi mathvariant="italic">l</mml:mi><mml:mi mathvariant="italic">a</mml:mi><mml:mi mathvariant="italic">t</mml:mi><mml:mi mathvariant="italic">i</mml:mi><mml:mi mathvariant="italic">o</mml:mi><mml:mi mathvariant="italic">n</mml:mi></mml:mrow><mml:mtext>&#xA0;</mml:mtext><mml:mrow><mml:mi mathvariant="italic">M</mml:mi><mml:mi mathvariant="italic">o</mml:mi><mml:mi mathvariant="italic">d</mml:mi><mml:mi mathvariant="italic">e</mml:mi><mml:mi mathvariant="italic">l</mml:mi></mml:mrow><mml:mo>&#x003A;</mml:mo><mml:mspace width="1em" /><mml:mi>&#x03B3;</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">W</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mi mathvariant="bold-italic">z</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-38"><mml:math id="mml-ieqn-38"><mml:mi mathvariant="bold-italic">z</mml:mi></mml:math></inline-formula> is a <inline-formula id="ieqn-39"><mml:math id="mml-ieqn-39"><mml:mi>J</mml:mi></mml:math></inline-formula> by 1 vector of observed indicators scores from one observation, <inline-formula id="ieqn-40"><mml:math id="mml-ieqn-40"><mml:mi>&#x03B3;</mml:mi></mml:math></inline-formula> is <inline-formula id="ieqn-41"><mml:math id="mml-ieqn-41"><mml:mi>P</mml:mi></mml:math></inline-formula> by 1 vector of components, <inline-formula id="ieqn-42"><mml:math id="mml-ieqn-42"><mml:mi mathvariant="bold-italic">A</mml:mi></mml:math></inline-formula> is a <inline-formula id="ieqn-43"><mml:math id="mml-ieqn-43"><mml:mi>P</mml:mi></mml:math></inline-formula> by <inline-formula id="ieqn-44"><mml:math id="mml-ieqn-44"><mml:mi>J</mml:mi></mml:math></inline-formula> matrix of factor loadings, <inline-formula id="ieqn-45"><mml:math id="mml-ieqn-45"><mml:mi mathvariant="bold-italic">B</mml:mi></mml:math></inline-formula> is a <inline-formula id="ieqn-46"><mml:math id="mml-ieqn-46"><mml:mi>P</mml:mi></mml:math></inline-formula> by <inline-formula id="ieqn-47"><mml:math id="mml-ieqn-47"><mml:mi>P</mml:mi></mml:math></inline-formula> matrix of structural path coefficients, <inline-formula id="ieqn-48"><mml:math id="mml-ieqn-48"><mml:mi mathvariant="bold-italic">W</mml:mi></mml:math></inline-formula> is a <inline-formula id="ieqn-49"><mml:math id="mml-ieqn-49"><mml:mi>J</mml:mi></mml:math></inline-formula> by <inline-formula id="ieqn-50"><mml:math id="mml-ieqn-50"><mml:mi>P</mml:mi></mml:math></inline-formula> matrix of weights for components. In addition, <inline-formula id="ieqn-51"><mml:math id="mml-ieqn-51"><mml:mi>&#x03F5;</mml:mi></mml:math></inline-formula> presents the residuals of indicators, which is expressed by a <inline-formula id="ieqn-52"><mml:math id="mml-ieqn-52"><mml:mi>J</mml:mi></mml:math></inline-formula> by 1 vector, and <inline-formula id="ieqn-53"><mml:math id="mml-ieqn-53"><mml:mi>&#x03B6;</mml:mi></mml:math></inline-formula> presents the residuals of components, which is expressed by a <inline-formula id="ieqn-54"><mml:math id="mml-ieqn-54"><mml:mi>P</mml:mi></mml:math></inline-formula> by 1 vector. The three equations can be merged into one equation as follows:</p>
<p><disp-formula id="eqn-5"><label>(5)</label><mml:math id="mml-eqn-5" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msup><mml:mi mathvariant="bold-italic">V</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mi>z</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mi mathvariant="bold-italic">C</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:msup><mml:mi mathvariant="bold-italic">W</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mi>z</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="bold-italic">e</mml:mi><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-55"><mml:math id="mml-ieqn-55"><mml:mi mathvariant="bold-italic">V</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mi mathvariant="bold-italic">I</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msup><mml:mi mathvariant="bold-italic">W</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable><mml:mo>]</mml:mo></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mi mathvariant="bold-italic">C</mml:mi><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:msup><mml:mi mathvariant="bold-italic">A</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msup><mml:mi mathvariant="bold-italic">B</mml:mi><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable><mml:mo>]</mml:mo></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo></mml:math></inline-formula> and <inline-formula id="ieqn-56"><mml:math id="mml-ieqn-56"><mml:mi mathvariant="bold-italic">e</mml:mi><mml:mo>=</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mi>&#x03F5;</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>&#x03B6;</mml:mi></mml:mtd></mml:mtr></mml:mtable><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
<p>While minimizing the residual term <inline-formula id="ieqn-57"><mml:math id="mml-ieqn-57"><mml:mi mathvariant="bold-italic">e</mml:mi></mml:math></inline-formula> in <xref ref-type="disp-formula" rid="eqn-5">Eq. (5)</xref>, the factor loadings, path coefficients, and weights are estimated via the alternating least square method. The details of the estimation procedure along with a matlab code can be found in Hwang et al. [<xref ref-type="bibr" rid="ref-9">9</xref>].</p>
</sec>
<sec id="s2_3">
<label>2.3</label>
<title>Fuzzy Clusterwise GSCA</title>
<p>As a similar fashion that Hwang et al. [<xref ref-type="bibr" rid="ref-30">30</xref>] applied fuzzy clustering into latent curve model [<xref ref-type="bibr" rid="ref-31">31</xref>], Ryoo et al. [<xref ref-type="bibr" rid="ref-8">8</xref>] applied fuzzy clustering to latent class model. By applying fuzzy clustering to GSCA in the both models, latent curve model and latent class model, the cluster-level heterogeneity can be taken into account. The distinguished feature between Hwang et al. [<xref ref-type="bibr" rid="ref-30">30</xref>] and Ryoo et al. [<xref ref-type="bibr" rid="ref-8">8</xref>] is that the former focused on continuous indicators whereas the latter focused on discrete/categorical indicators. The fuzzy clusterwise GSCA for LCA follows the four steps:
<list list-type="order">
<list-item><label>1.</label><p>(Step 1) Identify clusters and estimate membership probabilities, <inline-formula id="ieqn-58"><mml:math id="mml-ieqn-58"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, as the initial procedure, where the initial fuzzy clustering works based on the response data.</p></list-item>
<list-item><label>2.</label><p>(Step 2) Estimate the parameters of GSCA model by applying the optimal scaling, <inline-formula id="ieqn-59"><mml:math id="mml-ieqn-59"><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, and the residual sums of squares, <inline-formula id="ieqn-60"><mml:math id="mml-ieqn-60"><mml:mi>&#x03D5;</mml:mi></mml:math></inline-formula>, in <xref ref-type="disp-formula" rid="eqn-6">Eq. (6)</xref>.</p></list-item>
<list-item><label>3.</label><p>(Step 3) Update class membership probabilities <inline-formula id="ieqn-61"><mml:math id="mml-ieqn-61"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> by utilizing the estimates of GSCA and applying the objective function, <inline-formula id="ieqn-62"><mml:math id="mml-ieqn-62"><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, in <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref>.</p></list-item>
<list-item><label>4.</label><p>(Step 4) Step 2 and Step 3 are iteratively carried out until both <inline-formula id="ieqn-63"><mml:math id="mml-ieqn-63"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> and all GSCA parameters are no longer improved.</p></list-item>
</list></p>
<p>In addition to the fuzzy clusterwise GSCA by Hwang et al. [<xref ref-type="bibr" rid="ref-9">9</xref>], the fuzzy clusterwise GSCA for LCA employs the optimal scaling to preserve measurement characteristics of categorical indicators proposed by Young [<xref ref-type="bibr" rid="ref-32">32</xref>] in Step 2. To estimate the parameters of GSCA in Step 2, we minimize the residual sums of their squares weighted with fixed <inline-formula id="ieqn-64"><mml:math id="mml-ieqn-64"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> at Step 1 (or Step 3) as below:</p>
<p><disp-formula id="eqn-6"><label>(6)</label><mml:math id="mml-eqn-6" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mi>&#x03D5;</mml:mi><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo movablelimits="false">&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msubsup><mml:mi>S</mml:mi><mml:mi>S</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:msubsup><mml:mi>V</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>&#x2212;</mml:mo><mml:msubsup><mml:mi>C</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:msubsup><mml:mi>W</mml:mi><mml:mrow><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msubsup><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>)</mml:mo></mml:mrow><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>subject to the probabilistic condition, <inline-formula id="ieqn-65"><mml:math id="mml-ieqn-65"><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:math></inline-formula>. <inline-formula id="ieqn-66"><mml:math id="mml-ieqn-66"><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is the optimally scaled vector such as <inline-formula id="ieqn-67"><mml:math id="mml-ieqn-67"><mml:msub><mml:mi>s</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mo stretchy="false">(</mml:mo><mml:msub><mml:mi mathvariant="bold-italic">z</mml:mi><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo stretchy="false">)</mml:mo></mml:math></inline-formula> where <inline-formula id="ieqn-68"><mml:math id="mml-ieqn-68"><mml:mi>f</mml:mi></mml:math></inline-formula> is a optimal scaling function. In Step 3, we update <inline-formula id="ieqn-69"><mml:math id="mml-ieqn-69"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> by applying the objective function, <inline-formula id="ieqn-70"><mml:math id="mml-ieqn-70"><mml:msub><mml:mi>J</mml:mi><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, in <xref ref-type="disp-formula" rid="eqn-1">Eq. (1)</xref> for the fuzzy clustering. The fuzzifier, <inline-formula id="ieqn-71"><mml:math id="mml-ieqn-71"><mml:mi>m</mml:mi></mml:math></inline-formula>, of <inline-formula id="ieqn-72"><mml:math id="mml-ieqn-72"><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mi>m</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula> is often set up at 2 in practice [<xref ref-type="bibr" rid="ref-26">26</xref>], which is also used as a default in the R package, <bold>gscaLCA</bold>. The item response probabilities characterizing classes are estimated based on their membership within each class at the end of procedure. The standard error of estimation can also be calculated by the bootstrap method [<xref ref-type="bibr" rid="ref-33">33</xref>].</p>
</sec>
<sec id="s2_4">
<label>2.4</label>
<title>Model Evaluation in Fuzzy Clusterwise GSCA</title>
<p>In addition to <inline-formula id="ieqn-73"><mml:math id="mml-ieqn-73"><mml:msup><mml:mi>R</mml:mi><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> type of model evaluation tools, FIT and AFIT, from GSCA [<xref ref-type="bibr" rid="ref-9">9</xref>], the package gscaLCA computes the fuzziness performance index (FPI) and the normalized classification entropy (NCE) recommended by Roubens [<xref ref-type="bibr" rid="ref-34">34</xref>]. They are defined as follows:</p>
<p><disp-formula id="eqn-7"><label>(7)</label><mml:math id="mml-eqn-7" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mi>F</mml:mi><mml:mi>P</mml:mi><mml:mi>I</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>&#x2212;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mi>K</mml:mi><mml:mo>&#x22C5;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover><mml:msubsup><mml:mi>u</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:msubsup><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>K</mml:mi><mml:mo>&#x2212;</mml:mo><mml:mn>1</mml:mn></mml:mrow></mml:mfrac></mml:mstyle></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>and</p>
<p><disp-formula id="eqn-8"><label>(8)</label><mml:math id="mml-eqn-8" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mi>N</mml:mi><mml:mi>C</mml:mi><mml:mi>E</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mo>&#x2212;</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mn>1</mml:mn><mml:mi>N</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>N</mml:mi></mml:mrow></mml:munderover><mml:munderover><mml:mo>&#x2211;</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>K</mml:mi></mml:mrow></mml:munderover><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mi>K</mml:mi></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>.</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula></p>
<p>Both FPI and NCE applies the criterion of &#x201C;smaller is better&#x201D; between 0 and 1, which helps researchers decide the number of clusters in the process of fitting LCA [<xref ref-type="bibr" rid="ref-9">9</xref>].</p>
</sec>
</sec>
<sec id="s3">
<label>3</label>
<title>Fuzzy Clusterwise GSCA with Covariates</title>
<p>By adding covariates into latent class modeling, we are able to investigate how the covariates predict the latent class membership of individuals [<xref ref-type="bibr" rid="ref-21">21</xref>,<xref ref-type="bibr" rid="ref-35">35</xref>]. Two possible approaches in modeling LCA with covariates can be applied, either a one-step approach or a three-step approach. The one-step approach estimates the effect of covariates on the membership while estimating the class membership probabilities and item response probabilities [<xref ref-type="bibr" rid="ref-36">36</xref>,<xref ref-type="bibr" rid="ref-37">37</xref>]. Specifically, a multinomial regression of membership probabilities on covariates is fitted within the LCA modeling. That is, the combined model estimates all parameters simultaneously. The other approach estimates the effects of covariates by fitting multinomial regressions with partitioning based on the estimated class membership probability. This three-step approach fits LCA at the first step, assigns each subject based on the estimated membership at the second step, and then fits a logistic regression of the assigned membership on covariates at the third step, sequentially [<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-38">38</xref>] (see [<xref ref-type="bibr" rid="ref-39">39</xref>] for detailed explanation of three-step approach). Compared to the one-step approach, the three-step approach rarely encounters identification issues or convergence problems because of the separate and individualized steps. Considering these advantages, the package, <bold>gscaLCA</bold>, applies the three-step approach in the fuzzing clustering GSCA with covariates, denoted as gscaLCR in our package, by examining the covariate effects.</p>
<p>More specifically, the first step of the three-step approach is executing fuzzy clusterwise GSCA, which is explained in the previous section. For the second step, two types of partitioning are available. The partitioning methods are associated with class assignments: either a hard partitioning for mutually exclusive assignment or a soft partitioning based on membership probabilities [<xref ref-type="bibr" rid="ref-35">35</xref>,<xref ref-type="bibr" rid="ref-40">40</xref>]. The hard partitioning assigns each individual&#x0027;s class based on the highest membership probability of the individual. For example, if the estimated <inline-formula id="ieqn-74"><mml:math id="mml-ieqn-74"><mml:msub><mml:mi>u</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> for the individual <inline-formula id="ieqn-75"><mml:math id="mml-ieqn-75"><mml:mi>i</mml:mi></mml:math></inline-formula> are 0.3, 0.45, 0.25 for class 1, class 2, and class 3, respectively, then the individual is assigned as class 2. On the other hand, the soft partitioning focuses on the membership probabilities themselves for all classes. The estimated membership probabilities are used to assign individuals into each class proportionally. Thus, in the example above, the individual will be assigned class 2 with the probability of 0.45.</p>
<p>With the assignment in the second step, the third step fits either a multinomial or binomial logistic regression. In the case of the hard participating, the procedure is straightforward. A regression model is fitted with the assigned class of each individual as the dependent variable and with covariates as independent variables. The model can be expressed as</p>
<p><disp-formula id="eqn-9"><label>(9)</label><mml:math id="mml-eqn-9" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mi>log</mml:mi><mml:mo>&#x2061;</mml:mo><mml:mrow><mml:mo>(</mml:mo><mml:mstyle displaystyle="true" scriptlevel="0"><mml:mfrac><mml:mrow><mml:mover><mml:mi>u</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mi>k</mml:mi></mml:mrow><mml:mrow><mml:mover><mml:mi>u</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:msup><mml:mi>k</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>S</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>S</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
where <inline-formula id="ieqn-76"><mml:math id="mml-ieqn-76"><mml:mover><mml:mi>u</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover></mml:math></inline-formula> represents the assigned class with the hard partitioning, <inline-formula id="ieqn-77"><mml:math id="mml-ieqn-77"><mml:mi>k</mml:mi></mml:math></inline-formula> is a focal class and <inline-formula id="ieqn-78"><mml:math id="mml-ieqn-78"><mml:msup><mml:mi>k</mml:mi><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msup></mml:math></inline-formula> is a reference class. In addition <inline-formula id="ieqn-79"><mml:math id="mml-ieqn-79"><mml:mi>S</mml:mi></mml:math></inline-formula> represents the number of covariates.</p>
<p>For the binomial regression, the assignment is re-coded into dummy variables. By using each dummy variable as dependent variable, the binomial regression can be fitted for each focal class separately. The dummy variables can be denoted as <inline-formula id="ieqn-80"><mml:math id="mml-ieqn-80"><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mover><mml:mi>u</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> for the focal class <inline-formula id="ieqn-81"><mml:math id="mml-ieqn-81"><mml:mi>k</mml:mi></mml:math></inline-formula>, and the binomial regression can be expressed as follows:</p>
<p><disp-formula id="eqn-10"><label>(10)</label><mml:math id="mml-eqn-10" display="block"><mml:mtable rowspacing="4pt" columnspacing="1em"><mml:mtr><mml:mtd><mml:mrow><mml:mtext>logit</mml:mtext></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mrow><mml:mover><mml:mi>u</mml:mi><mml:mo accent="false">&#x00AF;</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>0</mml:mn><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mo>&#x22EF;</mml:mo><mml:mo>+</mml:mo><mml:msub><mml:mi>&#x03B2;</mml:mi><mml:mrow><mml:mi>S</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mi>c</mml:mi><mml:mi>o</mml:mi><mml:msub><mml:mi>v</mml:mi><mml:mrow><mml:mi>S</mml:mi><mml:mo>,</mml:mo><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>where <inline-formula id="ieqn-82"><mml:math id="mml-ieqn-82"><mml:mi>k</mml:mi></mml:math></inline-formula> can be <inline-formula id="ieqn-83"><mml:math id="mml-ieqn-83"><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mn>2</mml:mn><mml:mo>,</mml:mo><mml:mo>&#x2026;</mml:mo><mml:mo>,</mml:mo><mml:mi>K</mml:mi><mml:mo>,</mml:mo></mml:math></inline-formula> which is the number of classes. The selection of either multinomial or binomial regression is determined by researcher&#x0027; preference or intention. In the case of soft partitioning, researchers use the same logistic regression as in the hard participating because each participant holds one and only one membership based on the highest probability. However, the regression takes the degrees of their membership into account with weights of the estimated membership probabilities, i.e., soft partitioning. That is, through the weights, the contribution of units on each class within the regression is adjusted.</p>
</sec>
<sec id="s4">
<label>4</label>
<title>Package gscaLCA for LCA</title>
<p>The package, <bold>gscaLCA</bold>, enables to conduct a LCA based on fuzzy clusterwise GSCA by estimating the parameters of latent class prevalence and item response probability in LCA with a single command line. The fuzzy clusterwise GSCA model can be fitted with or without covariates. The two main functions of the <bold>gscaLCA</bold> packages, gscaLCA and gscaLCR, are described below, along with the key features of the results and visualizations it produces.</p>
<sec id="s4_1">
<label>4.1</label>
<title>Data Input and Sample Datasets</title>
<p>Data are the main input to the function gscaLCA, and they should be formatted as a data frame containing indicator variables and covariates. The function gscaLCA requires the indicator variables to be discrete or categorical. It, however, does not requires whether the categorical variables are integer or character. When any indicator variable is continuous, the function is still run by recognizing the type of variable as a categorical variable. Thus, a caution is necessitated. There is an option that a continuous indicator is assigned as a continuous. On the other hand, for the covariates, both discrete and continuous variable are available. When a covariate is categorical numeric variable, it is required to define this variable as a factor. Missing data should be coded as NA in <bold>gscaLCA</bold>. The missing values will be deleted for the analysis in the gscaLCA algorithm by applying a listwise deletion in the current version.</p>
<p>The package <bold>gscaLCA</bold> provides two sample datasets that are informative for exploring different situations: (1) categorical variables (binary and more than two categories) for indicators and (2) continuous or categorical variables for covariates.</p>
<sec id="s4_1_1">
<label>4.1.1</label>
<title>TALIS Data</title>
<p>These data provide 5 items from the 2,560 survey responses data of U.S. teachers from the Teaching and Learning International Survey (TALIS) 2018 [<xref ref-type="bibr" rid="ref-39">39</xref>]. Two items are from teacher&#x0027;s motivation, two items were from teaching pedagogy, and the last item is from teacher&#x2019;s satisfaction. The five items were coded as the ordinal responses from 1 (least) to 3 (most). Teachers&#x2019; responses are originally coded as four ordered categorical data. However, due to too small frequencies at the lowest levels at the five variables, we modified them into three ordered categories by merging the two lowest levels: (Not/low importance, moderate importance, and high importance) in motivation, (not at all/to some extent, quite a bit, and a lot) in pedagogy, and (strongly disagree/disagree, agree, and strongly agree) in satisfaction. Other missing codes were treated as a missing code, NA. The specific explanation about the categories and the corresponding question are presented in the manual, which is also accessible via the command of TALIS?</p>
</sec>
<sec id="s4_1_2">
<label>4.1.2</label>
<title>AddHealth Data</title>
<p>AddHealth data consist of 5,144 of the participants including their responses of five item variables about substance use such as: Smoking, Alcohol, Other Types of Illegal Drug (Drug), Marijuana, and Cocaine. The responses of the five variables are dichotomous as either &#x201C;Yes&#x201D; or &#x201C;No&#x201D; and treated the other missing codes as systematic missings. The AddHealth data additionally includes a randomly generated ID variable and two demographic variables: education level and gender, which can be used as covariates. Educational level consists of eight levels from not graduating high school to beyond master&#x0027;s degree. Gender has two levels, male and female. These data were obtained from the website (<uri xlink:href="https://www.cpc.unc.edu/projects/addhealth/documentation">https://www.cpc.unc.edu/projects/addhealth/documentation</uri>) of the National Longitudinal Study of Adolescent to Adult Health (Add Health) [<xref ref-type="bibr" rid="ref-41">41</xref>]. The study has mainly focused on the investigation of how health factors in young adulthood affect adult outcomes. Although full data collection includes four additional waves since 1994, in the package gscaLCA, only the data of &#x201C;specific section of substance use&#x201D; collected at the wave IV are provided, where participants were 24 to 32 years old.</p>
</sec>
</sec>
<sec id="s4_2">
<label>4.2</label>
<title>gscaLCA Command Line and Options</title>
<p>To estimate LCA based on fuzzy clusterwise GSCA with the gscaLCA algorithm, the default function of gscaLCA can be called with the following arguments:</p>
<p><inline-graphic xlink:href="CMES_19708-inline-1.png"/></p>
<p>The command gscaLCA requires main seven options to fit the fuzzy clusterwise GSCA that are specified:
<list list-type="bullet">
<list-item>
<p>dat: Dataset to be used to fit a model of gscaLCA.</p></list-item>
<list-item>
<p>varnames: A character vector. The names of columns to be used in the function gscaLCA.</p></list-item>
<list-item>
<p>ID.var: A character element. The name of ID variable. If ID variable is not specified, the function gscaLCA will try to search an ID variable in the given data. The ID of observations will be automatically generated as a numeric variable if the dataset does not include any ID variable. The default is NULL.</p></list-item>
<list-item>
<p>num.class: An integer element. The number of classes to be identified. When num.class is smaller than 2, gscaLCA terminates with an error message. The default is 2.</p></list-item>
<list-item>
<p>num.factor: Either EACH or ALLin1. EACH specifies the situation that each indicator is assumed to be its phantom latent variable. ALLin1 indicates that all variables are assumed to be explained by a common latent variable. The default is EACH. The specification here presents the relationship between indicators and latent variables, which is used for GSCA algorithm. These two options can be expressed with a diagram, presented in <xref ref-type="fig" rid="fig-1">Fig. 1</xref>.</p></list-item>
<list-item>
<p>Boot.num: An integer element. The number of bootstrap to be identified. The standard errors of parameters are computed from the bootstrap within the gscaLCA algorithm. The default is&#x00A0;20.</p></list-item>
<list-item>
<p>multiple.Core: A logical element. TRUE enables to use multiple cores for the bootstrap. The default is FALSE.</p></list-item>
</list></p>
<fig id="fig-1">
<label>Figure 1</label>
<caption>
<title>The example of diagrams of the options EACH and ALLin1 with five indicators. The circles represent the latent variables and the square represent indicator variables</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19708-fig-1.png"/>
</fig>
<p>When a model of the fuzzy clusterwise GSCA with covariates is fitted, the additional three arguments are required:
<list list-type="bullet">
<list-item>
<p>covnames: A character vector. The names of columns of the dataset that indicate covariates in the model fitted.</p></list-item>
<list-item>
<p>cov.model: A numeric vector. It is a vector of indicators of whether each covariate is used in specifying three sub-model in GSCA. The indicator is 1 if the covariate is involved in GSCA; and otherwise 0. Involving covariates in the GSCA model indicates that the covariates are used in defining the relation between indicators and latent variables.</p></list-item>
<list-item>
<p>multinomial.ref: A character element. Options of MAX, MIN, FIRST, and LAST are available for setting a reference group. The default is MAX, which indicates that the class whose prevalence is the highest is used for a reference class in fitting a multinomial regression. Contrary, MIN indicates that the class whose prevalence is the lowest is used for a reference class. FIRST and LAST indicates that the first class and last class are used as the reference class, respectively.</p></list-item>
</list></p>
<p>In these arguments, there are kind of generic arguments such as dat, Boot.num, and multiple.Core that utilizes functions in R. On the other hand, ID.var, num.class, num.factor, covnames, cov.model, and multinomial.ref are unique functionalities associated with the <bold>gscaLCA</bold> package. Thus, a user considers what those values might be and needs to specify them to run gscaLCA.</p>
</sec>
<sec id="s4_3">
<label>4.3</label>
<title>gscaLCA Output</title>
<p>The function gscaLCA returns an object involving many different elements. We have selected key features and recommend researchers to check all other arguments by using gscaLCA:
<list list-type="bullet">
<list-item>
<p>N: The number of observations used after applying listwise deletion when missing values exist.</p></list-item>
<list-item>
<p>N.origin: The number of observations before applying listwise deletion. This number is same as the number of observations of the input dataset.</p></list-item>
<list-item>
<p>LEVELs: The observed categories for each indicator.</p></list-item>
<list-item>
<p>all.levels.equal: The indicator whether all indicators used for analysis have the same answer categories. If it is FALSE, the program does not create a graph automatically.</p></list-item>
<list-item>
<p>num.class: The number of classes used for the analysis.</p></list-item>
<list-item>
<p>Boot.num: The number of bootstrap assigned by users.</p></list-item>
<list-item>
<p>Boot.num.im: The number of bootstrap implemented. Not all iterations of bootstrap is needed to estimate the standard error.</p></list-item>
<list-item>
<p>model.fit: The model fit indices. FIT, AFIT, FPI, and NCE are provided with the standard error and 95% credible interval lower and upper bounds.</p></list-item>
<list-item>
<p>LCprevalence: The latent class prevalence. The percent of class, the number of observation for each class, standard error, and 95% credible interval lower and upper bounds are provided.</p></list-item>
<list-item>
<p>RespProb: The item response probabilities for all variables used are reported as elements of a list. Each element consists of a table containing the probabilities with respect to the possible categories of each variable. The standard error and 95% credible interval with lower and upper bounds are also reported.</p></list-item>
<list-item>
<p>it.in: The number of iteration of in-loop. The in-loop is used for the estimation of the GSCA model.</p></list-item>
<list-item>
<p>it.out: The number of iteration of out-loop. The out-loop is used to update the membership probabilities of subjects.</p></list-item>
<list-item>
<p>membership: A data frame of the posterior probability for each subject with the predicted class membership.</p></list-item>
<list-item>
<p>plot: Graphs of item response probabilities within each category. For example, with two categories, each graph is stored as p1 and p2 in the list of plot. When the number of categories of indicators are different, the graphs are not provided.</p></list-item>
<list-item>
<p>A.mat: The estimated factor loading matrix of the GSCA model.</p></list-item>
<list-item>
<p>B.mat: The estimated path coefficient matrix of the GSCA model.</p></list-item>
<list-item>
<p>W.mat: The estimated weighted relation matrix of the GSCA model.</p></list-item>
<list-item>
<p>used.dat: The dataset that used for the analysis. When the input data include missing values, the data used are ones after applying listwise deletion.</p></list-item>
</list></p>
<p>When the covariates are involved in the analysis, the function gscaLCA returns eight additional elements:
<list list-type="bullet">
<list-item>
<p>cov_results.multi.hard: This is the main result of the multinomial regression with the hard partitioning.</p></list-item>
<list-item>
<p>cov_results_raw.multi.hard: This is the result of the multinomial regression with the hard partitioning, which is directly from the function nnet::multinom.</p></list-item>
<list-item>
<p>cov_results.bin.hard: This is the main result of binominal regression with the hard partitioning.</p></list-item>
<list-item>
<p>cov_results_raw.bin.hard: This is the result of the binominal regression with the hard partitioning, which is directly from the function stats::glm.</p></list-item>
<list-item>
<p>cov_results.multi.soft: This is the main result of the multinomial regression with the soft partitioning.</p></list-item>
<list-item>
<p>cov_results_raw.multi.soft: This is the result of the multinomial regression with the soft partitioning, which is directly from the function nnet::multinom.</p></list-item>
<list-item>
<p>cov_results.bin.soft: This is the main result of binominal regression with the soft partitioning.</p></list-item>
<list-item>
<p>cov_results_raw.bin.soft: This is the result of the binominal regression with the soft partitioning, which is directly from the function stats::glm.</p></list-item>
</list></p>
</sec>
<sec id="s4_4">
<label>4.4</label>
<title>gscaLCR Command Line and Options</title>
<p>In addition to the main function gscaLCA, the package gscaLCA provides a function which implements the second and third steps in the algorithm of fuzzy clusterwise GSCA with covariates (gscaLCR). The function is called gscaLCR. As aforementioned, fuzzy clusterwise GSCA with covariate includes three steps. Even if users fit the gscaLCA first without the covariates with the function gscaLCA, steps 2 and 3 of gscaLCR can be executed via the function gscaLCR.</p>
<p>R&#x003E; gscaLCR(results.obj, covnames, multinomial.ref &#x003D; "MAX")</p>
<p>The function gscaLCR requires three elements:
<list list-type="bullet">
<list-item>
<p>results.obj: The result object of gscaLCA.</p></list-item>
<list-item>
<p>covnames: A character vector of covariate names. The covariate variables have to be in the data that used to fit the gscaLCA model.</p></list-item>
<list-item>
<p>multinomial.ref: A character element. Options of MAX, MIN, FIRST, and LAST are available for setting a reference group. The default is MAX.</p></list-item>
</list></p>
<p>The output of gscaLCR is the same as gscaLCA.</p>
</sec>
</sec>
<sec id="s5">
<label>5</label>
<title>Example</title>
<p>To demonstrate the usage of the package <bold>gscaLCA</bold>, we fit two analyses in gscaLCA with empirical exemplars: the one is fitting gscaLCA model without covariate, and the other is the one with covariates. The former is demonstrated with the TALIS data, and the latter is demonstrated with the AddHealth data. It should be noted that although the results of this study were obtained with methodological rigor, they would be slightly different from other researchers&#x2019; results within the TALIS study due to lurking variables.</p>
<sec id="s5_1">
<label>5.1</label>
<title>An Example with gscaLCA without Covariate</title>
<p>For the TALIS data, we used the three-class model (num.class &#x003D; 3) with a single factor (num.factor &#x003D; "ALLin1"). It is the optimal model with these data that was found through the model comparison. Regarding the number of class, it is expected to have three factors, motivation, pedagogy, and satisfaction, as the variable name and explanation suggests (see <xref ref-type="sec" rid="s4_1_1">Section 4.1.1</xref>). Regarding the number of factors, all variables are supposed to be explained by a common latent variable, because TALIS data is a survey data is collected from U.S. teachers with a topic of teaching and learning. The following command with the function gscaLCA was implemented to fit the fuzzy clusterwise GSCA. Once the gscaLCA command is executed, it displays the degree of the completion process as percentages while gscaLCA is running. When the estimation completes, the summary function is available to display the results. The summary function print out the sample size for the analysis, model fit indices, estimated latent class prevalence, and item response probabilities.</p>

<p><inline-graphic xlink:href="CMES_19708-inline-2.png"/></p>
<p><inline-graphic xlink:href="CMES_19708-inline-3.png"/></p>

<p>The results report that 2,365 observations were used for the analysis, excluding 195 incomplete responses. FIT and AFIT were 0.5046 and 0.5033, respectively. With a single factor, FIT and AFIT are typically lower than the larger number of factors. The indices to evaluate the classification were relatively large (FPI &#x003D; 0.8623 and NCE &#x003D; 0.8742), but they are better than when the option num.factor is EACH for these data. The estimated latent class prevalences is 34.59%, 26.22%, and 39.20%. The conditional item response probabilities for each category per variable are also presented in a table. When the standard error and 95% credible interval of the model fit, the prevalence and conditional response probabilities are required, we can print out the objects through the following commands. These standard error was estimated by the bootstrap.</p>
<p><inline-graphic xlink:href="CMES_19708-inline-4.png"/></p>
<p>These response probabilities are used to define latent classes. In order to grasp the patterns of the probabilities, a visual representation of profiles based on the probabilities would be more helpful than numeric quantities in the output above. Once the gscaLCA function is executed, the graph is automatically created. When the command summary(T3) is executed, the resulting plot is displayed for all answer categories. When the plots for each answer category are required, it can also be printed by using the command T3$plot. <xref ref-type="fig" rid="fig-2">Fig. 2</xref> presents the conditional item response probabilities of the result, T3. Based on the patterns of responses in each class, we define &#x201C;Pedagogy focused teachers&#x201D;, &#x201C;Motivated teachers&#x201D;, and &#x201C;Balanced teachers&#x201D;. For example, for the second response, latent class 1 has relatively higher value in both Pdgg_1, and Pdgg_2 variables, but latent class 1 has relatively smaller value in other variables. Therefore, we define the first latent class as the &#x201C;Pedagogy focused teachers&#x201D;. For the second latent class, it has relatively higher value in both Mtv_1, and Mtv_2 variables, but has relatively smaller value in other variables. Therefore, we named the second latent class as the &#x201C;Motivated teachers&#x201D;. The third latent class has overall similar values across all five variables, thus, we called the third latent class as the &#x201C;Balanced teschers&#x201D;. Lastly, the membership probabilities of observations can be obtained from the membership of the saved objects, T3.</p>
<p><inline-graphic xlink:href="CMES_19708-inline-5.png"/></p>
<fig id="fig-2">
<label>Figure 2</label>
<caption>
<title>Profiles of three latent classes from fuzzy clusterwise GSCA by using TALIS data</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19708-fig-2.png"/>
</fig>
</sec>
<sec id="s5_2">
<label>5.2</label>
<title>An Example of gscaLCA with Covariates</title>
<p>In this example, we demonstrate how to fit the gscaLCA with covaritate by using the AddHealth data. We used the three-class model (num.class &#x003D; 3) with num.factor &#x003D; &#x201C;EACH&#x201D;. which was found in the previous research, Park et al. [<xref ref-type="bibr" rid="ref-10">10</xref>]. For this example, we considered 5 indicators (Smoking, Alcohol, Drug, Marijuana, and Cocaine) and the gender covariate. The gender covariate was involved in fitting the GSCA model (cov.model &#x003D; 1). Adding the covariate into the GSCA model depends on researchers&#x0027; decision in their fields, but we recommend to add the covariates into the GSCA model when the covarites affect not only the membership probabilities but also the relationship between indicators and latent variables. When multiple covariates are considered, the indicators of adding the covariates into the GSCA model are presented as a vector. For example, when researcher uses the second covariate out of three covariates in the GSCA model, the option cov.model with c(0, 1, 0) can be used. Fitting gscaLCA with covariates using the function gscaLCA produces the results as follows:</p>
<p><inline-graphic xlink:href="CMES_19708-inline-6.png"/>
<inline-graphic xlink:href="CMES_19708-inline-6a.png"/></p>
<p>The results report that 5,065 observations were used for the analysis after applying the listwise deletion. They also show that the model fit indices with the AddHealth data are acceptable. FIT and AFIT were 0.9993 and 0.9993, and they are close to 1. The indices to evaluate the classification were relatively low (FPI &#x003D; 0.4658 and NCE &#x003D; 0.5094). The estimated latent class prevalences are 54.30%, 20.23%, and 25.46%. These estimated prevalences are changed depending on whether we consider covariates into a GSCA model or not. The conditional item response probabilities are also presented for each category per variable. Like the example of the TALIS data, the results here provide the 95% standard error by using the following commands:</p>
<p><inline-graphic xlink:href="CMES_19708-inline-7.png"/></p>
<p>When a category of response is binary, a graph provided by this package shows the probabilities&#x2019; patterns of each category. In the AddHealth data, thus the graph is involved when the response is &#x201C;Yes&#x201D;, because the responses are binary (Yes or No),</p>
<p>From the plot shown in <xref ref-type="fig" rid="fig-3">Fig. 3</xref>, three latent class can be defined as &#x201C;the smoking and drinking class (Latent Class 1)&#x201D;, &#x201C;the binge drinking and heavy smoking class (Latent Class 3)&#x201D;, and &#x201C;the heavy substance user class (Latent Class 2)&#x201D; as previously shown in Park et al. [<xref ref-type="bibr" rid="ref-10">10</xref>]. For the case which needs the graphs for all categories of answers, plots for all categories are provided in the results of gscaLCA. Here, the outcome element of A3_1 contains the graphs for each response in plot. The graphs can be extracted by using the following two commands: the one for the category of &#x201C;Yes&#x201D; and the other for the category of &#x201C;No&#x201D;, respectively.</p>

<fig id="fig-3">
<label>Figure 3</label>
<caption>
<title>Profiles of three latent classes from fuzzy clusterwise GSCA in AddHealth data when the gender covariate is taken account into the GSCA model</title>
</caption>
<graphic mimetype="image" mime-subtype="png" xlink:href="CMES_19708-fig-3.png"/>
</fig>
<p><inline-graphic xlink:href="CMES_19708-inline-8.png"/></p>

<p>To examine the effects of covariates on the membership probabilities, the function summary can be used as demonstrated before. Other options in partitioning and fitting regression (multinomial.soft, binomial.hard, and binomial.soft) are available as well in the function summary. In addition, the following command also provides the results of regressions:</p>
<p><inline-graphic xlink:href="CMES_19708-inline-9.png"/></p>
<p>As the example of A3_1, other options are available for hard and soft partitioning and multinomial and binomial regression.</p>
</sec>
</sec>
<sec id="s6">
<label>6</label>
<title>Discussion</title>
<p>Latent class analysis and clustering analysis including fuzzy clustering are statistical tools to identify (dis-)similarity of data distribution as a model-based mixture model and a data-driven classification approach, respectively. We developed the package, <bold>gscaLCA</bold>, in R that utilizes fuzzy clusterwise GSCA and fit class analysis with covariates as implemented in the other LCA or clustering analysis packages. Moreover, the <bold>gscaLCA</bold> package allows researchers to consider a structure of underlying constructs by utilizing the GSCA framework. This feature was unique and possible by utilizing least squares estimate. On the other hand, both methods, LCA and clustering analysis, have pros and cons but have been used in various fields at the selection of one of two methods based on researcher&#x0027;s discretion. Based on the theoretical foundation (Ryoo et al. [<xref ref-type="bibr" rid="ref-8">8</xref>]) with its efficiency in estimation, fuzzy clusterwise GSCA is now applicable to latent class analysis within the <bold>gscaLCA</bold> package, and its extension to latent class regression as a three-step approach is also available.</p>
<p>The <bold>gscaLCA</bold> package is still undergoing active development until it is equipped with as many mixture models as in the maximum likelihood-based structural equation modeling. Its next journey of gscaLCA is 1) to implement additional options on missing data such as multiple imputation [<xref ref-type="bibr" rid="ref-42">42</xref>], 2) to extend to multiple group and/or multilevel analysis [<xref ref-type="bibr" rid="ref-21">21</xref>], and 3) to extend gscaLCA to longitudinal data so as to fit latent transition analysis [<xref ref-type="bibr" rid="ref-21">21</xref>].</p>
<p>The R package, <bold>gscaLCA</bold>, provides a unified framework of fitting an LCA model utilizing fuzzy clustering algorithm and generalized structured component analysis. Both dichotomized observed variables and ordered categorical observed variables can be used in the function gscaLCA. In addition, visual representation of results profiles are a key feature in <bold>gscaLCA</bold> that helps researchers identify characteristics of classes. It should also be noted that the capacities of GSCA [<xref ref-type="bibr" rid="ref-9">9</xref>] within <bold>gscaLCA</bold> will extend the application of <bold>gscaLCA</bold> in a variety of SEM modeling.</p>
</sec>
</body>
<back>
<fn-group>
<fn fn-type="other"><p><bold>Funding Statement:</bold> This research was supported by the Yonsei University Research Fund of 2021 (2021-22-0060).</p>
</fn>
<fn fn-type="conflict"><p><bold>Conflicts of Interest:</bold> The authors declare that they have no conflicts of interest to report regarding the present study.</p>
</fn>
</fn-group>
<ref-list content-type="authoryear">
<title>References</title>
<ref id="ref-1"><label>1.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Lazarsfeld</surname>, <given-names>P. F.</given-names></string-name></person-group> (<year>1950</year>). <chapter-title>The logical and mathematical foundation of latent structure analysis</chapter-title>. In: <source>Studies in social psychology in World War II</source>, vol. 4, pp. <fpage>362</fpage>&#x2013;<lpage>412</lpage>. <publisher-loc>Princeton, NJ</publisher-loc>: <publisher-name>Princeton University Press</publisher-name>.</mixed-citation></ref>
<ref id="ref-2"><label>2.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>McCutcheon</surname>, <given-names>A. L.</given-names></string-name></person-group> (<year>1987</year>). <source>Latent class analysis</source>. <publisher-loc>Thousand Oaks, California</publisher-loc>: <publisher-name>Sage</publisher-name>.</mixed-citation></ref>
<ref id="ref-3"><label>3.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dempster</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Laird</surname>, <given-names>N.</given-names></string-name>, <string-name><surname>Rubin</surname>, <given-names>D.</given-names></string-name></person-group> (<year>1977</year>). <article-title>Maximum likelihood from incomplete data via the EM algorithm</article-title>. <source>Journal of the Royal Statistical Society, Series B (Methodological)</source><italic>,</italic> <volume>39</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>1</fpage>&#x2013;<lpage>22</lpage>. DOI <pub-id pub-id-type="doi">10.1111/j.2517-6161.1977.tb01600.x</pub-id>.</mixed-citation></ref>
<ref id="ref-4"><label>4.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Gill</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>King</surname>, <given-names>G.</given-names></string-name></person-group> (<year>2004</year>). <article-title>What to do when your hessian is not invertible: Alternatives to model respecification in nonlinear estimation</article-title>. <source>Sociological Methods &#x0026; Research</source><italic>,</italic> <volume>33</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>54</fpage>&#x2013;<lpage>87</lpage>. DOI <pub-id pub-id-type="doi">10.1177/0049124103262681</pub-id>.</mixed-citation></ref>
<ref id="ref-5"><label>5.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Steinley</surname>, <given-names>D.</given-names></string-name>, <string-name><surname>Brusco</surname>, <given-names>M. J.</given-names></string-name></person-group> (<year>2011</year>). <article-title>Evaluating mixture modeling for clustering: Recommendations and cautions</article-title>. <source>Psychological Methods</source><italic>,</italic> <volume>16</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>63</fpage>&#x2013;<lpage>79</lpage>. DOI <pub-id pub-id-type="doi">10.1037/a0022673</pub-id>.</mixed-citation></ref>
<ref id="ref-6"><label>6.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Brusco</surname>, <given-names>M. J.</given-names></string-name>, <string-name><surname>Shireman</surname>, <given-names>E.</given-names></string-name>, <string-name><surname>Steinley</surname>, <given-names>D.</given-names></string-name></person-group> (<year>2017</year>). <article-title>A comparison of latent class, K-means, and K-median methods for clustering dichotomous data</article-title>. <source>Psychological Methods</source><italic>,</italic> <volume>22</volume><italic>(</italic><issue>3</issue><italic>),</italic> <fpage>563</fpage>&#x2013;<lpage>580</lpage>. DOI <pub-id pub-id-type="doi">10.1037/met0000095</pub-id>.</mixed-citation></ref>
<ref id="ref-7"><label>7.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Lubke</surname>, <given-names>G.</given-names></string-name>, <string-name><surname>Muthen</surname>, <given-names>B.</given-names></string-name></person-group> (<year>2005</year>). <article-title>Investigating population heterogeneity with factor mixture models</article-title>. <source>Psychological Methods</source><italic>,</italic> <volume>10</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>21</fpage>&#x2013;<lpage>39</lpage>. DOI <pub-id pub-id-type="doi">10.1037/1082-989X.10.1.21</pub-id>.</mixed-citation></ref>
<ref id="ref-8"><label>8.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ryoo</surname>, <given-names>J. H.</given-names></string-name>, <string-name><surname>Park</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Kim</surname>, <given-names>S.</given-names></string-name></person-group> (<year>2020</year>). <article-title>Categorical latent variable modeling utilizing fuzzy clustering generalized structured component analysis as an alternative to latent class analysis</article-title>. <source>Behaviormetrika</source><italic>,</italic> <volume>47</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>291</fpage>&#x2013;<lpage>306</lpage>. DOI <pub-id pub-id-type="doi">10.1007/s41237-019-00084-6</pub-id>.</mixed-citation></ref>
<ref id="ref-9"><label>9.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Hwang</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Takane</surname>, <given-names>Y.</given-names></string-name></person-group> (<year>2014</year>). <source>Generalized structured component analysis: A component-based approach to structural equation modeling</source>. <publisher-loc>Boca Raton</publisher-loc>: <publisher-name>CRC Press</publisher-name>.</mixed-citation></ref>
<ref id="ref-10"><label>10.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Park</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Kim</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Ryoo</surname>, <given-names>J. H.</given-names></string-name></person-group> (<year>2020</year>). <article-title>Latent class regression utilizing fuzzy clusterwise generalized structured component analysis</article-title>. <source>Mathematics</source><italic>,</italic> <volume>8</volume><italic>(</italic><issue>11</issue><italic>),</italic> <fpage>2076</fpage>&#x2013;<lpage>2031</lpage>. DOI <pub-id pub-id-type="doi">10.3390/math8112076</pub-id>.</mixed-citation></ref>
<ref id="ref-11"><label>11.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ryoo</surname>, <given-names>J. H.</given-names></string-name>, <string-name><surname>Park</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Kim</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Ryoo</surname>, <given-names>H. S.</given-names></string-name></person-group> (<year>2020</year>). <article-title>Efficiency of cluster validity indexes in fuzzy clusterwise generalized structured component analysis</article-title>. <source>Symmetry</source><italic>,</italic> <volume>12</volume><italic>(</italic><issue>9</issue><italic>),</italic> <fpage>1514</fpage>&#x2013;<lpage>1529</lpage>. DOI <pub-id pub-id-type="doi">10.3390/sym12091514</pub-id>.</mixed-citation></ref>
<ref id="ref-12"><label>12.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Muthen</surname>, <given-names>B.</given-names></string-name>, <string-name><surname>Muthen</surname>, <given-names>L.</given-names></string-name></person-group> (<year>1998</year>). <source>Mplus user&#x2019;s guide (eighth edition)</source>. <publisher-loc>Los Angeles, CA</publisher-loc>: <publisher-name>Muthen &#x0026; Muthen</publisher-name>.</mixed-citation></ref>
<ref id="ref-13"><label>13.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>van Horn</surname>, <given-names>M. L.</given-names></string-name>, <string-name><surname>Jaki</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Masyn</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Ramey</surname>, <given-names>S. L.</given-names></string-name>, <string-name><surname>Smith</surname>, <given-names>J. A.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2009</year>). <article-title>Assessing differential effects: Applying regression mixture models to identify variations in the influence of family resources on academic achievement</article-title>. <source>Developmental Psychology</source><italic>,</italic> <volume>45</volume><italic>(</italic><issue>5</issue><italic>),</italic> <fpage>1298</fpage>&#x2013;<lpage>1313</lpage>. DOI <pub-id pub-id-type="doi">10.1037/a0016427</pub-id>.</mixed-citation></ref>
<ref id="ref-14"><label>14.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>O&#x2019;Neill</surname>, <given-names>T. A.</given-names></string-name>, <string-name><surname>McLarnon</surname>, <given-names>M. J. W.</given-names></string-name>, <string-name><surname>Xiu</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Law</surname>, <given-names>S. J.</given-names></string-name></person-group> (<year>2016</year>). <article-title>Core self-evaluations, perceptions of group potency, and job performance: The moderating role of individualism and collectivism cultural profiles</article-title>. <source>Journal of Occupational and Organizational Psychology</source><italic>,</italic> <volume>89</volume><italic>(</italic><issue>3</issue><italic>),</italic> <fpage>447</fpage>&#x2013;<lpage>473</lpage>. DOI <pub-id pub-id-type="doi">10.1111/joop.12135</pub-id>.</mixed-citation></ref>
<ref id="ref-15"><label>15.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Linzer</surname>, <given-names>D. A.</given-names></string-name>, <string-name><surname>Lewis</surname>, <given-names>J. B.</given-names></string-name></person-group> (<year>2011</year>). <article-title>PoLCA: An R package for polytomous variable latent class analysis</article-title>. <source>Journal of Statistical Software</source><italic>,</italic> <volume>42</volume><italic>(</italic><issue>10</issue><italic>),</italic> <fpage>1</fpage>&#x2013;<lpage>29</lpage>. DOI <pub-id pub-id-type="doi">10.18637/jss.v042.i10</pub-id>.</mixed-citation></ref>
<ref id="ref-16"><label>16.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Schreiber</surname>, <given-names>J. B.</given-names></string-name></person-group> (<year>2017</year>). <article-title>Latent class analysis: An example for reporting results</article-title>. <source>Research in Social and Administrative Pharmacy</source><italic>,</italic> <volume>13</volume><italic>(</italic><issue>6</issue><italic>),</italic> <fpage>1196</fpage>&#x2013;<lpage>1201</lpage>. DOI <pub-id pub-id-type="doi">10.1016/j.sapharm.2016.11.011</pub-id>.</mixed-citation></ref>
<ref id="ref-17"><label>17.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Miranda</surname>, <given-names>V. P. N.</given-names></string-name>, <string-name><surname>dos Santos Amorim</surname>, <given-names>P. R.</given-names></string-name>, <string-name><surname>Bastos</surname>, <given-names>R. R.</given-names></string-name>, <string-name><surname>Souza</surname>, <given-names>V. G. B.</given-names></string-name>, <string-name><surname>de Faria</surname>, <given-names>E. R.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2019</year>). <article-title>Evaluation of lifestyle of female adolescents through latent class analysis approach</article-title>. <source>BMC Public Health</source><italic>,</italic> <volume>19</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>1</fpage>&#x2013;<lpage>12</lpage>. DOI <pub-id pub-id-type="doi">10.1186/s12889-019-6488-8</pub-id>.</mixed-citation></ref>
<ref id="ref-18"><label>18.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>van Rijnsoever</surname>, <given-names>F. J.</given-names>, </string-name><string-name><surname>Castaldi</surname>, <given-names>C.</given-names></string-name></person-group> (<year>2011</year>). <article-title>Extending consumer categorization based on innovativeness: Intentions and technology clusters in consumer electronics</article-title>. <source>Journal of the American Society for Information Science and Technology</source><italic>,</italic> <volume>62</volume><italic>(</italic><issue>8</issue><italic>),</italic> <fpage>1604</fpage>&#x2013;<lpage>1613</lpage>. DOI <pub-id pub-id-type="doi">10.1002/asi.21567</pub-id>.</mixed-citation></ref>
<ref id="ref-19"><label>19.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Xia</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Evans</surname>, <given-names>F. H.</given-names></string-name>, <string-name><surname>Spilsbury</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Ciesielski</surname>, <given-names>V.</given-names></string-name>, <string-name><surname>Arrowsmith</surname>, <given-names>C.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2010</year>). <article-title>Market segments based on the dominant movement patterns of tourists</article-title>. <source>Tourism Management</source><italic>,</italic> <volume>31</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>464</fpage>&#x2013;<lpage>469</lpage>. DOI <pub-id pub-id-type="doi">10.1016/j.tourman.2009.04.013</pub-id>.</mixed-citation></ref>
<ref id="ref-20"><label>20.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Lanza</surname>, <given-names>S. T.</given-names></string-name>, <string-name><surname>Dziak</surname>, <given-names>J. J.</given-names></string-name>, <string-name><surname>Huang</surname>, <given-names>L.</given-names></string-name>, <string-name><surname>Wagner</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Collins</surname>, <given-names>L. M.</given-names></string-name></person-group> (<year>2015</year>). <source>PROC LCA &#x0026; PROC LTA users&#x2019; guide version 1.3.2</source>. <publisher-loc>State College, PA</publisher-loc>: <publisher-name>The Methodology Center</publisher-name>.</mixed-citation></ref>
<ref id="ref-21"><label>21.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Collins</surname>, <given-names>L. M.</given-names></string-name>, <string-name><surname>Lanza</surname>, <given-names>S. T.</given-names></string-name></person-group> (<year>2009</year>). <source>Latent class and latent transition analysis: With applications in the social behavioral, and health sciences</source>, vol. 718. <publisher-loc>Hoboken, NJ</publisher-loc>: <publisher-name>John Wiley &#x0026; Sons</publisher-name>.</mixed-citation></ref>
<ref id="ref-22"><label>22.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Reynolds</surname>, <given-names>G. L.</given-names></string-name>, <string-name><surname>Fisher</surname>, <given-names>D. G.</given-names></string-name></person-group> (<year>2019</year>). <article-title>A latent class analysis of alcohol and drug use immediately before or during sex among women</article-title>. <source>The American Journal of Drug and Alcohol Abuse</source><italic>,</italic> <volume>45</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>179</fpage>&#x2013;<lpage>188</lpage>. DOI <pub-id pub-id-type="doi">10.1080/00952990.2018.1528266</pub-id>.</mixed-citation></ref>
<ref id="ref-23"><label>23.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Ryoo</surname>, <given-names>J. H.</given-names></string-name>, <string-name><surname>Wang</surname>, <given-names>C.</given-names></string-name>, <string-name><surname>Swearer</surname>, <given-names>S. M.</given-names></string-name>, <string-name><surname>Park</surname>, <given-names>S.</given-names></string-name></person-group> (<year>2017</year>). <article-title>Investigation of transitions in bullying/victimization statuses of gifted and general education students</article-title>. <source>Exceptional Children</source><italic>,</italic> <volume>83</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>396</fpage>&#x2013;<lpage>411</lpage>. DOI <pub-id pub-id-type="doi">10.1177/0014402917698500</pub-id>.</mixed-citation></ref>
<ref id="ref-24"><label>24.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hwang</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Takane</surname>, <given-names>Y.</given-names></string-name>, <string-name><surname>Jung</surname>, <given-names>K.</given-names></string-name></person-group> (<year>2017</year>). <article-title>Generalized structured component analysis with uniqueness terms for accommodating measurement error</article-title>. <source>Frontiers in Psychology</source><italic>,</italic> <volume>8</volume><italic>,</italic> <fpage>2137</fpage>&#x2013;<lpage>2148</lpage>. DOI <pub-id pub-id-type="doi">10.3389/fpsyg.2017.02137</pub-id>.</mixed-citation></ref>
<ref id="ref-25"><label>25.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Widaman</surname>, <given-names>K. F.</given-names></string-name></person-group> (<year>2007</year>). <chapter-title>Common factors versus components: Principals and principles, errors and misconceptions</chapter-title>. In: <person-group person-group-type="editor"><string-name><surname>Cudeck</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>MacCallum</surname>, <given-names>R. C.</given-names></string-name></person-group> (Eds.), <source>Factor analysis at 100: Historical developments and future directions</source>, pp. <fpage>177</fpage>&#x2013;<lpage>203</lpage>. <publisher-loc>Mahwah, NJ</publisher-loc>: <publisher-name>Lawrence Erlbaum Associates Publishers</publisher-name>.</mixed-citation></ref>
<ref id="ref-26"><label>26.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Bezdek</surname>, <given-names>J. C.</given-names></string-name></person-group> (<year>1981</year>). <source>Pattern recognition with fuzzy objective function algorithms; advanced applications in pattern recognition</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Plenum Press</publisher-name>.</mixed-citation></ref>
<ref id="ref-27"><label>27.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Mahata</surname>, <given-names>K.</given-names></string-name>, <string-name><surname>Sarkar</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Das</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Das</surname>, <given-names>S.</given-names></string-name></person-group> (<year>2017</year>). <chapter-title>Fuzzy evaluated quantum cellular automata approach for watershed image analysis</chapter-title>. In: <person-group person-group-type="editor"><string-name><surname>Bhattacharyya</surname>, <given-names>S.</given-names></string-name>, <string-name><surname>Maulik</surname>, <given-names>U.</given-names></string-name>, <string-name><surname>Dutta</surname>, <given-names>P.</given-names></string-name></person-group> (Eds.), <source>Quantum inspired computational intelligence</source>, pp. <fpage>259</fpage>&#x2013;<lpage>284</lpage>. <publisher-loc>Boston, MA</publisher-loc>: <publisher-name>Morgan Kaufmann</publisher-name>.</mixed-citation></ref>
<ref id="ref-28"><label>28.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Hastie</surname>, <given-names>T.</given-names></string-name>, <string-name><surname>Tibshirani</surname>, <given-names>R.</given-names></string-name>, <string-name><surname>Friedman</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2009</year>). <source>The elements of statistical learning: Data mining, inference, and prediction</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer</publisher-name>.</mixed-citation></ref>
<ref id="ref-29"><label>29.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Mulaik</surname>, <given-names>S.</given-names></string-name></person-group> (<year>2009</year>). <source>Foundations of factor analysis (2nd edition)</source>. <publisher-loc>Boca Raton</publisher-loc>: <publisher-name>CRC Press</publisher-name>.</mixed-citation></ref>
<ref id="ref-30"><label>30.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Hwang</surname>, <given-names>H.</given-names></string-name>, <string-name><surname>Desarbo</surname>, <given-names>W. S.</given-names></string-name>, <string-name><surname>Takane</surname>, <given-names>Y.</given-names></string-name></person-group> (<year>2007</year>). <article-title>Fuzzy clusterwise generalized structured component analysis</article-title>. <source>Psychometrika</source><italic>,</italic> <volume>72</volume><italic>(</italic><issue>2</issue><italic>),</italic> <fpage>181</fpage>&#x2013;<lpage>198</lpage>. DOI <pub-id pub-id-type="doi">10.1007/s11336-005-1314-x</pub-id>.</mixed-citation></ref>
<ref id="ref-31"><label>31.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Bollen</surname>, <given-names>K. A.</given-names></string-name>, <string-name><surname>Curran</surname>, <given-names>P. J.</given-names></string-name></person-group> (<year>2006</year>). <source>Latent curve models: A structural equation perspective</source>. <publisher-loc>Hoboken, NJ</publisher-loc>: <publisher-name>John Wiley &#x0026; Sons</publisher-name>.</mixed-citation></ref>
<ref id="ref-32"><label>32.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Young</surname>, <given-names>F. W.</given-names></string-name></person-group> (<year>1981</year>). <article-title>Quantitative analysis of qualitative data</article-title>. <source>Psychometrika</source><italic>,</italic> <volume>46</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>357</fpage>&#x2013;<lpage>388</lpage>. DOI <pub-id pub-id-type="doi">10.1007/BF02293796</pub-id>.</mixed-citation></ref>
<ref id="ref-33"><label>33.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Efron</surname>, <given-names>E.</given-names></string-name></person-group> (<year>1979</year>). <source>Bootstrap methods: Another look at the jackknife</source>. <publisher-loc>New York, NY</publisher-loc>: <publisher-name>Springer</publisher-name>.</mixed-citation></ref>
<ref id="ref-34"><label>34.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Roubens</surname>, <given-names>M.</given-names></string-name></person-group> (<year>1982</year>). <article-title>Fuzzy clustering algorithms and their cluster validity</article-title>. <source>European Journal of Operational Research</source><italic>,</italic> <volume>10</volume><italic>(</italic><issue>3</issue><italic>),</italic> <fpage>294</fpage>&#x2013;<lpage>301</lpage>. DOI <pub-id pub-id-type="doi">10.1016/0377-2217(82)90228-4</pub-id>.</mixed-citation></ref>
<ref id="ref-35"><label>35.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Vermunt</surname>, <given-names>J. K.</given-names></string-name></person-group> (<year>2010</year>). <article-title>Latent class modeling with covariates: Two improved three-step approaches</article-title>. <source>Political Analysis</source><italic>,</italic> <volume>18</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>450</fpage>&#x2013;<lpage>469</lpage>. DOI <pub-id pub-id-type="doi">10.1093/pan/mpq025</pub-id>.</mixed-citation></ref>
<ref id="ref-36"><label>36.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Vermunt</surname>, <given-names>J. K.</given-names></string-name></person-group> (<year>1997</year>). <source>LEM: A general program for the analysis of categorical data</source>. <publisher-loc>Netherlands</publisher-loc>: <publisher-name>Department of Methodology and Statistics, Tilburg University</publisher-name>.</mixed-citation></ref>
<ref id="ref-37"><label>37.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Yamaguchi</surname>, <given-names>K.</given-names></string-name></person-group> (<year>2000</year>). <article-title>Multinomial logit latent-class regression models: An analysis of the predictors of gender-role attitudes among Japanese women</article-title>. <source>American Journal of Sociology</source><italic>,</italic> <volume>105</volume><italic>(</italic><issue>6</issue><italic>),</italic> <fpage>1702</fpage>&#x2013;<lpage>1740</lpage>. DOI <pub-id pub-id-type="doi">10.1086/210470</pub-id>.</mixed-citation></ref>
<ref id="ref-38"><label>38.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Bolck</surname>, <given-names>A.</given-names></string-name>, <string-name><surname>Croon</surname>, <given-names>M.</given-names></string-name>, <string-name><surname>Hagenaars</surname>, <given-names>J.</given-names></string-name></person-group> (<year>2004</year>). <article-title>Estimating latent structure models with categorical variables: One-step versus three-step estimators</article-title>. <source>Political Analysis</source><italic>,</italic> <volume>12</volume><italic>(</italic><issue>1</issue><italic>),</italic> <fpage>3</fpage>&#x2013;<lpage>27</lpage>. DOI <pub-id pub-id-type="doi">10.1093/pan/mph001</pub-id>.</mixed-citation></ref>
<ref id="ref-39"><label>39.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><collab>OECD</collab></person-group> (<year>2019</year>). <article-title>Teachers and school leaders as lifelong learners. In: <italic>TALIS, 2018 results</italic></article-title>, vol. 1. <publisher-loc>Paris</publisher-loc>: <publisher-name>OECD Publishing</publisher-name>.</mixed-citation></ref>
<ref id="ref-40"><label>40.</label><mixed-citation publication-type="journal"><person-group person-group-type="author"><string-name><surname>Dias</surname>, <given-names>J. G.</given-names></string-name>, <string-name><surname>Vermunt</surname>, <given-names>J. K.</given-names></string-name></person-group> (<year>2008</year>). <article-title>A bootstrap-based aggregate classifier for model-based clustering</article-title>. <source>Computational Statistics</source><italic>,</italic> <volume>23</volume><italic>(</italic><issue>4</issue><italic>),</italic> <fpage>643</fpage>&#x2013;<lpage>659</lpage>. DOI <pub-id pub-id-type="doi">10.1007/s00180-007-0103-7</pub-id>.</mixed-citation></ref>
<ref id="ref-41"><label>41.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Harris</surname>, <given-names>K. M.</given-names></string-name>, <string-name><surname>Halpern</surname>, <given-names>C. T.</given-names></string-name>, <string-name><surname>Whitsel</surname>, <given-names>E.</given-names></string-name>, <string-name><surname>Hussey</surname>, <given-names>J.</given-names></string-name>, <string-name><surname>Tabor</surname>, <given-names>J.</given-names></string-name> <etal>et al.</etal></person-group> (<year>2009</year>). <source>The national longitudinal study of adolescent to adult health</source>. <publisher-loc>Chapel Hill, NC</publisher-loc>: <publisher-name>Carolina Population Center</publisher-name>.</mixed-citation></ref>
<ref id="ref-42"><label>42.</label><mixed-citation publication-type="book"><person-group person-group-type="author"><string-name><surname>Rubin</surname>, <given-names>D. B.</given-names></string-name></person-group> (<year>2004</year>). <source>Multiple imputation for nonresponse in surveys</source>. <publisher-loc>Hoboken, NJ</publisher-loc>: <publisher-name>John Wiley &#x0026; Sons</publisher-name>.</mixed-citation></ref>
</ref-list>
</back>
</article>
