<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>J. Baruník, T. Kley, Quantile coherency: A general measure for dependence between cyclical economic variables, The Econometrics Journal</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Spatial weighted robust clustering of multivariate time series based on quantile dependence with an application to mobility during COVID-19 pandemic</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ángel López-Oriona</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pierpaolo D'Urso</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>José A. Vilar</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Borja Lafuente-Rego</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Mathematics, Research Group MODES, Research Center for Information and Communication Technologies (CITIC), University of A Coruña</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Social Sciences and Economics, Sapienza University of Rome</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>22</volume>
      <issue>2019</issue>
      <fpage>0000</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>In this paper, a fuzzy clustering model for multivariate time series based on the quantile cross-spectral density and principal component analysis is improved. The extension consists of (i) a weighting system which assigns a weight to each principal component in accordance with its importance concerning the underlying clustering structure and (ii) a penalization term allowing to take into account the spatial information. The iterative solutions of the new model, which employs the exponential distance in order to gain robustness against outlying series, are derived. A simulation study shows that the introduction of the weighting system substantially enhances the efectiveness of the former approach. The behaviour of the extended model in terms of the spatial penalization term is also analysed. An application involving multivariate time series of mobility indicators concerning COVID-19 pandemic highlights the usefulness of the proposed technique.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;clustering</kwd>
        <kwd>multivariate time series</kwd>
        <kwd>principal component analysis</kwd>
        <kwd>quantile cross-spectral density</kwd>
        <kwd>weighting system</kwd>
        <kwd>spatial statistics</kwd>
        <kwd>COVID-19</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction and related work</title>
      <p>Clustering of time series is a central problem in data mining with applications in many fields [1, 2]. The objective
is to split a large set of unlabelled time series into homogeneous groups so that similar series are placed together
in the same group and dissimilar series are located in diferent groups. The majority of the proposed approaches
concern univariate time series (UTS) [3, 4, 5], while clustering of multivariate time series (MTS) has received much
less attention [6].</p>
      <p>This paper improves a fuzzy clustering model for MTS proposed in our previous work [7]. This model is based
on the so-called Quantile Cross-spectral Density (QCD) and the classical Principal Component Analysis (PCA).
Although highly successful, the model in [7] sufers from two major drawbacks. First, each principal component
contributes equally to the objective function of the clustering algorithm. This overlooks the fact that some principal
components could have more importance than others concerning the underlying clustering structure. Second,
no spatial information is considered in the clustering model, so it can not handle time series datasets containing
geographical information in a proper way. This paper proposes a new fuzzy clustering model aimed at overcoming
both previously mentioned issues.</p>
      <sec id="sec-1-1">
        <title>2.1. A fuzzy clustering model based on QCD and PCA</title>
        <p>Let {,  ∈ Z} = {(,1, . . . , ,),  ∈ Z} be a -variate real-valued strictly stationary stochastic process.
Denote by  the marginal distribution function of , ,  = 1, . . . , , and by  ( ) = − 1( ),  ∈ [0, 1], the
corresponding quantile function. Fixed  ∈ Z and an arbitrary couple of quantile levels (,  ′) ∈ [0, 1]2, consider
the cross-covariance of the indicator functions  {,1 ≤ 1 ( )} and  {+,2 ≤ 2 ( ′)} given by
 1,2 (, ,  ′) = Cov (︀  {,1 ≤ 1 ( )} ,  {︀ +,2 ≤ 2 ( ′)}︀)
for 1 ≤ 1, 2 ≤ . Taking 1 = 2 = , the function  , (, ,  ′), with (,  ′) ∈ [0, 1]2, so-called quantile
autocovariance function (QAF) of lag , generalizes the traditional autocovariance function.</p>
        <p>In the case of the multivariate process {,  ∈ Z}, we can consider the  ×  matrix</p>
        <p>Γ(, ,  ′) = (︀  1,2 (, ,  ′))︀ 1≤ 1,2≤  ,
which jointly provides information about both the cross-dependence (when 1 ̸= 2) and the serial dependence
(because the lag  is considered).</p>
        <p>In the same way as the spectral density is the representation in the frequency domain of the autocovariance
function, the spectral counterpart for the cross-covariances  1,2 (, ,  ′) can be introduced. Under suitable
summability conditions (mixing conditions), the Fourier transform of the cross-covariances is well-defined and the quantile
cross-spectral density is given by</p>
        <p>In Section 2, the procedure presented in [7] is concisely revised and its weighted version is introduced. A brief
numerical experiment is carried out in order to show why introducing a weighting system in the principal
components space is advantageous. Section 3 describes the new proposed model, in which spatial information is taken
into account. A simple simulation study is performed to assess the behaviour of the procedure. In Section 4, the
technique is applied to a dataset of mobility concerning the COVID-19 pandemic. Section 5 concludes. Finally, the
Appendix contains the derivation of the iterative solutions of the proposed algorithm.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Fuzzy clustering models based on the quantile cross-spectral density</title>
      <p>In this section, we first introduce a fuzzy clustering model based on QCD and PCA. Then, we present a modification
of this model in which a diferent weight is given to each principal component. A numerical experiment shows the
usefulness of the latter approach.</p>
      <p>∞
f1,2 (, ,  ′) = (1/2 ) ∑︁  1,2 (, ,  ′)− ,</p>
      <p>=−∞
for 1 ≤ 1, 2 ≤ ,  ∈ R and ,  ′ ∈ [0, 1]. Note that f1,2 (, ,  ′) is complex-valued.</p>
      <p>The quantile cross-spectral density contains information about the general dependence structure of a given
stochastic process. For a specific realization of the process, this quantity can be consistently estimated by means of
the so-called smoothed CCR-periodogram, ^1,,2 (, ,  ′), proposed by [8].</p>
      <p>
        Based on previous remarks, a simple dissimilarity measure between two realizations of the -variate process
(MTS) can be defined as follows. Given the -th MTS, (), consider the set () = {^1,,2 (, ,  ′), 1, 2 =
1, . . . , ,  ∈ Ω, ,  ′ ∈  }, where Ω is the set of Fourier frequencies and  = {0.1, 0.5, 0.9}. Let Ψ() be
the vector formed by concatenating separately the real and imaginary parts of the elements of the set (). A
dissimilarity measure between the series (1) and (2) is defined as the squared Euclidean distance between the
vectors Ψ(1) and Ψ(
        <xref ref-type="bibr" rid="ref1">2</xref>
        ). We call this dissimilarity .
      </p>
      <p>
        The distance measure  can be used as input to the traditional fuzzy -medoids algorithm to develop a
clustering procedure for MTS. However, the numerical simulations carried out in [7] have shown that performing
PCA as a preprocessing step frequently improves the efectiveness of the metric  in a clustering context. More
(1)
(
        <xref ref-type="bibr" rid="ref1">2</xref>
        )
(3)
whose goal is to find the subset of Ψ of size , Ψ̃︀ = { Ψ̃︀
 , solving the minimization problem
precisely, given a set of  MTS, {(1), . . . , ()}, the previously mentioned QCD-based features are extracted
from each series, thus providing the set Ψ = {Ψ(1), . . . , Ψ()}. These vectors are transformed by means of
PCA, obtaining the set of score vectors Ψ  = {Ψ(1), . . . , Ψ()}. The corresponding distance used in
the clustering algorithm is   ((1), (2)) = ⃦⃦⃦ Ψ(1) − Ψ(2)⃦⃦⃦ 2. From now on, we assume that this
is the distance being considered, although the subscript   is removed for the sake of simplicity. It is worth
remarking that the use of PCA has been shown to improve the accuracy of clustering algorithms in several contexts
[9, 10, 11, 12].
      </p>
      <p>The metric  is used to develop the so-called QCD-based Fuzzy -Medoids clustering model (QCD-FCMd),
(1) ()
, . . . , Ψ̃︀</p>
      <p>}, and the  ×  matrix of fuzzy coeficients,
mΨ̃︀ ,in ∑=︁1 ∑=︁1  ⃦⃦⃦ Ψ() − Ψ̃︀ ()⃦⃦⃦ 2</p>
      <p>with respect to ∑︁  = 1 ∀ and  ≥ 0,
=1
where  ∈ [0, 1] represents the membership degree of the -th series in the -th cluster, Ψ̃︀ is the vector of
QCD-based features with regards to the medoid series for the cluster , and  &gt; 1 is a parameter controlling the
fuzziness of the partition, usually referred to as fuzziness parameter.</p>
      <p>Several hyperparameters have to be selected in the model QCD-FCMd. In particular, the number of clusters,
, the fuziness parameter,  and the number of retained principal components, , need to be set in advance. We
propose to perform hyperparameter selection through a grid search by considering the four internal clustering
quality indexes given in [13]. It is worth mentioning that the model QCD-FCMd was compared in [7] with several
alternative procedures, clearly outperforming all of them.</p>
      <p>
        Despite its good behaviour, The QCD-FCMd model sufers from a major limitation: all the selected principal
components receive the same importance concerning the distances in (
        <xref ref-type="bibr" rid="ref4">4</xref>
        ), thus ignoring that some principal
components often contain more information than others about the underlying clustering structure. Solving the limitations
of the QCD-FCMd model is one of the main motivations of the present work.
()
      </p>
      <sec id="sec-2-1">
        <title>2.2. A weighted fuzzy clustering model based on QCD and PCA</title>
        <p>
          Consider the set of MTS {(1), . . . , ()} and denote by Ψ = {Ψ(1), . . . , Ψ()} the set of QCD-based features.
Assume that the subset of  first principal componens was selected. In this way, each Ψ() is a -dimensional vector,
Ψ() = (Ψ(1), . . . , Ψ()),  = 1, . . . , . We introduce now the Weighted QCD-based Fuzzy -Medoids clustering
model (W-QCD-FCMd), which intends to get around the limitations of the QCD-FCMd model by considering the
minimization problem
(
          <xref ref-type="bibr" rid="ref4">4</xref>
          )
(5)
subject to
min ∑︁ ∑︁  ∑︁ (︂ ︁( Ψ() − Ψ̃︀ ())︁ )︂ 2
Ψ̃︀ ,, =1 =1 =1
 
∑︁  = 1 and  ≥ 0, ∑︁  = 1,  ≥ 0, for  = 1, . . . , ,
=1 =1
where  = {1, . . . , } is a set of  weights and the remaining terms are the same as in (
          <xref ref-type="bibr" rid="ref4">4</xref>
          ). The W-QCD-FCMd
model associates the weight  to the -th principal component when computing the distance in (5). Since the
weights are variables in the objective function, the set  gets properly estimated. Basically, W-QCD-FCMd allows
for an appropriate tuning of the influence of the various principal components when computing the dissimilarity
between MTS. Hence, the algorithm can give more importance to the principal components containing the most
valuable information about the structure of the dataset. The iterative solutions of (5) can be easily obtained through
the Lagrange multipliers method and the procedure is quite similar to the one shown in the Appendix for computing
the solutions of the model presented in Section 3.
        </p>
        <p>In order to show the advantages of the W-QCD-FCMd model over QCD-FCMd model, we constructed a simple
simulation study as follows. We considered a scenario involving four diferent generating processes. 10 series
were simulated from each process and the corresponding set of 40 MTS was subject to clustering by means of both
models QCD-FCMd and W-QCD-FCMd. The generating processes were bivariate vector moving average processes
with vectorized matrices of coeficients (− 0.3, 0, 0, 0), (0.5, 0, 0, 0), (0, 0, 0.3, 0) and (0, 0, 0, 0.7), and Gaussian
innovations. The number of clusters was set to  = 4. The series length was set to  = 300. Several values for the
fuzziness parameter, , and the retained number of principal components, , were taken into account. Specifically,
we considered  = 1.4, 1.6, 1.8, 2 and  = 1, 2, 3, 4, 5.</p>
        <p>According to our clustering purpose, the partition defined by the generating processes was assumed to be the true
partition (the ground truth) and the assessment of both approaches was carried out with the fuzzy extension of the
Adjusted Rand Index [14], denoted by FARI. The simulation procedure was repeated 200 trials and average values of
FARI were calculated for each combination of  and . Table 1 contains the results. It is clear that, overall, the model
W-QCD-FCMd substantially outperformed QCD-FCMd in most of the settings. The only exception was when  = 1,
when both models are the same (only one weight is present in W-QCD-FCMd for the first principal component).
We have performed additional numerical experiments, achieving similar results. Therefore, we conclude that the
W-QCD-FCMd model substantially improves the QCD-FCMd model.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. The QCD-based exponential weighted fuzzy C-medoids clustering model with spatial component</title>
      <p>FCMd-S), whose aim is to find the subset of Ψ of size , Ψ̃︀ = { Ψ̃︀ , . . . , Ψ̃︀ }, the set of  weights  =
{1, . . . , }, and the  ×  matrix of fuzzy coeficients,  = (), solving the minimization problem:
  
min ∑︁ ∑︁  [︁1 − exp (︀ −  ∑︁ 2(Ψ() −
Ψ̃︀ ,, =1 =1 =1
  
Ψ̃︀ ())2)︀ ]︁ + 2 ∑︁ ∑︁  ∑︁ ∑︁ ′ ′′
where  &gt; 0 is a constant tuning the membership degrees to gain robustness to outlying series;  = (′ ) is
a symmetric  ×  matrix introducing information on spatial proximity of the statistical units, with ′ ≥ 0 for
 ̸= ′ and  = 0,  = 1, . . . , ;  ≥ 0 is a coeficient regulating the efect of the spatial proximity (such as
described below);  = {1, . . . ,  − 1,  + 1, . . . } denotes the set of all the clusters except cluster ; and the
remaining elements are the same as in (5).</p>
      <p>The objective function in (6) is formed by two specific terms, which are described in detail below.</p>
      <p>[1] The term ∑︁ ∑︁  [︁1 − exp (︀ −  ∑︁ 2(Ψ() − Ψ̃︀ ())2)︀ ]︁ corresponds, in essence, to the standard term in
=1 =1 =1
the classical fuzzy -medoids clustering algorithm but including two main modifications. The Euclidean distance is
replaced by an exponential-type distance to endow the clustering procedure with higher robustness against outliers
[15]. The exponential distance assigns small weights to anomalous points and larger weights to elements from
welldefined and compact clusters (for more details, see [15, 16]). Moreover, this term also considers diferent weights for
each principal component in order to improve the accuracy of the standard non-weighted version such as argued
in Section 2.</p>
      <p>[2] The second term,  ∑︁ ∑︁  ∑︁ ∑︁ ′ ′′ , is a spatial penalty term, which can be seen as a
regulariza2 =1 =1 ′=1 ′∈
tion element. The matrix  incorporates the spatial information of the  units where the MTS were recorded and
it is often restricted to indicate if two diferent units are or not contiguous, i.e.,  is usually constructed as follows:
′ = 1 if MTS  is contiguous to MTS ′, 0 otherwise.
(7)</p>
      <p>The goal of the spatial regularization term is penalizing the fact that two contiguous units are located in diferent
clusters with high membership degrees, since it is expected that neighbouring sites tend to have similar features.
The constant  regulates the trade-of between internal cohesion based on the set Ψ and the fact that the clusters
are constituted by adjacent MTS. In this way, large values of  force neighbouring spatial units to belong to the
same cluster, whereas when  = 0, the spatial information is not taken into account.</p>
      <p>The optimal iterative solutions of the minimization problem (6) are given in the Appendix.</p>
      <p>Remark 1 (Selection of  ). An appropriate choice of the hyperparameter  is totally crucial for a good performance
of the EW-QCD-FCMd-S procedure. The value of  is determined in this work as the inverse of the variability in
the data, which is a common choice in the literature. For more details, see [16].</p>
      <p>Remark 2 (Selection of  ). The selection of the optimal value for  is a complex problem and frequently relies on
heuristic procedures. To that aim, we adapt a procedure given in [17], which is based on maximizing the so-called
within cluster spatial autocorrelation.</p>
      <p>A brief simulation study was carried out in order to assess the behaviour of the EW-QCD-FCMd-S model. The
simulation mechanism involved two vector autoregressive processes or order 1, denoted by VAR1 and VAR2, with
vectorized matrices of coeficients (− 0.2, 0.2, 0.2, 0.2) and (0.3, − 0.2, 0, 0), respectively, and Gaussian
innovations. Ten diferent series were generated from each process. The first 8 MTS and MTS 19 and 20 pertained to
VAR1. On the other hand, MTS from 9 to 18 were associated with VAR2. We encapsulated the spatial information
by means of a matrix  indicating that the first 10 MTS belonged to the same area, the same holding for the last
10 MTS. The matrix  was constructed as
 = 1, ,  = 1, . . . , 10,  ̸= ,  = 1, ,  = 11, . . . , 20,  ̸= .
(8)</p>
      <p>The remaining entries of the matrix  were set to zero. The number of clusters was set to  = 2. The series
length was set to  = 500. The hyperparameter  was chosen to be  = 1 for illustrative purposes. Concerning
the number of principal components, , we selected the optimal value of  ∈ {2, . . . , 6} according to the FARI
attained when  = 0 and the true partition is given by the true generating processes, which resulted  = 2.</p>
      <p>The simulation mechanism was performed for all pairs (,  ), with  = 1.4, 1.6, 1.8, 2 and  = 0, 0.01, 0.05,
0.1, 0.2, 0.5. Note that, when  = 0, there is no spatial penalization term, whereas when increasing the values of  ,
more importance is given to this term, which is, the clusters are more required to be formed by series in the same
area to a higher extent.</p>
      <p>To gain insights into the behaviour of EW-QCD-FCMd-S from a spatial point of view, we recorded, for each pair
(,  ), the sum of the membership degrees of the MTS 9, 10, 19 and 20 with respect to the cluster in which they
must be located according to the matrix  . In other words, we took the membership degrees of MTS 9 and 10 in
the cluster where MTS 1-8 showed a higher membership degree. Analogously, we took the membership degrees of
MTS 19 and 20 in the cluster where MTS 11-18 showed a higher membership degree. The average of the 4 previous
quantities was taken into account. Note that this measure indicates to what extent the spatial requirements defined
by  are fulfilled. Results are given in Table 2 concerning 200 simulation trials. As expected, when  = 0, the
corresponding quantities are close to 0 as the grouping is made uniquely based on underlying generating processes.
When the value of  increases, the MTS 9, 10, 19 and 20 are more afected by the penalization term, constituting
clusters along with MTS in the same areas, and the corresponding measures also increase. Indeed, for  = 0.5, the
spatial requirement gets so stringent that the series are located in their spatial clusters with very high membership
degrees.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Application. Spatial clustering of Spanish regions based on mobility time series during COVID-19 pandemic</title>
      <p>In this section we develop a study case related to the non-supervised classification of geographical zones in terms
of their temporal records of some mobility indicators concerning the COVID-19 pandemic.</p>
      <p>In particular, we considered mobility time series observed during COVID-19 outbreak in the 17 autonomous
communities of Spain. Specifically, we recorded daily data about mobility in (i) retail and recreation places, (ii)
groceries and pharmacies and (ii) transit stations. The three variables are measured as a percentage of change with
respect to a baseline level associated with the pre-pandemic stage. The sample period spans from 15th February 2020
to 22nd May 2021, thus resulting serial realizations of length  = 1096. Each community is described by means of
a trivariate series indicating the daily percent of change in mobility with respect to the mentioned locations. Note
that, from an epidemological point of view, it is reasonable to think that the joint behaviour of these three quantities
is diferent depending on the geographical location of each region.</p>
      <p>The EW-QCD-FCMd-S model was applied to the set of 17 series. Note that 5 hyperparameters had to be chosen,
namely , , ,  and  . To choose ,  and , we applied a grid search based on the four internal clustering
quality indexes presented in [13]. To this aim, we considered the W-QCD-FCMd clustering model, which does not
require the selection of  and  . The optimal combination was (, , ) = (2, 1.7, 4). Given the selected value
 = 2, the hyperparameter  was chosen according to the procedure indicated in [18]. The corresponding value
was  = 0.432. The matrix  was constructed so that  = 1 if communities  and  are adjacent and  =0
otherwise. In this way, the only communities lacking neighbours are the Canary Islands and the Balearic Islands.
The optimal value of the tuning parameter  was chosen according to the procedure described in [17]. The optimal
value of  was  = 1.76. The optimal weights were 0.82 and 0.18 for the first and second principal component,
respectively.</p>
      <p>Table 3 shows the membership degrees concerning the 4 clustering solution given by EW-QCD-FCMd-S. For
each community, the highest value for the membership degree was highlighted in bold. According to Table 3, there
exist a small cluster, 4, including the autonomous communities of Asturias (AS) and Cantabria (CA). None of the
remaining communities show a high membership degree in this cluster. There are three more clusters, 1, 2
and 3, each one containing 5, 4, and 6 communities, respectively. Interestingly, the community of Castile and
León (CL) presents membership degrees which are spread out between the 4 clusters. Hence, this region could be
considered an outlier.</p>
      <p>In order to clarify the results in Table 3, we have depicted in Figure 1 a map of Spain where the communities
were coloured according to the underlying crisp clustering partition (blue, red, yellow and green colours for
Clusters 1, 2, 3 and 4, respectively). The abbreviation of each region given in Table 3 was incorporated. It is
clear from Figure 1 that the resulting partition is highly interpretable from a spatial point of view, suggesting that
the climate and other geographical conditions could have influenced how people moved during the COVID-19
pandemic. Clusters 3 and 4 are formed mainly by communities located in the northern part of the country, with
prototypes Galicia (GA) and AS, respectively. On the other hand, Cluster 1 is constituted by communities close
to the Mediterranean Sea. The medoid of Cluster 1 is the community of Castile-La Mancha (CM). It is worth</p>
      <p>Community 1
Andalusia (AN) 0.71
Aragon (AR) 0.25</p>
      <p>Asturias (AS) 0.00
Balearic Islands (BI) 0.77
Basque Country (BC) 0.21
Canary Islands (CI) 0.00</p>
      <p>Cantabria (CB) 0.12
Castile and León (CL) 0.26
Castile-La Mancha (CM) 1.00</p>
      <p>Catalonia (CT) 0.62
Community of Madrid (MD) 0.09</p>
      <p>Extremadura (EX) 0.04</p>
      <p>Galicia (GA) 0.00
La Rioja (RI) 0.03</p>
      <p>Navarre (NC) 0.06</p>
      <p>Region of Murcia (MC) 0.23
Valencian Community (VC) 0.76
remarking that the only Mediterranean area not included in Cluster 1 from a crisp point of view, the region of
Murcia (MC), shows a relatively high membership in this group, 0.23. Finally, Cluster 2 is the more heterogeneous
cluster from a geographical perspective. This cluster is composed by a northern community, Aragon (AR), a central
community, Madrid (MD), a southern zone, Murcia (MC), and the insular area of the Canary Islands (CI), which
corresponds to the medoid. These islands are located far away from the Iberian Peninsula although they are shown
near Andalusia (AN) in Figure 1 for the sake of simplicity. It is important to emphasize that, although the elements
in Cluster 2 are not spatially connected, the strong similarity shown by the corresponding MTS ofsets the spatial
penalization term and provokes the formation of this group. Further investigations based on epidemiological data
will be carried out so as to explain the formation of this group.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Concluding remarks</title>
      <p>In this work, we have extended a fuzzy clustering model for multivariate time series based on QCD and PCA.
The former model employs a set of features obtained through the smoothed CCR-periodogram and performs the
traditional fuzzy -medoids clustering algorithm by considering the squared Euclidean distance in the principal
components space.</p>
      <p>The proposed extension is based on two tools. On the one hand, a system of weights is introduced in the objective
function so that more importance is given to the principal components having more discriminative ability in terms
of the clustering structure. On the other hand, a penalization term which permits to take into account spatial
information is considered. The resulting model uses the exponential distance, thus attaining robustness against
outlying series.</p>
      <p>We have showed that the consideration of the weighting system is totally advantageous in terms of clustering
efectiveness. The behaviour of the model according to the spatial penalization term was also analysed by means of
a brief example. Finally, the proposed technique was applied to cluster a dataset containing COVID-19 data where
the MTS come from diferent geographic locations, leading to illuminating conclusions.</p>
    </sec>
    <sec id="sec-6">
      <title>Appendix</title>
      <p>( , ,  ,  ) =
Here we derive the iterative solutions of the constrained minimization problem (6) via the Lagrangian multiplier method.</p>
      <p>First, consider the corresponding Lagrangian function taking the form:</p>
      <p>1, . . . ,  } and  stand for the Lagrange multipliers concerning the constraints of the membership degrees and
By fixing</p>
      <p>and setting equal to zero the partial derivatives of  with respect to  and  , for arbitrary  ∈ {1, . . . , } and
the weights, respectively.
 ∈ {1, . . . , }, we obtain that
is equivalent to

=1
which yields
 =
︂(   ︂) − 1 [︁</p>
      <p>1

︂(   ︂) 1− 1 ∑︁ [︁</p>
      <p>′=1
︂(   ︂) − 1</p>
      <p>1

=
[︃  [︂
∑︁
′=1
(11)
(12)
(13)
(14)
− 1[︁1
− exp (︀ −  ∑︁ 2(Ψ() −
− exp (︀ −  ∑︁ 2(Ψ() −
′=1 ′′∈

′′′
, 2 =</p>
      <p>′′′
On the other hand, the optimal weights  can be determined in a similar way. Now, we fix
 and set equal to zero the
partial derivatives of  with respect to  and the Lagrange multiplier  so that an equivalence occurs between
. Expression (14) gives the iterative solution for the
member
From the first equation in (15), we can write
By replacing (16) in the second equation of (15), we have</p>
      <p>︂[
′=1</p>
      <p>∑︁ ∑︁  (Ψ
()</p>
      <p>By plugging in (16) the expression for − / 2 in (18), we obtain the iterative solution for the weights, which takes the form:
 =
[︃

∑︁
′=1
∑︀
∑︀
=1
=1
∑︀
∑︀
=1
=1
 (Ψ −</p>
      <p>()
 (Ψ
()
′ −
,
where for the sake of simplicity we have used the notation (1, 2) = exp (︀
−
 ∑︀
=1</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>[2] [3] [6] [7] [11] [12] [13] [14] [15] [16] [17]</source>
          [1]
          <string-name>
            <given-names>T. W.</given-names>
            <surname>Liao</surname>
          </string-name>
          ,
          <article-title>Clustering of time series data-a survey</article-title>
          ,
          <source>Pattern recognition 38</source>
          (
          <year>2005</year>
          )
          <fpage>1857</fpage>
          -
          <lpage>1874</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>S.</given-names>
            <surname>Aghabozorgi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Shirkhorshidi</surname>
          </string-name>
          , T. Y. Wah,
          <article-title>Time-series clustering-a decade review</article-title>
          ,
          <source>Information Systems</source>
          <volume>53</volume>
          (
          <year>2015</year>
          )
          <fpage>16</fpage>
          -
          <lpage>38</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>B.</given-names>
            <surname>Lafuente-Rego</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Vilar</surname>
          </string-name>
          ,
          <article-title>Clustering of time series using quantile autocovariances</article-title>
          ,
          <source>Advances in Data Analysis and classification 10</source>
          (
          <year>2016</year>
          )
          <fpage>391</fpage>
          -
          <lpage>415</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Vilar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Lafuente-Rego</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          <article-title>D'Urso, Quantile autocovariances: a powerful tool for hard and soft partitional clustering of time series</article-title>
          ,
          <source>Fuzzy Sets and Systems</source>
          <volume>340</volume>
          [5]
          <string-name>
            <surname>P. D'Urso</surname>
            ,
            <given-names>E. A.</given-names>
          </string-name>
          <string-name>
            <surname>Maharaj</surname>
          </string-name>
          ,
          <article-title>Autocorrelation-based fuzzy clustering of time series</article-title>
          ,
          <source>Fuzzy Sets and Systems</source>
          <volume>160</volume>
          (
          <year>2009</year>
          )
          <fpage>3565</fpage>
          -
          <lpage>3589</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>P. D'Urso</surname>
            ,
            <given-names>E. A.</given-names>
          </string-name>
          <string-name>
            <surname>Maharaj</surname>
          </string-name>
          ,
          <article-title>Wavelets-based clustering of multivariate time series</article-title>
          ,
          <source>Fuzzy Sets and Systems</source>
          <volume>193</volume>
          (
          <year>2012</year>
          )
          <fpage>33</fpage>
          -
          <lpage>61</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>