<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Methods of Acceleration of Term Correlation Matrix Calculation in the Island Text Clustering Method</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Faculty of Applied Mathematics, NTUU "Igor Sikorsky Kyiv Polytechnic Institute"</institution>
          ,
          <addr-line>Kyiv, Peremohy ave., 37</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <fpage>0000</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>This paper considers the task of accelerating the text clustering by the island clustering method. Based on the consideration of the basic implementation of this method, the step of calculating the term correlation matrix is highlighted for further acceleration. This step has quadratic complexity from the number of terms since it considers each combination of them in pairs. In the framework of this study, it is proposed to use parallel computing and memoization as methods to accelerate the calculation of the term correlation matrix. The experiment was carried out using a specially developed deterministic algorithm for generating test data, which is also described. Based on the results of the data obtained analysis, it is shown that the use of the proposed methods accelerates the calculation of the matrix by 8-10 times on a virtual machine with 4 physical cores and 8 logical ones.</p>
      </abstract>
      <kwd-group>
        <kwd>text clustering</kwd>
        <kwd>memoization</kwd>
        <kwd>parallelism</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The clustering of natural language textual data is widely used both as one of the
variants of automatic systematization and as a proper analysis tool [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Clustering of a
natural-language text corpus allows to: understand its structure; highlight the most
atypical texts in this specific corpus that will not belong to any cluster; reduce the size
of the text corpus before its further processing or storing, discarding the most similar
texts, and more [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        Thus, due to the constant increase in the volume of text corpora and complication
of their structure, formulation of new and improvement of existing methods of
automatic (unsupervised) clustering of text documents becomes an urgent task [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        One of the existing effective methods of text clustering, which provides clarity to
the process of obtaining clusters for humans, is the method of island clustering [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        The basis of this method is the implementation of a two-step procedure: the first
stage provides the clustering of terms (hereinafter, “term” is used in general
information retrieval meaning – any word/expression) that documents consist of; the
second stage – construction of document clusters, based on the clusters of terms
received in the first stage. Below is a more detailed list of steps of this method [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]:
      </p>
      <p>Stage I. Clustering of terms:
1. pre-processing of texts from the input collection of documents: removal of stop
words, lemmatization, etc.;
2. selection from the texts the set of terms which they consist of;
3. if necessary, filtering the resulting set of terms (for example, in situations where
the initial centroids of the clusters are known or the resulting set is too large);
4. construction of a graph of terms correlation;
5. pre-processing of the graph and obtaining its approximation (optional);
6. clustering of the obtained approximation of the graph;</p>
    </sec>
    <sec id="sec-2">
      <title>Stage II. Construction of document clusters:</title>
      <p>7. splitting documents into clusters based on received term clusters.</p>
      <p>One of the biggest cost steps is to build a graph of terms correlation. This procedure
lays in calculation of the graph distances matrix (for this method of clustering the
measure of terms correlation serves as a distance, so the names «distance matrix» and
«correlation matrix» are identical for this research), which requires a pairwise
consideration of all terms and leads to the quadratic complexity of the calculation process.
Theoretically, the original clustering procedure for the obtained approximation of the
graph is also quadratic (since it considers all edges of the obtained graph), but when
applying the effective procedures of this very approximation, the complexity of this
step is almost linear in practice.</p>
      <p>Regardless of the mentioned above, as will be shown in Section 2 of this paper,
none of the papers dealing with the island clustering method considers the
acceleration of the term correlation matrix calculation. From this, we can conclude that the
development of methods to accelerate the calculation of the term correlation matrix is
a pressing issue, the successful solution of which will increase the speed of
implementation of island clustering of texts as a whole.</p>
      <p>Thus, the main objective of the research is to accelerate the calculation of the term
correlation matrix in the method of island clustering of texts. The object of the study
is the process of automatic clustering of natural language text data. The subject of the
study is the methods of calculation acceleration and applying them in the context of
the island text clustering method. According to the stated objective, the following
tasks were set and solved:
1. research of existing methods to accelerate the calculations;
2. reasonable choice of methods to accelerate the calculations for their use in the
method of island clustering of texts when calculating the term correlation matrix;
3. software implementation of the calculation of the term correlation matrix using the
selected methods of accelerating the calculations;
4. analysis of the efficiency of the proposed means by the criterion of the speed of
calculation.</p>
      <sec id="sec-2-1">
        <title>Related Works</title>
        <p>
          The island clustering method was first described in [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. One objective of this study
was to develop a method, the computational complexity of which has no more than a
log-linear dependence on the number of texts. The island clustering method meets this
requirement since the calculation time of the correlation matrix linearly depends on
the number of documents. In the section devoted to the experiment, it was described
that clustering took about 9 minutes on a pre-indexed subset (consisting of 17545
documents, longer than 100 characters) of the standard Reuters-21578 collection.
        </p>
        <p>
          [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] is devoted to the improvement of the island clustering method, describing the
method of essential clustering. The new method did not set any new requirements for
clustering performance, leaving the requirement only for the dependence of
complexity on the number of texts. Since the main objective of this work was to increase
accuracy, there is no information about clustering performance in the work.
        </p>
        <p>However, in practice, more important is the dependence of performance on the
volume of the text corpus, which is more natural to consider as the number of terms.
Described related papers did not describe any techniques for improving this
calculation.
3</p>
      </sec>
      <sec id="sec-2-2">
        <title>An Existing Algorithm for Calculating the Term Correlation</title>
      </sec>
      <sec id="sec-2-3">
        <title>Matrix in the Method of Island Clustering of Texts</title>
        <p>
          The term correlation matrix is a symmetric matrix containing, as elements, the
correlation value between the corresponding pairs of terms [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. Let be
the number of unique terms (the dimension of the matrix), then to fill in the
correlation matrix, it is necessary to calculate the correlation value between
        </p>
        <p>
          pairs of terms. The correlation value between the terms and is determined as
follows. Let be the total number of terms in all documents, – be the number of
terms in the documents that meet the term , – the total number of term
occurrences in all documents, and – the number of term occurrences in documents
containing the term . Then the probability that in documents containing term ,
is found or more of terms , can be used as a basis for calculating the measure of
the correlation of terms and . This probability can be calculated by the formula of
the binomial distribution (1) [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. The lower the probability obtained, the more terms
are correlated with each other.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Since the probability of is not symmetric, in practice as a measure of the correlation of terms. (1) is used</title>
      <p>As a result, we can build matrix</p>
      <p>based on the obtained values and this matrix will
look like
. Such a matrix can be effectively
stored in memory in the form of a vector with length</p>
      <p>
        Thus, the basic algorithm for calculating the term correlation matrix
the following steps [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]:
.
consists of
1. in the cycle with counter from 1 to
2. in the cycle with counter from
3. find , and calculate ;
4. assign and .
      </p>
      <p>to
with step 1 perform pp.2-4;
with step 1 perform pp.3-4;
This algorithm is a test of the statistical hypothesis of pairwise independence of the
presence of terms in documents. The consequence of this is the fact that for randomly
generated texts this algorithm will not find any significant relationship between the
terms, and therefore, no clusters will be received at the output.</p>
      <p>
        In practice, matrix also can be additionally filtered to remove insignificant
relations between terms (e.g. it can be done by use absolute threshold for term correlation
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]) – p.5 in the island clustering algorithm. Such a procedure will increase the quality
of obtained term clusters.
4
4.1
      </p>
      <sec id="sec-3-1">
        <title>Methods of Accelerating the Calculations That Can Be Used in Calculating the Term Correlation Matrix</title>
        <sec id="sec-3-1-1">
          <title>Parallel Implementation of Cycles of the Correlation Matrix Calculation</title>
          <p>In the algorithm of the term correlation matrix calculation, it is obvious that each of
the iterations is independent of the others. Thus, instead of
sequentially executing iterations of cycles, their parallel execution is possible.</p>
          <p>Let be the time of steps 3-4 of the algorithm (calculating the probabilities,
finding the maximum and assigning it to the elements of the matrix). Then – the
duration of the sequential implementation of the algorithm will be
.</p>
          <p>
            Let all the iterations be divided into parts of approximately equal size, which will
be processed by independent threads in parallel. Then – the minimum
duration of running a parallel algorithm implementation, which is
can theoretically be achieved. [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ]
          </p>
          <p>Thus, comparing and we can conclude that it is theoretically
possible to achieve an acceleration of the matrix calculation by times (with a theoretical
maximum in at ). However, the use of
parallel implementation of the algorithm in practice requires additional data sharing
and flow management costs, so that the practical acceleration will be less than .
4.2</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>Memoization of Interim Results</title>
          <p>
            Memoization is one of the methods of optimizing calculations, which consists of
storing the results of a function execution in a lookup table to prevent repeated
calculations [
            <xref ref-type="bibr" rid="ref7">7</xref>
            ]. When using this technique, before performing the memoized function, it is
checked whether this function has already been performed for the passed parameter
set. If the lookup table already saves the result of this function execution for a given
set of parameters – it is used without performing calculations; otherwise, the
calculation is performed, and the given result is recorded in the table [
            <xref ref-type="bibr" rid="ref7">7</xref>
            ]. In this way,
memoization differs from the pure use of lookup tables in a way that when used, the
lookup table is not pre-calculated but filled out if relevant.
          </p>
          <p>In practice, memoization can only be applied to pure features that do not create
side effects (such as file entries, database reads, etc.). Also, memoization cannot be
applied if the calculation results have only a certain validity period, after which they
must be re-calculated (in this case, more sophisticated techniques such as caching
should be used).</p>
          <p>The additional cost of memoization is only the need for additional memory to store
the lookup table (its size depends on the set of valid values of the memoized function
parameters).</p>
          <p>In the case of calculating the term correlation matrix, memoization can be applied
at the third point of the algorithm when calculating the parameters and from
formula (1) of the calculation of . This will only require of additional
memory, and if you use hash tables as lookup tables for memoization, the difficulty of
getting the calculated result will be only .
5
5.1</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Experiment and Results</title>
        <sec id="sec-3-2-1">
          <title>Test Data Generation Algorithm</title>
          <p>To test the proposed methods of accelerating the calculation of the term correlation
matrix, in the framework of this research an own algorithm for generating a textual
data corpus was developed.</p>
          <p>The developed algorithm has the following properties:
1. accepts only from the input parameters – the number of unique terms in the
corpus texts;
2. is fully deterministic – the same result will be obtained for startups with the same
input parameter value;
3. the generated corpus of texts contains pairs of terms with different degrees of
correlation – fully correlated, partially, not at all correlated.</p>
          <p>Due to the deterministic nature of this algorithm, it is possible to compare the
obtained experimental data (the rate of calculation of the term correlation matrix) for the
same matrix dimension with the use of different combinations of the proposed
methods.</p>
          <p>The developed algorithm for generating a textual data corpus (with a term
correlation matrix of a given dimension) consists of the following steps:
1. based on the input value, calculate the number of texts in the corpus being
generated, as well as the boundaries of the range of texts in which the next term
will occur. Formulas (2) and (3) are used to calculate these two parameters;
2. for each from to :
(a) obtain the row value of the next term by converting to hexadecimal;
(b) calculate the indexes of the texts in which the next term will occur according to
formula (4). In case if any index of the resulting set is negative or exceeds the
index of the last text – this index is reduced to valid values;
(c) calculate the number of occurrences of the term to the texts found in p.b:
(i) the total number of occurrences of the term to the corpus is calculated by
the formula (5);
(ii) the number of occurrences of term to the central text of the range is ;
(iii) the rest of the occurrences are distributed equally among the other
texts of the range;
(d) the term according to the number of occurrences calculated is added to the
texts in the range.
5.2</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>Features of Implementation of the Software Used for Testing</title>
          <p>The software based on C# and the .NET Core 2.2 platform has been developed for
experimental studies of the proposed methods to accelerate the calculation of the term
correlation matrix. The source code for the developed software is fully open and
available at https://github.com/yakivyusin/IslandClusteringAcceleration.</p>
          <p>
            The measurement of the performance of the calculation of the term correlation
matrix at different dimensions of the matrix and with different combinations of
acceleration methods used was performed using the BenchmarkDotNet library (v.0.11.5) [
            <xref ref-type="bibr" rid="ref8">8</xref>
            ].
          </p>
          <p>
            The launch of the developed software was performed in Google Cloud Platform
virtual machine [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ]: machine type – n1-highcpu-8 [
            <xref ref-type="bibr" rid="ref10">10</xref>
            ], the operating system installed
– Windows Server 2012 R2 Datacenter, .NET runtime – .NET Core 2.2.6 (CoreCLR
4.6.27817.03, CoreFX 4.6.27818.02), 64bit RyuJIT.
(2)
(3)
(4)
(5)
5.3
          </p>
        </sec>
        <sec id="sec-3-2-3">
          <title>Results Obtained</title>
          <p>Testing the speed of calculating the term correlation matrix was performed for 8
dimensions: 49, 169, 245, 371, 441, 569, 659, and 785 terms. According to the
developed algorithm for generating a textual data corpus, this set of dimensions provided
the generation of corpora with sizes from 2 to 5 texts.</p>
          <p>For each dimension of the test set, the term correlation matrix was calculated in
twelve different ways: with sequential execution of cycles, with parallel execution of
an external cycle, with parallel execution of both cycles, with memoization of
calculation , with memoization of calculation , with memoization of calculation of both
parameters.</p>
          <p>The test results are shown in Table 1, in which the following notations are used to
denote a matrix calculation variant:
 S – sequential execution of cycles;
 P – parallel execution of the external cycle only;
 P2 – parallel execution of both cycles;
 - – memoization of parameter calculation is absent (at the second position in the
variant code corresponds to the parameter at the third position – parameter );
 + – memoization of parameter calculation is present (at the second position in the
variant code corresponds to the parameter , at the third position – parameter ).
The present results have high accuracy with a low variance – the ratio of the standard
deviation to mean is in the range from 0.01% to 6.7% (standard error is also within
the range from 0.02% to 0.92%). All measurements (except one) where this ratio
exceeds the mean value of the dataset are related to parallel implementations. This
may be caused by the fact that parallel implementations are more sensitive to random
changes in the workload of the operating system with multitasking, while sequential
implementations just utilize one core and work on it. Also, the process of garbage
collection might influence the variance of results without memoization, but on the
experimental dataset this process is almost deterministic due to dataset volume.</p>
          <p>
            The charts of the obtained results for matrices with dimensions of 659 and 785
terms are shown in Fig. 1 and Fig. 2 respectively. Interactive charts for all test data
are created using the Highcharts library [
            <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
            ] and are
https://yakivyusin.github.io/IslandClusteringAcceleration/plots.html.
available
at
          </p>
          <p>From the obtained results, it can be concluded that the use of all the proposed
methods to accelerate the calculation of the term correlation matrix, increases the
performance by 8-10 times in total, provided there are 8 virtual cores (depending on the
dimension of the matrix and the number of texts; shown in Fig. 3), compared to the
«naive» implementation of calculations.
As we can see, the main objective of the present paper is achieved –all proposed
methods improve correlation matrix calculation performance in any given
combination of used methods. The combination of all proposed methods (enabled parallelism
and memoization of both calculation parameters) allows getting acceleration more
than 8 times on machine with 4 physical cores / 8 virtual cores. Also, from the results
obtained, the following dependencies can be emphasized:
1. clean use of parallelism leads to acceleration improvement more than 4 times. Such
a result is because the test virtual machine has only 4 physical cores with 8 virtual
cores;
2. memoization of parameter has a much greater effect on acceleration than
memoization of parameter (average 2 times against average 1.05). We relate it
to the fact that the calculation of is a more complex process and requires read
collection of all corpus terms;
3. variants of parallelism implementation with one cycle and two almost do not differ
from each other (P2 variant on average is faster on 4-6%). The internal features of
.NET Core parallelism implementation might influence this result and also that
even in case of used one parallel cycle the number of chunks exceeds free cores
count.
6</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>Conclusion</title>
        <p>In the present paper, we have described methods of acceleration of term correlation
matrix calculation in the island text clustering method. Proposed methods are based
on two techniques: parallel execution of iterations and memoization of interim results
(coefficients in correlation calculation formula). According to experimental results,
the performance of calculation has been improved more than 8 times on 8 virtual core
machine within the range of different text corpora (test data was generated by a
described algorithm which was developed especially for this paper). The obtained result
exceeds a theoretical limit of parallel execution (which is equal to the number of
cores) due to memoization using.</p>
        <p>Future work on this subject may be in the following areas:
1. experimental testing on different machines (with different number of cores) to
better determine the dependence between improvement ratio and the degree of
parallelism;
2. porting an implementation to a programming language without a garbage collector
to reduce the number of factors that affect the results.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Reddy</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kinnicutt</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Text Document Clustering: The Application of Cluster Analysis to Textual Document</article-title>
          . In: Arabnia H.,
          <string-name>
            <surname>Deligiannidis</surname>
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            <given-names>M</given-names>
          </string-name>
          . (eds.)
          <source>INTERNATIONAL CONFERENCE ON COMPUTATIONAL SCIENCE AND COMPUTATIONAL INTELLIGENCE</source>
          , pp.
          <fpage>1174</fpage>
          -
          <lpage>1179</lpage>
          . IEEE Computer Society, USA (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Allahyari</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pouriyeh</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Assefi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Safaei</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Trippe</surname>
            ,
            <given-names>E.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gutierrez</surname>
            ,
            <given-names>J.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kochut</surname>
            ,
            <given-names>K.J.:</given-names>
          </string-name>
          <article-title>A Brief Survey of Text Mining: Classification, Clustering and Extraction Techniques</article-title>
          . ArXiv, abs/1707.02919 (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Berry</surname>
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Survey of text mining: clustering, classification, and retrieval</article-title>
          . Springer, New York (
          <year>2002</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Shmulevich</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiselev</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pivovarov</surname>
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>The method of clustering texts, taking into account the joint occurrence of key terms, and its application to the analysis of the thematic structure of the news flow, as well as its dynamics</article-title>
          .
          <source>In: INTERNET MATHEMATICS</source>
          , pp.
          <fpage>412</fpage>
          -
          <lpage>435</lpage>
          . Yandex, Moscow (
          <year>2005</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Shmulevich</surname>
            <given-names>M.:</given-names>
          </string-name>
          <article-title>The method of automatic clustering of texts, based on the extraction of object names from the texts and the subsequent construction of graphs for the joint occurrence of key terms</article-title>
          .
          <source>Ph.D. Thesis</source>
          , MFTI, Moscow, Russia (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Amdahl</surname>
            ,
            <given-names>G.M.:</given-names>
          </string-name>
          <article-title>Validity of the single-processor approach to achieving large scale computing capabilities</article-title>
          .
          <source>In: AFIPS CONFERENCE PROCEEDINGS</source>
          vol.
          <volume>30</volume>
          (
          <string-name>
            <surname>Atlantic City</surname>
            ,
            <given-names>N.J.</given-names>
          </string-name>
          ,
          <source>Apr. 18-20)</source>
          , pp.
          <fpage>483</fpage>
          -
          <lpage>485</lpage>
          . AFIPS Press, Reston,
          <string-name>
            <surname>Va.</surname>
          </string-name>
          (
          <year>1967</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Michie</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Memo Functions and Machine Learning</article-title>
          .
          <source>Nature</source>
          <volume>218</volume>
          ,
          <fpage>19</fpage>
          -
          <lpage>22</lpage>
          (
          <year>1968</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8. Overview | BenchmarkDotNet, https://benchmarkdotnet.org/articles/overview.html,
          <source>last accessed</source>
          <year>2019</year>
          /10/27.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9. Cloud Computing Services | Google Cloud, https://cloud.google.com/,
          <source>last accessed</source>
          <year>2019</year>
          /10/27.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10. Machine types | Compute Engine Documentation | Google Cloud, https://cloud.google.com/compute/docs/machine-types,
          <source>last accessed</source>
          <year>2019</year>
          /10/27.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <article-title>Interactive JavaScript charts for your webpage</article-title>
          | Highcharts, https://www.highcharts.com/,
          <source>last accessed</source>
          <year>2019</year>
          /10/27.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12. Box plot | Highcharts, https://www.highcharts.com/demo/box-plot,
          <source>last accessed</source>
          <year>2019</year>
          /10/27.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>