<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Methodology for Preprocessing Semi-Structured Data for Making Managerial Decisions in the Healthcare* 1</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>E. Makarova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>D. Lagerev</string-name>
          <email>LagerevDG@mail.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Bryansk State Technical University</institution>
          ,
          <addr-line>Bryansk</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes the process of supporting management decisionmaking in healthcare based on data mining. The authors described various problems and specifics of data in medical information systems, leading to the complexity of their analysis and integration, such as: the presence of a large number of specific abbreviations, errors in the data and their poor structure. The paper demonstrates an approach to the search and further disclosure of abbreviations in texts, built on a combination of machine and human processing. A method for extracting features from semi-structured fields using an expert in the subject area and using various visualizations is proposed. The proposed abbreviation search and disclosure methods, based on a hybrid approach combining the strengths of processing with the help of a machine and an expert, can increase the number of abbreviations found automatically and significantly reduce the time spent by experts on processing the remaining reductions. In addition, the method for automated feature extraction during integration can significantly increase the amount of useful input data, while reducing the time of the expert.</p>
      </abstract>
      <kwd-group>
        <kwd>Natural Languages Processing</kwd>
        <kwd>Data Integration</kwd>
        <kwd>Healthcare</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The digitalization of Russian medicine poses new challenges for managers and
engineers - implementation, security and support issues, large data storage and processing
systems. But the collection of this data in digital format, in turn, opens up new
opportunities for researchers and healthcare managers through the use of data analysis
technologies.</p>
      <p>Over the years of informatization of various cities and regions of the Russian
Federation, in medical information systems (hereinafter referred to as MIS), more and more
data has been accumulating on various aspects of the work of medical organizations
from medical histories and prescriptions of specific patients to various aspects related
to providing medical institutions with necessary medicines and supplies materials.</p>
      <p>The reported study was funded by RFBR, project № 20-04-60185.</p>
      <p>
        Thus, the improvement of analysis technologies and the filling of databases with
medical data makes it relevant to use data mining for tasks related to healthcare
management [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], such as: planning material and human resources in healthcare, forecasting
statistical indicators, optimal and timely provision of resources, tracing outbreaks and
spreading of diseases.
      </p>
      <p>These processes explain the relevance of solving the problem of creating an
automated system to support managerial decision-making in healthcare and solving side
problems associated with the implementation of this system.</p>
      <p>The emphasis is on the need for both strategic (development of the healthcare sector
in the region) and operational (decisions at the level of a medical organization, response
to outbreaks of diseases, etc.)
2</p>
    </sec>
    <sec id="sec-2">
      <title>Management decision-making in healthcare</title>
      <p>
        The forecasting task is very relevant for healthcare: it is necessary to predict the
incidence rate, assess the required resources to maintain the effective operation of the
system, etc. Researchers devote much attention to the problems of disease prognosis in
specific regions [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        Forecasting is a task that can be solved in many ways - from classical statistical
methods to models based on machine learning technology. Recently, the neural network
approach has become widespread. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. For example, recent research show that the
accuracy of predicting many diseases using CNN is greater than, for example, the KNN
and NB algorithms. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        In addition, a research that used deep learning to process textual medical data
(various models of embedded representations of words were used) showed an increase in
forecasting accuracy [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. To predict a number of diseases, classic time series are best
used, using only numerical values (for example, injuries, SARS, etc.). However, in
which groups of diseases it is necessary to apply more complex approaches to the
analysis. For example, when predicting malignant neoplasms, it is necessary to understand
not only the general characteristics, but also the number of patients at different stages
of cancer. In a research on the prediction of breast cancer using machine learning,
patients were divided into cohorts depending on the stage of the disease and other
parameters, which allowed better identification of factors contributing to patient survival [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>To make a decision on the distribution of resources between various medical
institutions, the regional health department needs to make a forecast about the development
of certain diseases and act according to long-term planning. The process of creating
such a forecast can simplify the development of an automated system to support
management decisions. (Fig 1)</p>
      <p>The integration and setting up of this process requires the investment of certain
human resources, however, given the need to regularly make such management decisions
and constant updating of data in the regional information system (RIS), in the long term,
these labor costs will be justified. The general scheme of this process is shown in
Figure 2.</p>
      <p>Methodology for Preprocessing Semi-Structured Data for Making Managerial Decisions… 3
where: R - available resources (budget for all medical institutions);</p>
      <p>M - budget for a specific medical organization (MO) for various articles (equipment,
maintenance of an inpatient hospital, procurement of medicines, rates of health
workers, etc.);</p>
      <p>D - effectiveness of the development of this budget;
Z - requests for resources (the current need for various MO in them);
I - information available for analysis and forecasting resource requirements.</p>
      <p>
        The most time-consuming step in data mining is still the process of collecting,
cleaning and pre-processing data before analysis. According to various researchers, this
process takes from 60% to 80% of the time [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        In previous works of the authors [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] concerning the process of collecting and
processing semi-structured data, much attention was paid to the “hybrid” approach in the
field of developing data analysis systems. In this approach, human expertise is used in
conjunction with automatic analysis methods, which allows, on the one hand, to
improve the quality of the system on tasks that cannot be solved without human
intervention, and on the other, to relieve the expert from solving typical, routine tasks. This was
achieved using various methodologies for pre-processing and data visualization, which
helped the expert make faster decisions about the inclusion / exclusion of a particular
data source [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>Data collection is only the first step in the data preparation process for use in
ensembles of data mining models. In addition to the general problems for all subject areas
arising at the stage of data preprocessing, when constructing analytical models for the
analysis of biomedical data, researchers and developers encounter a number of
problems specific to the described data, which will be discussed in more detail in the next
section.</p>
    </sec>
    <sec id="sec-3">
      <title>Pre-processing of medical data for analysis</title>
      <p>also an important part of solving the integration problem. For example, the same
indicator in one database is called “growth”, and in another - “body length”.</p>
      <p>The problem of poorly structured data is that the ICD-10 classification does not
contain details on the diagnosis, doctors must fill out this information on their own
according to pre-created forms in the system or in free form.</p>
      <p>Also in the field of free entry is usually indicated: degree (stage) of the disease, form,
prescribed medications and other. For example, regarding the classification of
malignant neoplasms according to ICD-10, the classifier does not reflect the stage of the
disease. Usually it is indicated in natural language in another field. However, the
availability of these data could constitute a better prognostic model for the stages of the
disease. So, for example, for patients with a malignant formation of the first stage, there
is a significant risk under certain circumstances, an article by patients with a second
stage of cancer, second - third are at the second - go to the third, etc.</p>
      <p>Here are a few examples of such an ICD-10 diagnosis uncertainty. For example, the
ICD-10 code “S82.6” (fracture of lateral malleolus) should have at least an explanation
of the right leg or left leg (which can be expressed in free form as «левая», «слева»,
«левый», «лев», «л.», etc.), but also an indication of whether the fracture is closed or
not, complete or incomplete.</p>
      <p>For example, Table 1 presents some examples of how certain important terms are
indicated. In addition, in some cases it is not clear that this is an abbreviation, a specific
term or word spelled out with errors.</p>
      <p>The database also contains specific grammatical constructions that make it difficult
to extract features. For example, when identifying symptoms by standard methods, the
phrase from the patient’s history “did not have hepatitis, tuberculosis”, information
about hepatitis and tuberculosis without mentioning “not sick” could fall into the
patient’s model, which would worsen the quality of the prediction models.</p>
    </sec>
    <sec id="sec-4">
      <title>A hybrid approach to finding and revealing abbreviations, incorrect spelling of words</title>
      <p>
        Based on the available data, various approaches to finding abbreviations have been
tried, from a standard approach based on regular expressions and a dictionary of
commonly used abbreviations. Since many abbreviations in the sample are specific, the
combined methods based on a combination of heuristic, vocabulary, and statistical
approaches gave the greatest increase in accuracy. A detailed description of this approach
and the results of its use are described in a previous work of the authors [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>However, it is not yet possible to reveal specific abbreviations absolutely precisely
in a fully automated mode. One way or another, when solving this problem, you will
have to turn to knowledge by a competent person. By analyzing the context of
abbreviations, it is possible to significantly reduce the degree of expert intervention, if we
train the word embedding model on the available data and predict the overall meaning
of abbreviations depending on the context. There is enough data when training the
model to establish syntagmatic and paradigmatic relationships.</p>
      <p>The semantic similarity between linguistic units is calculated as the distance between
vectors. In studies on distributive semantics, the most often used cosine measure, which
is calculated by the formula</p>
      <p>where a, b are the number of elements in the “стд” token and the “стадия” token,
respectively;
c is the number of common elements in the "std" and "stage" tokens.</p>
      <p>In the standard setting, k is taken to be large 0.5 for words with a non-matching
beginning and ending (“стд” and “стадия”) and 0.35 for words with a matching
beginning and ending and containing a hyphen (“бр-ть” and “беременность”). In previous
where A and B are the vectors of words, the distance between which is calculated.</p>
      <p>
        In this study, vectorization was implemented using the Bag of Words method [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ],
but other methods are also possible. For this sample, a sufficiently high threshold is a
value of 0.7
      </p>
      <p>In addition, in order to automatically recognize the word appropriate context
abbreviation, three conditions are necessary:
the abbreviation coincides with the beginning of a semantically close word;
the found word is not an abbreviation and is present in the dictionary of used words
of the Russian language;</p>
      <p>the found word is the only one satisfying the first and second conditions in the range
of semantic proximity from 0.7 to 0.99.</p>
      <p>
        For words that do not coincide with the abbreviation but are close in terms of the
cosine measure, the Tanimoto coefficient is used with a match value greater than or
equal to 0.5 [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. For example, we calculate the syntactic proximity of the abbreviation
"стд" and the word "стадия":
sim =
n
 (Ai  Bi )
i=1
n n
 (Ai )2   (Bi )2
i=1 i=1
      </p>
      <p>,
sim ≥ 0,7
k =</p>
      <p>c
a + b − c ,
(2)
(3)
(4)</p>
      <p>Methodology for Preprocessing Semi-Structured Data for Making Managerial Decisions… 7
work of the authors was presented word embedding visualization technique for these
tasks analytics, which was also implemented in this case.</p>
      <p>Of the available sample of depersonalization records from integrated electronic
medical records (IEMR) of approximately 1.4 million residents of the Bryansk region,
a sample of 60,000 records was created, balanced by diagnosis and length of text
description, of which 3,000 records were similarly selected. Each of abbreviations was
manually specified to verify the developed methodology. Next, a comparison was made
of the results of a fully manual approach, a fully automated and the hybrid approach
described above. Results are presented in table 2.
The interface of a specialist in the subject area for marking up data when implementing
manual (expert) control is presented in Figure 3. The proposed approach checks not
only all cases that do not go beyond the boundaries of automatic marking, but also
510% of instances automatically classified by the system in order to verify the correct
operation of the algorithms and their settings, if necessary. The large spread in labor
costs for the expert when marking up the data is explained by the number of checked
examples and the severity of the thresholds for automatic marking depends enough. In
the described experiment, the choice of the percentage of data considered by experts
and trusted by the system depends on the accuracy requirements and is limited by
available resources.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Visual interface for data extraction</title>
      <p>The solution to the problem of bringing various abbreviations with a single value is
only one way to reduce the number and improve the quality of features that are input to
analytical models. Using a large number of features (some of which will be duplicate,
some will be useless) is irrational. There are problems such as overfitting, an increase
in processing time and the presence of "noise" and "garbage data".</p>
      <p>In addition to selecting features that will go into the model, they often need to be
additionally extracted from poorly structured data. In this case, we also use a hybrid
approach (Fig. 4).</p>
      <p>In this approach, word2vec models are used (to select contextually close tokens) and
a visual editor (Fig. 5). The main visualization metric here is coverage of features. It is
calculated on a weighted limited sample. In this case, 3000 records to quickly recount
the results and provide interactive visualizations.</p>
      <p>As the main extraction method used rules based on the principles of regular
expressions. A user who is an expert in this field, but does not understand regular expressions
and word processing, is provided with a visual editor of these expressions and
instructions for use. After re-calculation, several random entries are also presented for manual
evaluation of the created rules. The results of these evaluations are saved for automatic
validation when the rules change.</p>
      <p>Table 3 presents the results of an experiment conducted on a sample of oncological
diagnoses, where important metadata, such as the "stage", were described in the free
entry fields.
To effectively manage healthcare resources, it is necessary to collect, save and analyze
data received from all regions of the Russian Federation. At the moment, one of the
main methods of data integration in the USISH project is integration through
documents. Since documents are a poorly structured source of information, with such
integration there are problems associated with the interpretation of various text data, as well
as problems of their quality: the presence of specific abbreviations, errors, difficulties
in extracting various features, etc. The presence of a large number of noise, duplicates,
and incorrect features degrades the quality of data analysis models.</p>
      <p>The proposed abbreviation search and disclosure methods, based on a hybrid
approach combining the strengths of processing with the help of a machine and an expert,
can increase the number of abbreviations found automatically by 21%, as well as detect
in automated mode up to 55% of cases (with a probability of correctness higher 70%)
and significantly reduce the time spent by experts on processing the remaining
reductions.</p>
      <p>In addition, the method for automated feature extraction during integration can
significantly increase the amount of useful input data, while reducing the time of the
expert.</p>
      <p>Using a hybrid approach to preprocessing poorly structured data increases the
efficiency of managerial decisions in the field of healthcare by increasing the reliability of
data mining models and reducing the time spent by experts on their creation and
support. A further line of work in this area will be directed to the development of methods
for the semi-automatic selection of features for analytical models.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Zakharova</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lagerev</surname>
            ,
            <given-names>D. G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Podvesovskii</surname>
            ,
            <given-names>A. G.</given-names>
          </string-name>
          <article-title>: Multi-level Model for Structuring Heterogeneous Biomedical Data in the Tasks of Socially Significant Diseases Risk Evaluation</article-title>
          .
          <source>In: 3rd Conference on Creativity in Intelligent Technologies and Data Science, CIT and DS</source>
          <year>2019</year>
          , pp.
          <fpage>461</fpage>
          -
          <lpage>473</lpage>
          ,
          <string-name>
            <surname>Volgograd</surname>
          </string-name>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Choporov</surname>
            ,
            <given-names>O.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zolotuhin</surname>
            ,
            <given-names>O.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bolgov</surname>
            ,
            <given-names>S.V.</given-names>
          </string-name>
          :
          <article-title>Algoritmizaciya intellektual'nogo analiza dannyh o rasprostranennosti zabolevanij na regional'nom i municipal'nom urovnyah</article-title>
          .
          <source>In: Modelirovanie, optimizaciya i informacionnye tekhnologii № 2</source>
          (
          <issue>9</issue>
          ), (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Lazarenko</surname>
            ,
            <given-names>V.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Antonov</surname>
            ,
            <given-names>A.E.</given-names>
          </string-name>
          :
          <article-title>Diagnostika i prognozirovanie veroyatnosti vozniknoveniya holecistita na osnove nejrosetevogo analiza faktorov riska</article-title>
          .
          <source>In: Issledovaniya i praktika v medicine. №4</source>
          (
          <issue>4</issue>
          ), pp.
          <fpage>67</fpage>
          -
          <lpage>72</lpage>
          . (
          <year>2017</year>
          ) https://doi.org/10.17709/
          <fpage>2409</fpage>
          -2231-2017-4
          <issue>-4</issue>
          -7
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Dahiwade</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patle</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meshram</surname>
          </string-name>
          , E.:
          <article-title>Designing Disease Prediction Model Using Machine Learning Approach</article-title>
          .
          <source>In: 2019 3rd International Conference on Computing Methodologies and Communication (ICCMC)</source>
          , pp.
          <fpage>1211</fpage>
          -
          <lpage>1215</lpage>
          , Erode, India (
          <year>2019</year>
          ) https://doi.org/10.1109/ICCMC.
          <year>2019</year>
          .8819782
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Christensen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frandsen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Glazier</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Humpherys</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <article-title>: Machine Learning Methods for Disease Prediction with Claims Data</article-title>
          .
          <source>In: 2018 IEEE International Conference on Healthcare Informatics (ICHI)</source>
          , pp.
          <fpage>467</fpage>
          -
          <lpage>474</lpage>
          , New York, NY (
          <year>2018</year>
          ). https://doi.org/10.1109/ICHI.
          <year>2018</year>
          .00108
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Shukla</surname>
            ,
            <given-names>N</given-names>
          </string-name>
          , Hagenbuchner,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Win</surname>
          </string-name>
          ,
          <string-name>
            <surname>T. K.</surname>
          </string-name>
          :
          <article-title>Breast cancer data analysis for survivability studies and prediction</article-title>
          . In: Computer Methods and Programs in Biomedicine (
          <year>2017</year>
          ) https://doi.org/10.1016/j.cmpb.
          <year>2017</year>
          .
          <volume>12</volume>
          .011
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Lohr</surname>
            ,
            <given-names>S.: For</given-names>
          </string-name>
          <string-name>
            <surname>Big-Data</surname>
            <given-names>Scientists</given-names>
          </string-name>
          , 'Janitor Work' is Key Hurdle to Insights, http://www.nytimes.com/
          <year>2014</year>
          /08/18/technology/for-big
          <article-title>-data-scientists-hurdle-to-insights-is-janitorwork</article-title>
          .
          <source>html?_r=0 Last accessed 14 July 2020</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Makarova</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lagerev</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lozbinev</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Approaches to visualizing big text data at the stage of collection and pre-processing</article-title>
          .
          <source>In: Scientific Visualization N. 11.4</source>
          , pp.
          <fpage>13</fpage>
          -
          <lpage>26</lpage>
          , (
          <year>2019</year>
          ). https://doi.org/10.26583/sv.11.4.
          <fpage>02</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Karpov</surname>
            ,
            <given-names>O.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gavrishev</surname>
            ,
            <given-names>M.YU.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>SHishkanov</surname>
          </string-name>
          , D.V.:
          <article-title>Integraciya medicinskoj informacionnoj sistemy i sistemy administrativno-hozyajstvennoj deyatel'nosti kak instrument optimizacii processov medicinskoj organizacii. Otdel'nye problemy i puti ih resheniya</article-title>
          .
          <source>In: Sovremennye naukoemkie tekhnologii. № 9-1</source>
          . pp.
          <fpage>46</fpage>
          -
          <lpage>50</lpage>
          . (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <article-title>Portal of operational interaction of USISH participants http://portal</article-title>
          .egisz.
          <source>rosminzdrav.ru/materials Last accessed 14 July 2020</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Kreuzthaler</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oleynik</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Avian</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schulz</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Unsupervised Abbreviation Detection in Clinical Narratives</article-title>
          . In:
          <article-title>Studies in Health Technology and Informatics</article-title>
          . v.
          <volume>245</volume>
          , pp.
          <fpage>539</fpage>
          -
          <lpage>543</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Lagerev</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Makarova</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <article-title>Features of preliminary processing of semi-structured medical data in Russian for use in ensembles of data mining models</article-title>
          .
          <year>2020</year>
          . Т.
          <volume>17</volume>
          , № 7. pp.
          <fpage>43</fpage>
          -
          <lpage>53</lpage>
          . https://doi.org/ 10.14489/vkit.
          <year>2020</year>
          .
          <volume>07</volume>
          .pp.
          <fpage>043</fpage>
          -
          <lpage>053</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Zellig</surname>
            ,
            <given-names>S. H.</given-names>
          </string-name>
          :
          <article-title>Distributional Structure</article-title>
          . v.
          <volume>10</volume>
          . pp.
          <fpage>146</fpage>
          -
          <lpage>162</lpage>
          ,
          <string-name>
            <surname>Word</surname>
          </string-name>
          (
          <year>1954</year>
          ), https://doi.org/10.1080/00437956.
          <year>1954</year>
          .11659520
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Tanimoto</surname>
          </string-name>
          , T.T.:
          <source>IBM Internal Report 17th Nov. IBM. Corp</source>
          , New York (
          <year>1957</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>