<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Quantitative Characteristics of Key Words in Texts of Scientific Genre (on the Material of the Ukrainian Scientific Journal)</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Lviv Polytechnic National University</institution>
          ,
          <addr-line>12 Bandera street, Lviv, Ukraine, 79013</addr-line>
        </aff>
      </contrib-group>
      <fpage>0000</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>The key words in a corpus linguistics statistically deals with the words or word combinations that occur in text more frequently than the others. Each style of the language is characterized by different set of words that makes it stand out among others and easily be referred to. The reason behind this is to convey the relevant information. The scientific style of the Ukrainian language is not the exception in such a respect. It possesses the specific vocabulary that allows conveying the needed information in the most objective, accurate and justified way. In the scientific style, text is most representative language unit as all researches and recent findings are presented in the form of text, so it serves a rich source for a research.</p>
      </abstract>
      <kwd-group>
        <kwd>frequency</kwd>
        <kwd>key words</kwd>
        <kwd>corpus linguistics</kwd>
        <kwd>scientific texts</kwd>
        <kwd>system</kwd>
        <kwd>analysis</kwd>
        <kwd>research</kwd>
        <kwd>scientific journal</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Key words are considered to be the most frequently occurring words in text. Their
main function is to indicate the ‘aboutness’ of a particular text or corpus [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and refer
text to a particular style of genre. They could provide a comprehensive understanding
of the topic, that is why recently they have been of a high interest of the applied
linguistics, corpus linguistics and stylistics.
      </p>
      <p>
        According to Kintsch and van Dijk [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] any statement that is referred to repetitively
should be of a higher importance than the others. Such scholars as Berber [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], Scott
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], Tribble [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], Lazinski [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], Tarasheva [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] study the notion of key words. The
researchers believe that if the particular word often repeats in text, it means that it has
more importance than the other words. And it is also believed that key words can
reveal more information than the other words or word combinations that occur in text.
That is why a list of key words has been an inevitable part of any scientific article. It
is given right after the abstract aiming to briefly introduce the topic to the reader as
well as to facilitate the categorization of articles in general. The object of the research
is key words in the scientific texts. The subject is the investigation of the frequency of
key words in a scientific paper. The aim is to find out what are the most frequent
words used by the Ukrainian scientists and check whether these the most frequent
words are included in abstract and the key words section of the scientific paper to
properly convey the aboutness of a research. All calculations presented in the paper
were made with AntConc [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and a set of author’s programs written in Python.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>The research corpus</title>
      <p>The material of the research is the scientific journal ‘Information systems and
networks’ published by the department of the Lviv Polytechnic National University
‘Information systems and networks’ from 2009 to 2016. Since 2009 this scientific
journal has received a standard structural formatting and division according to subject
headings. A total of 391 articles were processed, which were divided into three
sections: ‘Information systems, networks and technologies’ (further in the text ISN) –
235 articles, which generally contain about 742 000 words, ‘Computational and
mathematical linguistics’ (CML) – 101 articles, which generally contain close 352 000
words, ‘Project and program management’ (PPM) – 55 articles, which in total contain
about 171 000 words.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Key words in the structural units of the research corpus</title>
      <p>The first step towards the organization and preparation of the research corpus, was
decided to divide the selection into structural units such as: text, abstract, the list of
key words. That was the most natural division as all the researched articles had such a
structure. The next step was to identify the key words in each structural unit. The
percentage of the most frequent words that have been encountered in the structural
units of each group of texts is presented below:
It is no surprise that the highest percentage of the most frequent words is presented in
the list of the key words and in abstract. The number of the most frequent words in
text itself is at least 3 times lower than in the abovementioned sections. From one
side, it can be concluded that the most frequent words are really statistically the most
frequent in those sections where they had to be (taking into account the overall
number of words presented in all three structural units in general), although it should be
noted that the expected results for the list of key words were higher than the actual
ones. So, probably the authors, who were publishing the results of their researches in
the scientific journal ‘Information systems and networks’, when it came to the key
words list compilation, were not thorough enough to make sure they provided the
specific key words where they were meant to be.</p>
      <p>The next step of the research was to define the most frequent word for each group
of articles in general with the reference to each structural unit.
For the group of texts ISN the most frequent word is a word система (system). It is
the most frequent one in all structural units. It means that the authors accurately
reflected the key word of the research both in abstract and key words list.
In the second group of texts, the most frequent in the word слово (word) for the
structural unit text, метод (method) for the structural unit abstract and інформаційний
(informational) for the structural unit key words. Hence, it can be concluded that the
texts under the topic ‘Computational and mathematical linguistics’, the key word is
not that obvious as in the first group of texts. It differs in different structural units. In
this case it is also interesting to trace the aboutness of the articles of this group as
computational and structural linguistics is about words, methods and information.</p>
      <sec id="sec-3-1">
        <title>Word</title>
        <p>дані (data)
інформаційний
(informational)
інформаційний
(informational)</p>
      </sec>
      <sec id="sec-3-2">
        <title>Quatity</title>
        <p>1237
42
29</p>
      </sec>
      <sec id="sec-3-3">
        <title>Frequency</title>
        <p>0,74%
1,40%
4,50%
In the third group of texts, the tendency is quite different as we can see that the key
word in the structural unit text is дані (data) while this word cannot be observed in
such structural units as abstract and key words what can lead to the conclusions that
the authors do not convey an accurate analysis of their key words and their abstracts
and key words lists may lack accurate information.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>The intersections of the sets of the most frequent words</title>
      <p>The next step was to find out the most frequent words in each group of texts and how
they coincide following structural units:
 intersection of text with abstract (BiAi)
 intersection of text with key words section (BiKi)
 intersection of abstract with key words section (AiKi)
 intersection of text with abstract with key words section (BiAiKi)
The aim of this research was to find out if the words that experimentally proved to be
the most frequent words are the key words in the group of texts ‘Information Systems
and Networks’ (ISN group), ‘Computer Science’ (CML group) and ‘Project and
program management’ (PPM group).</p>
      <p>The experiment was carried out in the following way:
1. The most frequent words for each group of texts were counted.
2. Each of them was assigned a rank from 1 to 25. Ranks and frequencies of words
are inversely proportional, meaning that the most frequent word has a rank 1 while
the least frequent one has rank 25.
3. 25 the most frequent words were taken to illustrate the experiment.
4. The intersection of 25 most frequent words in the sets of text and abstract (BiAi),
text and key words (BiKi), abstract and key words (AiKi), and text and abstract
and key words were found (BiAiKi).
5. Such manipulation was done for three group of texts.</p>
      <p>The results of the first group of texts is presented in the table below:
From the results above it can be concluded that not all the most frequent words that
were found experimentally are really the most frequent in a scientific text.</p>
      <p>For example, only 12 words among 25 the most frequent both occur in text and
abstract (in other words, the intersection BiAi is the set that contains all elements of B
that also belong to A) and they have different ranks meaning that they have different
frequency. These words are presented in the table in terms of their ranks.</p>
      <p>The ideal picture would be if all cells from 1 to 12 are colored. Instead it can be
observed that those words that occur in a intersection of sets (be it text and abstract or
abstract and key words or others) are not necessarily the most frequent words in the
selection.</p>
      <p>The same experiment was done for all three group of text and the following results
were obtained:
From the results, it can be concluded that for the 1st group of texts ISN, approximately
half of the most frequent words occur in both text and abstract, if to be more precise
only 48%. Only 36% of the most frequent words occur in both text and key words
section, only 48% – in abstract and key words sections. And only 32% of all the most
frequent words fount experimentally are really the most frequent in text, abstract and
key words section. Quite the same tendencies are observed in the 2d group of texts
CML and 3d group PPM.
5</p>
    </sec>
    <sec id="sec-5">
      <title>The most frequent words and their collocations</title>
      <p>After finding out the most frequent words in each group of texts and in all
intersections described above, it was decided to look at the collocations these the most
frequent words occur in the texts. The most frequent words from the whole selection is
presented below:
система (system)
дані (data)
аналіз (analysis)
інформація (information)
модель (model)</p>
      <p>Frequency
From the list of the most frequent words in the whole selection it is seen that the word
система (system) and дані (data) come to the fore. They have relatively high
occurrence comparing to the other most frequent words what shows that the researchers
publish their findings in the ‘Information systems and networks’ describe some
system and data, then they conduct the analysis, work with information, model or
content. To go further and discover what collocations they use with the most frequent
words, it was decided to find the total number of collocations with the most frequent
words. The final stage of the research was to identify the collocations that frequently
occur with the top 5 frequent words of the selection. The results can be seen in the
tables below:
To obtain the most comprehensive results, all the collocations (occurring both on the
left and right sides) of the given word was taken into account and shown:</p>
      <sec id="sec-5-1">
        <title>Rank Wordform</title>
        <p>total
499
290
193
160
151
149
142
132
121
117
138
130
1
103
lef
96
86
86
76
right
15
2
13
6
0
21
6
0
1
7
6
2
1
115
0
2
2
1
(search)
отримання
(receive)
необхідної
(needed)
подання
(presentation)
обміну
(exchange)
джерел
(sources)
даних (data)
інформаційної
(informational)
системи
(systems)
математичні
(mathematical, pl)
математичної
(mathematical,
sing)
концептуальної
(conceptual)
оцінки
(assessment)
математичну
(mathematical)
побудови
(building)
предметної
(subject)
останніх
(recent)
даних (data)
total
125
59
58
46
39
39
38
37
37
37
total
231
51
45
46
43
9
1
0
1
0,61%
0,60%
0,46%
0,46%
0,44%
%
1,65%
0,78%
0,76%
0,61%
0,51%
0,51%
0,50%
0,49%
0,49%
0,49%
%
2,81%
отриманих
(received)
контенту
(content)
основі (basis)
результатів
(results)
кластерного
(cluster)
системи
(system)
тексту (text)
морфологічного
(morphological)
69
The most frequent collocations are the following: electronic system, management
system, information system, database, data repository, text information, data model,
recent analysis and received analysis.
6</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusions</title>
      <p>The main question the paper aims to answer is if the most frequent words in text
could be considered as key words revealing the aboutness of this text, and whether the
authors use these words in abstract and key words sections. The results showed that
the list of key words they compile for their articles do not accurately reflect the
aboutness of their researches, because these are not the most frequent words of their
texts. Quite the same situation was observed with abstract. There are some
discrepancies in terms of the most frequent words between the abstracts, key words lists and
text itself. It can be concluded that when the author writes a scientific paper, he or
she, of course, uses some words more frequently than others. This is a natural process
as they outline some narrow issue. When they write the abstract and compile the list
of key words, they do not necessarily remember about the importance to use the most
frequent words from their researches in abstract and key words sections.</p>
      <p>Generally, abstract and key words section must convey the meaning of the article,
their function is to convey paper’s aboutness, this is why it is of high importance to
include the most frequent words there, and therefore consider these words as key
words. Such tendency was only observed partially allowing to conclude that only
every 2d or 3d word in the abstract and list of key words are really paper’s the most
frequent.</p>
      <p>Hence, the modern Ukrainian researchers when describing their results, use slightly
different key words in abstracts, the list of key words and text itself. They are still
within the same topic, but these are different words what leads to the assumption that
when dealing with abstract or the list of key words, the authors does not carefully
analyze their text and make decision of which lexical units to use in abstract or the list
of key words based on their feelings of the subject area.</p>
      <p>
        Another thing raised in the paper is what are the most frequent words of the
analyzed scientific journal and what is the hidden meaning to use such words. The most
frequent word used in the whole selection is the word система (system). According to
the definition [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], the word система (system) means order, or a set of principles or
procedures according to which somethings is done; an organized methods or scheme.
So, it can be concluded that the Ukrainian researchers who work in the fields of
information systems, computational and mathematical linguistics and project and
program management, tend to order and systematize things they investigate.
      </p>
      <p>It is also interesting to mention that the second most frequent word is дані (data)
assuming that the object of their researches in the most cases is data. So, further to
conclude is that the modern Ukrainian scientists basically work with data. The word
аналіз (analysis) is the third most frequent word allowing concluding that the
preliminary goal of the scientists is to carry out an analysis of something. Among the most
frequent words are also such as інформація (information), модель (model), контент
(content), час (time), використання (use), кількість (quantity), значення (meaning).
Although the results are already representative, the further work is definitely to be
done on a larger selection and with a reference corpus.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>1. AntConc: http://www.laurenceanthony.net/software/antconc/release s/AntConc343/help.pdf.</mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Berber</given-names>
            <surname>Sardinha</surname>
          </string-name>
          ,
          <string-name>
            <surname>T.</surname>
          </string-name>
          :
          <article-title>Wordsets, keywords, and text contents: an investigation of text topic on the computer</article-title>
          ,
          <source>Delta</source>
          <volume>141</volume>
          -
          <fpage>149</fpage>
          . (
          <year>1999</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Kintsch</surname>
            , W., van Dijk,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Toward a model of text comprehension and production</article-title>
          .
          <source>In: Psychological Review</source>
          , vol.
          <volume>85</volume>
          (
          <issue>5</issue>
          ),
          <fpage>363</fpage>
          -
          <lpage>394</lpage>
          . (
          <year>1978</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Lazinski</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Key words in semantics and statistics</article-title>
          .
          <source>In: Biuletyn Polskiego Towarzystwa Językoznawczego</source>
          , vol.
          <volume>62</volume>
          ,
          <fpage>57</fpage>
          -
          <lpage>68</lpage>
          . (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Medvedev</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pashchenko</surname>
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Teoriia imovirnostei ta metematychna statystyka</article-title>
          , Lira, Kyiv,
          <volume>536</volume>
          . (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Merriam-Webster Dictionary</surname>
          </string-name>
          : https://www.merriam-webster.com/dictionary/system.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Scott</surname>
          </string-name>
          , M.:
          <article-title>PC Analysis of key words-and key key words</article-title>
          .
          <source>In: System</source>
          , vol.
          <volume>25</volume>
          ,
          <fpage>233</fpage>
          -
          <lpage>245</lpage>
          . (
          <year>1997</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Tarasheva</surname>
          </string-name>
          , E.: Repetitions of Word Forms in Texts, Cambridge Scholars Publishing. (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Tribble</surname>
          </string-name>
          , Ch.,
          <string-name>
            <surname>Scott</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Textual Patterns: Key Words and Corpus Analysis</article-title>
          .
          <source>In: Language Education</source>
          , vol.
          <volume>11</volume>
          . (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>