<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Quantitative Properties of Russian Adjective-Noun Collocations across Dictionaries and Corpora*</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Khokhlov</string-name>
          <email>m.khokhlova@spbu.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>St. Petersburg State University</institution>
          ,
          <addr-line>7/9 Universitetskaya nab., 199034 St. Petersburg</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The paper discusses the differences between collocations extracted from a number of Russian dictionaries paying attention to their frequency characteristics based on corpora. The aim of the study was, first, to analyze how collocations and set expressions are described in Russian explanatory and specialized dictionaries and to what extent their data coincide with each other, and, secondly, to investigate how collocations presented in dictionaries are reflected in text corpora. This will make it possible to examine the interrelation between the “manually” collected data and modern corpora (the Russian National Corpus and ruTenTen). We tested the following hypothesis, i.e. high collocation frequencies correspond to the fact that the item is represented in several dictionaries. In our paper we considered 180 collocations built according to the “adjective / participle + noun” model. The results show the heterogeneity of the dictionary data while the choice of lexical items does not coincide with its frequency characteristics: the examples are low-frequency and about 34% are absent in the disambiguated subcorpus. Explanatory dictionaries and collocation dictionaries show the smallest overlap.</p>
      </abstract>
      <kwd-group>
        <kwd>Collocations</kwd>
        <kwd>Russian Language</kwd>
        <kwd>Dictionaries</kwd>
        <kwd>Corpora</kwd>
        <kwd>Statistics</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Our project deals with the process of building a database that will represent Russian
collocations extracted from dictionaries and corpora [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. The results of this research
can be used in various NLP tasks and also in different fields of theoretical and applied
linguistics, i.e. Russian lexicology, morphology and syntax or teaching the Russian
language. Data about Russian collocability can be valuable for machine translation,
clustering of words and word combinations, sentiment analysis, text summarization,
disambiguation etc. It is expected that the collocations extracted from dictionaries will
be used for the evaluation of machine learning algorithms dealing with automatic text
processing, since today there is no single standard that would include verified
information and at the same time in sufficient quantity.
      </p>
      <p>Since dictionaries are an important source of data about collocations, it is
important to analyze them in order to understand the possible distinctions. At the same
time, it can be tricky to compare different dictionaries.</p>
      <p>In the paper we consider collocations which were extracted from six Russian
dictionaries, analyzing how they are reflected in corpora of the Russian language. Within
the framework of our research, we dwell on two tasks. First, to analyze how recurrent
word combinations are presented in different dictionaries and how much they
coincide with each other. Secondly, to investigate the extent to which collocations that are
reflected in dictionaries can be found in corpora and, therefore, trace the intersection
between “manually” collected data and modern corpora.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>The tradition of Russian lexicography has a rich history, however, there are not many
projects dealing with collocability in Russian and moreover based on corpus data or
created by automatic methods. At the same time corpora are seen as the main source
of language data in Western lexicography and up-to-date projects implement them
(Macmillan Dictionary).</p>
      <p>
        The issue of selecting collocations is crucial in lexicography, and not only for
monolingual dictionaries, representing “the most controversial and vulnerable part of
almost every bilingual dictionary” [2, p. 61]. Atkins and Rundell [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] point out the
difficulty of selecting examples from the corpus and suggest using collocation lists for
this task. As some authors note [
        <xref ref-type="bibr" rid="ref12 ref4">4, 12</xref>
        ] the issue of differentiating phrases of various
types is still controversial, which leads to the fact that “specific cases of idiomatic
combinations often do not receive an unambiguous qualification, which is reflected,
in particular, in dictionaries” [12, p. 2].
      </p>
      <p>
        What kind of phrases to include in a computer dictionary is discussed, for example,
in [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Multi-word expressions can be seen as an umbrella term and it is true for
lexicographic resources. Authors list different word combinations in dictionary entries
calling them idioms, phrasemes, collocations etc. A detailed overview of the available
Russian dictionaries was given in the paper [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. The paper [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] describes the machine
learning procedure of selecting coefficients for searching collocations based on the
examples selected from the dictionaries.
      </p>
      <p>
        The idea of comparison between dictionaries and corpora attracted attention from
scholars a few decades ago. The interest was focused on automatic extraction (either
rule-based or statistical one) from the sources and their further evaluation. The results
of the analysis complement each other and can be applied for constructing NLP
lexicons [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The research presented later in [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] puts emphasis on collocations in
machine translation. Their implementation improved the quality of the analyzed systems.
      </p>
      <p>
        There is also a number of works involving comparison between text corpora (for
example, [
        <xref ref-type="bibr" rid="ref10 ref17">17, 10</xref>
        ]). Most of them deal with building frequency lists and calculating
statistical metrics based on several corpora trying to find a suitable test “supporting
the comparison of small and large corpora” [10, p. 258].
      </p>
    </sec>
    <sec id="sec-3">
      <title>Experiment: Description</title>
      <p>The merging of dictionaries from different sources implies not only a single
lexicographic format, but also raises the question of data relevance. When describing
collocations, a lexicographer needs to select examples taking into account their
representativeness in corpus, coverage in dictionaries, and also suitability for language users and
their purposes.</p>
      <p>
        We identified items from several Russian dictionaries of different types:
1. Explanatory dictionaries, i.e. the Dictionary of the Russian Language (DRL [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]);
the Large Explanatory Dictionary of the Russian Language (LEDR [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]);
2. Collocation dictionaries [
        <xref ref-type="bibr" rid="ref16 ref18 ref3">18, 16, 3</xref>
        ];
3. Online dictionary [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>
        In our research, we tested the following hypothesis: high collocation frequencies in
the corpus correspond to high values of the dictionary index introduced in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. This
index is understood as the number of dictionaries in which the item is recorded. That
is, we expect to see a directly proportional relationship between lexicographic and
corpus data and, therefore, a positive correlation between dictionaries and corpora.
None of the collocations was recorded in all six dictionaries we examined, so the
maximum value of the dictionary index turned out to be 4.
      </p>
      <p>To assess the extracted data across text corpora, we randomly selected 20
collocations from groups with dictionary indices 2, 3 and 4 (see Table 1 for the examples)
resulting in 60 collocations. During the statistical analysis we implemented
nonparametric Friedman and Kruskal-Wallis tests for the comparison between corpora.</p>
      <p>We also analyzed 20 collocations that are present only in one dictionary (i.e., their
dictionary index was 1) resulting in 120 collocations.</p>
      <p>
        As a material for our study we used the Russian National Corpus (RNC), i.e. a
disambiguated subcorpus of 6 mln tokens and the main corpus of 321 mln tokens. The
given subcorpora represent a “classical” (traditional) approach to corpus building
compared to an automatic one but are not so large, as opposed to other Russian
corpora. Therefore we also consulted ruTenTen corpus of more than 18 bln tokens that was
crawled automatically [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>Collocation</p>
      <p>Borisova 1995
1. adskaya
bol’ ‘hellish
pain’
2. glubokaya
drevnost’
1 1 and 0 stand for presence or absence of the collocation in the dictionary.</p>
      <p>LEDR</p>
      <p>Dictionary index
0
0
2
3</p>
      <p>an‘great
tiquity’
3. goryachaya
lyubov’
‘burning
love’
4. ostraya
diskussiya
‘heated
discussion’
5. vysokoye
masterstvo
‘superior
skill’
6. zverinaya
zhestokost’
‘monstrous
cruelty’
1
1
0
0
1
1
1
1
1
1
1
1
0
0
0
0
1
1
1
0
0
0
0
0
4
4
3
2
In our study we will also address to the following questions: 1) can we use corpora of
a smaller volume for collocation analysis or in tasks dealing with their automatic
processing? 2) do “traditional” and large web corpora produce the same results? 3) do
dictionaries present homogeneous data, i.e. collocations of the same language nature
demonstrate similar quantitative features?
4
4.1</p>
    </sec>
    <sec id="sec-4">
      <title>Experiment: Results</title>
      <sec id="sec-4-1">
        <title>Dictionary index 4</title>
        <p>Even for the high value of dictionary index, the results are not uniform (see examples
in Table 2). 4 collocations are absent in the disambiguated subcorpus, although they
are present in several dictionaries. Spearman’s rank correlation coefficient is about
0.81 for all three pairs of corpora, which indicates a strong positive relationship
between them and their similar ranking of collocations.
The average frequency in the subcorpus of the RNC was 1.34, the main corpus of the
RNC and ruTenTen showed 1.93 and 2.01 respectively, but the differences between
the data are statistically insignificant (p &gt; 0.05 according to the Friedman test).
Hence, the frequencies are homogeneous but two collocations with the lexeme
bol’shoy ‘large’ (bol’shoye znacheniye ‘great meaning’ and bol’shoy uspekh ‘great
success’) show outliers.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Dictionary index 3</title>
        <p>The frequencies of this group of collocations (see Table 3) show lower correlation
(Spearman’s rank correlation coefficient varies between 0.62 and 0.79), while the
differences between them in corpora are significant (p &lt; 0.05 according to the
Friedman test).
For collocations selected from two dictionaries, we see that 12 units out of 20 (60%)
were not found in the smallest corpus, and 3 of them were not recorded in the main
RNC corpus either (see Table 4 for the results).
14.
15.
16.
17.</p>
        <p>Thus, in the case of the above given examples that are present in three dictionaries,
the frequencies tend to reveal more diversity.
Spearman’s correlation coefficient increased up to 0.92 for ruTenTen and the main
RNC corpus while it decreased to 0.53 for ruTenTen and the disambiguated RNC
subcorpus. Hence we can register differences in ranking in the latter case. But the
fluctuations in the frequencies between all three corpora are not significant (p &gt; 0.05
according to the Friedman test). This result enables us to suggest that collocations
found only in two dictionaries are rare.
4.4</p>
      </sec>
      <sec id="sec-4-3">
        <title>Dictionary index 1</title>
        <p>
          Despite the fact that the online dictionary of idiomatic expressions [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] was compiled
on the basis of the RNC, only half of the collocations were recorded in the
disambiguated subcorpus. Collocations extracted from the dictionary are characterized by
extremely low frequencies in all three corpora and show minimal values compared to
other lexicographic resources. This is the poorest result among collocations obtained
for all dictionaries. The differences in frequency values between corpora are
insignificant (p &gt; 0.05 according to the Friedman test), and the standard deviation values are
also low, i.e. one can assume some homogeneity of noun collocations in this
dictionary.
        </p>
        <p>
          The collocations extracted from the dictionary of lexical intensifiers [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] are also
characterized by low frequencies in corpora and insignificant differences.
        </p>
        <p>
          The dictionary of set expressions [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] shows the highest results for the collocation
frequencies (the differences between corpora are also insignificant, p &gt; 0.05 according
to the Friedman test), i.e. it can be concluded that this lexicographic resource reflects
more frequent collocations. For example, vyssheye obrazovaniye ‘higher education’
and dukhovnaya zhizn’ ‘spiritual life’.
        </p>
        <p>
          Analysis of the data in the collocations dictionary [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] suggests that the selected
items occupy an intermediate position according to their frequency characteristics, i.e.
8 items are not recorded in the disambiguated subcorpus, and 5 collocations have only
1 occurrence in the main RNC corpus. But nevertheless the collocations extracted
from the given dictionary prove to be the only ones showing significant differences
between corpora.
        </p>
        <p>
          The results for collocations from both explanatory dictionaries suggest that the
sources differ to a certain degree in how they represent unique phrases. For units from
the LEDR, the distribution of frequencies in three corpora is characterized by outliers
and a large range of values (for example, aktsionernoye obscestvo ‘joint stock
company’, organicheskoye veschestvo ‘organic matter’, pochtovyy yaschik ‘letterbox’).
Collocations from the DRL have smaller deviations from the mean values. Both
explanatory dictionaries show very little overlap with other lexicographic sources. This can
be explained by the fact that dictionaries are aimed at describing different vocabulary:
for example, the dictionary [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] represents only phrases with the meaning of high
intensity, while DRL and LEDR are aimed at a more complete presentation of
vocabulary and dictionary entries list phraseological units.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Discussion</title>
      <p>The analysis shows that in total there are no significant differences between corpora
in frequencies of collocations with the same dictionary indexes (or from the same
dictionary). Thus the analyzed items prove to be rare units. About 34% of the
considered collocations are absent in the RNC disambiguated subcorpus, i.e. it can be
assumed that the volume of 6 mln tokens is not enough to study collocability. About
12% of the analyzed collocations yield less than 0.01 occurrences per million even in
the largest ruTenTen corpus.</p>
      <p>It should be mentioned that collocation frequencies in corpora are steadily
decreasing with the decrease of dictionary index (the differences are statistically significant,
p &lt; 0.05 according to the Kruskal-Wallis test), and the value of the Spearman’s rank
correlation coefficient also decreases. Collocations represented in four dictionaries
tend to be more widespread in corpora but also have low frequencies.</p>
      <p>
        It is worth noting that with the rise of corpus volume the unique collocations (with
dictionary index equal to 1) tend to show more diversity in their frequencies. Only
the collocation dictionaries [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] demonstrated significant differences on the
smallest RNC corpus while the main RNC corpus could exemplify one more pair, e.g.
the dictionary of set expressions [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] and the dictionary of lexical intensifiers [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>
        The ruTenTen corpus proved to have the largest number of pairs with significant
differences in frequencies (here we can name additionally, firstly, the dictionary of
idiomatic expressions [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and the collocations dictionary [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], and secondly, the
former [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and LEDR [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]). This can suggest that the dictionary of set expressions [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]
includes more frequent phrases compared to other sources, while the dictionary of
Russian idiomatics [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] contains the least recurrent units.
      </p>
      <p>
        With the exception of a few collocations (bol’shoye znacheniye ‘great importance’,
bol’shoy uspekh ‘great success’, vyssheye obrazovaniye ‘higher education’ and
nervnaya sistema ‘nervous system’ the examples turn out to be low-frequency in all three
corpora. The hypothesis is confirmed that with the decrease of the dictionary index,
the relative frequencies of collocations in the corpus decrease (with the exception of
unique collocations in the dictionary [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], whose frequencies, on the contrary, exceed
the others). The presence of collocations in several dictionaries indicates their higher
frequencies and hence possible prediction by automatic methods.
6
      </p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>In our study we examined the Russian collocations which were extracted from six
dictionaries. Their quantitative characteristics obtained on corpora of different
volumes show that the analyzed examples turn out to be low-frequency and demonstrate
their ambiguous nature. The overwhelming majority of dictionary collocations are
unique, i.e. presented in only one dictionary; hence such items are difficult to be
identified in corpora by using automatic methods.</p>
      <p>The issue of data volume deserves much more attention, and the very phenomenon
of collocability must be investigated in larger corpora as small volume does not show
any occurrences for a number of collocations. Automatically crawled large corpora
reveal more fascinating findings as well as peculiarities obscured in smaller text
collections. Hence we can assume that machine learning algorithms that process word
combinations should be based on large datasets counting several bln tokens.</p>
      <p>The number of dictionary collocations depends on the dictionaries used.
Unfortunately, despite the processing of several sources, the volume of the extracted data is
still insufficient, therefore it is important to analyze other dictionaries and
lexicographic sources and extract examples from them. Explanatory dictionaries may
contain set expressions in other parts of dictionary entries (in the texts of quotations or
illustrative examples), therefore, their further analysis is necessary, which will be
performed at the next stage of our work. In future we plan to consider other resources
and to study collocations based on other syntactic models.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Atkins</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rundell</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          : The Oxford Guide to Practical Lexicography. Oxford U.P.,
          <string-name>
            <surname>Oxford</surname>
          </string-name>
          (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Berkov</surname>
          </string-name>
          , V.:
          <article-title>Bilingual lexicography [Dvuyazychnaya leksikografiya]. 2nd edition</article-title>
          . AST, Moscow (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Borisova</surname>
          </string-name>
          , E.:
          <article-title>A Word in a Text. A Dictionary of Russian Collocations with EnglishRussian Dictionary of Keywords [Slovo v tekste. Slovar' kollokatsiy (ustoychivykh sochetaniy) russkogo yazyka s anglo-russkim slovarem klyuchevykh slov]</article-title>
          .
          <source>Filologiya</source>
          , Moscow (
          <year>1995</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Calzolari</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fillmore</surname>
          </string-name>
          , Ch.,
          <string-name>
            <surname>Grishman</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ide</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lenci</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>MacLeod</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zampolli</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Towards Best Practice for Multiword Expressions in Computational Lexicons</article-title>
          .
          <source>In: Proceedings of LREC - 2002</source>
          ,
          <fpage>1934</fpage>
          -
          <lpage>1940</lpage>
          (
          <year>2002</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <article-title>5. The Dictionary of the Russian Language [Slovar' russkogo jazyka v 4 tomakh]. Yevgen'yeva, A</article-title>
          . P. (ed.-
          <source>in-chief)</source>
          . Vol.
          <article-title> 1-4, 2nd edition, revised and supplemented</article-title>
          .
          <source>Russkij jazyk</source>
          , Moscow (
          <year>1981</year>
          -
          <fpage>1984</fpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Fontenelle</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Collocation acquisition from a corpus or from a dictionary: a comparison</article-title>
          .
          <source>In: Proceedings I-II Papers submitted to the 5th EURALEX International Congress on Lexicography in Tampere</source>
          ,
          <volume>221</volume>
          -
          <fpage>228</fpage>
          (
          <year>1992</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Jakubíček</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kilgarriff</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kovář</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rychlý</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suchomel</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>The TenTen Corpus Family</article-title>
          .
          <source>In: Proceedings of the 7th International Corpus Linguistics Conference CL</source>
          <year>2013</year>
          ,
          <article-title>the United Kingdom</article-title>
          ,
          <year>July 2013</year>
          ,
          <fpage>125</fpage>
          -
          <lpage>127</lpage>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Khokhlova</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Building a Gold Standard for a Russian Collocations Database</article-title>
          .
          <source>In: Proceedings of the XVIII EURALEX International Congress: Lexicography in Global Contexts. Ljubljana</source>
          ,
          <volume>863</volume>
          -
          <fpage>869</fpage>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Khokhlova</surname>
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Collocations in Russian Lexicography and Russian Collocations Database</article-title>
          .
          <source>In: Proceedings of The 12th Language Resources and Evaluation Conference. Marseille, France. European Language Resources Association</source>
          ,
          <fpage>3191</fpage>
          -
          <lpage>3199</lpage>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Kilgarriff</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          : Comparing Corpora.
          <source>International Journal of Corpus Linguistics</source>
          ,
          <volume>6</volume>
          (
          <issue>1</issue>
          ),
          <fpage>97</fpage>
          -
          <lpage>133</lpage>
          (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Klyshinsky</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khokhlova</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          : In Search of Lost Collocations:
          <article-title>Combining Measures to Reach the Top Range</article-title>
          .
          <source>In: Internet and Modern Society: Proceedings of the International Conference IMS-2017 (St. Petersburg; Russian Federation</source>
          ,
          <fpage>21</fpage>
          -
          <lpage>24</lpage>
          June 2017). Radomir V.
          <string-name>
            <surname>Bolgov</surname>
          </string-name>
          ,
          <string-name>
            <surname>Nikolai</surname>
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Borisov</surname>
          </string-name>
          ,
          <string-name>
            <surname>Leonid</surname>
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Smorgunov</surname>
            ,
            <given-names>Irina I. Tolstikova</given-names>
          </string-name>
          , Victor P. Zakharov (eds.). ACM International Conference Proceeding Series,
          <fpage>160</fpage>
          -
          <lpage>163</lpage>
          . ACM Press, N.Y. (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Kustova</surname>
          </string-name>
          , G.:
          <article-title>Dictionary of Russian Idiomatic Expressions [Slovar' russkoyj idiomatiki</article-title>
          .
          <source>Sochetaniya slov so znacheniyem vysokoy stepeni]</source>
          (
          <year>2008</year>
          ), http://dict.ruslang.ru,
          <source>last accessed</source>
          <year>2020</year>
          /10/14.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <article-title>The Large Explanatory Dictionary of the Russian Language [Bol'shoy tolkovyy slovar' russkogo yazyka]</article-title>
          . Kuznetsov, S. (ed.). Norint, St.
          <source>Petersburg</source>
          (
          <year>1998</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Lukashevich</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dobrov</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chuyko</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Selecting word phrases for an automatic text processing system dictionary [Otbor slovosochetaniy dlya slovarya sistemy avtomaticheskoy obrabotki tekstov]</article-title>
          .
          <source>In: Computational linguistics and intellectual technologies: Proceedings of Int. Conf. “Dialog-2008”</source>
          ,
          <fpage>339</fpage>
          -
          <lpage>344</lpage>
          . RSUH, Moscow (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15. Macmillan Dictionary, https://www.macmillandictionary.com,
          <source>last accessed</source>
          <year>2020</year>
          /10/14.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Oubine</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Dictionary of Russian and English Lexical Intensifiers [Slovar' usilitel'nykh slovoso-chetaniy russkogo I angliyskogo yazykov]</article-title>
          .
          <source>Russian Language</source>
          , Moscow (
          <year>1987</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Rayson</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garside</surname>
          </string-name>
          , R.:
          <article-title>Comparing corpora using frequency profiling</article-title>
          .
          <source>In: Proceedings of the workshop on Comparing Corpora. Association for Computational Linguistics</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          (
          <year>2000</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Reginina</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tjurina</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shirokova</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Set Expressions of the Russian Language. A Reference Book for Foreign Students [Ustoychivye slovosochetaniya russkogo yazyka: Uchebnoye posobiye dlya studentov-inostrantsev]</article-title>
          .
          <string-name>
            <surname>Shirokova</surname>
            ,
            <given-names>L. I</given-names>
          </string-name>
          . (ed.). Moscow (
          <year>1980</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19. Russian National Corpus, http://ruscorpora.ru,
          <source>last accessed</source>
          <year>2020</year>
          /10/14.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Wehrli</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Seretan</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nerima</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Russo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Collocations in a rule-based MT system: A case study evaluation of their translation adequacy</article-title>
          .
          <source>In: Proceedings of the 13th Annual Conference of the EAMT, Barcelona</source>
          ,
          <fpage>128</fpage>
          -
          <lpage>135</lpage>
          (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>