<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Information Technology for Extraction Fragments from Scientometric Databases</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Svitlana Petrasova</string-name>
          <email>svetapetrasova@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nina Khairova</string-name>
          <email>nina_khajrova@yahoo.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>National Technical University “Kharkiv Polytechnic Institute”</institution>
          ,
          <addr-line>2, Kyrpychova str., Kharkiv, 61002</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The paper considers the information technology for extraction of Ukrainian and English coherent text fragments from scientometric systems. The up-to-date methods of Data Mining, in particular, statistical models of Information Extraction and Machine Learning for identifying text fragments are analyzed. Based on intelligent data processing tools, the technology will make it possible to determine common information spaces of authors' scientific interaction due to identification of statistically significant collocations that display the topic of texts. The technology includes the distributive-statistical model and Natural Language Processing tools for identification and extraction coherent text fragments. The distributional semantic model MI is applied at the stage of probable collocation identification. Based on POS-tagging formalism, regular expressions, developed in accordance with the grammar of a particular language, allow extracting grammatically correct constructions, i.e. potential substantive, verb and adjective collocations. To identify coherence between these text fragments, dependency parsing is applied. The advantage of the developed technology is that both the grammatical structure and frequency of collocations are taken into account. The corpora of Ukrainian and English scientific texts are built on the basis of abstracts from articles indexed in Google Scholar and ScienceDirect scientometric databases. The effectiveness of the developed technology is assessed and exceeds the results of analogs. The use of the proposed technology could improve the quality of natural language processing. The solution to the problem of automatic extraction of coherent text fragments can be employed to monitor the development of scientific directions, extend research fronts, identify texts of the same domain, extract facts, etc.</p>
      </abstract>
      <kwd-group>
        <kwd>1 Scientometric database</kwd>
        <kwd>coherent text fragment</kwd>
        <kwd>collocation</kwd>
        <kwd>Data Mining</kwd>
        <kwd>distributionstatistical model</kwd>
        <kwd>POS-tagging</kwd>
        <kwd>dependency parsing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>are insufficient in the development of information support for libraries, electronic catalogs, computer
bibliography, systems for automated import of documents, etc. Therefore, it is necessary to apply
novel technologies for extraction of an implicitly expressed connection between text fragments from
scientometric databases that will make it possible to provide relevant search and access to research
works performed on the similar topics.</p>
      <p>Such short text fragments, inter alia, can be collocations that represent a non-random syntactic and
semantic combination of two or more lexemes and provide more specific semantic information than
certain words. So, automatic extraction of collocations that display the topic of text documents can be
an additional tool to identify common information spaces of authors’ scientific interaction.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Works</title>
      <p>Currently, documents, as online multimedia information with hyperlinks, have become a means of
monitoring, influencing and communicating. One of the ways of storing such unstructured data,
namely text documents, is a scientometric system that is a bibliographic and abstract database with
tools for checking the citation of articles published in scientific journals. As the result, it is possible to
identify the state-of-the-art directions of scientific research.</p>
      <p>
        However, to form information spaces of scientific communities adequately, the level of
automation of unstructured data (text) processing needs increasing [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], in particular, by solving the
problems of intelligent analysis.
      </p>
      <p>
        In general, intelligent analysis of data or Data Mining is defined as a decision support process
based on searching hidden patterns (information patterns). Data Mining technologies are one of the
most promising tools to extract valuable knowledge from massive amounts of data, discovering
structures, relations, and interconnectedness of data [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Data Mining methods are at the integration of
several areas (Fig. 1): Data Mining, Web Mining, Machine Learning, Information Retrieval,
Information Extraction, etc.
      </p>
      <p>Thus, we can say that the challenging task of extraction of hidden information patterns such as
collocations is highlighted at the intersection of Information Extraction and Machine Learning
methods.</p>
      <p>
        To identify collocations, distributive-statistical models are considered to be most applicable, the
sense of which is the statistical analysis of the probable co-occurrence of variables (lexemes) in large
data streams. Traditionally, these types of algorithms [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ] use means of mathematical statistics and
algebra, without linguistic information.
      </p>
      <p>Statistical metrics or association measures are based on the frequencies of collocates (words)
included in collocations to calculate the stability of lexical units. In total, there are more than
80 measures to assess the strength of connectivity. The most commonly used measures are MI, PMI,
t-score, log-likelihood, probability coefficient, Pearson’s chi-squared test and others. However,
statistical methods extract noisy data and ignore syntactic links between words.</p>
      <p>At the same time, the analysis of Data Mining methods and statistical models of distributive
semantics shows that collocation extraction requires the use of morphological and syntactic tools in
addition to statistical models. In this case, identification of collocations allows us not only to take into
account the probability of co-occurrence of collocates, but also to formalize grammatical
dependencies between the main and dependent components of text fragments.</p>
      <p>
        To solve the problem, we propose the productive combination of a statistical measure,
in particular, the model of mutual information MI that compares dependent context-related
frequencies with independent ones, and Natural Language Processing tools based on the Universal
Dependency (UD) formalism [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], describing coherence between components of collocations.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Information Technology</title>
      <p>Forming common information spaces of scientific interaction of authors, the developed technology
for extraction of Ukrainian and English coherent text fragments includes four stages (Fig. 2).</p>
      <p>At the first stage of extraction of abstracts from articles indexed in Google Scholar (a freely
accessible scientometric database) and ScienceDirect (a database of scientific publications of the
Dutch publisher Elsevier), the scientific text corpora were designed. Each corpus contains 350
abstracts of research articles on Artificial Intelligence in the Ukrainian and English languages over a
period of 2017-2019. The dataset is of 70 000 (Ukrainian) and 65 000 (English) words.</p>
      <p>
        Since the research on collocations in Natural Language Processing focuses on two related tasks:
collocation identification and collocation extraction [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], collocation identification is implemented at
the second stage of our technology (to discover the collocation tokens in the corpus) while collocation
extraction is carried out at the third stage of the technology (to find the collocation types).
      </p>
      <p>Consequently, based on the model of distributive semantics MI, we identify potential candidates
for the main and dependent components of collocations at the second stage. The metric makes it
possible to single out rare colocations and identify terminology and other constructions where
frequency rates of collocates are very small:
 ( ,  ) =</p>
      <p>( ,  ) × 
2  ( ) ×  ( )
,
(1)
where n is a main word; c is a collocate; f (n, c) is the frequency of the main word n in pairs with the
collocate c; f(n), f(c) are absolute (independent) frequencies of the main word n and the collocate c in
the corpus; N is the total number of words in the corpus.</p>
      <p>The result of MI computation is represented by the list of statistically significant bigrams and
trigrams, probable collocations, that display the topic of processed texts identified in corpora.</p>
      <p>The next stages are devoted to establishing the linguistic correctness of the identified candidates,
namely identification of the types of potential collocations and syntactic dependency between
collocates via Stanza part-of-speech (POS) tagging and dependency parsing.</p>
      <p>
        Stanza [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] is a Python Natural Language Processing toolkit that features a language-agnostic fully
neural pipeline for text analysis, including, in particular, part-of-speech tagging, dependency parsing,
using the UD formalism.
      </p>
      <p>
        The application of Stanza POS-tagging was conducted to extract 3 types of collocations:
substantive, verb and adjective. According to the corpus-oriented approach, these types of
collocations represent significant text fragments that are most often found in corpora [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>In this way, at the third stage, we define the following POS-tags in accord with the main types of
collocations:
1. Substantive collocations:
• &lt;NOUN&gt; &lt;NOUN&gt; (e.g. input process, scrambling of data, “налаштування пристрою”
device setting);
• &lt;ADJ&gt; &lt;NOUN&gt; (e.g. user-centered approach, “програмне забезпечення” - software);
2. Verb collocations:
• &lt;VERB&gt; &lt;NOUN&gt; (e.g. to meet requirements, “нести відповідальність” - to bear
responsibility);
• &lt;VERB&gt; &lt;ADP&gt; &lt;NOUN&gt; (e.g. to be caused by changes, “використовувати в описах” - to
use in manuals);
• &lt;ADV&gt; &lt;VERB&gt; (e.g. to process iteratively, “уважно ознайомитися” - to see thoroughly);
3. Adjective collocations:
• &lt;ADV&gt; &lt;ADJ&gt; (e.g. highly accurate, “абсолютно симетричний” - absolutely symmetric).</p>
      <p>Using Stanza DepparseProcessor at the final stage of the technology, each sentence in the output is
parsed into the UD structure. As the result, syntactic dependency relations between extracted
collocates are defined, representing coherence of the text fragments, i.e. collocations:
• nmod: nominal modifier shows nominal dependents of another noun or a noun phrase,
corresponds to a genitive complement (in Ukrainian). The nmod relation is used in Substantive
collocations (&lt;NOUN&gt; &lt;NOUN&gt;);
• compound: compound relation is used for noun compounds (in English) in Substantive
collocations (&lt;NOUN&gt; &lt;NOUN&gt;);
• amod: adjectival modifier of a noun serves to modify the noun. The relation is applied in
Substantive collocations (&lt;ADJ&gt; &lt;NOUN&gt;);
• obj: object of a verb denotes the entity acted upon, the object is marked by the accusative case
(in Ukrainian). This core argument of a verb is represented in Verb collocations (&lt;VERB&gt;
&lt;NOUN&gt;);
• obl: oblique nominal relation functionally corresponds to the prepositional construction in the
double object construction, as well as temporal and locational modifiers or nominal modifiers for
the agent of a passive verb. It occurs in Verb collocations with prepositions (&lt;VERB&gt; &lt;ADP&gt;
&lt;NOUN&gt;);
• advmod: adverbial modifier functions like adverbs and serves to modify predicates in Verb
collocations (&lt;ADV&gt; &lt;VERB&gt;) or modifier words like adjectives in Adjective collocations
(&lt;ADV&gt; &lt;ADJ&gt;).</p>
      <p>Thus, collocations are considered as coherent text fragments if grammatical characteristics of
collocates, identified via MI model, satisfy POS-tagging and syntactic dependency parsing features.
The collocations extracted from scientific text corpora allow not only identifying the missing
combinations of words in Ukrainian and English linguistic sources, but also calculating the statistical
indicators of their stability.</p>
      <p>The developed technology is implemented as software applications to extract English (Fig. 3) and
Ukrainian (Fig. 4) coherent text fragments from large data streams. Software implementation allows
you to download text corpora, run POS-tagging via Stanza POSProcessor and dependency parsing of
texts via Stanza DepparseProcessor, previously processed by the TokenizeProcessor, MWTProcessor,
and LemmaProcessor, and display all the extracted collocations with examples of their contexts in
descending order of MI coefficient values.</p>
      <p>Running the software program, coherent text fragments are identified that display the topic of
scientific texts and form a common information space (a scientific front) of researchers’ interaction.
Applying the identification results subsequently makes it possible to get new abstracts of the same
topic and extend scientific fronts.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Experimental Results</title>
      <p>To assess the effectiveness of the developed technology and carry out comparative analysis with
existing technologies for extracting collocations from large data streams, the precision score was
calculated.</p>
      <p>
        To compare the obtained results with another existing software implementation dealing with a
similar problem, the web application Sketch Engine [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] was chosen. This application includes the
function of identification of word sketches, i.e. typical phrases that are defined, on the one hand, by
syntax that limits the combination of words in a particular language, and on the other one by the
probability connected to the frequency of usage of words. Unfortunately, Sketch Engine does not
support corpora in the Ukrainian language.
      </p>
      <p>Table 2 shows the comparative analysis of the precision coefficient, determined by the ratio of the
number of relevant decisions made by the system (Correctly Extracted Collocations) to the total
number of collocations extracted by the system (All the Extracted Collocations). The expert evaluated
the extracted collocation as relevant if its morphological and syntactic dependencies were identified
correctly and false otherwise.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>Based on the statistical model of distributive semantics and Natural Language Processing tools, the
information technology is proposed for determining common information spaces of authors’ scientific
interaction due to identification of statistically significant collocations that display the topic of texts.
The effectiveness of the technology is assessed and exceeds the results of analogs.</p>
      <p>The developed software implementation will improve the quality of collocation extraction, and can
also be proposed as an advanced model for existing text processing systems. In instance, it can be
applied to monitor the development of scientific directions, extend research fronts, classify texts of
the same topic, extract facts, etc.</p>
      <p>In future studies, we intend to broaden the scope of our research and focus on a more complex
challenging problem of distant multi-gram collocations extraction. Additionally, our further work will
extend the domain of the texts studied. In prospect, we intend to spread our dataset for free access to
fulfil similar approaches.</p>
    </sec>
    <sec id="sec-6">
      <title>6. References</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>C.</given-names>
            <surname>King</surname>
          </string-name>
          , Research Fronts The Hottest Areas in Science,
          <year>2016</year>
          . URL: http://stateofinnovation.com/research-fronts
          <article-title>-2016-the-hottest-areas-in-science</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L.</given-names>
            <surname>Wen</surname>
          </string-name>
          , Ya. Lu,
          <string-name>
            <given-names>Hui</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Long</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <source>Detecting of Research Front Topic in Artificial Intelligence Based on SciVal, in: Proceedings of the 2nd International Conference on Artificial Intelligence and Advanced Manufacture (AIAM2020)</source>
          .
          <article-title>Association for Computing Machinery</article-title>
          , New York, USA,
          <year>2020</year>
          , pp.
          <fpage>145</fpage>
          -
          <lpage>149</lpage>
          . doi: https://doi.org/10.1145/3421766.342179
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>H.</given-names>
            <surname>Sasaki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Fugetsu</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Sakata</surname>
          </string-name>
          ,
          <article-title>Emerging Scientific Field Detection Using Citation Networks and Topic Models-A Case Study of the Nanocarbon Field</article-title>
          , in: Applied System Innovation.
          <year>2020</year>
          ,
          <volume>3</volume>
          (
          <issue>3</issue>
          ):
          <fpage>40</fpage>
          . doi: https://doi.org/10.3390/asi3030040
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Yun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ahn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Young</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <article-title>Return to basics: Clustering of scientific literature using structural information</article-title>
          ,
          <source>Journal of Informetrics, 14</source>
          <volume>4</volume>
          (
          <year>2020</year>
          ). doi: https://doi.org/10.1016/j.joi.
          <year>2020</year>
          .
          <volume>101099</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.V.</given-names>
            <surname>Petrasova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.F.</given-names>
            <surname>Khairova</surname>
          </string-name>
          ,
          <article-title>Using a Technology for Identification of Semantically Connected Text Elements to Determine a Common Information Space</article-title>
          , in: Cybernetics and
          <string-name>
            <given-names>Systems</given-names>
            <surname>Analysis</surname>
          </string-name>
          ,
          <year>2017</year>
          ,
          <volume>53</volume>
          (
          <issue>1</issue>
          ), pp.
          <fpage>115</fpage>
          -
          <lpage>124</lpage>
          . doi: https://doi.org/10.1007/s10559-017-9912-z
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Ya. Zhao</surname>
            , Ch. Zhang, Yi. Zhang,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>A review of data mining technologies in building energy systems: Load prediction, pattern identification, fault detection and diagnosis</article-title>
          ,
          <source>in: Energy and Built Environment</source>
          ,
          <volume>1 2</volume>
          (
          <issue>2020</issue>
          )
          <fpage>149</fpage>
          -
          <lpage>164</lpage>
          . doi: https://doi.org/10.1016/j.enbenv.
          <year>2019</year>
          .
          <volume>11</volume>
          .003
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Lenci</surname>
          </string-name>
          , Distributional Models of Word Meaning, in: Annual Review of Linguistics,
          <year>2018</year>
          , 4. pp.
          <fpage>151</fpage>
          -
          <lpage>171</lpage>
          . doi: https://doi.org/10.1146/annurev-linguistics-
          <volume>030514</volume>
          -125254
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dinu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Dinu</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Sorodoc</surname>
          </string-name>
          ,
          <article-title>Aggregation methods for efficient collocation detection</article-title>
          ,
          <source>in: Proceedings of the Ninth International Conference on Language Resources and Evaluation</source>
          ,
          <year>2014</year>
          , pp.
          <fpage>4041</fpage>
          -
          <lpage>4045</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>M.-C. de Marneffe</surname>
            , T. Dozat,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Silveira</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Haverinen</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Ginter</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Nivre</surname>
          </string-name>
          , Ch.D. Manning, Universal Stanford Dependencies:
          <article-title>A cross-linguistic typology</article-title>
          ,
          <source>in: Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14)</source>
          , Reykjavik, Iceland,
          <year>2014</year>
          , pp.
          <fpage>4585</fpage>
          -
          <lpage>4592</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Huang</surname>
          </string-name>
          , Zh. Yin,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <article-title>Recognition of Collocation Frames from Sentences</article-title>
          ,
          <source>in: IEICE Transactions on Information and Systems</source>
          ,
          <year>2019</year>
          , E102.
          <source>D(3)</source>
          , pp.
          <fpage>620</fpage>
          -
          <lpage>627</lpage>
          . doi: https://doi.org/10.1587/TRANSINF.2018EDP7255
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>P.</given-names>
            <surname>Qi</surname>
          </string-name>
          , Yu. Zhang, Yu. Zhang, J. Bolton, Ch.D. Manning, Stanza:
          <string-name>
            <given-names>A Python</given-names>
            <surname>Natural Language Processing</surname>
          </string-name>
          <article-title>Toolkit for Many Human Languages</article-title>
          ,
          <source>in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>101</fpage>
          -
          <lpage>108</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2020</year>
          .acl-demos.
          <fpage>14</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S.</given-names>
            <surname>Petrasova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Khairova</surname>
          </string-name>
          , W. Lewoniewski,
          <article-title>Building the semantic similarity model for social network data streams</article-title>
          ,
          <source>in: Proceedings of 2018 IEEE Second International Conference on Data Stream Mining &amp; Processing (DSMP)</source>
          ,
          <year>Lviv</year>
          ,
          <year>2018</year>
          , pp.
          <fpage>21</fpage>
          -
          <lpage>24</lpage>
          , doi: 10.1109/DSMP.
          <year>2018</year>
          .8478480
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kilgarriff</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Baisa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bušta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jakubíček</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kovář</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Michelfeit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rychlý</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Suchomel</surname>
          </string-name>
          ,
          <article-title>The Sketch Engine: ten years on</article-title>
          , in: Lexicography,
          <string-name>
            <surname>ASIALEX</surname>
          </string-name>
          ,
          <year>2014</year>
          ,
          <volume>1</volume>
          (
          <issue>1</issue>
          ):
          <fpage>7</fpage>
          -
          <lpage>36</lpage>
          . doi:
          <volume>10</volume>
          .1007/s40607-014-0009-9
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>