<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>An overview of Lithuanian Internet media n-gram corpus</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ieva Bumbuliene˙</string-name>
          <email>ieva.bumbuliene@bpti.lt</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Loïc Boizou</string-name>
          <email>l.boizou@hmf.vdu.lt</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tomas Krilavicˇius</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Vilnius University, Lithuania Baltic Institute of Advanced Technology</institution>
          ,
          <country country="LT">Lithuania</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Vytautas Magnus University</institution>
          ,
          <country country="LT">Lithuania</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Vytautas Magnus University, Lithuania Baltic Institute of Advanced Technology</institution>
          ,
          <country country="LT">Lithuania</country>
        </aff>
      </contrib-group>
      <fpage>24</fpage>
      <lpage>28</lpage>
      <abstract>
        <p>-This paper describes construction and properties of the open 70 million words Lithuanian Internet media n-gram corpus. Due to copyright limitations often contemporary media based resources availability is restricted, while n-grams corpora (e.g., Google N-gram viewer/corpus) solve the problem. Lithuanian language is under-resourced, hence n-gram corpus of Lithuanian media is designed to contribute to publicly available ready-to-use lexical resources. In this paper we report corpus construction procedure, preprocessing, corpus statistics and possible areas of application.</p>
      </abstract>
      <kwd-group>
        <kwd>corpus</kwd>
        <kwd>Internet media</kwd>
        <kwd>Lithuanian</kwd>
        <kwd>n-grams</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>I. INTRODUCTION</title>
      <p>
        This paper describes the construction, and properties of
the Lithuanian Internet media n-gram corpus. N-grams have
been used in the variety of Natural Language Processing
(NLP) tasks. Besides, n-gram language models have become
popular as a general resource for data-driven algorithms in
many areas such as speech recognition, text tagging, spelling
correction, named entity recognition, word sense
disambiguation, lexical substitution [1] and especially statistical machine
translation [2], [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. A bit less common application of n-grams
is sentiment analysis or polarity mining, i.e., identifying and
extracting subjective information from text [4], developing a
grammar checker without grammar rules [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Also, n-grams
were reported to improve classification tasks, e.g., email-act
classification, where n-grams provide contextual information
for better characterization of distinct classes [6], legal text
classification [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], [8]. This versatile applicability of n-grams
is one of the main motivations for construction of Lithuanian
Internet media n-gram corpus.
      </p>
      <p>
        A number of papers (e.g., [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]) report that effectiveness
of NLP techniques and methods is sensitive to the data size.
Therefore there have been significant efforts to create larger
datasets and thus at least several large n-gram corpora became
available. Among others, Google Books N-gram Corpus is
worth mentioning in this context. Primarily designed for
building better language models for machine translation and other
NLP applications, Google Books N-gram Corpus is a database
of n-grams up to 5 words with the frequency distribution of
each sequence/unit in each year from 1500 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. In preparation
of the corpus, n-grams were filtered by eliminating unigrams
Justina Mandravickaite˙
      </p>
      <p>DATA SOURCE AND INITIAL CORPUS</p>
      <p>Initial corpus is constructed of articles from a popular
Lithuanian news portal delfi.lt and covers time period from
2014 March to 2016 November, inclusive. Articles and their
meta-information were crawled and stored to a database. The
database consists of the following fields:
1)
2)
1delfi.lt
category,
date,
3)
4)
5)
6)
7)
1)
2)
3)
4)
5)
6)
7)
8)
9)
10)
11)
title,
author,
link (url),
word count and
article itself.</p>
      <p>Articles were crawled from the following 11 categories:
people,
projects,
science,
auto,
sport,
life,
news,
citizen,
business,
fit,
other category, which consists of articles that did not
fit into any of the other mentioned categories or which
category was not recognized.</p>
      <p>Video category was ignored because the articles it contained
were too short. Thus initial corpus is constructed of 190000
articles, it has 72 million words and requires for a database
of 0:48 GB. Such initial corpus was used as a base to build
an n-gram (where n ranges from 1 to 4) corpus consisting of
4 n-gram collections. Development and characteristics of the
corpus are described in the following sections.</p>
      <p>III. PREPROCESSING AND N-GRAM GENERATION
Preprocessing and n-gram generation were separated into
several phases:
1)
2)
3)
4)
documents symbols analysis,
sentence splitting,
n-grams generation according to the rules,
n-grams frequency calculation.</p>
      <p>During the first step of preparation for n-gram generation,
analysis of symbols that form news articles was performed.
It was found out that articles are constructed not only from
Lithuanian letters, digits and main punctuation marks but
also of other miscellaneous symbols, e.g., letters of different
foreign languages (e.g., Cyrillic for Russian, some Hebrew,
Greek letters, etc.), currency symbols, degree symbol, accented
symbols, etc. Moreover, some of the symbols were similar and
had similar purpose, but were encoded differently, e.g., there
were 14 types of dashes and 8 types of double quotes. While
the usage of acronyms in the articles was not analysed, it was
decided to keep words of foreign languages during n-gram
generation process as they are important for the meaning of a
sentence. Discarding these words might be incorrect action
because the sequence of words in the sentence cannot be
changed while generating n-grams.</p>
      <p>The second step was sentence splitting. Each file was split
into sentences by a rule-based Haskell segmentation library
developed at the Centre of Computational Linguistics
(Vytautas Magnus University) since 2013 in order to process rough
or morphologically annotated data. Older versions have been
used to prepare data for Sketch Engine and for the training
of the morphological analyser of semantika.lt web service. In
1)
2)
3)
4)
5)
6)
7)
8)
1)
2)
3)
4)
5)
late 2016, the current version of the library contributed to the
processing of Lithuanian Parseme data.</p>
      <p>The sentence limits in the files used to generate n-gram
lists are marked by empty lines. Tokenisation rules for n-grams
were created based on the analysis of symbols in the previous
step. Thus n-gram splitting consisted of the following steps:</p>
    </sec>
    <sec id="sec-2">
      <title>Every sentence was split by spaces. All text was converted to lowercase characters. Analysis of each created token was performed. a)</title>
      <p>b)
a)
b)
c)</p>
      <p>If a token consisted only of letters or digits,
it was declared a word.</p>
      <p>A token of one symbol which was not a
letter and not a digit was ignored as it was
considered to be a punctuation mark or some
symbol that can be thrown out.</p>
      <p>Predefined patterns where symbols should not be
deleted were searched:</p>
      <p>URL addresses,
email addresses,
float numbers and a special word structure
consisting of a number, symbol and a part
of word, e. g. “32-tasis (the 32nd)”,
“101osiose (in the 101st)”, “1965-ais (in the year
1965)”.</p>
      <p>
        A dictionary of abbreviations [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] was used to find
shortened forms of the words in order to keep such
tokens as words without discarding full stops because
in the later stages all full stops were removed.
      </p>
      <p>When a token in format “digit:digit” occurred, it was
split into two digits which were then saved as two
separate tokens. This type of token, e.g., could refer
to a result of a sport game match.</p>
      <p>All the punctuation marks that occurred at the start
or at the end of a token were removed while the
remaining part was saved.</p>
      <p>Finally, tokens that contained no letters and no digits
were removed. The analysis of the removed content
revealed it to be various meaningless sequences of
symbols that mostly were punctuation notations, e.g.,
&lt;...&gt;.</p>
      <p>After creating and preprocessing tokens, n-grams were
generated for each sentence and stored in database. Additionally,
some meta-information about n-grams was saved as well. It
included:
separate words that formed n-gram,
information about the document n-gram was extracted
from,
sentence number,
index of every word,
length of every word.</p>
      <p>The result was 4 collections of n-grams (unigrams, bigrams,
trigrams and tetragrams) in the database. These collections
were aggregated to count frequencies of different n-grams and
construct final corpora. Statistics of constructed n-gram corpus
is presented in the next section.</p>
      <p>N-GRAM CORPUS BASIC STATISTICS
In total
unigrams</p>
      <p>CORPUS STATISTICS</p>
      <p>Summary of the basic statistics of n-gram corpus is
presented in the Table I. The percent of unique n-grams varies
from 1.45 for unigrams to 87.82 for 4-grams. In addition,
the percent of n-grams which occur in corpus once (hapax
legomena) is 0.62 for unigrams and 82.05 for 4-grams whereas
values for 2-grams and 3-grams are distributed in mentioned
intervals. On average, unique n-grams consist of 45.13% and
n-grams with frequency 1 make 39.29% of all the n-grams in
the corpus.</p>
      <p>Figures 1-4 present variability of n-gram frequencies in the
corpus. X axis shows frequencies of every n-gram that occurs
in the corpus whereas y axis represents a count of different
ngrams that are in the corpus with a specified frequency. Thus
Figure 1 shows that when a value of unigram frequency is
higher than 200, the amount of unigrams with the same
frequency in the n-gram corpora is less than 100.</p>
      <p>Meanwhile Figure 2 shows that counts of different
2grams with frequency higher than 50 are much smaller
in comparison to other n-grams with lower frequency. For
3grams and 4-grams frequencies of different n-grams continue
to drop as well. These values are approximately 25 for 3-grams
and 10 – for 4-grams (see 3 and 4).</p>
      <p>V.</p>
      <p>CONCLUSION</p>
      <p>This paper describes construction and properties of the
70 million Lithuanian Internet media n-gram corpus. All
important steps of preprocessing, and statistical details of
the corpus are discussed. Initial texts were collected from
the Lithuanian news portal delfi.lt (190 thousands articles;
72 million words). Then analysis of symbols that comprise
extracted texts was performed for generation of tokenization
rules. Later sentence splitting and tokenization using additional
resources and patterns were executed. After that, n-grams (uni,
bi, tri and tetragrams) were generated for every sentence. The
following statistics were calculated for the newly constructed
Lithuanian Internet media n-gram corpus for each n-gram
collection: total number, number of unique n-grams, percentage
of unique n-grams, percentage of n-grams with frequency 1,
and frequency counts. N-gram corpus of Lithuanian Internet
media is designed to contribute to publicly available
readyto-use lexical resources for various NLP, linguistic, etc. tasks,
and soon will be available for all users.</p>
    </sec>
    <sec id="sec-3">
      <title>ACKNOWLEDGMENT This research was funded by the Research Council of Lithuania (No. LIP-027/2016)), see www.mwe.lt for more details.</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Flor</surname>
          </string-name>
          , “
          <article-title>A fast and flexible architecture for very large word ngram datasets,” Natural Language Engineering</article-title>
          , vol.
          <volume>19</volume>
          (
          <issue>1</issue>
          ), pp.
          <fpage>61</fpage>
          -
          <lpage>93</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>C.</given-names>
            <surname>Buck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Heafield</surname>
          </string-name>
          ,
          <string-name>
            <surname>B. Van Ooyen</surname>
          </string-name>
          , “
          <article-title>N-gram counts and language models from the common crawl,” in LREC</article-title>
          , vol.
          <volume>2</volume>
          . Citeseer, p.
          <fpage>4</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Pauls</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Klein</surname>
          </string-name>
          , “
          <article-title>Faster and smaller n-gram language models,” in Proc. of the 49th Annual Meeting of the ACL: Human Language Technologies-Volume 1</article-title>
          . ACL, pp.
          <fpage>258</fpage>
          -
          <lpage>267</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>D.</given-names>
            <surname>Bespalov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Bai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Qi</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Shokoufandeh, “
          <article-title>Sentiment classification based on supervised latent n-gram analysis</article-title>
          ,
          <source>” in Proceedings of the 20th ACM international conference on Information and knowledge management. ACM</source>
          , pp.
          <fpage>375</fpage>
          -
          <lpage>382</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>R.</given-names>
            <surname>Nazar</surname>
          </string-name>
          , I. Renau, “
          <article-title>Google books n-gram corpus used as a grammar checker</article-title>
          ,”
          <source>in Proceedings of the Second Workshop on Computational Linguistics and Writing (CLW</source>
          <year>2012</year>
          )
          <article-title>: Linguistic and Cognitive Aspects of Document Creation and Document Engineering</article-title>
          . ACL, pp.
          <fpage>27</fpage>
          -
          <lpage>34</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>V. R.</given-names>
            <surname>Carvalho</surname>
          </string-name>
          , W. W. Cohen, “
          <article-title>Improving email speech acts analysis via n-gram selection</article-title>
          ,”
          <source>in Proceedings of the HLT-NAACL 2006 Workshop on Analyzing Conversations in Text and Speech. ACL</source>
          , pp.
          <fpage>35</fpage>
          -
          <lpage>41</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>V.</given-names>
            <surname>Mickevicˇius</surname>
          </string-name>
          , T. Krilavicˇius, V. Morkevicˇius, “
          <article-title>Classification of short legal lithuanian texts</article-title>
          ,
          <source>” BSNLP</source>
          <year>2015</year>
          , p.
          <fpage>106</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>V.</given-names>
            <surname>Mickevicˇius</surname>
          </string-name>
          , T. Krilavicˇius, V. Morkevicˇius, A. Mackute˙- Varoneckiene˙, “
          <article-title>Automatic thematic classification of the titles of the seimas votes,”</article-title>
          <source>in Proc. of the 20th Nordic Conf. of Computational Linguistics, NODALIDA 2015, May 11-13</source>
          ,
          <year>2015</year>
          , Vilnius, Lithuania,
          <string-name>
            <given-names>B.</given-names>
            <surname>Megyesi</surname>
          </string-name>
          , Ed. Linköping University Electronic Press / ACL,
          <year>2015</year>
          , pp.
          <fpage>225</fpage>
          -
          <lpage>231</lpage>
          . [Online]. Available: http://aclweb.org/anthology/W/W15/ W15-1828.pdf
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>P.</given-names>
            <surname>Norvig</surname>
          </string-name>
          , “
          <article-title>Statistical learning as the ultimate agile development tool,” in ACM 17th Conf</article-title>
          .
          <article-title>on Information and Knowledge Management Industry Event (CIKM-</article-title>
          <year>2008</year>
          ),
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Banko</surname>
          </string-name>
          , E. Brill, “
          <article-title>Mitigating the paucity-of-data problem: Exploring the effect of training corpus size on classifier performance for natural language processing</article-title>
          ,”
          <source>in Proceedings of the first international conference on Human language technology research. ACL</source>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>T.</given-names>
            <surname>Hawker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gardiner</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Bennetts, “
          <article-title>Practical queries of a massive n-gram database,”</article-title>
          <source>in Proc. of the Australasian Language Technology Workshop</source>
          , pp.
          <fpage>40</fpage>
          -
          <lpage>48</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S.</given-names>
            <surname>Evert</surname>
          </string-name>
          , “
          <article-title>Google web 1t 5-grams made easy (but not for the computer),” in Proc. of the NAACL HLT 2010 sixth web as corpus workshop</article-title>
          . ACL, pp.
          <fpage>32</fpage>
          -
          <lpage>40</lpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>J.-B. Michel</surname>
            ,
            <given-names>Y. K.</given-names>
          </string-name>
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>A. P.</given-names>
          </string-name>
          <string-name>
            <surname>Aiden</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Veres</surname>
            ,
            <given-names>M. K.</given-names>
          </string-name>
          <string-name>
            <surname>Gray</surname>
            ,
            <given-names>J. P.</given-names>
          </string-name>
          <string-name>
            <surname>Pickett</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Hoiberg</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Clancy</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Norvig</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Orwant</surname>
          </string-name>
          et al., “
          <article-title>Quantitative analysis of culture using millions of digitized books,” Science</article-title>
          , vol.
          <volume>331</volume>
          (
          <issue>6014</issue>
          ), pp.
          <fpage>176</fpage>
          -
          <lpage>182</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>K.</given-names>
            <surname>Gulordava</surname>
          </string-name>
          , M. Baroni, “
          <article-title>A distributional similarity approach to the detection of semantic change in the google books ngram corpus</article-title>
          ,”
          <source>in Proceedings of the GEMS 2011 Workshop on GEometrical Models of Natural Language Semantics. ACL</source>
          , pp.
          <fpage>67</fpage>
          -
          <lpage>71</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>P.</given-names>
            <surname>Juola</surname>
          </string-name>
          , “
          <article-title>Using the google n-gram corpus to measure cultural complexity,” Literary and linguistic computing</article-title>
          , vol.
          <volume>28</volume>
          (
          <issue>4</issue>
          ), pp.
          <fpage>668</fpage>
          -
          <lpage>675</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.-I.</given-names>
            <surname>Chiu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Hanaki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hegde</surname>
          </string-name>
          , S. Petrov, “
          <article-title>Temporal analysis of language through neural language models</article-title>
          ,
          <source>” ACL</source>
          <year>2014</year>
          , p.
          <fpage>61</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17] “
          <article-title>Abbreviations dictionary of lithuanian language</article-title>
          ,” https://github. com/tokenmill/ltlangpack/blob/master/tokenizer/abbr-dictionary.xml,
          <source>accessed: 24 02</source>
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>