<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Linguistic processing in lattice-based taxonomy construction</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Anastasia Novokreshchenova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maria Shabanova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dmitry Zaytsev</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nina Belyaeva</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Higher School of Economics</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <fpage>350</fpage>
      <lpage>355</lpage>
      <abstract>
        <p>Building a lattice-based taxonomy over a text corpus with formal concept analysis (FCA) methods requires preliminary text processing that would enable construction of a context. We consider several natural language processing methods aimed at automatic attribute acquisition from texts. In particular, we derive attributes of three types: frequent words, latent topics and named entities. Afterwards, we construct a context for each type taking documents in the corpus as a set of objects. Then the corresponding concept lattices are built and pruned with the help of stability index in order to improve the readability of the diagrams. The proposed technique is illustrated on a collection of 26 texts in English dealing with political domain. In this case, the technique serves as a tool for deeper understanding of the interests of di erent political actors producing political texts by clarifying the connections between notions they use in them.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Constructing a taxonomy over a text corpus requires preliminary text
processing. In this paper we explore capabilities of several natural language processing
methods for extracting a set of attributes from the documents each taken as
an object. The rst method is keyword extraction which is based on the
Vector Space Model [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The second technique is Latent Dirichlet Allocation, an
extension of probabilistic latent semantic analysis (pLSA) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], which allows one
to extract latent topics from word distributions. These two methods are based
on the \bag-of-words" assumption|that the order of the words in documents
can be beglected. The third technique we use for attribute extraction is Named
Entity Recognition (NER), which involves linguistic analysis and
part-of-speechtagging (POS-tagging).
      </p>
      <p>Here, we apply these three methods to a collection of 26 texts dealing with
political domain. After constructing an object-attribute matrix and building a
lattice, tering based on the stability index is applied in order to improve the
readability of the diagrams.
?? This work is supported by project 10-04-0017 of the Scienti c Foundation of the State
University-Higher School of Economics \Discrete mathematical models for political
analysis of democratic institutions and human rights".</p>
      <p>Our primary data consists of 26 texts corresponding to speeches of European
and American leaders and spokepersons during the 2007-2010 period addressing
their relations with Russia. These documents were collected by experts in
political science as part of an interdisciplinary research project. From the viewpoint
of political studies, the aims were (1) to de ne the context in which Russia is
addressed in the speeches of Western leaders and international organizations,
(2) to analyze the role and importance of democracy and human rights agenda.
However, this text corpus is used here primarily as the testing ground for di
erent methods of text analysis, and will be expanded in further research in order
to reach a higher level of validity for the conclusions.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Building a context with frequent words</title>
      <p>
        In The Vector Space Model (VSM) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] documents are represented as vectors of
features consisting of the weights of the terms that occur within the collection.
The term weighing we perform here is based on word frequences. This factor is
called term frequency of a term j within a document i:
tfij = Pk nik
      </p>
      <p>;
nij
where nij is the number of occurrences of the word j in document i; the
denominator is the total number of words in the document i.</p>
      <p>
        Constructing the Context and Stability-based Pruning Formal
Concept Analysis (FCA) methods have been proven useful for representing a
meaningful structure of a given knowledge community (such as a scienti c community
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]) in a form of a lattice-based taxonomy (see [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] for FCA preliminaries). While
constructing the context we let the documents be a set of objects and words that
are most frequently used within each document be a set of attributes. The
context construction involved two common techniques: stemming (using the Porter
stemmer [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]) and elimination of stop words.
      </p>
      <p>Using the database of 26 political texts we have built a context made of
documents and terms mentioned in each document most frequently. In particular,
we took 20 most popular terms for each document according to tf measure. The
resulting context contains 26 documents and 249 terms which yields a lattice of
453 formal concepts.</p>
      <p>
        The number of concepts is too large to be shown in a diagram. In order to
obtain intelligible diagrams, we apply the pruning technique based on the notion
of concept stability [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. For a formal context K = (G; M; I) and a formal concept
(A; B) of the context K the stability index is de ned as follows:
(A; B) = jfC
      </p>
      <p>A j C0 = Bgj
2jAj</p>
      <p>The basic stability-based pruning method is to remove all concepts with
stability below a xed threshold. Of course, stable concepts (i.e., satisfying the
chosen stability threshold) do not always form a lattice. However, this fact does
not in uence interpretation of results.</p>
      <p>
        The reduced substructure featuring the 31 most stable concepts is presented
in Fig. 1. The diagrams were produced with Concept Explorer.[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] From this
diagram it is possible to provide the following description concerning the area
of discourse by European countries and the US the situation with Russia. The
term \russia" is obviously a central issue|this word is one of the most frequent
in 20 documents out of 26. It is also a parent for several associated subtopics
such as concepts with intents f\security", \russia"g and f\european", \russia"g.
On the whole, it may be concluded that according to word frequencies security
issues and the relationships between Russia and Europe, as well as some global
problems, are the most discussed topics.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Building a context with latent topics</title>
      <p>
        In this section we address an issue of probabilistic modeling of text. Its basic
idea is that documents are represented as random mixtures over latent topics,
where each topic is characterized by a distribution over words. Such probabilistic
topic approach to document modeling was introduced as probabilistic Latent
Semantic Analysis (pLSA) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Here we apply an extended probabilistic model
Latent Dirichlet Allocation (LDA) [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] to our dataset of 26 political texts taking
the number of latent topics equal to 20.1 As previously, before running the model,
all words were stemmed with Porter algorithm and stop-words were eliminated.
1 The experiments were done
http://psiexp.ss.uci.edu/research.
      </p>
      <p>with</p>
      <p>Matlab</p>
      <p>Topic</p>
      <p>Modeling</p>
      <p>Toolbox:
right nation georgia russia global secur europ
human unit russian russian polici today european
govern nuclear intern interest institut member union
peopl america georgian energy issu challeng idea
democraci american territori medvedev e ort afganistan point
work interest south issu respons strateg thing
women futur order rule achiev face bring
democrat weapon process trust approach matter global</p>
      <p>After applying the LDA model to the text corpora we construct a context
taking topics as attributes for the documents|a topic is assigned to a document
if the total number of times that words in this document were assigned to a
particular topic exceeds a xed threshold. The resulting substructure of the
corresponding pruned lattice is presented in Fig. 2. Lists of words in squares
represent the rst six words assigned to a topic with the highest probability.</p>
      <p>According to this lattice the most actual topics are those connected with
European Union (topic represented by terms \europ", \european", \union", \idea"
and assigned to 12 documents), global problems (\global", \polici", \institut",
\issu", \e ort", 14 documents) and security issues (\secur", \today", \member",
\challeng", \afganistan", \starteg", 14 documents), as well as energy resources
(\russia", \relat", \russian", \interest", \energi", \medvedev", 12 documents).
There is a concept corresponding to the topic of Russian-Georgian con ict which
contains 5 documents. In addition there is an isolated concept related to
economic development of China (\peopl", \work", \commit", \econom", \help",
\china", 9 documents).
4</p>
    </sec>
    <sec id="sec-4">
      <title>Building a Context with Named Entities</title>
      <p>Name Entity Recognition (NER) is the process of nding mentions of xed
types of information in human language. From the 26 texts, we extracted 38
paragraphs that touch issues related to Russia. The paragraphs were processed
with GATE2 system and three types of named entities were extracted|names of
persons, organizations and geographical objects. We construct a context taking
paragraphs as objects and organizations, persons and locations mentioned in a
certain paragraph more than a xed number of times as attributes. The resulting
pruned lattice is presented in Fig. 3.</p>
      <p>From this lattice it can be noticed that Europe and European Union are the
most discussed topics as it appeared in previous results. It is worth noticing that
2 GATE: General Architecture for Text Engineering: http://www.gate.ac.uk/
the United Nations (UN) is mentioned only in the context of Afghanistan, which
in its turn is mentioned solely in the context of NATO. Topics corresponding to
China and Ukraine form isolated concepts.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>The rst method based on frequent words allowed us to identify what questions
are raised most frequently by European and American leaders while talking
about Russia, whereas latent topic modeling allowed us to specify and describe
these issues more thoroughly. The results obtained with named-entity lattice are
rather similar. We concluded that it could be more useful to merge NER with
the LDA model, for instance, by taking named entities as a set of objects and
topics as a set of attributes. Expanding the corpus of the texts and combining
NER with the LDA model can be useful in testing the hypotheses addressed in
political studies.</p>
      <p>On the whole, building lattice-based taxonomies using various language
processing techniques is promissing for obtaining additional knowledge regarding
open and hidden intentions of political actors. What is important for social
sciences is that the knowledge is obtained without collecting any additional
information and by a \neutral" tool|through deriving deeper connections between
notions used by the actors in their texts.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Salton</surname>
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wong</surname>
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Yang</surname>
            <given-names>C.</given-names>
          </string-name>
          <article-title>A Vector Space Model for Automatic Indexing</article-title>
          .
          <source>Communications of the ACM</source>
          , vol.
          <volume>18</volume>
          , pp.
          <volume>613</volume>
          {
          <issue>620</issue>
          (
          <year>1975</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Steyvers</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gri</surname>
          </string-name>
          ts T.
          <article-title>Probabilistic Topic Models. Latent Semantic Analysis: A Road to Meaning</article-title>
          , Laurence
          <string-name>
            <surname>Erlbaum</surname>
          </string-name>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Jurafsky</surname>
            <given-names>D.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Martin</surname>
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Speech</surname>
            and
            <given-names>Language</given-names>
          </string-name>
          <string-name>
            <surname>Processing</surname>
          </string-name>
          .
          <article-title>An Introduction to Natural Language Processing</article-title>
          , Computational Linguistics and
          <string-name>
            <given-names>Speech</given-names>
            <surname>Recognition</surname>
          </string-name>
          , Prentice
          <string-name>
            <surname>Hall</surname>
          </string-name>
          (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Ganter</surname>
            <given-names>B.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Wille R. Formal Concept</surname>
            <given-names>Analysis</given-names>
          </string-name>
          ,
          <source>Mathematical Foundations</source>
          . Springer Verlag, Berlin (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Roth</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Obiedkov</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kourie</surname>
            ,
            <given-names>D.G.</given-names>
          </string-name>
          <article-title>Towards concise representation for taxonomies of epistemic communities</article-title>
          . In: Yahia,
          <string-name>
            <given-names>S.B.</given-names>
            ,
            <surname>Nguifo</surname>
          </string-name>
          , E.M. (eds.)
          <source>Proc. CLA 4th Intl. Conf. on Concept Lattices and their Applications</source>
          .
          <source>LNCS/LNAI</source>
          , vol.
          <volume>4923</volume>
          , pp.
          <volume>240</volume>
          {
          <fpage>255</fpage>
          . Springer (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Porter</surname>
            <given-names>M.</given-names>
          </string-name>
          <article-title>An algorithm for su x stripping</article-title>
          .
          <source>Program</source>
          , vol.
          <volume>14</volume>
          , pp.
          <volume>130</volume>
          {
          <issue>137</issue>
          (
          <year>1980</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Kuznetsov</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Obiedkov</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roth</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <article-title>Reducing the representation complexity of lattice-based taxonomies</article-title>
          . In: Priss,
          <string-name>
            <given-names>U.</given-names>
            ,
            <surname>Polovina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Hill</surname>
          </string-name>
          ,
          <string-name>
            <surname>R</surname>
          </string-name>
          . (eds.) Conceptual Structures:
          <article-title>Knowledge Architectures for Smart Applications: 15th Intl Conf on Conceptual Structures</article-title>
          ,
          <string-name>
            <surname>ICCS</surname>
          </string-name>
          <year>2007</year>
          ,
          <article-title>She eld</article-title>
          ,
          <source>UK. LNCS/LNAI</source>
          , vol.
          <volume>4604</volume>
          , pp.
          <volume>241</volume>
          {
          <fpage>254</fpage>
          . Springer (July
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Yevtushenko</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <article-title>System of data analysis \Concept Explorer". (In Russian)</article-title>
          .
          <source>Proceedings of the 7th national conference on Arti cial Intelligence KII-2000</source>
          , p.
          <volume>127</volume>
          {
          <issue>134</issue>
          ,
          <string-name>
            <surname>Russia</surname>
          </string-name>
          ,
          <year>2000</year>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Blei</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ng</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jordan</surname>
            <given-names>M. Latent</given-names>
          </string-name>
          <article-title>Dirichlet allocation</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          , vol.
          <volume>3</volume>
          : pp.
          <volume>993</volume>
          {
          <issue>1022</issue>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>