<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards Concept Identification using a Knowledge-Intensive Approach</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>O´ scar Mun˜oz-Garc´ıa</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andr´es Garc´ıa-Silva</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Oscar Corcho</string-name>
          <email>ocorcho@fi.upm.es</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Havas Media Group</institution>
          ,
          <addr-line>Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Ontology Engineering Group, Departamento de Inteligencia Artificial Universidad Polit ́ecnica de Madrid</institution>
          ,
          <addr-line>Madrid</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2013</year>
      </pub-date>
      <volume>1019</volume>
      <fpage>45</fpage>
      <lpage>49</lpage>
      <abstract>
        <p>This paper presents a method for identifying concepts in microposts and classifying them into a predefined set of categories. The method relies on the DBpedia knowledge base to identify the types of the concepts detected in the messages. For those concepts that are not classified in the ontology we infer their types via the ontology properties which characterise the type.</p>
      </abstract>
      <kwd-group>
        <kwd>concept identification</kwd>
        <kwd>microposts</kwd>
        <kwd>dbpedia</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>Spotting concepts</title>
      <p>The concept spotting stage analyses the micropost for extracting the keywords
that are candidates for being concepts, or can serve as context for disambiguating
the concepts. The stage is executed in three steps, namely:</p>
      <sec id="sec-2-1">
        <title>1. Text normalisation. 2. Part-of-speech tagging. 3. Keyword selection.</title>
      </sec>
      <sec id="sec-2-2">
        <title>Next, each of the steps is described.</title>
        <sec id="sec-2-2-1">
          <title>Text Normalisation</title>
          <p>
            The text normalisation step converts the text of the micropost, that often
includes metalanguage elements, to a syntax more similar to the usual natural
language. Previous results demonstrate that this normalisation step improves
the accuracy of the part-of-speech tagger [
            <xref ref-type="bibr" rid="ref3">3</xref>
            ]. Specifically, we have implemented
several rules for syntactic normalization of Twitter messages (some of them have
been described in [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ]). The rules executed are the following:
– Transform to lower-case the text completely written with upper-case
characters.
– Delete the sequence of characters “RT” followed by a mention to a Twitter
user (marked by the symbol “@”) and, optionally, by a colon punctuation
mark.
– Delete mentions to users that are not preceded by a coordinating or
subordinating conjunction, a preposition, or a verb.
– Delete the word “via” followed by a mention to a user at the end of the
tweet.
– Delete the hashtags found at the end of the tweet.
– Delete the “#” symbol from the hasthtags that are maintained.
– Delete the hyperlinks contained within the tweet.
– Delete ellipses that are at the end of the tweet, followed by a hyperlink.
– Delete characters that are repeated more than twice (e.g., “yeeeeeessss” is
transformed to “yes”).
– Transform underscores to blank spaces.
– Divide camel-cased words in multiple words (e.g., “AnalyticsTools” is
converted to “Analytics Tools”).
2.2
          </p>
        </sec>
        <sec id="sec-2-2-2">
          <title>Part-of-speech Tagging</title>
          <p>
            After normalising the micropost text, we execute the part-of-speech analysis of
the normalised text. For doing so, we make use of Freeling [
            <xref ref-type="bibr" rid="ref7">7</xref>
            ].
2.3
          </p>
        </sec>
        <sec id="sec-2-2-3">
          <title>Keyword Selection</title>
          <p>Once the part-of-speech tagging is obtained, the keyword selection step is
executed. For each sentence within the micropost text we extract all the possible
n-grams. In this case a gram is a word in the sentence. After that we select only
the n-grams that satisfy the following criteria:
– The n-gram contains at least one noun.
– The n-gram is not contained in a set of stop words.
– If the number of words included in the n-gram is greater than one, the n-gram
is included in the set of Wikipedia article titles.
– The n-gram is not contained in another n-gram that has been added to the
keyword set (longer n-grams prevail).</p>
          <p>
            To speed-up the process of querying the millions of Wikipedia article titles
we have uploaded the list of titles (available at [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ]) to a Redis store [
            <xref ref-type="bibr" rid="ref8">8</xref>
            ].
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Semantics of the Concepts</title>
      <p>
        To identify the semantics of the keywords we tap into the DBpedia knowledge
base [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] to elicit the types of the concepts to which the keywords correspond.
DBpedia contains knowledge from Wikipedia for close to 3.5 million resources;
1.6 million resources are classified in a cross domain ontology containing 272
classes. DBpedia strengths include its large coverage and the fact that its data
are exposed in RDF allowing to query them using SPARQL queries through the
available endpoint [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>Our process starts by identifying for each keyword the DBpedia resource
which represents its intended meaning. Once we have the corresponding resource
we query in DBpedia its classes, whenever they are available, or infer them
through the identification of specific resource properties, so that we can identify
the types defined in the challenge.
3.1</p>
      <sec id="sec-3-1">
        <title>From keywords to DBpedia resources</title>
        <p>First we query DBpedia for a resource with a label matching the keyword. We
use exact string matching between the resource label and the keyword which has
been previously modified to fit the style of article titles in Wikipedia. The output
resource of this query represents the most frequent meaning of the keyword
defined by Wikipedia editors. We call this resource default sense of a keyword.
If the resource is not related to a disambiguation resource, we consider that
the term is not ambiguous and therefore we use the default sense as the one
representing the keyword meaning.</p>
        <p>In case we do not find a match between the keyword and a resource label
we use an spelling service that suggests similar titles of Wikipedia articles. This
spelling service3 compares the n-grams based on characters of both keywords and
article titles, and takes into account the popularity of the articles in Wikipedia
(i.e., the times that an article has been linked from other articles) when
producing the final ranking of suggestions. We use the most similar suggestion, above
a given threshold, for searching for the DBpedia resource.</p>
        <p>If the resource is related to a disambiguation resource, then we have to select
the proper sense among the candidates. To do so we leverage the correspondence
of DBpedia resources with Wikipedia articles to obtain textual descriptions of
each resource. Thus, we calculate similarity between each resource and the term
by comparing the resource textual description with the term context. The most
similar resource is selected as the resource representing the term meaning. By
context we mean the set of keywords identified in the same sentence.</p>
        <p>
          To calculate similarity between the keyword context and the resource
description we use a vector space model. The components of the vectors are the
most frequent terms of the Wikipedia articles related to each candidate DBpedia
resource. To populate the vectors representing resources we use term frequency
3 The spelling service was built upon Lucene [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] spell checker.
and inverse document frequency (TF-IDF) as term weighting scheme. IDF is
calculated using only the set of textual descriptions corresponding to the candidate
resources. We calculate the cosine of the angle between the vector representing
the keyword context with each of the vectors representing candidate resources.
The candidate with the highest cosine is selected as the resource to represent
the keywords. Details of this procedure can be found in [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ].
        </p>
        <p>In short for ambiguous keywords if there is not context and there is a default
sense we select the DBpedia resource corresponding to the default sense. If there
is context and default sense, and the context do not overlap with any of the
candidate vectors we use the DBpedia resource corresponding to the default
sense too. If there is overlap between the context and candidate vectors we use
the most similar candidate.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Identifying concept types</title>
        <p>We manually select the classes from DBpedia and linked ontologies that allow
us to identify the types of the concepts defined in the challenge. For instance,
– dbpedia-owl:Person and foaf:Person are the classes for People;
– dbpedia-owl:Place is the class for Location;
– dbpedia-owl:Organisation, dbpedia-owl:Company, and umbel:Organization are
the classes for Organizations;
– dbpedia-owl:ProgrammingLanguage, umbel:SoftwareObject, dbpedia-owl:Film,
dbpedia-owl:TelevisionShow, dbpedia-owl:Award, and dbpedia-owl:SportsEvent,
are the classes for the Miscellaneous type.</p>
        <p>Therefore, for each DBpedia resource we obtain its class from the ontology
and classify it according to the challenge types.</p>
        <p>However, many DBpedia resources are not classified in the ontology. For
those resources we infer its type from certain properties which are characteristic
of the type. For instance, from a triple
&lt;subject&gt; dbpedia-owl:birthPlace &lt;object&gt;</p>
        <p>we can infer that object is a location given that it is the birth place of the
subject described in the triple. The same rationale can be used with predicates
such as dbpedia-owl:hometown and dbpedia-owl:location. Similarly, from a triple
&lt;subject&gt; dbpprop:mvp &lt;object&gt;</p>
        <p>we can infer that the subject is an sport event since it has a most
valuable player. Other predicates used for identifying sport events include
dbpprop:menDraw, dbpprop:teams, dbpprop:sport, and dbpprop:referee.</p>
        <p>Finally, in case we cannot identify the concept type using DBpedia, we use
a list of concepts and their types which have been collected from the training
data set. From this list we take the first type associated with that concept.</p>
        <p>We have not included evaluation results in this extended abstract since the
only available source of annotated data for the evaluation in this challenge was
the the training data set (the test data set was not annotated). Given that our
approach uses a list of concepts gathered from the training set is not fair to
report evaluation results on this data set.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgements</title>
      <p>This research is partially supported by the Spanish Centre for the Development
of Industrial Technology under the CENIT program, project CEN-20101037,
“Social Media” (http://www.cenitsocialmedia.es).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Apache</given-names>
            <surname>Software</surname>
          </string-name>
          <article-title>Foundation: Apache Lucene</article-title>
          . http://lucene.apache.org (
          <year>2013</year>
          ), [Online; accessed 23-May-2013]
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kobilarov</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Becker</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cyganiak</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hellmann</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>DBpedia - A crystallization point for the Web of Data</article-title>
          .
          <source>Journal of Web Semantic</source>
          <volume>7</volume>
          (
          <issue>3</issue>
          ),
          <fpage>154</fpage>
          -
          <lpage>165</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Codina</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Atserias</surname>
          </string-name>
          , J.:
          <article-title>What is the text of a tweet? In: Proceedings of @NLP can u tag #user generated content?! via lrec-conf.org</article-title>
          . ELRA, Istanbul, Turkey (May
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. DBpedia:
          <article-title>DBpedia SPARQL endpoint</article-title>
          . http://dbpedia.org/sparql (
          <year>2013</year>
          ), [Online; accessed 23-May-2013]
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Garc</surname>
          </string-name>
          <article-title>´ıa-</article-title>
          <string-name>
            <surname>Silva</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cantador</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corcho</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Enabling folksonomies for knowledge extraction: A semantic grounding approach</article-title>
          .
          <source>International Journal on Semantic Web and Information Systems (IJSWIS) 8</source>
          (
          <issue>3</issue>
          ),
          <fpage>24</fpage>
          -
          <lpage>41</lpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. Kaufmann,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Jugal</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.</surname>
          </string-name>
          :
          <article-title>Syntactic normalization of twitter messages</article-title>
          .
          <source>In: Proceedings of the International Conference on Natural Language Processing (ICON-2010)</source>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. Padro´,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Stanilovsky</surname>
          </string-name>
          , E.:
          <article-title>Freeling 3.0: Towards wider multilinguality</article-title>
          .
          <source>In: Proceedings of the Language Resources and Evaluation Conference (LREC</source>
          <year>2012</year>
          ). ELRA, Istanbul, Turkey (May
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Sanfilippo</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Redis. http://redis.io (
          <year>2009</year>
          ), [Online; accessed 23-May-2013]
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9. Wikipedia: Wikipedia:
          <article-title>Database download</article-title>
          . http://en.wikipedia.org/wiki/ Wikipedia:Database_download (
          <year>2013</year>
          ), [Online; accessed 23-May-2013]
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>