<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The TALP-UPC approach to Tweet-Norm 2013</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>TALP Research Center Universitat Politecnica de Catalunya</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes the methodology used by the TALP-UPC team for the SEPLN 2013 shared task of tweet normalization (Tweet-Norm). The system uses a set of modules that propose di erent corrections for each out-of-vocabulary word. The nal correction is chosen by weighted voting according to each module accuracy.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The increasing use of social networks to
brie y express opinions and facts is
leading to large amounts of text written with
misspellings and neologisms, such as the
case of tweets. The SEPLN 2013
TweetNorm shared task focuses on the
evaluation of approaches useful for normalizing
out-of-vocabulary words occurring in
Spanish tweets, similar to the previous works such
as
        <xref ref-type="bibr" rid="ref1">(Han and Baldwin, 2011)</xref>
        for English and
        <xref ref-type="bibr" rid="ref3 ref4">(Mosquera, Lloret, and Moreda, 2012)</xref>
        for
English and Spanish. In this paper we
describe the UPC system for this task and the
results achieved.
      </p>
    </sec>
    <sec id="sec-2">
      <title>Our approach</title>
      <p>The UPC system for SEPLN 2013
TweetNorm shared task consists of a collection of
expert modules, each of which proposes
corrections for out-of-vocabulary (OOV) words.
The nal decision is taken by weighted
voting according to each expert accuracy on the
development corpus.</p>
      <p>First, a preprocessing step is applied,
where consecutive occurrences of the same
letter are reduced to one (except valid
Spanish digraphs like rr or ll). We generate
also a version of the OOV with those
repetitions reduced to two occurrences (to
capture cases such as coordinar, leed, accion,
etc.). In this way, we obtain three di erent
OOV versions (original, reduction to one
repeated letter, reduction to two repeated
letters) that will be checked against dictionaries
and gazetteers as described below.</p>
      <p>
        All expert modules are implemented
using FreeLing
        <xref ref-type="bibr" rid="ref3 ref4">(Padro and Stanilovsky, 2012)</xref>
        library facilities for dictionary access,
multiword detection, or PoS tagging. Some
experts use string edit distance (SED) measures
to nd words in a dictionary similar to the
target OOV. FOMA library
        <xref ref-type="bibr" rid="ref2">(Hulden, 2009)</xref>
        is used in these cases for fast retrieval of
candidates.
      </p>
      <p>The used expert modules can be divided
in three classes:</p>
      <p>Regular-expression experts:
Experts in this class are regular
expression collections that propose corrections
for recurring patterns or words, such as
smileys, laughs (e.g. jajjaja, jeje,
etc.), frequent abbreviations (e.g. TQM !
te quiero mucho, xq ! porque, etc), or
frequent mistakes (e.g. nose ! no se).
Experts in this category propose a xed
solution for each case.</p>
      <p>Single-word experts: Each module
belonging to this class uses a speci c
single-word lexical resource and a set
of string edit distance (SED) measures
to nd candidates similar to the target
word. The three SED measures
specifically used for the task are: character
distance (the conventional edit distance
metric between strings), phonetic
distance (transformations according to
similarity in pronunciation) and keyboard
distance (transformations due to
possible errors when typing).</p>
      <p>Multi-word experts: Modules in this
category take into account the context
where an OOV is located to select the
best candidate among those proposed by
the other experts. We used three di
erent experts in this category. First, the
multiword dictionary module takes into
account proposals of the single-word
experts that use di erent distances over
a dictionary consisting only of tokens
that appear in known multiwords. All
combinations of possible candidates for
the OOV and its context are checked
against the multiwords dictionary, and
those matching an entry are suggested
as corrections. Second, the PoS tagger
expert takes into account all proposals
of all single-word experts, retrieves the
possible PoS tags for each of them, and
creates a virtual token with a
morphological ambiguity class including all
obtained categories. Then, a PoS tagger is
applied, and the best category for each
OOV is selected. The module lters out
all proposals not matching the
resulting tag, and produces as candidates only
those with the selected category.
Finally, the glued words expert, which
consists of a FSM that recognizes the
language L( L)+, where L is the language
of all valid words in the Spanish
dictionary. Using foma-based SED search on
this FSM with an appropriate cost
matrix, we can obtain, for instance, that
lo siento is the word in the FSM
language closer to losiento, and propose
it as a candidate correction.</p>
      <p>All the resources used in these experts are
brie y enumerated in Section 3.</p>
      <p>After all experts have been applied, a
selection function is used on the set of resulting
candidates. This selection function takes into
account the SED distance of each proposal to
the original OOV, the number of experts that
proposed it, and the precision, recall, and F1
of each expert on the development corpus to
perform a weighted voting and select the nal
correction. Di erent experiments with di
erent functions are reported in Section 4.
3</p>
    </sec>
    <sec id="sec-3">
      <title>The set of lexical resources</title>
      <p>In this section, we describe the di erent
lexical resources we have employed in order to
provide correct candidates for OOV words.
Some of these resources are merely looked
up with exact search, whilst for others we
have considered useful to perform
approximate search as well, using the SED metrics
described in the previous section. Note, in
addition, that the unnanotated tweets
provided by the organization have not been used
to enrich our resources.</p>
      <p>Next, we list each of these resources as
well as the types of searches for which they
are used.
3.1</p>
      <sec id="sec-3-1">
        <title>Resources for regular-expression experts</title>
        <p>Gazetteer of acronyms: List of di
erent sorts of acronyms. It also includes
several abbreviations frequently used in
tweets and other short messages. Used
only for exact searches.</p>
        <p>Gazetteer of emoticons: List of
emoticons, some of them expressed as
regular expressions. We just deal with those
emoticons which are not composed by
only punctuation signs, since the other
ones would be accepted by FreeLing,
and consequently they will not be OOV
words.</p>
        <p>Gazeteer of onomatopoeias: List of
onomatopoeias, many of them expressed
as regular expressions. The RAE
dictionary is used as the reference for the
correct normalization of each candidate
onomatopoeia found.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Resources for single-word experts</title>
        <p>Spanish dictionary: List of Spanish
words, according to FreeLing dictionary.
The three types of SED metrics are
performed on it.</p>
        <p>English dictionary: List of English
words, according to FreeLing dictionary.
The three types of SED metrics are also
performed on it.</p>
        <p>Spanish dictionary expanded with
morphological derivates: A set of
morphological derivates has been
generated for the words in the Spanish
dictionary. The speci c derivates have
been applied according to the PoS of
each word. Concretely, superlatives
and diminutives have been generated for
nouns, adjectives, adverbs and
participles, and enclitic pronouns have been
su xed to in nitive and imperative
verbal forms, as well as to gerunds.
However, due to the high volume of
generated alternatives, a previous lter based
on commonness and length of the words
has been performed on the dictionary
and only the derivates for the resultant
words have been generated. On this
resource, the exact search and the three
types of SED metrics are performed.
Gazetteer of names: List of person
names (including also certain
diminutives). Both the exact and the SED
metrics are carried out on it.</p>
        <p>Uniwords NE gazetteer: It comprises
a far from exhaustive list of proper nouns
such as di erent types of locations,
companies, artists and other personalities,
TV channels and programs, products,
newspapers and media groups or even
shopping centers. As mentioned in the
previous section, this gazetteer is used
both as a preliminary search for the
elements of the multiwords gazetter and as
a gazetter in itself. In the latter case,
exact search and SED metrics for
approximate searches are performed on it.
3.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Resources for multi-word experts</title>
        <p>Multiwords NE gazetteer: It
comprises a far from exhaustive list of proper
nouns composed by more than one word,
belonging to the same categories
mentioned for the uniwords gazetteer.
Neither exact search nor SED metrics are
performed directly on it. As mentioned
in the previous section, they are
performed in the uniwords gazetteer.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experiments on di erent functions for the best candidate</title>
      <p>Combining the experts from Section 2 and
the lexical resources described in Section 3,
we obtain a total of 32 di erent producers
that are integrated in our tweet normalizer.
Additionally, we add a 33rd producer that
always proposes to leave the target OOV as it
is. The combined outputs of these producers
yield several hundreds of spelling alternatives
for the OOV words, therefore we need a
principled method to choose the best one among
them, including to leave the original word as
it is given. This strategy is able to propose
the correct spelling alternative to 89.42% of
the OOVs found in the development corpus,
therefore, this is the upper-bound accuracy
of our system.</p>
      <p>Using the development corpus, we have
computed the precision, recall and F1 of each
producer. Since the producers yield a list of
spelling alternatives that are sortable
according to the SED metrics, we have devised three
di erent levels where we can measure its
condence:</p>
      <p>TopN: At this level, we check only if the
producer produces the correct correction
anywhere in the alternatives list.</p>
      <p>Top1: This level checks how many times
the correct correction has the smallest
SED in the whole list of alternatives (i.e.,
it is on the front of the list). In this
case, the precision is computed against
the total number of proposed corrections
having the smallest SED.</p>
      <p>Top0: This measures how many times
the correct correction has a SED
distance of zero over the total number of
proposals at distance zero. Note that
all exact searches (regular-expression
experts and look up dictionaries) yield
alternatives with distance zero.</p>
      <p>We compute precision, recall and F1 for
each producer for all three levels of measure.</p>
      <p>To produce a proposal for each OOV, we
implement a voting scheme. Each producer
votes for each of their proposed corrections
using the suitable TopN, Top1 or Top0 scores
as their vote weight. The possible
corrections are pooled together and the one with
the largest total score is our nal proposal.</p>
      <p>Note that a proposed correction in the
Top0 position is also in the Top1 and the
TopN positions. Therefore, we can choose
if the weight of a producer vote is just the
score of it's best measure (e.g. Top1 instead
of TopN) or the addition of all suitable
measures (e.g. Top1 plus TopN for a proposal in
Top1). We have experimented with these two
scheme
single
additive</p>
      <p>single
additive</p>
      <p>single
additive
single
single
single
additive
voting schemes that we call single or additive
and with using precision (P), recall (R) or F1
as the actual vote weight. We have also
considered the possibility of squaring the weights
in order to strengthen the relevance of high
precision producers. Table 4 shows the
results achieved using the development corpus
for testing and estimating the weights. We
have set up some baselines giving xed weight
to the votes. We have the scheme w111,
which gives a weight of 1 to TopN, Top1 and
Top0; the scheme w110 gives a weight of 0 to
TopN and of 1 to the other two, and so on.</p>
      <p>The experiments show that using squared
precision as the con dence measure within an
additive scheme yields the best results: a
precision of 69.81% on the development corpus.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <p>The o cial result of our single run is an
accuracy of 65.26% on a test corpus of 500 tweets
containing roughly 700 OOVs. In this run we
use the weights estimated from the
development corpus. This results is 4.5 points behind
what we obtained on the development
corpus, suggesting that our estimation method
is reasonable but may be over tting.</p>
      <p>To elucidate this issue, we have repeated
our experiment using the test gold standard
to estimate the vote weights (instead of using
the development data). With this setup and
identical voting scheme, the precision is
increased by 0.76 points, less than four points
behind the 69.81% we got in the development
set. Additionally, we have used the gold
standard to calculate the upper-bound of our
producers as we did with the development data.
We are able to propose the correct word for
85.47% of the OOVs, which is 4 poins behind
the 89.42% for development data.</p>
      <p>Since little improvement is obtained when
using the test data, this suggests that our
strategy of estimating each producers'
precision is not over tting. Additionally, we can
see how the drop in the system's upper-bound
matches its accuracy drop. Therefore, we
believe that the nature and distribution of
OOVs in Twitter streams may vary over time
more than it is represented on the the
development set, thus, our strategy as a whole is
more suited to this particular set of
development data than to the test data.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusions and future work</title>
      <p>In this paper, we have described our
approach for the SEPLN 2013 Tweet-Norm
shared task, which consists in normalizing a
set of prede ned out-of-vocabulary words
occurring in tweets. Our system is based on a
voting schema that combines 33 di erent
experts on OOV candidate selection, each one
using a speci c viewpoint de ned by a
particular pair of edit distance similarity metric
and lexical resource. This approach achieved
a precision of 65.26% in the test corpus,
ranking our system in the 3rd best place among
the participants. This result show the
appropriateness of our approach for the task.
However, it is far to achieve the upper-bound
results (i.e., from the 69.81% achieved for the
development corpus to the upper-bound of
85.47% achievable in that corpus). This fact
shows that there is room enough to improve
our system.</p>
      <p>In order to get improvement, main lines in
our future work involve enriching the lexical
resources with OOV words occurring in the
unnanotated tweets provided by the
organizers, using a richer context of the OOV words
to drop out false candidates, tuning the costs
of the edit distances operators, and
considering other alternative voting schemes.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This research has been partially funded by
the Spanish research project SKATER
(TIN2012-38584-C06-01) and the EU FP7
Programme project XLike (FP7-ICT-2011.4.2).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Han</surname>
            , Bo and
            <given-names>Timothy</given-names>
          </string-name>
          <string-name>
            <surname>Baldwin</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Lexical normalisation of short text messages: Makn sens a #twitter</article-title>
          .
          <source>In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics (ACL</source>
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Hulden</surname>
          </string-name>
          , Mans.
          <year>2009</year>
          .
          <article-title>Fast approximate string matching with nite automata</article-title>
          .
          <source>Procesamiento del Lenguaje Natural</source>
          , (
          <volume>43</volume>
          ):
          <volume>57</volume>
          {
          <fpage>64</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Mosquera</surname>
            , Alejandro,
            <given-names>Elena</given-names>
          </string-name>
          <string-name>
            <surname>Lloret</surname>
            , and
            <given-names>Paloma</given-names>
          </string-name>
          <string-name>
            <surname>Moreda</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Towards facilitating the accessibility of web 2.0 texts through text normalisation</article-title>
          .
          <source>In Proceedings of the LREC 2012 Workshop: Natural Language Processing for Improving Textual Accessibility (NLP4ITA).</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Padro</surname>
            , Llu s
            <given-names>and Evgeny</given-names>
          </string-name>
          <string-name>
            <surname>Stanilovsky</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Freeling 3.0: Towards wider multilinguality</article-title>
          .
          <source>In Proceedings of the Language Resources and Evaluation Conference (LREC</source>
          <year>2012</year>
          ). ELRA.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>