<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Attribute Selection Techniques for Classi cation of Aggressive Tweets</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Gabriela Ram rez-de-la-Rosa</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Esau Villatoro-Tello</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hector Jimenez-Salazar</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Language and Reasoning Research Group, Information Technologies Department, Universidad Autonoma Metropolitana Unidad Cuajimalpa</institution>
          ,
          <addr-line>Mexico City</addr-line>
          ,
          <country country="MX">Mexico</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <fpage>515</fpage>
      <lpage>519</lpage>
      <abstract>
        <p>This paper describes the participation of our Research Group in the shared task of MexA3T 2019. We evaluated the impact of using as principal features, the set of terms identi ed as discriminant by distinct feature selection strategies. Our main goal was to test if a condensed set of words can be indicative of the aggressiveness of a short text written in a very informal setting (i.e., tweets). Our experiments indicate that di erent feature selection techniques favor di erent aspects of the aggressiveness in a short text.</p>
      </abstract>
      <kwd-group>
        <kwd>Feature selection Aggressiveness identi cation Text classi cation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Hate speech, harassment, and cyberbullying are a few examples of how social
media can negatively a ect groups of people online. According to [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], aggression
in social media is targeted to a particular person or group aiming to damage their
identity or lowering their prestige. Previous research proposes that aggression
is often expressed in two ways: directly expressed or hidden in the posts [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ],
resulting in a very challenging task to be performed automatically. Additionally,
aggressiveness depends heavily on cultural aspects, as well as the local context,
hence, the necessity of building automatic systems for languages di erent than
English, and culturally oriented is becoming more relevant.
      </p>
      <p>
        Accordingly, in this paper, we describe our proposed system for aggressiveness
identi cation in Mexican-Spanish tweets. Speci cally, we describe our
participation in the second edition of the MexA3T 2019 challenge [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. For this particular
task, we were given a set of 7700 training tweets, labeled as being aggressive
and non-aggressive. Our approach replicates the traditional pipeline of a
nonthematic text classi cation system. However, contrastingly to previous research
using the same dataset, our main goal was to evaluate the impact of distinct
feature selection strategies. Thus, we evaluated if a highly condensed set of words
are capable of providing to the learning algorithms with valuable and
discriminant information solving the posed task. An initial analysis of the obtained
features indicates that a distinct aspect of the aggressiveness can be detected by
each of the feature selection strategies.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>System description</title>
      <p>Our main goal was to evaluate if a condensed set of terms, obtained through a
feature selection strategy, would provide enough information to identify
aggressiveness in a tweet. To that purpose, we used the traditional pipeline for text
classi cation. First step is the preprocessing, we lowercased all texts, removed
numbers and symbols (except for exclamation points), we used a common word
to indicate a hashtag, and all emojis where described1 in text format. Second,
we extracted a set of attributes to generate a representation based on the vector
model. In this phase we tested the three strategies described below for selecting
a relative small set of attributes. Third, we use a learning algorithm to classify
tweets into two classes: aggressive and non-aggressive.</p>
      <p>Next, we describe three di erent strategies for attribute selection. Two of
them are widely used in many text classi cation task, however the third one,
to the best of our knowledge, has not been used as feature selection technique
before.
2.1</p>
      <sec id="sec-2-1">
        <title>Document frequency</title>
        <p>Document frequency is a very simple approach for selecting informative terms,
aiming at lowering the dimensionality of vector representations. The idea behind
this technique is to ignore terms that are used in very few documents (hapax) or
used in almost all documents (known as corpus-speci c stop words), assuming its
importance is locally to only those texts or they are not useful for distinguishing
between classes because those terms appeared in all documents, respectively;
hence, we ignore terms not tting on one of this criteria.</p>
        <p>For our experiments we used a grid search to nd the best value for these
thresholds. We search for both high and low thresholds (high indicate we ignore
terms appearing in more than h% of all documents; and low indicate that we
ignore terms appearing in less than l% of all documents). The best empirical
result was obtained when h = 80% and l = 0%.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Mutual information</title>
        <p>
          From information theory, mutual information (MI) of two random variables x
and y can be describe as the reduction in the uncertainty of variable x due to
1 We used an emoji library for python: https://github.com/carpedm20/emoji/
the knowledge of y [
          <xref ref-type="bibr" rid="ref3 ref6">3, 6</xref>
          ]. In other words, MI measures the relevance of a feature
x to predict the class y.
        </p>
        <p>Thus, we computed the MI value for each feature, and those with the greater
MI value were used as attributes. For our experiments we selected n top features
with greater MI, the best performance was obtained with n = 1000.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Lexical availability</title>
        <p>
          This technique is a linguistically motivated approach aiming to identify those
lexical markers that represent the words springing to mind in response to a
speci c topic. Lexical Availability (LA) measures the ease with which a word is
generated in a given communicative situation, and allows to obtain the mental
lexicon which represents the vocabulary ow usable of a person [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. The terms
with greater LA can be seen as the most important ones for a set of tweets from
the same class. Thus, we computed the mental lexicon for each class, and then,
we used the resulting set of combining both lexicons as features.
        </p>
        <p>In our experiments, for combining the mental lexicon for each class we used
the union, intersection and symmetric di erence of both lists. However, the best
performance was obtained when we used the union of both lexicons to generate
the nal vocabulary. We also search for the n top attributes with the higher
value of lexical availability before combining the lists. The best performance was
obtained with n = 1000 and n = 3000.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>
        During the development phase of the MexA3T challenge, a training set of tweets
was released. The total number of instances for the aggressive class was 2727,
and 4973 for no-aggressive tweets, for a total of 7700 tweets in the training set.
A better description of the data can be found in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>For validating our experiments in this stage we performed a 10 fold cross
validation technique. We used the F-score to measure the best performances;
although we use the macro average score we also reported the F-score for the
positive class (aggressive) and for the negative class (non-aggressive).</p>
      <p>Furthermore, we used several classi ers along the development phase but we
found consistent performances with Nave Bayes classi er. Thus, for all of our
reported results we used the scikit-learn2 implementation of this classi er with
the default parameters.</p>
      <p>Table 1, under the validation phase column, shows the best performance for
all the con gurations tested, using the three di erent feature selection techniques
with some variations as described in Section 2. We also included the bag-of-word
(BOW) with the top 1000 most frequent features. As we can see from the table,
all feature selection strategies showed similar performances. A slighter better
Fmacro was obtained with document frequency (DF) and lexical availability (LA
2 https://scikit-learn.org/stable/modules/naive bayes.html
with n=3000). However, note that while DF performed similar to LA (n=3000)
the size of the vocabulary for DF was approximately 15000 for each fold; while
the vocabulary size for LA with n=3000 was approximately 4500 for each fold.</p>
      <p>Validation phase</p>
      <p>Test phase
BOW (n=1000)
DF (h=80%)
MI (n=1000)
LA (n=1000)
LA (n=3000)
Ensamble</p>
      <p>F+</p>
      <p>For the evaluation stage of the MexA3T shared task a set of 3156 tweets were
given. The nal results are shown in Table 1 under the test phase column. For
this stage, we also included an ensemble con guration, which assigns the class
corresponding to the majority vote of the ve submissions. As we can see, similar
as in the validation stage, we obtained the best f-macro with document frequency
(DF) and lexical availability (LA). However, the performance for the positive
class (aggressive) decreased, for all submissions, in about 30% in comparison to
the validation data.</p>
      <p>To have a better idea of the amount of unique information given by the three
feature selection strategies, Table 2 indicates the size of unique sets of attributes
when comparing pairs of strategies. To make a fair judgment, we examined sets
with the same sizes (i.e. 1000). The last row in this table shows, on the one
hand, that lexical availability (LA) has 191 unique attributes compared with the
set obtained by the bag-of-words (BOW) vocabulary. On the other hand, only
109 attributes are unique in the set of BOW compared against the LA method.
Similarly, when LA is compared against the MI technique, the total number
of unique attributes is 212 for LA but only 131 for MI. This initial analysis
illustrates that the LA method is capable of nding more diverse attributes in
general.
0
179
191</p>
      <p>MI</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusion and future work</title>
      <p>This paper describes our participation in the shared task of MexA3T 2019
for identify aggressive tweets written in Mexican-Spanish. We used the general
pipeline for text classi cation varying the strategy used for feature selection. Our
goal was to test if a condensed set of words could be indicative of the
aggressiveness of a short text. Our experiments indicated that di erent feature selection
techniques favor di erent aspects of the aggressiveness in a short text.</p>
      <p>As feature work we plan to performed a deeper analysis of the set of attributes
selected by the tested strategies. Our initial analysis indicates that lexical
availability can extract set of attributes with a di erent linguistic meaning as opposed
to bag-of-words or mutual information.</p>
      <sec id="sec-4-1">
        <title>Acknowledgements.</title>
        <p>We thank UAM Cuajimalpa and CONACyT (project grant CB-2015-01-258588)
for their support. The rst author thanks INAOE for the facilities given in her
research visit.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Mario</given-names>
            <surname>Ezra</surname>
          </string-name>
          <string-name>
            <surname>Aragon</surname>
          </string-name>
          ,
          <article-title>Miguel A Alvarez-Carmona, Manuel Montes-y-</article-title>
          <string-name>
            <surname>Gomez</surname>
          </string-name>
          ,
          <article-title>Hugo Jair Escalante, Luis Villasen~or-Pineda, and Daniela Moctezuma. Overview of MEX-A3T at IberLEF 2019: Authorship and aggressiveness analysis in mexican spanish tweets</article-title>
          .
          <source>In Notebook Papers of 1st SEPLN Workshop on Iberian Languages Evaluation Forum (IberLEF)</source>
          , Bilbao, Spain, September,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Rosa</given-names>
            <surname>Mar</surname>
          </string-name>
          <article-title>a Jimenez Catalan. Lexical availability in English and Spanish as a second language</article-title>
          , volume
          <volume>17</volume>
          . Springer,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Thomas</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Cover</surname>
          </string-name>
          and Joy A. Thomas.
          <source>Elements of Information Theory. WileyInterscience</source>
          , New York, NY, USA,
          <year>1991</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Jonathan</given-names>
            <surname>Culpeper</surname>
          </string-name>
          .
          <article-title>Impoliteness: Using language to cause o ence</article-title>
          , volume
          <volume>28</volume>
          . Cambridge University Press,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Sreekanth</given-names>
            <surname>Madisetty</surname>
          </string-name>
          and
          <article-title>Maunendra Sankar Desarkar</article-title>
          .
          <article-title>Aggression detection in social media using deep neural networks</article-title>
          .
          <source>In Proceedings of the First Workshop on Trolling, Aggression and Cyberbullying (TRAC-2018)</source>
          , pages
          <fpage>120</fpage>
          {
          <fpage>127</fpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Brian C Ross</surname>
          </string-name>
          .
          <article-title>Mutual information between discrete and continuous data sets</article-title>
          .
          <source>PloS one</source>
          ,
          <volume>9</volume>
          (
          <issue>2</issue>
          ):e87357,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>