<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>UACh-INAOE at HASOC 2019: Detecting Aggressive Tweets by Incorporating Authors' Traits as Descriptors</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marco Casavantes</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roberto Lopez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luis Carlos Gonzalez-Gurrola</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Manuel Montes-y-Gomez</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Facultad de Ingenier a, Universidad Autonoma de Chihuahua (UACh)</institution>
          ,
          <country country="MX">Mexico</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Laboratorio de Tecnolog as del Lenguaje, Instituto Nacional de Astrof sica</institution>
          ,
          <addr-line>Optica y Electronica (INAOE)</addr-line>
          ,
          <country country="MX">Mexico</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we describe our participation for the Aggressiveness Detection Track in English texts for HASOC 2019. We evaluate di erent strategies for text classi cation, including classi ers such as Logistic Regression and Support Vector Machines trained on n-grams (words and characters) and word embeddings for clustering techniques. We also study the incorporation of contextual characteristics to explore whether people verbally attack di erently depending on their traits and environment.</p>
      </abstract>
      <kwd-group>
        <kwd>English text classi cation Aggressiveness Detection Twitter</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>As people increasingly communicate online through social media, they may deal
with negative experiences such as being targets of cyberbullying or expose
themselves to hateful and vulgar content. These problems have become more relevant
in the past few years, as they pose several challenges to preserve the freedom
of speech and sharing of ideas over these communication channels. The growth
in the volume of the messages that are posted on social media on a daily basis
demands more e cient means to detect and moderate the spread of o ensive
content and hate speech. Furthermore, administrators of social media platforms
could prevent abusive behavior and harmful experiences. It is crucial to address
the importance of early identi cation of users that promote hate speech, as this
could enable important outreach programs, to prevent an escalation from speech
to action [11]. Moreover, considering the high levels of aggressiveness and hostile
behaviour of certain users towards particular groups or individuals, more serious
real-life issues, like self-harm or suicide, could actually be prevented.</p>
      <p>In the last years, several shared tasks have been organized with the purpose
of attracting attention to these problems [14, 12, 7, 13]. Take for instance the
second edition of MEX-A3T [4]. In that event, our participation focused on
detecting aggressive tweets in a Mexican Spanish dataset, by incorporating traits
of authors (e.g., occupation, location). Therefore, by participating in HASOC [9]
(Sub-task A for English), we aimed to test our approach on a di erent collection
of tweets, tweaking our system to face this new challenge.</p>
      <p>In this study, we evaluate common strategies such as lexical feature
engineering through term frequency representations (e.g., bag of words through t df ),
along with di erent approaches with the aim to enhance features by adding
context to each document. Furthermore, we also advanced our research by including
the authors' traits, and using the outcome of unsupervised methods as potential
useful features.</p>
      <p>The hypothesis behind our approach is that o ensive messages could be
better recognized by analyzing not only the message but the user pro le. The rest
of this document is organized as follows: in section 2 we describe our approach;
in section 3, the results attained are detailed and analyzed; nally, in section 4
we state our conclusions and delineate some future work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Proposed Method</title>
      <p>Similar to our participation in MEX-A3T 2019 [5], we aim to enrich the
classication of aggressive tweets by including a possible theme to which each tweet
belongs, being this the main experiment that attempt to support our
hypothesis. This section gives a complete description of the changes and adaptation of
features that we propose in our approach.
2.1</p>
      <sec id="sec-2-1">
        <title>Data Pre-processing</title>
        <p>Once the text les were loaded using UTF-8 encoding, we conducted our
experiments in a custom version of the dataset where:
{ All words are made lowercase.
{ Emojis are converted into their text representation.</p>
        <p>(e.g., \:face with tears of joy:")
{ Tweets are stripped from non-alphanumeric characters excluding some
relevant symbols (#, @ and ).
{ Every URL (occurrence of the sequence \http") was replaced with \weblink"
to evenly represent references to external sources.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Features</title>
        <p>We conducted our research using the following features:
Lexical: We use both word n-grams (n=1, 2) and char n-grams (n=2, 3, 4),
however this collection of terms was only weighted with its term frequency.
Document Embeddings: Using only the text available in both the train and
test set, we employed a representation of the tweets through Word Embeddings
[8] to feed di erent clustering strategies.</p>
        <p>Grouping tweets by theme: We use di erent clustering methods (an
implementation of Self Organizing Maps [1], K-Means and A nity Propagation) to
generate new features based on thematic terms in each tweet.</p>
        <p>{ The SOM allowed us to locate each tweet on a two-dimensional plane, taking
the coordinates as new features.
{ Using K-Means and A nity Propagation we calculate, for every sample, the
distance between itself and the rest of clusters.</p>
      </sec>
      <sec id="sec-2-3">
        <title>Flesch Reading Ease and Flesch Kincaid Grade scores: Based on [6], we</title>
        <p>wanted to capture the quality of each tweet by getting the Flesch Reading Ease
and Kincaid Grade scores using textstat [3]. In our experiments the number of
sentences is also xed at one.</p>
      </sec>
      <sec id="sec-2-4">
        <title>Named Entity Recognition (NER) counters: Upon manual inspection of</title>
        <p>frequent tokens (Table 4), we observed that a big part of the dataset included
references to people like Donald Trump (current president of USA), Boris
Johnson (current Prime Minister of the United Kingdom), Mahendra Singh Dhoni
(indian international cricketer) and organizations like ICC (International Cricket
Council). Based on this information we decided to incorporate counters of how
many persons, organizations and locations were mentioned in each text using
polyglot [2].
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Experiments and Results</title>
      <p>The datasets were provided by the HASOC-2019 organization team. Table 1
shows the distribution of training and test partitions for English tweets.</p>
      <p>We started our research by recreating our baselines used in MEX-A3T 2019,
this time focusing on the word unigrams and bigrams baseline, as it holds the
best performance in this task in comparison to the character n-grams baseline.
In order to generalize our results for the test set, we evaluated our experiments
using two di erent con gurations, a single strati ed train-validation split and a
5-Fold Cross Validation.</p>
      <p>We trained Linear Support Vector Machines and a Logistic Regression classi er
for this task, and we decided to use both of them to submit our predictions:
{ Run 1 consists of a LinearSVM trained with the best 800 features from
a Bag of Words of range=(1,2) considering the term frequency of all the
tokens. The feature selection was done by a chi-squared statistics test on a
70-30% train-validation split.
{ Run 2 is the same as Run 1, but in this case the top 1250 features were
selected from a strati ed 5-fold cross validation on the train set, specifying
a 20% split for the validation set.
{ Run 3 is the result of creating an ensemble of two Logistic Regression
classi ers, one trained with a Bag of Words and the other one with a Bag of
Character n-grams. The predictions were assigned by choosing the model
with the highest probability for each tweet.
As stated before, a Linear Support Vector Machine was chosen as our system's
classi er adding Named Entity Recognition counters for runs 1 and 2, and a
Logistic Regression classi er ensemble was used to submit run 3. Table 3 lists
the results of our three submissions for the English Hate Speech and O ensive
Content Identi cation Sub-task A for HASOC 2019, more information of all
results of the contest is available at [9].
3.2</p>
      <sec id="sec-3-1">
        <title>Analysis</title>
        <p>We analyzed our participation in HASOC'19 in two ways. The rst analysis
focuses on observing what are the 10 most frequent n-grams (excluding stopwords)
at word level (separated by length) in the Hate-O ensive class, these are shown
in Table 4. We also exhibit in Table 5 the best word n-grams per class according
to the Logistic Regression classi er (LRC) trained with the whole training set.
In our nal con guration, it was easier for an o ensive tweet to be
missclassi ed as non-aggressive, and despite running several experiments, most of our
attempts to improve classi cation in this task by adding new features trying to
give context to the tweets unfortunately a ected the results negatively. After
inspection, we observed that this could have happened because:
{ The clustering techniques that we used didn't add anything new since the
tweets were kind of grouped from the beginning, as some main topics can be
spotted (e.g., Trump, Dhoni/ICC and "DoctorsFightBack" protest related
tweets).</p>
        <p>The second analysis addresses the performance of our proposal, regarding
F1-score and contrasted against the rest of the competitors. Fig. 1 presents two
box plots for the complete distribution of competitors in terms of Macro F1 and
Weighted F1. This analysis suggests that the outcome achieved by our proposal is
competitive, practically been located within the rst quartile for all participants.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusions and Future Work</title>
      <p>In this paper, we describe our strategy to classify o ensive and non-o ensive
tweets in a relatively new English collection of tweets. Regarding our
experiments for this task we can conclude that, in our best performing system, term
frequency matrices of words and character n-grams complement each other in
an ensemble of Logistic Regression classi ers. After seeing that the NER
counters were basically the only useful features in the validation stage and the fact
that we could not improve our classi cation scores with our current approach
on providing context to tweets motivates the idea of future work focusing on
nding new features to help us in our goal to see if it's possible to di erentiate
an o ensive text from a non o ensive one based on the message's underlying
properties and the author's attributes.</p>
      <p>References
[1] Github - justglowing/minisom: Minisom is a minimalistic implementation
of the self organizing maps. https://github.com/JustGlowing/minisom.
(Accessed on 06/03/2019).
[2] polyglot PyPI. https://pypi.org/project/polyglot/. (Accessed on
09/09/2019).
[3] textstat PyPI. https://pypi.org/project/textstat/. (Accessed on
09/09/2019).
[4] Aragon, M. E., Alvarez-Carmona, M. A., Montes-y Gomez, M., Escalante,
H. J., Villasen~or-Pineda, L., and Moctezuma, D. (2019). Overview of
MEXA3T at IberLEF 2019: Authorship and aggressiveness analysis in Mexican
Spanish tweets. In Notebook Papers of 1st SEPLN Workshop on Iberian
Languages Evaluation Forum (IberLEF), Bilbao, Spain, September.
[5] Casavantes, M., Lopez, R., and Gonzalez, L. C. (2019). UACh at MEX-A3T
2019 : Preliminary Results on Detecting Aggressive Tweets by Adding Author
Information Via an Unsupervised Strategy.
[6] Davidson, T., Warmsley, D., Macy, M. W., and Weber, I. (2017).
Automated hate speech detection and the problem of o ensive language. CoRR,
abs/1703.04009.
[7] Kumar, R., Ojha, A. K., Malmasi, S., and Zampieri, M. (2018).
Benchmarking aggression identi cation in social media. In Proceedings of TRAC.
[8] Le, Q. and Mikolov, T. (2014). Distributed representations of sentences
and documents. In Proceedings of the 31st International Conference on
International Conference on Machine Learning - Volume 32, ICML'14, pages
II{1188{II{1196. JMLR.org.
[9] Modha, S., Mandl, T., Majumder, P., and Patel, D. (2019). Overview of the
HASOC track at FIRE 2019: Hate Speech and O ensive Content Identi cation
in Indo-European Languages. In Proceedings of the 11th annual meeting of
the Forum for Information Retrieval Evaluation.
[10] Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel,
O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J.,
Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E. (2011).
Scikit-learn: Machine learning in Python. Journal of Machine Learning
Research, 12:2825{2830.
[11] Waseem, Z. and Hovy, D. (2016). Hateful symbols or hateful people?
predictive features for hate speech detection on twitter. pages 88{93.
[12] Wiegand, M., Siegel, M., and Ruppenhofer, J. (2018). Overview of the
germeval 2018 shared task on the identi cation of o ensive language.
[13] Zampieri, M., Malmasi, S., Nakov, P., Rosenthal, S., Farra, N., and Kumar,
R. (2019a). Predicting the Type and Target of O ensive Posts in Social Media.</p>
      <p>In Proceedings of NAACL.
[14] Zampieri, M., Malmasi, S., Nakov, P., Rosenthal, S., Farra, N., and Kumar,
R. (2019b). Semeval-2019 task 6: Identifying and categorizing o ensive
language in social media (o enseval). In Proceedings of the 13th International
Workshop on Semantic Evaluation, pages 75{86.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>