<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>IRLab@IITBHU at HASOC 2019: Traditional Machine Learning for Hate Speech and O ensive Content Identi cation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Indian Institute of Technology (BHU) India</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Varanasi</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>anitas.rs.cse</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>rajeshkm.rs.cse</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>spal.cseg@iitbhu.ac.in</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>In this paper, the results obtained from the Support Vector Machine, XGBoost method by IRLab@IIT(BHU) on HASOC shared task-organized at FIRE-2019 are reported. The HASOC shared task has three subtasks, namely Hate speech identi cation, O ensive language identi cation and Fine-grained classi cation for the English, Hindi and German languages. The best result for English is obtained after applying Support Vector Machine, XGBoost with a frequency-based feature for hate speech and o ensive content identi cation.</p>
      </abstract>
      <kwd-group>
        <kwd>O ensive</kwd>
        <kwd>Hate Speech</kwd>
        <kwd>Language</kwd>
        <kwd>Social Media</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Hate speech is a type of communication of verbal expression to attack a human
or group based on characteristics such as caste, religion, ethnic origin, sexual
orientation, disability or gender [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Hate speech and o ensive content in
IndoEuropean languages have become a common phenomenon in the social media.
Recent years have seen the spread of o ensive language on social media
platforms such as Facebook and Twitter. With the freedom of privilege of
expression granted to social media users, it became easy to spread disrespect or hatred
against individuals or groups. Automated hate language and o ensive content
detection systems may contain the spread inhibition of toxin textual material.
Beyond psychological harm, such toxic online content can give rise to real hate
crimes [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]; which justi es the need to automatically detect abusive language
and o ensive content shared on social media platforms.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>RELATED</title>
    </sec>
    <sec id="sec-3">
      <title>WORK</title>
      <p>Over the last few years, several studies on hate speech and o ensive content
identi cation have been published. The literature has explored di erent o ensive
and abusive language identi cation problems ranging from aggression to cyber
bullying, hate speech, poisonous comments and o ensive language. We brie y
discuss each of them in this section.
2.1</p>
      <sec id="sec-3-1">
        <title>Aggressive content identi cation</title>
        <p>
          The rst shared task on aggression identi cation is Trolling, Aggression and
Cyberbullying (TRAC-1) at COLING 2018. In this task Aggressive Language
Identi cation on Facebook and Twitter data targeted using word embeddings
and sentiment features [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. Moreover, the best result were obtained through
sentiment features with Random Forest (RF) and Support Vector Machine (SVM)
with 0.5830, 0.5074 accuracies, respectively. Later, e orts went to develop a
classi er that could discriminate between overly aggressive, hidden aggressive,
and non-aggressive text. Long short-term memory (LSTM), Convolutional
Neural Network (CNN)-LSTM, Bidirectional LSTM with Glove embeddings, the
combination of the Passive-Aggressive (PA) and SVM classi ers with
characterbased n-gram where n is from 1 to 5, TF-IDF as feature representation were
used for aggression identi cation. The best system explained above to achieve a
weighted F-score of 0.64 on the Facebook test set entitled as English and Hindi,
and the best scores for the surprise set were 0.60 and 0.50 for Hindi and English
respectively [
          <xref ref-type="bibr" rid="ref1 ref16 ref18 ref19 ref8">8, 16, 1, 18, 19</xref>
          ].
2.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Bullying content identi cation</title>
        <p>
          Bullying, also known as peer victimization, has been recognized as a serious
national health issue by the White House (2011), the American Academy of
Pediatrics (2009), and the American Psychological Association (2004) [
          <xref ref-type="bibr" rid="ref2 ref5 ref6">5, 2, 6</xref>
          ].
The growing research into cyberbullying in online social networks have catalyzed
by the widespread and profound consequences of abuse. Earlier research works
on automatic cyberbullying detection have mainly focused on using
(sophisticated) text-based methods [
          <xref ref-type="bibr" rid="ref12 ref15 ref4">4, 12, 15</xref>
          ]. Expanded the text-based identi cation
approach to model the use of hashtags, simultaneously with the emotions the
spatio-temporal cyberbullying measures to understand and explore. [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ].
2.3
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>Hate speech identi cation</title>
        <p>
          Hate speech is a statement of intent to o end another and use cruel or abusive
language based on actual or perceived membership to another group [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
Established a lexical baseline for discriminating between profane and hate speech on
the standard dataset this is the main aim of the paper [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. The authors adopted
a linear support vector machine classi er with three groups of extracted features
for these tests: word skip-grams, surface n-gram and Brown cluster.
2.4
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>O ensive language identi cation</title>
        <p>
          User-generated content on social media platforms such as Twitter often includes
a high level of rude, o ensive or sometimes hateful language [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]. Increasing
vulgarity in online discussions and user comment sections have recently been
discussed as relevant issues in society as well as in science [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ], and identi ed
o ensive tweets with an accuracy of 83.14 %, f1-score 0.7565 on the real test
data for the classi cation of o ensive vs non-o ensive.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>METHODOLOGY</title>
      <p>In this paper, we focus on hate, o ensive, and profane exclusively, for English. We
participated in the competition using the team name IRLAB@IITBHU. Figure 1
shows the methodology of the paper.
3.1</p>
      <sec id="sec-4-1">
        <title>Data</title>
        <p>
          The dataset was created from Twitter, Facebook and distributed in tab-separated
format. We have participated for all three sub-tasks of English language [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. The
size of training and testing data is 5852 and 1153 posts for the English language,
respectively. In Sub-task A the HOF containing Hate, o ensive, and profane
posts are 288, and NOT not containing any Hate speech, o ensive content posts
are 865. In Sub-task B, HATE Hate speech posts are 124, and NONE posts are
865. O ensive posts are 71, and Profane posts are 93. In Sub-task C, NONE
posts are 865, and TIN (Targeted Insult) posts are 245, and UNT (Untargeted)
posts are 43.
3.2
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Pre-processing</title>
        <p>First the data were cleaned using the tweet preprocessing library4. We got the
cleaned data after removing the Retweets Symbols (RT), Hashtag, URL's,
Twitter Mentions, Emoji's and Smileys. The preprocessed data also exclude the
English stop words (available in NLTK5) while tokenizing the sentences for the
extraction of frequency-based feature extraction. The Hate speech, O ensive
and Profane have been predicted through TF-IDF feature.
3.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Classi er</title>
        <p>We use two machine learning classi ers Support Vector Machine (SVM) and
XGBoost (XGB) classifying for classi cation of Hate speech, O ensive and Profane.
The input for both the classi er is in the form of TF-IDF feature matrix and
output is a label for the categorical result. Both the classi ers give a di erent
score, as classi ers have di erent specialities.
4 https://pypi.org/project/tweet-preprocessor/
5 https://www.nltk.org/
4</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>RESULTS</title>
      <p>We start by investigating the accuracy of our TF-IDF features based on machine
learning method for this task. We rst train the classi er, with each of them using
a type of TF-IDF feature. The results of these experiments are listed in Table 1
and Table 2. In Sub-task A, accuracy of XGBoost is 81% better as compared to
SVM 73%. The Sub-task B and Sub-task C accuracy is 80% the same for the
XGBoost.</p>
    </sec>
    <sec id="sec-6">
      <title>CONCLUSION</title>
      <p>In this paper we used text classi cation techniques to recognise among hate
speech, profane and o ensive posts. As a baseline we use XGBoost and a SVM
classi er. The best result was achieved by XGBoost achieving 81% accuracy.
The results displayed in this paper showed that identi cation of profanity from
abusive language is a very challenging task.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Aroyehun</surname>
            ,
            <given-names>S.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gelbukh</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Aggression detection in social media: Using deep neural networks, data augmentation, and pseudo labeling</article-title>
          .
          <source>In: Proceedings of the First Workshop on Trolling, Aggression and Cyberbullying (TRAC-2018)</source>
          . pp.
          <volume>90</volume>
          {
          <issue>97</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Association</surname>
            ,
            <given-names>A.P.</given-names>
          </string-name>
          , et al.:
          <article-title>Apa resolution on bullying among children and youth</article-title>
          . Washington, DC: American Psychological Association pp.
          <volume>1</volume>
          {
          <issue>4</issue>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Britannica</surname>
          </string-name>
          , E.: Britannica academic.
          <source>Encyclop dia Britannica Inc</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Dinakar</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Havasi</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lieberman</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Picard</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Common sense reasoning for detection, prevention, and mitigation of cyberbullying</article-title>
          .
          <source>ACM Transactions on Interactive Intelligent Systems (TiiS) 2</source>
          (
          <issue>3</issue>
          ),
          <volume>18</volume>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>House</surname>
            ,
            <given-names>W.:</given-names>
          </string-name>
          <article-title>Background on white house conference on bullying prevention (</article-title>
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. Committee on Injury,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Prevention</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          , et al.:
          <article-title>Role of the pediatrician in youth violence prevention</article-title>
          .
          <source>Pediatrics</source>
          <volume>124</volume>
          (
          <issue>1</issue>
          ),
          <volume>393</volume>
          {
          <fpage>402</fpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Johnson</surname>
          </string-name>
          , N.,
          <string-name>
            <surname>Leahy</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Restrepo</surname>
            ,
            <given-names>N.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velasquez</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zheng</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manrique</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Devkota</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wuchty</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>Hidden resilience and adaptive dynamics of the global online hate ecology</article-title>
          .
          <source>Nature</source>
          pp.
          <volume>1</volume>
          {
          <issue>5</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Kumar</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ojha</surname>
            ,
            <given-names>A.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Malmasi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zampieri</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Benchmarking aggression identi cation in social media</article-title>
          .
          <source>In: Proceedings of the First Workshop on Trolling, Aggression and Cyberbullying (TRAC-2018)</source>
          . pp.
          <volume>1</volume>
          {
          <issue>11</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Malmasi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zampieri</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Detecting hate speech in social media</article-title>
          .
          <source>arXiv preprint arXiv:1712.06427</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Matsuda</surname>
            ,
            <given-names>M.J.:</given-names>
          </string-name>
          <article-title>Public response to racist speech: Considering the victim's story</article-title>
          .
          <source>In: Words That Wound</source>
          , pp.
          <volume>17</volume>
          {
          <fpage>51</fpage>
          .
          <string-name>
            <surname>Routledge</surname>
          </string-name>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Modha</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mandl</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Majumder</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patel</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Overview of the HASOC track at FIRE 2019: Hate Speech and O ensive Content Identi cation in Indo-European Languages</article-title>
          .
          <source>In: Proceedings of the 11th annual meeting of the Forum for Information Retrieval Evaluation</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Nahar</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          , Y.:
          <article-title>Cyberbullying detection based on textstream classi cation</article-title>
          .
          <source>In: The 11th Australasian Data Mining Conference (AusDM</source>
          <year>2013</year>
          ) (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Orasan</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Aggressive language identi cation using word embeddings and sentiment features</article-title>
          .
          <source>In: Proceedings of the First Workshop on Trolling, Aggression and Cyberbullying (TRAC-2018)</source>
          . pp.
          <volume>113</volume>
          {
          <issue>119</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Ramakrishnan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zadrozny</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tabari</surname>
          </string-name>
          , N.:
          <article-title>UVA wahoos at SemEval-2019 task 6: Hate speech identi cation using ensemble machine learning</article-title>
          .
          <source>In: Proceedings of the 13th International Workshop on Semantic Evaluation</source>
          . pp.
          <volume>806</volume>
          {
          <fpage>811</fpage>
          . Association for Computational Linguistics, Minneapolis, Minnesota, USA (Jun
          <year>2019</year>
          ). https://doi.org/10.18653/v1/
          <fpage>S19</fpage>
          -2141, https://www.aclweb.org/anthology/S19- 2141
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Reynolds</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kontostathis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Edwards</surname>
            ,
            <given-names>L.:</given-names>
          </string-name>
          <article-title>Using machine learning to detect cyberbullying</article-title>
          .
          <source>In: 2011 10th International Conference on Machine learning and applications and workshops</source>
          . vol.
          <volume>2</volume>
          , pp.
          <volume>241</volume>
          {
          <fpage>244</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Risch</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krestel</surname>
          </string-name>
          , R.:
          <article-title>Aggression identi cation using deep learning and data augmentation</article-title>
          .
          <source>In: Proceedings of the First Workshop on Trolling, Aggression and Cyberbullying (TRAC-2018)</source>
          . pp.
          <volume>150</volume>
          {
          <issue>158</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Sui</surname>
          </string-name>
          , J.:
          <article-title>Understanding and ghting bullying with machine learning</article-title>
          .
          <source>Ph.D. thesis, Ph. D. dissertation, The Univ. of Wisconsin-Madison</source>
          ,
          <string-name>
            <surname>WI</surname>
          </string-name>
          , USA (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Wiegand</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Siegel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ruppenhofer</surname>
          </string-name>
          , J.:
          <article-title>Overview of the germeval 2018 shared task on the identi cation of o ensive language (</article-title>
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Zampieri</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Malmasi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nakov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosenthal</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Farra</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kumar</surname>
          </string-name>
          , R.:
          <article-title>Predicting the Type and Target of O ensive Posts in Social Media</article-title>
          .
          <source>In: Proceedings of NAACL</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Zampieri</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Malmasi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nakov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosenthal</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Farra</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kumar</surname>
          </string-name>
          , R.:
          <article-title>Semeval-2019 task 6: Identifying and categorizing o ensive language in social media (o enseval)</article-title>
          .
          <source>In: Proceedings of the 13th International Workshop on Semantic Evaluation</source>
          . pp.
          <volume>75</volume>
          {
          <issue>86</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>