<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>DEEP at HASOC2019 : A Machine Learning Framework for Hate Speech and O ensive Language Detection</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Hamada A. Nayel</string-name>
          <email>hamada.ali@fci.bu.edu.eg</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shashirekha H. L.</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science Faculty of Computers and Arti cial Intelligence Benha University</institution>
          ,
          <country country="EG">Egypt</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Computer Science Mangalore University</institution>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we describe the system submitted by our team for Hate Speech and O ensive Content Identi cation in Indo-European Languages (HASOC) shared task held at FIRE 2019. Hate speech and o ensive language detection have become an important task due to the overwhelming usage of social media platforms in our daily life. This task has been applied for three languages namely, English, Germany and Hindi. The proposed model uses classical machine learning approaches to create classi ers that are used to classify the given post according to di erent subtasks.</p>
      </abstract>
      <kwd-group>
        <kwd>Multi-task Classi cation</kwd>
        <kwd>Multi-lingual Text Analysis</kwd>
        <kwd>Hate</kwd>
        <kwd>Speech and O ensive Detection 3</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Wide spread use of internet is giving more freedom to people to express their
thoughts freely and anonymously in di erent forms such as blogs, business
networks, forums and social media such as Twitter and Facebook. Internet has
enabled the interaction among people from di erent culture, race, religion,
origin, gender and nationality. It has also paved the way for the online information
to reach the mass within a matter of seconds. Social media is generating lot of
content as it has opened up a new galaxy of opportunities for people to express
their opinions online resulting in the exchange of ideas in a positive way as well as
contributing to the propagation of hate speech in a negative way. The anonymity
of users provided by the internet is also impacting the society in generating the
web contents in a negative way by the usage of profane words, hate speech,
3 Copyright c 2019 for this paper by its authors. Use permitted under Creative
Commons License Attribution 4.0 International (CC BY 4.0). FIRE 2019, 12-15 December
2019, Kolkata, India.
derogatory terms, racists slurs, toxic, abusive and o ensive language. This
negative content may be aggressive and potentially harmful lowering the self-esteem
of people leading to mental illness and suicidal attempts and has also forced
people to deactivate their social media accounts. The increase of cyberbullying
and cyberterrorism, and the use of hate speech content on the Internet, make
the identi cation of hate speech an essential ingredient for anti-bullying policies
of social media [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Many terrorist activities, which are related to hate speech,
are gaining huge attention in social media in terms of posts and suggestions [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>Hate speech and other o ensive and objectionable content online are posing
threats and challenges to the society. A lot of countries prohibit hate speech
in social media subject to the condition that it should not target any group or
trigger any crime. As hate speech acts as a opinion builder, many online forums
such as YouTube, Face book and Twitter, have their own policies to remove hate
speech content or anything which is impacting the society in a negative way.</p>
      <p>
        It is therefore important to take preventive measures to cope up with hate
speech and objectionable content or remove such content. Detecting hate speech
is a di cult task as there is no clear de nition for hate speech and can vary
according to the context. For example, speech which contain subtle and nuance
sentences are di cult to detect as hate speech. Further, the usage of slang terms
may be found neutral, but may trigger hate crime or o ense in some context.
Although there is no universal de nition for hate speech, the most accepted
de nition is provided by Nockleby [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]: \any communication that disparages a
target group of people based on some characteristic such as race, colour,
ethnicity, gender, sexual orientation, nationality, religion, or other characteristic" [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>Detection and removal of hate speech and objectionable content manually is
a tedious task due to the massiveness of the web and the increasing number of
online users. Further, the task becomes di cult due to the anonymity of users
on the internet. Hence, there is a great demand for tools and techniques that
automatically detect the hate speech and objectionable content quickly on the
web to reduce the spread of hate speech.</p>
      <p>
        In this paper, we describe the system submitted by our team for the shared
task of Hate Speech and O ensive Content Identi cation in Indo-European
Languages (HASOC) track at FIRE 2019 [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Rest of the paper is organized as
follows: In Section 2 we review the related work on hate speech and o ensive
language detection. Description of the task and a summary of dataset is given
in Section 3. Our proposed model is discussed in Section 4 followed by the
experiments and results in Section 5. Finally, we conclude the paper in Section
6.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related work</title>
      <p>
        Though, de ning and understanding hate speech is di cult, several online
forums, IT industries and researchers have explored many algorithms to detect
hate speech, o ensive and abusive content on the web [
        <xref ref-type="bibr" rid="ref15 ref2 ref3 ref4 ref5 ref6">5, 2, 6, 3, 15, 4</xref>
        ]. Thomas
et.al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] trained a multi-class classi er to classify tweets into one of three
categories, namely, hate speech, o ensive but not hate speech, neither o ensive and
nor hate speech. They performed experiments using logistic regression, Nave
Bayes, decision trees, random forests, and linear SVM classi ers with 5-fold
cross validation and obtained an overall precision 0.91, recall of 0.90, and F1
score of 0.90 for the best performing model. Gibert et.al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], constructed a
manually labeled hate speech dataset composed of thousands of sentences obtained
from Stormfront, a white supremacist online forum which contains hateful and no
hateful sentences. Their model was built using Support Vector Machines (SVM),
Convolutional Neural Networks (CNN), Recurrent Neural Networks with Long
Short term Memories to annotate the test data.
      </p>
      <p>
        Sean MacAvaney et.al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], examines the challenges faced by online
automatic approaches for hate speech detection in text in addition to the existing
approaches and various datasets for hate speech detection. They also have
proposed a multi-view SVM approach that achieves near state-of-the-art
performance, while being simpler and producing more easily interpretable decisions
than neural methods. The experiments were conducted on HatebaseTwitter,
Stanfront and TRAC datasets. Avishek Garain and Arpan Basu [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], presents a
description of their system to detect o ensive language in Twitter submitted to
\SemEval-2019 Task 6 [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]" which uses bidirectional LSTM, a neural network
based model to capture information from both the past and future context. Ziqi
and Lei [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], discusses about the typical features responsible for classi cation
between hate speech and no hate content. In their work they proposed Deep
Neural Network structures serving as feature extractors that are particularly
e ective for capturing the semantics of hate speech. Their methods were
evaluated on the largest collection of hate speech datasets based on Twitter, and are
shown to be able to outperform the best performing method by up to 5
percentage points in macro-average F1, or 8 percentage points in the more challenging
case of identifying hateful content. Aditya et.al., [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] proposes an approach to
automatically classify tweets on Twitter into three classes: hateful, o ensive and
clean by considering features like n Grams and TF/IDF values. They performed
a comparative analysis of the models considering several values of n in n-grams
and TF/IDF normalization methods on Twitter dataset and achieved 95.6%
accuracy. Areej Al-Hassan and Hmood Al-Dossari [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] in their survey paper, gives
detailed information about hate speech in terms of insu cient dataset,
complexities and challenges and the complete knowledge about text mining and Natural
Language Processing techniques to nd and classify hateful and no-hateful
contents.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Task Description and Corpus</title>
      <p>Hate Speech and O ensive Content Identi cation in Indo-European Languages
(HASOC) shared task aims at identifying the hate speech and o ensive content in
three Indo-European languages namely English, Germany and code-mixed Hindi.
The shared task is divided into three sub-tasks to identify the hate speech and
o ensive content posts and classify them into various prede ned categories. The
rst sub-task, Sub-task A is a coarse-grained binary classi cation that focuses
on the identi cation of hate speech and o ensive language in the given posts
and is o ered for English, German and Hindi language posts. Sub-task B is a
ne-grained multi-class classi cation that classi es the hate speech and o ensive
language posts into one of the three prede ned categories, namely hate speech,
o ensive and profane. This subtask is o ered for English, German and Hindi
language posts. The third sub-task, Sub-task C aims at checking the type of
o ense as targeted insult or un-targeted insult and is o ered for English and
Hindi. The corpus consists of posts collected from Twitter tweets and Facebook
comments in English (EN), German (GR) and Hindi (HI) languages. Each post
contains three labels for Sub-task A, Sub-task B and Sub-task C respectively.
Table 1 gives a glimpse of the shared task and the associated categories in each
sub-task.
A detailed description of our model and the classi ers used are given in this
section.
Given a set of posts P = fp1; p2; :::; png, where each post is composed of a set of
words pi = fw1; w2; :::; wkg, in each language L = fEN; GR; HIg and the classes
A = fN OT; HOF g, B = fHAT E; OF F N; P RF N g and C = fT IN; U N T g for
the sub-tasks A, B and C respectively, the shared task is to formulate a
multilabel classi cation problem, where an unlabeled post/instance is assigned with
multiple class labels one from each class A, B and C. ie., given a unlabeled post
pk, multi-label classi cation problem will assign the triple &lt; a; b; c &gt; such that,
a 2 A, b 2 B and c 2 C.
4.2</p>
      <p>Model
The general structure of the proposed model is given in Fig. 1. This model creates
sub-models for each subtask and then assigns multiple class labels, one for each
subtask, by combining the output of each sub-model. The model consists of the
following phases:</p>
      <p>I. Preprocessing Preprocessing also known as corpus cleaning is a crucial and
rst step in building any classi ers. Each post pk has been tokenized into a set
of words to get n-gram (n = 2) bag of words. In this step, all un-informative
tokens such as urls, digits and special characters have been removed from all the
posts.</p>
      <p>
        II. Feature Extraction In this phase, a TF/IDF vector has been computed
for all the posts in the training set. This vector will be used as an input for
training the classi er. TF/IDF has been calculated as described in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
III. Training the Classi er In this phase, the features that have been
extracted in previous phase are used as input for training the classi ers. Three
classi cation algorithms namely, Linear classi er, SVM and Multilayer
Perceptron (MLP) have been used separately to train the proposed model. Linear
classi er is a simple classi er that uses a set of linear discriminant functions to
distinguish between di erent classes [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. SVM is a kind of linear classi er which
uses the samples close to the classes's boundaries for training and these samples
are known as support vectors [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. SVM has been used e ectively in di erent
NLP tasks [
        <xref ref-type="bibr" rid="ref8 ref9">9, 8</xref>
        ]. MLP is a deep learning approach that uses back propagation
for training the neural network [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. It is characterized by several layers of input
nodes connected as a directed graph between the input and output layers.
      </p>
      <p>Three di erent classi ers namely, Linear classi er, SVM and MLP are created
for each subtask. Classi ers for Sub-task A will classify the unlabeled instance
into one of the two prede ned categories as mentioned in Table-1. While the
entire data is used to construct classi ers for Sub-task A, a sub-corpus (only posts
labelled as HOF in Sub-task A) that contain only the hate speech and o ensive
language posts is used as training corpus to create classi ers for Sub-tasks B and
C. Classi ers for Sub-task B will classify the unlabeled instance into one of the
three prede ned categories and classi er for Sub-task C will classify the same
instance into one of the two prede ned categories as mentioned in Table 1.
5</p>
    </sec>
    <sec id="sec-4">
      <title>Experiments and Results</title>
      <p>
        In our proposed model, the Stochastic Gradient Descent (SGD) optimization
algorithm has been used for optimizing the parameters of linear classi er, while
SVM uses Lagrange multipliers to solve the optimization problem [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. The loss
function used in linear classi er was "Hinge" loss function. Linear kernel has
been used for SVM classi er. In MLP classi er the logistic function has been
used as activation function using 20 neurons in the hidden layer.
      </p>
      <p>
        We have used cross-validation approach with ve folds to train di erent
submodels for each subtask and the outputs of each sub-task have been combined to
generate the nal output. The organizers of the shared task used Macro-averaged
F1-score (M-f1) and weighted F1-score (W-f1) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] as performance evaluation
metrics for all sub-tasks. The dataset comprises of a set of tweets and Facebook
comments in the three languages labelled with di erent categories. Statistics of
the dataset is shown in Table 2.
      </p>
      <p>Table 3 shows the M-f1 and W-f1 of our submissions for all sub-tasks over
English, Germany and Hindi language posts. From the table, it is clear that SVM
outperforms other classi ers for all subtasks for English. For German language
posts, MLP outperforms other classi ers for Sub-tasks A and B. It is worth to
note that, Sub-task C is not applied for Germany. For Hindi language posts,
MLP outperforms other classi ers for Subtask A, while SVM reported better
results than MLP and linear classi er for subtask B.</p>
      <p>There is a gab between M-f1 and W-f1 for all subtasks except for the subtask
A for Hindi. This is due to the fact that the performance of the subtasks B and
C depends on the performance of the subtask A. If the prediction of subtask A
is wrong then by default subtasks B and C will also have wrong prediction.
Sub-task Class Labels English German Hindi
In this paper, a machine learning approaches have been used for creating a
model for detecting hate speech and o ensive language content. Proposed model
achieved good results compared to its simplicity. Extension of our work includes
using deep learning approach to build the classi er and test it on much bigger
dataset.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Al-Hassan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Al-Dossari</surname>
          </string-name>
          , H.:
          <article-title>Detection of Hate Speech in Social Networks: A Survey on Multilingual Corpus</article-title>
          .
          <source>In: Proceedings of 6th International Conference on Computer Science and Information Technology (CoSIT</source>
          <year>2019</year>
          ). pp.
          <volume>83</volume>
          {
          <fpage>100</fpage>
          .
          <string-name>
            <surname>Dubai</surname>
          </string-name>
          ,
          <source>UAE (February</source>
          <volume>23</volume>
          -24
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Davidson</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Warmsley</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Macy</surname>
            ,
            <given-names>M.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weber</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Automated Hate Speech Detection and the Problem of O ensive Language</article-title>
          .
          <source>In: Proceedings of the Eleventh International Conference on Web and Social Media</source>
          ,
          <string-name>
            <surname>ICWSM</surname>
          </string-name>
          <year>2017</year>
          , Montreal, Quebec, Canada, May
          <volume>15</volume>
          -18,
          <year>2017</year>
          . pp.
          <volume>512</volume>
          {
          <fpage>515</fpage>
          . AAAI Press (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Garain</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basu</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>The Titans at SemEval-2019 Task 6: O ensive Language Identi cation, Categorization and Target Identi cation</article-title>
          . In: May,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Shutova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Herbelot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            ,
            <surname>Apidianaki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Mohammad</surname>
          </string-name>
          , S.M. (eds.)
          <source>Proceedings of the 13th International Workshop on Semantic Evaluation, SemEval@NAACL-HLT</source>
          <year>2019</year>
          ,
          <article-title>Minneapolis</article-title>
          , MN, USA, June 6-7,
          <year>2019</year>
          . pp.
          <volume>759</volume>
          {
          <fpage>762</fpage>
          .
          <article-title>Association for Computational Linguistics (</article-title>
          <year>2019</year>
          ), https://www.aclweb.org/anthology/S19-2133/
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Gaydhani</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Doma</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kendre</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bhagwat</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Detecting Hate Speech and O ensive Language on Twitter using Machine Learning: An N-gram and TFIDF based Approach</article-title>
          .
          <source>CoRR</source>
          (
          <year>2018</year>
          ), http://arxiv.org/abs/
          <year>1809</year>
          .08651
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>de Gibert</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perez</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garc</surname>
            a-Pablos,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cuadros</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Hate Speech Dataset from a White Supremacy Forum</article-title>
          .
          <source>In: Proceedings of the 2nd Workshop on Abusive Language Online (ALW2)</source>
          . pp.
          <volume>11</volume>
          {
          <fpage>20</fpage>
          . Association for Computational Linguistics, Brussels, Belgium (Oct
          <year>2018</year>
          ). https://doi.org/10.18653/v1/
          <fpage>W18</fpage>
          -5102
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>MacAvaney</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yao</surname>
            ,
            <given-names>H.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Russell</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goharian</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frieder</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Hate speech detection: Challenges and solutions</article-title>
          .
          <source>PLoS ONE</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Modha</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mandl</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Majumder</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patel</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Overview of the HASOC track at FIRE 2019: Hate Speech and O ensive Content Identi cation in Indo-European Languages</article-title>
          . In:
          <article-title>Proceedings of the 11th annual meeting of the Forum for Information Retrieval Evaluation (December</article-title>
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Nayel</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shashirekha</surname>
            ,
            <given-names>H.L.</given-names>
          </string-name>
          :
          <article-title>Improving NER for Clinical Texts by Ensemble Approach using Segment Representations</article-title>
          .
          <source>In: Proceedings of the 14th International Conference on Natural Language Processing (ICON-2017)</source>
          . pp.
          <volume>197</volume>
          {
          <fpage>204</fpage>
          . NLP Association of India, Kolkata,
          <source>India (December</source>
          <year>2017</year>
          ), http://www.aclweb.org/anthology/W/W17/W17-7525
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Nayel</surname>
            ,
            <given-names>H.A.</given-names>
          </string-name>
          :
          <article-title>NAYEL@APDA: Machine Learning Approach for Author Pro ling and Deception Detection in Arabic Texts</article-title>
          . In:
          <article-title>Proceedings of the 11th annual meeting of the Forum for Information Retrieval Evaluation (December</article-title>
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Nayel</surname>
            ,
            <given-names>H.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shashirekha</surname>
            ,
            <given-names>H.L.</given-names>
          </string-name>
          : Mangalore University INLI@FIRE2018:
          <article-title>Arti cial Neural Network and Ensemble based Models for INLI</article-title>
          . In: Mehta,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Majumder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Mitra</surname>
          </string-name>
          , M. (eds.) Working Notes of FIRE 2018 -
          <article-title>Forum for Information Retrieval Evaluation, Gandhinagar</article-title>
          , India, December 6-
          <issue>9</issue>
          ,
          <year>2018</year>
          .
          <source>CEUR Workshop Proceedings</source>
          , vol.
          <volume>2266</volume>
          , pp.
          <volume>110</volume>
          {
          <fpage>118</fpage>
          .
          <string-name>
            <surname>CEUR-WS.org</surname>
          </string-name>
          (
          <year>2018</year>
          ), http://ceurws.org/Vol-
          <volume>2266</volume>
          /
          <fpage>T2</fpage>
          -10.pdf
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Nayel</surname>
            ,
            <given-names>H.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shashirekha</surname>
            ,
            <given-names>H.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shindo</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matsumoto</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Improving MultiWord Entity Recognition for Biomedical Texts</article-title>
          . CoRR abs/
          <year>1908</year>
          .05691 (
          <year>2019</year>
          ), http://arxiv.org/abs/
          <year>1908</year>
          .05691
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Nockleby</surname>
          </string-name>
          , J.T.:
          <article-title>Hate Speech</article-title>
          .
          <source>In: Encyclopedia of the American Constitution</source>
          . pp.
          <volume>1277</volume>
          {
          <issue>1279</issue>
          (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Theodoridis</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koutroumbas</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Chapter 3 - Linear Classi ers</article-title>
          .
          <source>In: Pattern Recognition (Fourth Edition)</source>
          , pp.
          <volume>91</volume>
          {
          <fpage>150</fpage>
          . Academic Press, Boston, fourth edition edn. (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Zampieri</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Malmasi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nakov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosenthal</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Farra</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kumar</surname>
          </string-name>
          , R.: SemEval
          <article-title>-2019 Task 6: Identifying and Categorizing O ensive Language in Social Media (O ensEval)</article-title>
          .
          <source>In: Proceedings of the 13th International Workshop on Semantic Evaluation</source>
          . pp.
          <volume>75</volume>
          {
          <issue>86</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Luo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Hate Speech Detection: A Solved Problem? The Challenging Case of Long Tail on Twitter</article-title>
          . CoRR abs/
          <year>1803</year>
          .03662 (
          <year>2018</year>
          ), http://arxiv.org/abs/
          <year>1803</year>
          .03662
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>