<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Journal
of King Saud University</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.1109/78.650093</article-id>
      <title-group>
        <article-title>Hate Speech: ATLANTIS for Eficient Hate Span Detection</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Partha Pakray</string-name>
          <email>partha@cse.nits.ac.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Niyar R Barman</string-name>
          <email>barmanniyar@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Krish Sharma</string-name>
          <email>iamkrish9090@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yashraj Poddar</string-name>
          <email>yash.raj.poddar.yp@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Advaitha Vetagiri</string-name>
          <email>advaitha21_rs@cse.nits.ac.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Hate Speech Detection, Named Entity Recognition (NER)</institution>
          ,
          <addr-line>Sequence Labeling, Natural Language Process-</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>National Institute of Technology</institution>
          ,
          <addr-line>Silchar, Assam</addr-line>
          ,
          <country country="IN">India -</country>
          <addr-line>788010</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>1</volume>
      <fpage>15</fpage>
      <lpage>18</lpage>
      <abstract>
        <p>Hate speech poses significant challenges to maintaining healthy online conversations, and automated systems are crucial for its accurate detection and mitigation. In this paper, we (CNLP-NITS-PP) introduce ATLANTIS (Attentive Transformer-LSTM for Named Entity and Token Identification System), a robust model designed to address the pervasive issue of hate speech in online social media platforms. ATLANTIS focuses on hate span identification within sentences labeled as hate speech, framed as a sequence labeling task using BIO notation. Leveraging a Hate dataset enriched with Named Entity Recognition (NER) tags, ATLANTIS efectively identifies hate speech spans within the text by combining contextualized representations and sequential modeling. The empirical results showcase ATLANTIS's efectiveness in isolating explicit signs of hate from a contextual backdrop, ofering a promising solution for creating safer online environments. We achieve a macro F1 score of 0.488 on the public test set and 0.508 on the private test set. This work not only lays the foundation for future advancements in hate-span detection but also emphasizes the importance of model eficiency, interpretability, and expanded training data that encompass diverse linguistic nuances and evolving hate speech trends. Code is available at https://github.com/niyarrbarman/hasoc23 htp:/ceur-ws.org CEUR Workshop Proceedings (CEUR-WS.org) ISN1613-073</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        Social media platforms like Twitter and Facebook have become commonplace in modern life,
giving people worldwide easy access to voice their thoughts and connect. However, the open
nature of these platforms also allows harmful content like hate speech, harassment, and threats
aimed at vulnerable groups to spread [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. This has created an urgent need for automated systems
that accurately recognise abusive language to maintain healthy online conversations [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
LGOBE
CEUR
Workshop
Proceedings
      </p>
      <p>
        A significant hurdle is that ofensive content can take many linguistic forms, necessitating
context-aware models to pinpoint the specific snippets of text that render a post hateful or
abusive [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Furthermore, implicit forms of hate speech, like veiled insults, require deducing
pragmatic implications [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] rather than just spotting explicit derogatory terms [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. This has
driven recent research into models for singling out spans of text that communicate hateful
intent within a given post [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        This paper tackles the problem of hate span identification within sentences labelled as hate
speech in the HASOC 2023 [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] shared task [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. In this paper, we delve into the challenges and
innovations of the HASOC subtrack at FIRE 2023, focusing on the ’Detection of Hate Spans and
Conversational Hate-Speech,’ as outlined by Satapara et. al [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Given an English social media
sentence already deemed hateful, the goal is to pinpoint contiguous spans of tokens that relay
its hateful purpose. This is framed as a sequence labelling task using BIO notation, where each
token is tagged as the Beginning (B), Inside (I), or Outside (O) of a hate span [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>
        The HASOC dataset provides ground truth BIO tag sequences for abusive sentences from
public hate speech sources [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Participants construct models to predict these spans in test
sentences without extra preprocessing to avoid incongruities. This focused evaluation enables
the systematic development of context-aware models and techniques for fine-grained hate
speech analysis, moving beyond the binary classification of posts [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>
        We present our proposed model design and tactic for the hate span identification task,
harnessing contextualised representations and sequential modelling [12]. Results showcase
our techniques’ eficacy in isolating explicit signs of hate from a contextual backdrop. By
classifying specific linguistic cues and semantic relationships that encode hate, our method
provides insights into the underlying fabric of abusive language [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>2. Application and Target Audience</title>
      <p>The research presented in this paper holds significant promise in tackling the pervasive problem
of hate speech on online social media platforms. ATLANTIS, the hate span detection system
that has been developed, carries practical implications for content moderation, user safety, and
the improvement of online discussions. By precisely identifying and extracting hate spans
from hateful sentences, ATLANTIS equips social media platforms to more eficiently filter and
eliminate hateful content, thereby promoting a safer and more inclusive online environment.
Furthermore, this technology can serve as a valuable tool for gaining insights into the prevalence
and dynamics of hate speech, assisting researchers and policymakers in formulating
evidencebased strategies to combat online hatred.</p>
      <p>This research paper is intended for a diverse audience encompassing various stakeholders
concerned with the detection and mitigation of hate speech. Content moderators and social
media platform administrators will find valuable insights and methodologies within as they
work towards maintaining respectful and secure online communities. Researchers in the
ifelds of natural language processing (NLP) and machine learning will appreciate the detailed
methodology and architecture of the ATLANTIS model, which represents an advancement in
state-of-the-art hate span detection. Policymakers and organizations focused on addressing
online hate speech will also gain valuable insights into the potential of machine learning-based
solutions for addressing this pressing issue. Furthermore, educators and students studying NLP,
machine learning, and technology ethics can utilize this paper as a resource for understanding
the development and application of advanced models for hate speech detection. Ultimately,
this research paper aims to engage a broad and diverse audience, fostering collaboration and
innovation in the ongoing efort to create safer online spaces.</p>
    </sec>
    <sec id="sec-4">
      <title>3. Objective</title>
      <p>The primary objective of this research is to create a hate span detection system capable of
pinpointing and extracting uninterrupted sequences of tokens found within hateful sentences,
which we refer to as “hate spans”. These hate spans are characterized as consecutive sets of
tokens within a sentence that collectively expresses explicit hatefulness. The aim of this shared
task is to automatically identify and extract all such hateful spans from preprocessed sentences.
The hate span detection task is approached as a sequence labeling problem, wherein each token
in a sentence is labeled with a specific tag to indicate its association with a hateful span. The
labeling follows the BIO notation, with ‘B’ signifying the beginning of a hate span, ‘I’ denoting
the continuation of a hate span, and ‘O’ indicating all other tokens that are not part of any hate
span within the sentence.</p>
      <p>The goal is to develop a machine-learning model to accurately predict the correct sequence
of BIO tags for each token in a given sentence, efectively detecting and delineating hate spans
within the text.</p>
    </sec>
    <sec id="sec-5">
      <title>4. Proposed Methodology</title>
      <p>The methodology employed to address the issue of hate speech at scale through the ATLANTIS
model comprises a systematic approach encompassing data preprocessing, tokenization, model
architecture, and the classification process. Leveraging the HateNorm23 dataset, which features
text samples paired with Named Entity Recognition (NER) tags categorizing each word as ‘B’
(signifying the start of a hate span), ‘I’ (indicating inclusion within a hate span), or ‘O’ (denoting
other), we conduct word-level tokenization to segment the text into meaningful units. A custom
tokenizer is then fine-tuned on the dataset to tailor tokenization for hate span detection. The
ATLANTIS model adopts a multi-stage architecture, initially processing tokenized text through a
custom transformer section followed by a bidirectional long short-term memory (Bi-LSTM) [13]
section. The transformer captures contextual information and relationships, while the Bi-LSTM
captures sequential dependencies. Subsequently, fused representations from these sections
traverse fully connected layers for the conclusive classification task. Detailed insights into
the architecture, hyperparameters, and experimental findings will be presented to substantiate
ATLANTIS’s eficacy in mitigating hate speech at scale.</p>
      <p>ATLANTIS consists of three primary components:
Transformer Encoder Block: The Transformer [14] block is a foundational component for
capturing contextual relationships within sequences. Its self-attention mechanism enables the
model to weigh the significance of each word in relation to others, allowing it to understand
complex dependencies and semantic connections. This block excels at learning hierarchical
features from the input data, providing a solid basis for understanding the underlying patterns
in the sequential data, which is particularly crucial in NLP tasks.</p>
      <p>BiLSTM Layer: The BiLSTM layer complements the Transformer’s strengths by efectively
capturing sequential dependencies in the data. By incorporating a BiLSTM layer, the model
can capture fine-grained temporal relationships and contextual nuances that might be missed
by the Transformer alone. This is especially valuable for NER, where identifying entities often
relies on sequential patterns.</p>
      <p>Sequential Block with FC Layers: The Sequential Block, containing Fully Connected layers,
serves as a vital element for transforming the enriched features from the preceding blocks into
a suitable format for making predictions. These FC layers allow for nonlinear transformations
and higher-level abstractions, enabling the model to learn complex mappings from the learned
representations to the target NER labels.</p>
      <p>Engineering Decisions: We aimed to identify a solution that excels in performance and
eficiency. Our approach led us to employ a sequence of six transformer blocks. Upon extending
the number of blocks, we observed a period during which the F1 score plateaued, roughly
around 9 to 10 blocks. Subsequently, the score rapidly declined, indicative of overfitting taking
hold.</p>
      <p>Regarding the BiLSTM layers, we integrated a single BiLSTM layer for the ultimate modeling
phase. Elevating the count of BiLSTM layers increased the model’s complexity, rendering it
more challenging to train and subsequently slowing down inference processes.</p>
      <p>We settled on a configuration of num_heads = 4 for the transformer block. Introducing
additional num_heads led to a stage of diminishing returns. Given the limited size of our dataset,
the model tended to memorize the training data rather than exhibiting the capacity to generalize
to novel data. This phenomenon, in turn, resulted in overfitting or diminished performance.</p>
      <p>Adam was used as the optimizer with learning_rate = 1e-3</p>
    </sec>
    <sec id="sec-6">
      <title>5. Dataset</title>
      <p>The dataset [15] comprises a total of 2421 data points. We partitioned this dataset into an
80:10:10 ratio, allocating segments for training, validation, and testing purposes. Within the
dataset, a sum of 8165 distinct words can be found. The visualization of the dataset is presented
in Figure 2. Notably, hate speech constitutes 17.422% of the entire dataset.</p>
    </sec>
    <sec id="sec-7">
      <title>6. Results and Analyses</title>
      <p>In this section, we present the results of our experiments, organized into three subsections:
Baseline Methods, Intrinsic Results, and Extrinsic Results. We discuss the models we used in
the Baseline Methods section and provide details on the intrinsic and extrinsic performance of
our approach.</p>
      <sec id="sec-7-1">
        <title>6.1. Baseline Methods</title>
        <p>To establish a benchmark for our experiments and assess the efectiveness of our proposed
method, we employed the following baseline models:
Pretrained BERT: BERT [12] has shown remarkable success in various natural language
processing tasks, and we included it as a reference to evaluate the performance of our approach
against a state-of-the-art model.</p>
        <p>Transformer Encoder: The incorporation of the Transformer [14] Encoder, in our study
serves a dual purpose. Firstly, it provides a reference point for evaluating the performance of
our approach. Secondly, it underscores the efectiveness of the encoder layers, equipped with
self-attention mechanisms, which play a key role in the remarkable success of BERT and similar
models across various natural language processing tasks.</p>
        <p>BiLSTM: BiLSTM [13] networks have been widely used for sequence labeling tasks, and we
included this baseline to evaluate our approach against a more traditional sequence labeling
model.</p>
      </sec>
      <sec id="sec-7-2">
        <title>6.2. Intrinsic Results</title>
        <p>In this subsection, we present the intrinsic results of our approach to the validation set. We
discuss the performance of our model and provide a detailed analysis of the results.</p>
        <p>Our model’s performance on the validation set was evaluated using various metrics, including
precision, recall and F1-score. They have been presented in Table 2.</p>
      </sec>
      <sec id="sec-7-3">
        <title>6.3. Extrinsic Results</title>
        <p>In this subsection, we present the extrinsic results of our approach to the competition test set.
We report public and private test scores, commonly used in Kaggle competitions to evaluate
model performance on unseen data. Table 3 summarizes our model’s public and private test
scores and compares them with the baseline models.</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>7. Related Work</title>
      <p>In this section, we review several relevant studies that contribute to the understanding and
development of hate speech detection, ofensive language detection, and related natural
language processing tasks. These works collectively provide insights into various approaches and
techniques employed in this field.</p>
      <p>In Qian et al.’s (2019)[16] study [14], a new challenge called generative hate speech
intervention was introduced. The authors augmented their research with two comprehensive
datasets obtained from Reddit and Gab, which contained intervention responses collected from
crowdsourcing. The assessment of three generative models, specifically Seq2Seq, VAE, and RL,
revealed areas where hate speech intervention methods could be enhanced.</p>
      <p>In the work conducted by Alshalan et al. [17], they tackled the problem of hate speech in
the Arabic Twittersphere. They introduced a dataset consisting of 9316 tweets categorized into
hate speech, abuse, and normalcy. Their assessment encompassed various models, including
CNN, GRU, CNN + GRU, and BERT. Among these models, CNN emerged as the most efective,
achieving superior performance with an F1-score of 0.79 and an AUROC of 0.89.</p>
      <p>In the research conducted by Elalami et al. [18], they introduced a transfer learning strategy
for detecting ofensive language in multiple languages. This approach leveraged several BERT
models, such as BERT, mBERT, and AraBERT. Their results were outstanding, surpassing
the performance of current leading methods that employ joint-multilingual and
translationbased approaches. This study underscored the robustness of BERT models in the context of
Multilingual Ofensive Language Detection.</p>
      <p>Ozler et al. [19] explored the application of BERT for multi-label and multi-domain incivility
detection tasks. They successfully established a new state-of-the-art performance across various
datasets. The study suggested that direct data combination from multiple domains yielded
superior results compared to more intricate training methods.</p>
      <p>The study by Hoang et al. [20] introduced ViHOS, a novel Vietnamese dataset for hate
and ofensive span detection, containing 26,467 annotated spans in 11,056 comments. Baseline
models, including XLM-RBase, XLM-RLarge, PhoBERTBase, and PhoBERTLarge, were evaluated,
with the XLM-RLarge model leading with an F1-score of 0.7770. The study found that detecting
multiple spans outperformed single-span detection in Vietnamese hate speech.</p>
      <p>
        Lample et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] introduced a discriminative parsing-based approach for nested named
entity recognition, demonstrating strong performance on top-level and nested entities. However,
the study acknowledged a limitation in terms of speed compared to conventional flat techniques.
The paper advocated for reconsidering the exclusion of embedded entities in NER corpora,
highlighting the substantial information loss incurred by this design choice.
      </p>
      <p>Ma (2016) [21] presented a neural network architecture for sequence labeling, representing
an end-to-end model without needing task-specific resources, feature engineering, or data
preprocessing. The study attained state-of-the-art performance on two linguistic sequence
labeling tasks, outperforming prior state-of-the-art systems.</p>
      <p>Peters et al. (2017) [22] proposed a simple semi-supervised approach using pre-trained neural
language models to enhance token representations in sequence tagging models. Their approach
consistently outperformed state-of-the-art models in NER and Chunking datasets. Notably,
the study showed that including both forward and backward language models consistently
improved performance.</p>
      <p>These related works collectively contribute valuable insights and methodologies that inform
the development of hate speech detection and associated natural language processing tasks,
showcasing the advancements and challenges in this field.</p>
    </sec>
    <sec id="sec-9">
      <title>8. Conclusion and Future Scope</title>
      <p>In this research, we have presented ATLANTIS (Attentive Transformer-LSTM for Named Entity
and Token Identification System), a robust model designed to combat hate speech at scale.
Leveraging a Hate dataset with detailed Named Entity Recognition (NER) tags, ATLANTIS
efectively identifies hate speech spans within textual content. Our multi-stage architecture,
comprising a custom transformer and bidirectional LSTM, captures contextual information
and sequential dependencies, facilitating precise hate span classification. Empirical results
demonstrate ATLANTIS’s efectiveness in this critical task. As we continue to address the
pressing issue of hate speech in digital spaces, ATLANTIS ofers a promising solution for safer
online environments.</p>
      <p>The work presented here lays the foundation for future advancements in hate span detection.
Further improvements in model eficiency and interpretability, along with expanded training
data encompassing diverse linguistic nuances and evolving hate speech trends, hold promise.
Investigating the integration of real-time monitoring and incorporating user-specific context
may enhance the model’s capabilities in dynamically changing online environments.
Additionally, exploring multilingual and cross-platform hate speech detection is vital for broader
impact. As technology evolves, ATLANTIS and its successors are poised to play a pivotal role
in fostering safer, more inclusive digital spaces.</p>
    </sec>
    <sec id="sec-10">
      <title>Acknowledgments</title>
      <p>We wish to extend our appreciation to the Computer Science and Engineering Department of
the National Institute of Technology Silchar for granting us the opportunity to carry out our
research and experiments. We are grateful for the support, resources, and research environment
ofered by the CNLP &amp; AI Lab at NIT Silchar.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>B.</given-names>
            <surname>Vidgen</surname>
          </string-name>
          , L. Derczynski, (
          <year>2020</year>
          ),
          <article-title>Directions in abusive language training data, a systematic review: Garbage in, garbage out</article-title>
          .
          <source>PloS one 15</source>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P.</given-names>
            <surname>Fortuna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Nunes</surname>
          </string-name>
          ,
          <article-title>A survey on automatic detection of hate speech in text, ACM Computing Surveys (CSUR) 51 (</article-title>
          <year>2018</year>
          )
          <fpage>1</fpage>
          -
          <lpage>30</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vetagiri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. K.</given-names>
            <surname>Adhikary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Pakray</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Das</surname>
          </string-name>
          ,
          <string-name>
            <surname>“CNLP-NITS at</surname>
          </string-name>
          SemEval-2023
          <source>Task</source>
          <volume>10</volume>
          :
          <article-title>Online sexism prediction</article-title>
          , PREDHATE!“,
          <source>In the 17th International Workshop on Semantic Evaluation SemEval 2023 Toronto, Canada July 9-14</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Jurgens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Hemphill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Chandrasekharan</surname>
          </string-name>
          ,
          <article-title>A just and comprehensive strategy for using NLP to address online abuse</article-title>
          ,
          <source>Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vetagiri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. K.</given-names>
            <surname>Adhikary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Pakray</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Das</surname>
          </string-name>
          , “
          <article-title>Leveraging GPT-2 for Automated Classification of Online Sexist Content“</article-title>
          , In Exist 2023 Lab at CLEF 2023:
          <article-title>Conference and Labs of the Evaluation Forum</article-title>
          ,
          <source>September 18-21</source>
          ,
          <year>2023</year>
          , Thessaloniki, Greece,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>B.</given-names>
            <surname>Mathew</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Saha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Yimam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Biemann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Mukherjee,</surname>
          </string-name>
          (
          <year>2021</year>
          ),
          <article-title>Hatexplain: A benchmark dataset for explainable hate speech detection</article-title>
          .
          <source>Proceedings of the AAAI Conference on Artificial Intelligence</source>
          <volume>35</volume>
          (
          <year>2021</year>
          )
          <fpage>14867</fpage>
          -
          <lpage>14875</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Masud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Khan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Akhtar</surname>
          </string-name>
          , T. Chakraborty,
          <article-title>Overview of the HASOC Subtrack at FIRE 2023: Identification of Tokens Contributing to Explicit Hate in English by Span Detection</article-title>
          , in: Working Notes of FIRE 2023 -
          <article-title>Forum for Information Retrieval Evaluation</article-title>
          ,
          <string-name>
            <surname>CEUR</surname>
          </string-name>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mandl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Modha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Majumder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Patel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dave</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Mandlia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Patel</surname>
          </string-name>
          ,
          <source>Overview of the HASOC track at FIRE</source>
          <year>2019</year>
          :
          <article-title>Hate speech and ofensive content identification in IndoEuropean languages</article-title>
          ,
          <source>Proceedings of the 11th Forum for Information Retrieval Evaluation</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Satapara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Masud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Madhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Khan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Akhtar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Chakraborty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Modha</surname>
          </string-name>
          , T. Mandl,
          <article-title>Overview of the HASOC subtracks at FIRE 2023: Detection of hate spans and conversational hate-speech, in: Proceedings of the 15th Annual Meeting of the Forum for Information Retrieval Evaluation</article-title>
          ,
          <string-name>
            <surname>FIRE</surname>
          </string-name>
          <year>2023</year>
          , Goa,
          <source>India. December 15-18</source>
          ,
          <year>2023</year>
          , ACM,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>G.</given-names>
            <surname>Lample</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ballesteros</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Subramanian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kawakami</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Dyer</surname>
          </string-name>
          ,
          <article-title>Neural architectures for named entity recognition</article-title>
          ,
          <source>arXiv preprint arXiv:1603</source>
          (
          <year>2016</year>
          )
          <fpage>01360</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , L. Luo,
          <article-title>Hate speech detection: A solved problem? The challenging case of long tail on Twitter</article-title>
          ,
          <source>Semantic Web</source>
          <volume>10</volume>
          (
          <year>2019</year>
          )
          <fpage>925</fpage>
          -
          <lpage>945</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>