<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>A. S. Kumar); bharathib@ssn.edu.in
(B. Bharathi); bhuvanaj@ssn.edu.in (J. Bhuvana); mirnalineett@ssn.edu.in (T.T Mirnalinee)
~ https://www.ssn.edu.in/staf-members/dr-b-bharathi/ (B. Bharathi);
https://www.ssn.edu.in/staf-members/dr-j-bhuvana/ (J. Bhuvana);
https://www.ssn.edu.in/staf-members/dr-t-t-mirnalinee/ (T.T Mirnalinee)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Abusive and Threatening Language Detection in Native Urdu Script Tweets Exploring Four Conventional Machine Learning Techniques and MLP</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>A. Karthikraja</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aarthi Suresh Kumar</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>B. Bharathi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jayaraman Bhuvana</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>T.T Mirnalinee</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of CSE Sri Sivasubramaniya Nadar College of Engineering</institution>
          ,
          <addr-line>Chennai, Tamil Nadu</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>The lack of clarity in rules imposed on discussions on social media and the lack of critical eyes on discussions in regional languages, unlike the languages with a greater audience like English, the vulgarity of most of the discussions go unnoticed. This demands an automated model to classify abusive and threatening messages to maintain decorum in social media platforms. Here in this work we have used classic models from Sklearn library to classify the data given in task HASOC 2021 - Abusive and Threatening language detection in Urdu. It has been observed that the best model for abusive classification was MLP with paraphrase multilang v1 encoding and for threatening language dataset, the best model observed was an nu-SVM.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Abusive language identification</kwd>
        <kwd>Threatening language detection</kwd>
        <kwd>Tf-Idf Vectorization</kwd>
        <kwd>Sentence Transformers</kwd>
        <kwd>Hate Speech Detection</kwd>
        <kwd>Text Classification</kwd>
        <kwd>MLP</kwd>
        <kwd>SVM</kwd>
        <kwd>Urdu</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>The boom in social media platforms has led to a inevitable freedom of self-expression, especially
among communities sharing the same native language and script.This rise of access of various
native language communities to the means of self-expression via the Internet raised the need for
detecting threatening and abusive language in their native scripts. The myriad of variations in
the meaning for the same scripts in a diferent language removes the possibility for one model
classifier for all languages. This creates a need for classifiers in each language,Roman Urdu,
where Urdu is written in English script has seen a lot of input. Here we have created a model to
classify the threatening and abusive nature of a sentence in native Urdu script. Independent
models based on Multi layer perceptron, Logistic Regression, SVM, KNN were used to perform
this classification task.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Literature Survey</title>
      <p>
        Identifying hate speech in social media is one of the most essential tasks to prevent spreading
hatred nowadays. Detection of such hate speech is tedious and the organisations hosting social
media platforms are taking steps to prevent the hate speech spreading through their platforms[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
Several works have been carried out to identify abusive and hate speech in diferent languages.
This section gives the background and state of the current approaches to identify hate speech
in Urdu.
      </p>
      <p>
        Variety of machine learning algorithms such as, Linear regression, SVM, Random Forest,
Naive Bayes and SGD classifier are applied on custom Roman Urdu [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] dataset with a 10 fold
cross-validation. Among all the listed approaches SVM has reported to give 0.774% of accuracy.
      </p>
      <p>
        Diferent models with n-gram pre-processing have been used for Ofensive classification in
Urdu sentences[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In their experiment character trigram preprocessing and logistic regression
proved to be the best model
      </p>
      <p>
        Propaganda Spotting in Online Urdu Language (ProSOUL) is designed to identify the sources
of propaganda in Urdu language [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] . Psycho-linguistic features were extracted using Linguistic
Inquiry and Word Count. NEws LAndscape (NELA) along with TF-IDF, N-grams, Word2Vec
and BERT features are fed to CNN and Logistic regression classifiers. It is reported that out of
all the listed features Word2Vec has outperformed BERT. In [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], describes automatic Abusive
Language Detection in Urdu Tweets.
      </p>
      <p>
        Hate Speech Roman Urdu 2020 (HS-RU-20) corpus has been created in order to classify the
Roman Urdu tweets into three classes namely, Neutral - Hostile, Simple - Complex, and Ofensive
classes [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Both conventional machine learning classifiers and CNN have been applied to detect
the ofensive ones, where an F1-score of 0.90 has been achieved with a Logistic Regression
model.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Datasets[7]</title>
      <p>Classification</p>
      <p>Train set</p>
      <p>Test set
Abusive
Not Abusive
Total</p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ], the overview of the shared task on threatening and abusive detection in Urdu.
Threatening
Non Threatening
Total
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Implementation and Experiments</title>
      <sec id="sec-4-1">
        <title>4.1. MLP Classifier</title>
        <p>Multilayer perceptron (MLP) classifier is a multilayer neural network. It uses back propagation
to tune its weights and learns from the loss function in each iteration. MLP works for even
linearly-unseparable problems. The model used for both the tasks contains 2 hidden layers with
256, 128 neurons respectively, and the neural weights were adjusted through 300 epochs with
learning rate of 0.001. The default ReLU activation function was used for all layers.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Logistic Regression</title>
        <p>
          Logistic Regression is a classical statistical analysis approach that relies on prior observation.
Logistic regression is usually used for classification sort of problems. It uses the sigmoid on the
given parameters to perform binary classification.[
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]
ℎ () =
        </p>
        <p>1
1 + −   
(1)</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Support Vector Machine</title>
        <p>nu-SVC a Support Vector Machine (SVM) classification algorithms was tested and trained with
the given datasets.The best F1-score was achieved for a RBf kernel (where the classes marked
on higher dimensions are governed by a Gaussian radial basis function) for regularization
parameter value in the range 0.3 to 0.5.
4.4. KNN
K Nearest Neighbors (KNN) classification is a clustering algorithm that labels a point with
the class of majority of its neighbours. Our model analyses 50 neighbours for every node to
predict the class of the node. Since the training set was not biased towards any one class,
this algorithm gave reasonable results for both the the classification tasks. The KNN model
requires high dimensional vectors as input which was derived using Tf-Idf vectorization method
available in the sklearn library version 0.0, default version provided in google colab[11]. The
parameters passed were n_neighbors=50, weights=’uniform’, algorithm=’auto’.The same model
configuration was used for both the classification tasks.</p>
      </sec>
      <sec id="sec-4-4">
        <title>4.5. Feature Extraction</title>
        <p>4.5.1. Embeddings for MLP
Two embeddings from the SentenceTransformers module , a Python framework for
state-ofthe-art sentence, text and image embeddings were used to fetch corresponding embedding of the
training set tweets[12]. The train data was lemmatized using the lemmatizer in urduhack, an
NLP library for Urdu language. The lemmatized sentences are transformed to similar embedding
using the below mentioned transformers: ’distiluse-base-multilingual-cased-v2’ and
’paraphrasexlm-r-multilingual-v1’ and trained separately . The results of training the model on encodings of
the lemmatized versions of the sentences was better than just training on raw data. This might
be because of an internal working of the pretrained models used for tokenizing the sentences.
4.5.2. Tf-Idf for KNN, LR, SVM
Tf-Idf Vectorization was used to vectorize the sentences along with a character 10-gram with
a max features of 50000. Character n-gram gave better results than word n-gram. It might be
because of the complex morphology of Urdu that character n-gram works better extracting
features from the samples.[13] While vectorizing, the sentences are converted to lower case in
order to avoid the confusion caused by the case of words in learning the context of the sentence.
Lemmatizing the words did not help in improving the accuracy. It might be because the root
word of any Urdu word does not necessarily have the same meaning in the context or the
derived word has a more precise meaning which helps the model understand the context of the
sentence better. So only the Tf-Idf vector with 10-word gram was fed as input to these models.
4.6. Hardware specification and link to computation
The training and testing of the models was done in google colab. A general purpose RAM size
of 8GB was allotted with a 2.3GHz Intel Xenon CPU was used for training of the above models.
Python note books associated with the abusive and Threatening tasks are given in the link. 1</p>
        <p>The above algorithms with the aforementioned extracted features are trained with k fold
cross validation and the best models and their parameters are tabulated below for each task.</p>
      </sec>
      <sec id="sec-4-5">
        <title>4.7. Performance analysis</title>
        <p>The performance of the proposed system for abusive language detection using training data are
tabulated in Table 3 and training data performance of threatening language detection is shown
in Table 4.</p>
        <p>The training performance shows that MLP-paraphrase and Nu-SVC have performed well for
Abusive language detection and threatening language detection with 89% and 93.2% respectively.
The MLP models trained on the 2 diferent encodings gave almost similar results on training
but paraphrase-xlm-r-multilingual-v1 encoding worked better than the rest on the test data.</p>
        <p>The performance of the proposed system for abusive language detection using test data are
tabulated in Table 5, threatening language detection performance is tabulated in Table 6.
1https://github.com/ask-1710/Abusive-and-Threatening-Language-Detection-Task-in-Urdu</p>
        <p>Model
distiluseMLP
KNN
Nu-SVC
LR</p>
        <p>MLP-paraphrase
Model
MLP-distiluse
KNN
Nu-SVC kernel:RBF
LR
MLP-paraphrase</p>
        <p>F1-Score
0.82
0.82
0.84
0.83
0.81
distiluseMLP
KNN
Nu-SVCrbf
LR</p>
        <p>MLP-paraphrase
Model
distiluseMLP
KNN
Nu-SVC, kernel:RBF
LR
MLP-paraphrase
private F1
private ROC_AUC
public F1
public ROC_AUC</p>
        <p>From Table 5 and Table 6, it has been noted that for both abusive and threatening language
detection task, paraphrase-xlm-r-multilingual-v1 embeddings with MLP models produces better
results than other approaches.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>Spreading hatred to the community on the basis of ethnicity, race, religion and gender is a
menace to the society. Social media applications nowadays serve as a unintended medium for
enabling transfer of such abusive and hatred messages. Techniques have to be developed to
curtail such abusive messages from spreading. In this work, five machine leaning approaches
have been explored to detect the abusive and threatening Language Task in Urdu Language.</p>
      <p>Classical machine learning were able to come close to the MLP for both tasks in terms of
F1scores. This work can be enhanced further by exploring the linguistic features of Urdu and also
other deep learning approaches can be employed with fine tuned parameters for this task.
[11] F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel,
P. Prettenhofer, R. Weiss, V. Dubourg, et al., Scikit-learn: Machine learning in python,
Journal of machine learning research 12 (2011) 2825–2830.
[12] N. Reimers, I. Gurevych, Making monolingual sentence embeddings multilingual using
knowledge distillation, arXiv preprint arXiv:2004.09813 (2020). URL: http://arxiv.org/abs/
2004.09813.
[13] M. P. Akhter, Z. Jiangbin, I. R. Naqvi, M. Abdelmajeed, M. T. Sadiq, Automatic detection
of ofensive language for urdu and roman urdu, IEEE Access 8 (2020) 91213–91226.
doi:10.1109/ACCESS.2020.2994950.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Laub</surname>
          </string-name>
          ,
          <article-title>Hate speech on social media: Global comparisons (</article-title>
          <year>2019</year>
          ). URL: https://www.cfr. org/backgrounder/hate
          <article-title>-speech-social-media-global-comparisons.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>T.</given-names>
            <surname>Sajid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hassan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Gillani</surname>
          </string-name>
          ,
          <article-title>Roman urdu multi-class ofensive text detection using hybrid features and SVM</article-title>
          , in: 2020
          <source>IEEE 23rd International Multitopic Conference (INMIC)</source>
          , IEEE,
          <year>2020</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Akhter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Jiangbin</surname>
          </string-name>
          , I. Naqvi,
          <string-name>
            <given-names>M.</given-names>
            <surname>Abdelmajeed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. T.</given-names>
            <surname>Sadiq</surname>
          </string-name>
          ,
          <article-title>Automatic detection of ofensive language for urdu and roman urdu, IEEE Access PP (</article-title>
          <year>2020</year>
          )
          <fpage>1</fpage>
          -
          <lpage>1</lpage>
          . doi:
          <volume>10</volume>
          .1109/ ACCESS.
          <year>2020</year>
          .
          <volume>2994950</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Kausar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Tahir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Mehmood</surname>
          </string-name>
          ,
          <article-title>Prosoul: a framework to identify propaganda from online urdu content</article-title>
          ,
          <source>IEEE Access 8</source>
          (
          <year>2020</year>
          )
          <fpage>186039</fpage>
          -
          <lpage>186054</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Amjad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Noman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Grigori</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Alisa</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-H. Liliana</surname>
          </string-name>
          , G. Alexander,
          <article-title>Automatic abusive language detection in urdu tweets</article-title>
          ,
          <source>Acta Polytechnica Hungarica</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>M. M. Khan</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Shahzad</surname>
            ,
            <given-names>M. K.</given-names>
          </string-name>
          <string-name>
            <surname>Malik</surname>
          </string-name>
          ,
          <article-title>Hate speech detection in roman urdu</article-title>
          ,
          <source>ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP) 20</source>
          (
          <year>2021</year>
          )
          <fpage>1</fpage>
          -
          <lpage>19</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Amjad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ashraf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zhila</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sidorov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zubiaga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gelbukh</surname>
          </string-name>
          ,
          <article-title>Threatening language detection and target identification in urdu tweets</article-title>
          ,
          <source>IEEE Access 9</source>
          (
          <year>2021</year>
          )
          <fpage>128302</fpage>
          -
          <lpage>128313</lpage>
          . doi:
          <volume>10</volume>
          .1109/ACCESS.
          <year>2021</year>
          .
          <volume>3112500</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Amjad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Alisa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Oxana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Sabur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Hamza</given-names>
            <surname>Imam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Grigori</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Alexander, Overview of the shared task on threatening and abusive detection in urdu at fire 2021</article-title>
          , CEUR Workshop Proceedings (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Amjad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Alisa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Oxana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Sabur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Hamza</given-names>
            <surname>Imam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Grigori</surname>
          </string-name>
          , G. Alexander, Urduthreat@ fire2021:
          <article-title>Shared track on abusive threat identification in urdu, Forum for Information Retrieval Evaluation (</article-title>
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mannor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <article-title>Robust logistic regression and classification</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>27</volume>
          (
          <year>2014</year>
          )
          <fpage>253</fpage>
          -
          <lpage>261</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>