<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Detection of deceptions in Twitter and News Headlines written in Arabic</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Francisco Eros Blazquez del Rio</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Manuel Conde Rodr guez</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jose M. Escalante</string-name>
          <email>3jmescalantefernadez@gmail.com</email>
        </contrib>
      </contrib-group>
      <abstract>
        <p>In this work we present an model to detect deceptive texts (twitters and news headlines) written in Arabic. To develop the model we have focused on the several characteristics to di erentiate them from the regular texts, such as: numbers, special characters or n-grams, which could be a new type of deception indicator. The used classi er method is Support Vector Machine, due to it has excellent performance in di erent low and high-dimensional classi cation tasks.</p>
      </abstract>
      <kwd-group>
        <kwd>Support Vector Machine</kwd>
        <kwd>Stop of Words</kwd>
        <kwd>N-grams</kwd>
        <kwd>Special</kwd>
        <kwd>Characters</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Currently, the amount of information exchanged by users on the network has
reached dimensions that at the beginning of century were considered
unimaginable. There are more than 2.5 quintillion bytes of data created every day[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
Over the last two years alone 90 percent of the data in the world was generated
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. A large amount of the data generated is due to the explosion of social
networks, such as Facebook, Twitter or Instagram [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], which has been favored by
the exponential increase in internet speed in the last decade [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] together with
the development of wireless communication technology [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Among the jungle
of di erent social networks, Twitter (microblogging network of just 140
characters) has been the fastest expanding social network considering its simplicity.
Its strength lies in its instantaneity to access information [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Due to the
instantaneity of this social network, it has a high vulnerability to the propagation of
hoaxes, opinions of doubtful credibility and deceptive news.
      </p>
      <p>
        Although currently we have in the era of information. The misinformation
oods our daily life [
        <xref ref-type="bibr" rid="ref6 ref7">6,7</xref>
        ]. This fact makes mandatory the veri cation of the
veracity of a news or comments posted. The ability of these deceptions (hoaxes
and fake news) to shape people's opinion is somewhat worrisome [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], doing that
the propagation of hoax on Twitter and the fake news made are highly correlated
due to the characteristic of this social network mentioned above.
      </p>
      <p>
        There is an increasing interest in deception detection in the context of
intelligenceservices, law enforcement or intra-organizational monitoring, but also from
a civil point view [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Though it is expected that human capability to indulge
in deception would be balanced by the ability to detect deception, it has been
borne out in studies [
        <xref ref-type="bibr" rid="ref10 ref11 ref12">10,11,12</xref>
        ] that humans are not naturally good in deception
detection. In fact, even trained personnel correctly identify deception at rates
only a little better than chance (52 percent accuracy) [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>In contrast, software systems based on number-crunching algorithms and
correlation analysis often provide surprisingly accurate results. This is especially
true in real-time applications where computerized systems can quickly and e
ciently parse huge datasets, agging objects of interest that can then be sent for
nal analysis and scrutiny by human judges.</p>
      <p>
        Psychological studies [
        <xref ref-type="bibr" rid="ref12 ref13 ref14">12,13,14</xref>
        ] in the area of deception detection have shown
that changes in behaviour such as the body posture, facial expressions, speech
rhythm and pitch are closely correlated with deception. Markers of deception
that are under conscious control (like the content of the deceptive `story') may
be modi ed by those who wish to deceive, but fortunately most of these cues
are a result of sub conscious processes and therefore even awareness of vigilance
and scrutiny does not make deception any easier.
      </p>
      <p>
        In the context of deception in text, though communication is stripped to
its essentials and no non-verbal deception-cues are transmitted, the linguistic
manifestations of deception remain consistent across most domains. It has been
empirically shown [
        <xref ref-type="bibr" rid="ref11 ref13 ref15">13,11,15</xref>
        ] that deception leaves a linguistic signature caused
largely due to the high demands that indulging in deception generates on a
person's cognitive capabilities.
      </p>
      <p>
        Many research groups in the eld of psychology [
        <xref ref-type="bibr" rid="ref15 ref16 ref17">16,15,17</xref>
        ] have constructed
sets of linguistic markers for deception detection in text using di erent
methodologies. James Pennebaker et al. at the Department of Psychology, University of
Texas [
        <xref ref-type="bibr" rid="ref14 ref18">14,18</xref>
        ] have constructed an empirical deception model based on cue-word
usage-frequency pro les. Though this linguistic signature of deception is not easy
for humans to detect directly, it is easily detected by software. According to the
model, deception in text is marked by:
{ Decreased frequency of rst-person pronouns (I, mine, myself etc.) { a
subconscious attempt by the author to disassociate from the deceptive content.
{ Decreased frequency of exclusive words (or, but, without etc.) { to keep the
content simple, concrete and without abstractions in order to avoid faltering
while being repeatedly interrogated.
{ Increased frequency of negative emotion words (anger, abandon, hate etc.) {
a re ection of the subconscious feeling of guilt involved with being deceptive.
{ Increased frequency of action verbs (move, run, lead etc.) { as a form of
distraction to keep the `story' moving while the basic content remains simple
and insigni cant.
      </p>
      <p>
        In this work, we try another approach much simpler. The kind of texts
considered (twitters and news headlines) have other special characteristics that
differentiate them from a regular text. Generally, in this kind of texts it is used
numbers, special character or combination of words (n-grams, 2-grams and
4grams) to capture the attention of readers, so the study will be focused in this
kind of structures as main signatures of deception[
        <xref ref-type="bibr" rid="ref19">19</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>Support Vector Machines model</title>
      <p>
        Support Vector Machine (SVM) model is a supervised learning approach
introduced by Vapnik in 1995 for solving two-class pattern recognition problem [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
It is based on the Structural Risk Minimization principle for which error-bound
analysis has been theoretically motivated by the works [
        <xref ref-type="bibr" rid="ref20 ref21">20,21</xref>
        ]. The method is
de ned over a vector space where the problem is to nd a decision surface that
best separates the data points in two classes. This model has a excellent
performance in di erent classi cation task [
        <xref ref-type="bibr" rid="ref22 ref23">22,23</xref>
        ] comparing with other methods, such
as: k-Neares Neighbor (k-NN), Neural Networks (NNet), Linear Leas-Square Fit
(LLSF) or Naive Bayes (NV). Figure a shows the training ans test schema of
the model.
(a)
      </p>
      <p>(b)</p>
    </sec>
    <sec id="sec-3">
      <title>Methodology</title>
      <p>The code, written in R-code initially generates a bag of words (BoW) (n = 1000,
number of BoW) and uses a cross validation method (k = 3 and r = 1, where
k and r are parameter to control the cross validation method performance of
R-code package) to train a SVM model (method="svmlinear").</p>
      <p>After several computation with di erent values of n, k and r, the chosen
values were: n = 100, k = 10 and r = 5, achieving a balance between
computation time neccesary and the accuracy obtained. Although the reduction in the
number of BoW is important, the accuracy is not strongly a ected, going from a
accuracy of 0.73 to 0.62, but with an important reduction of computation time.</p>
      <p>It could be thought that the length of BoW is too short, but taking into
account the classes of texts considered (no more 140 characters in the longest
case) with n = 100 it is enough. This is supported by Fig. 1b, where it is
plotted an histogram the frequencies of every word of BoW. We can see that
approximately the rst 25 words (n=25) concentrate the highest frequencies of
appearance in the texts considered.</p>
      <p>
        The methodology followed is as follows:
1. Initially, we train the two variant of SVM model: SVMlinear and
SVMlinear3 [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ], with the BoW generated.
2. Next, we train the two models with a new BoW where the numbers, special
character and Stop Words 1 have been removed.
3. Next, we train the two models with the N-grams(2 and 4-grams)
generated using the texts without removing numbers, special character and Stop
Words.
4. Finally, the models are trained with the N-grams generated after removing
numbers, special character and Stop Words.
      </p>
      <p>
        The results of accuracies obtained are summarized on the Tables 1 (Twitter)
and 2 (News Headlines).
1A set of Stop Words is any set of words can be chosen as the stop words for a given
purpose [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. In our case this set of words does not provide any relevant information
to decide if a text is true or false. In this work we have use a library of RStudio called
"arabicStemR"[
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] to detect and remove this words.
{ BoW-SVM/SVM3! The model SVM Linear/SVM Linear 3 has been just
trained with BoW.
{ BoW-C-SVM/SVM3! The model SVM Linear/SVM Linear 3 has been
trained with BoW after removing from text numbers, special characters
and Stop Words.
{ BoW-Ng-SVM/SVM3! The model SVM Linear/SVM Linear 3 has been
just trained with BoW and N-grams without removing from texts numbers,
special characters and Stop Words.
{ BoW-C-Ng-SVM/SVM3! The model SVM Linear/SVM Linear 3 has
been just trained with BoW and N-grams removing from texts numbers,
special characters and Stop Words.
      </p>
      <p>In general we observe that SVM Linear 3 works better than SVM Linear.
Also, in general, we observe that the accuracies are better for Twitter (less
for case BoW-C-Ng-SVM) than for News Headlines. It could be due to News
Headline text length are shorter than Twitter text length, this gives us a training
corpus of words very di erent.</p>
      <p>On the other hand, while we observe a certain growth trend in News Headlines
with the methods used to train the model, for Twitter case this tendency is not
so clear. Therefore, we should use di erent methodologies, adapting them to the
characteristic of the text.</p>
      <p>Finally, we observe, in a generalized way, that the accuracy is a ected when
Data Cleaning is applied. We think that this e ect comes from the use of library
called "arabicStemR" to remove Stop Words. Maybe, it could be solve using
a set of Stop Words more adapted to Arabic language used in Twitter.</p>
    </sec>
    <sec id="sec-4">
      <title>Forward works</title>
      <p>Considering the results obtained above, we have detected several de ciencies
which have to be improved, such as:
{ Study much deeper the e ect of Data Cleaning: detect character more
common in twitter written in Arabic and new set of Stop Word.
{ Adjust much better the di erent parameter of SVM model in R-code.
{ Study deeper the e ect of N-gram.
{ Adapt the methodologies to the characteristics of the text.</p>
      <p>On the other hand, we propose some future steps to improve the model, such
as:
{ Detection and check the URLS.</p>
      <p>This kind of structure is very common in Twitter, so if the twitter share an
URL of a fake news o fake website we can consider the twitter as deceptive.
{ Detection and checking the hashtags and Twitter accounts found on the
texts, trying to see if them are related with fake hashtags and Twitter
accounts.</p>
      <p>This kind of structure are very common in Twitter, so if the twitter retweet
a hashtag which is fake, we can consider the twitter as deceptive. The same
for Twitter accounts.
{ The study of N-grams di erent of 2-grams and 4-grams.</p>
      <p>It would be interesting to study N-grams with other lengths.</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>We have observed that the process of cleaning data can be a ected negatively
the accuracy of the model since we could remove relevant information. Mainly,
News Headlines are a ected stronger than Twitter. The explanation can be base
on the number of words in Twitter is bigger.</p>
      <p>Considering our approach, we have observe that just the use of N-gram
improves the model, mainly in News Headlines, showing that there are structure
of two and four word which can be indicator of deceptive text for this type of
texts. In other words, this would be able to show that N-grams structures have
a special characteristics in deceptive text for this type of texts.</p>
      <p>On the other hand, we see that the methodologies have to be adapted to the
characteristics of the text.</p>
      <p>Finally we have proposed several improvements which can give a way to new
jobs in the same line as this work, such as: detection and check of fake URLs,
hashtag or Twitter accounts.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Bernard</given-names>
            <surname>Marr</surname>
          </string-name>
          ,
          <article-title>"How Much Data Do We Create Every Day? The Mind-Blowing Stats Everyone Should Read"</article-title>
          , https://www.forbes.com/
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Jamie</surname>
          </string-name>
          ,
          <article-title>"84 thoughts on \65+ Social Networking Sites You Need to Know About"</article-title>
          , https://makeawebsitehub.com/social-media-sites/
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Brian</given-names>
            <surname>Patrick</surname>
          </string-name>
          <string-name>
            <surname>Eha</surname>
          </string-name>
          ,
          <article-title>"An Accelerated History of Internet Speed (Infographic)"</article-title>
          , https://www.entrepreneur.com/
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Rajiv</surname>
          </string-name>
          ,
          <article-title>"Evolution of wireless technologies 1G to 5G in mobile communication"</article-title>
          , https://www.rfpage.com/
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Maria</given-names>
            <surname>Valera</surname>
          </string-name>
          , "Historia de Twitter:
          <article-title>de un comienzo brillante a los rumores sobre su futuro incierto"</article-title>
          , https://marketing4ecommerce.net/
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Lautaro</given-names>
            <surname>Rubbi</surname>
          </string-name>
          ,
          <article-title>"The Age of Disinformation"</article-title>
          , https://www.thebubble.com/
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Rasha</given-names>
            <surname>El Hallak</surname>
          </string-name>
          ,
          <article-title>"In an era of information abundance, how much disinformation can we handle?"</article-title>
          , https://blog.usejournal.com/
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Meredith</given-names>
            <surname>Wilson</surname>
          </string-name>
          ,
          <article-title>"Disinformation is changing the way we view the world"</article-title>
          , https: //emergentriskinternational.com/
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Jaume</given-names>
            <surname>Masip</surname>
          </string-name>
          ,
          <article-title>"Deception detection: State of the art and future prospects"</article-title>
          <source>Psicothema</source>
          , Vol.
          <volume>29</volume>
          , No.
          <issue>2</issue>
          , pp.
          <fpage>149</fpage>
          -
          <lpage>159</lpage>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>P.</given-names>
            <surname>Ekman</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. O'Sullivan</surname>
            , and
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Frank</surname>
          </string-name>
          .
          <article-title>A Few Can Catch A Lair</article-title>
          .
          <source>In Psychological Science</source>
          , pages
          <volume>10</volume>
          :
          <fpage>263</fpage>
          {
          <fpage>266</fpage>
          (
          <year>1999</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>M.L. Newman</surname>
            ,
            <given-names>J.W.</given-names>
          </string-name>
          <string-name>
            <surname>Pennebaker</surname>
            ,
            <given-names>D.S.</given-names>
          </string-name>
          <string-name>
            <surname>Berry</surname>
            , and
            <given-names>J.M.</given-names>
          </string-name>
          <string-name>
            <surname>Richards</surname>
          </string-name>
          . Lying Words:
          <article-title>Predicting Deception from Linguistic Styles</article-title>
          .
          <source>In Personality and Social Psychology Bulletin</source>
          , pages
          <volume>29</volume>
          :
          <fpage>665</fpage>
          {
          <fpage>675</fpage>
          (
          <year>2003</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>D.P.</given-names>
            <surname>Twitchell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Jr J.F.</given-names>
            <surname>Nunamaker</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.K.</given-names>
            <surname>Burgoon</surname>
          </string-name>
          .
          <article-title>Using Speech Act Pro ling for Deception Detection</article-title>
          .
          <source>In Intelligence and Security Informatics: Second Symposium on Intelligence and Security Informatics</source>
          ,
          <string-name>
            <surname>ISI</surname>
          </string-name>
          <year>2004</year>
          , pages
          <fpage>403</fpage>
          {
          <fpage>410</fpage>
          (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>B.M. DePaulo</surname>
            ,
            <given-names>J.J.</given-names>
          </string-name>
          <string-name>
            <surname>Lindsay</surname>
            ,
            <given-names>B.E.</given-names>
          </string-name>
          <string-name>
            <surname>Malone</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Muhlenbruck</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Charlton</surname>
            , and
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Cooper</surname>
          </string-name>
          . Cues to Deception. In Psychology Bulletin, pages
          <volume>9</volume>
          :
          <fpage>74</fpage>
          {
          <fpage>118</fpage>
          (
          <year>2002</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <given-names>P.</given-names>
            <surname>Ekman</surname>
          </string-name>
          .
          <article-title>Why Lies Fail and What Behaviors Betray A Lie</article-title>
          . In J.C. Yuille (Ed.) Credibility Assessment, pages
          <volume>71</volume>
          {
          <fpage>81</fpage>
          (
          <year>1989</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15. L.
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>D.P.</given-names>
          </string-name>
          <string-name>
            <surname>Twitchell</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Qin</surname>
            ,
            <given-names>J.K.</given-names>
          </string-name>
          <string-name>
            <surname>Burgoon</surname>
            , and
            <given-names>Jr J.F.</given-names>
          </string-name>
          <string-name>
            <surname>Nunamaker</surname>
          </string-name>
          .
          <article-title>An Exploratory Study into Deception Detection in Text-Based Computer Mediated Communication</article-title>
          .
          <source>In Proceedings of the 36th Hawaii International Conference on Systems Science</source>
          (
          <year>2003</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <given-names>D.P.</given-names>
            <surname>Biros</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sakamoto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.F.</given-names>
            <surname>George</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Adkins</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kruse</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.K.</given-names>
            <surname>Burgoon</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>Jr. J.F.</given-names>
            <surname>Nunamaker</surname>
          </string-name>
          .
          <article-title>A quasi-experiment to determine the impact of a computer based deception detection training system: The use of Agent 99 trainer in the US military</article-title>
          .
          <source>In Proceedings of the 38th Hawaii International Conference on Systems Science</source>
          , volume
          <volume>1</volume>
          (
          <year>2005</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17. L.
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>J.K.</given-names>
          </string-name>
          <string-name>
            <surname>Burgoon</surname>
            ,
            <given-names>Jr J.F.</given-names>
          </string-name>
          <string-name>
            <surname>Nunamaker</surname>
            , and
            <given-names>D.P.</given-names>
          </string-name>
          <string-name>
            <surname>Twitchell</surname>
          </string-name>
          .
          <article-title>Automating Linguistic-Based Cues for Detecting Deception in Text-Based Asynchronous Computer-Mediated Communication</article-title>
          . In Group Decision and Negotiation Vol.
          <volume>13</volume>
          , pp:
          <volume>81</volume>
          {
          <fpage>106</fpage>
          (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <given-names>J.W.</given-names>
            <surname>Pennebaker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.E.</given-names>
            <surname>Francis</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.J.</given-names>
            <surname>Booth</surname>
          </string-name>
          .
          <article-title>Linguistic Inquiry and Word Count (LIWC)</article-title>
          .
          <source>Technical report</source>
          , Erlbaum Publishers,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Char</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zaghouani</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ghanem</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Snchez-Junquera</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Overview of the track on author pro ling and deception detection in arabic</article-title>
          . In: Mehta P.,
          <string-name>
            <surname>Rosso</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Majumder</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mitra</surname>
            <given-names>M</given-names>
          </string-name>
          . (Eds.)
          <article-title>Working Notes of the Forum for Information Retrieval Evaluation (FIRE 2019)</article-title>
          . CEUR Workshop Proceedings. In: CEUR-WS.org, Kolkata, India, December
          <volume>12</volume>
          -
          <fpage>15</fpage>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20. V.
          <article-title>Vapnic, "The Nature of Statistical Learning Theory"</article-title>
          ,
          <source>Edt</source>
          . Springer (
          <year>1995</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <given-names>C.</given-names>
            <surname>Cortez</surname>
          </string-name>
          and
          <string-name>
            <given-names>V.</given-names>
            <surname>Vapnic</surname>
          </string-name>
          ,
          <article-title>"Support Vector Networks"</article-title>
          ,
          <source>Machine Learning</source>
          Vol.
          <volume>20</volume>
          , pp:
          <fpage>273</fpage>
          -
          <lpage>297</lpage>
          (
          <year>1995</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Yang</surname>
            <given-names>Y. M.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Liu</surname>
            <given-names>X.</given-names>
          </string-name>
          ,
          <article-title>"A Re-Examination of Text Categorization Methods"</article-title>
          ,
          <source>Proceeding of SIGIR-99,22 nd ACM International Conference on Research and Development in Information Retrieval</source>
          (
          <year>1999</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Hu</surname>
            <given-names>Zhang</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shan-de</surname>
            <given-names>WEei</given-names>
          </string-name>
          ,
          <article-title>Hong-ye Tan and Jia-heng Zheng, "A Study on Deception Detection Based on Classi cation for Chinese Text"</article-title>
          ,
          <source>Journal of Computational Information Systems</source>
          Vol.
          <volume>3</volume>
          No.
          <issue>5</issue>
          , pp:
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24. A.nand Rajaraman and
          <string-name>
            <given-names>J.</given-names>
            <surname>Ullman</surname>
          </string-name>
          ,
          <article-title>"Mining of Massive Datasets"</article-title>
          , Ed.
          <source>Cambrige 2nd Edition</source>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <article-title>Documentation of"arabicStemR"</article-title>
          , https://www.rdocumentation.org/packages/ arabicStemR/versions/1.2
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <given-names>Rahul</given-names>
            <surname>Saxena</surname>
          </string-name>
          ,
          <article-title>"Support Vector Machine Classi er Implementation in R with caret package"</article-title>
          , https://dataaspirant.com/
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>