<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Deception Detection in Arabic Texts Using N-grams Text Mining</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jorge Cabrejas</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jose Vicente Martí</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antonio Pajares</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Víctor Sanchis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>jorcabpe</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>jvmao</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>apajares</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>vicsanig}@gmail.com</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universitat Politècnica de València</institution>
          ,
          <addr-line>Camino de Vera s/n, 46022 Valencia</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we present our team participation at Author Profiling and Deception Detection in Arabic (APDA) - Task 2: Deception Detection in Arabic Texts. Our analysis shows that the accuracy of a unigram method outperforms both bigrams and hybrid modeling based on unigram and bigrams. We show that the accuracy of the unigram modeling can achieve performance above 76%.</p>
      </abstract>
      <kwd-group>
        <kwd>Arabic</kwd>
        <kwd>deception</kwd>
        <kwd>n-grams</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Social networks such as Facebook, Twitter or Instagram are excellent
communication channels between customers and diferent brands. According to statistics,
approximately 92% of business-to-business marketers in North America use
social networks as part of their marketing tactics [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. For companies, this is the
paramount importance as a means of customer acquisition. On the other hand,
customers generally send thousands of messages to show their personal
opinions in social networks either to criticize or recommend, for instance, a hotel
or restaurant. Many times these views are unbiased and help companies to
improve their mistakes. However, many others are full of hatred and resentment
and sometimes they cause damage to the reputation of the company.
      </p>
      <p>
        Being aware of the importance of social networks in the commercialization
of the companies, it is key to detect when a comment in social networks is real
or when it is suspected of deception and tries to damage the reputation of the
brand. This problem is constant throughout the world, regardless of the
country. However, for instance, the Arabic language sufers from a lack of Natural
Language Processing (NLP) techniques that does not happen in the English
language. In particular, one of the first contributions can be found in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] where
authors try to detect spam opinions in the Yahoo!-Maktoob social network through
dictionaries of linguistic polarity and machine learning classification techniques.
However, there have only been some articles in the literature that aim to detect
the irony [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], analyze the sentiment [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] or profile authors in Arabic texts [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        In this paper, we present our participation in Author Profiling and
Deception Detection in Arabic (APDA) - Task 2 carried out together with the FIRE
2019 Forum for Information Retrieval Evaluation 12-15 December 2019, Kolkata,
India [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. In the article, we show that word n-grams (hereinafter referred to
ngrams) are appropriate methods to detect when a tweet or a new is false or
true.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Task Description</title>
      <p>Given a set of Arabic messages on Twitter or new headlines, the goal of the task
is to predict when a given message is deceptive when it was written trying to
sound authentic to mislead the reader. Therefore, the target labels are Truth or
Lie. The organizers provided 1 1443 new headlines
(APDA@FIRE.Qatar-Newscorpus.training.csv) and 532 tweets (APDA@FIRE.Qatar-Twitter-corpus.traini
ng.csv) to train our model. Both files have identically the same fields, that is,
an ID (identifier of each text), a label (to describe if a text is true or false),
and a text (words to analyze). It is important to say that some tweets or news
were repeated throughout the file and we decided to delete them so as not to
overtrain the model with those messages. Similarly, organizers also provided
370 tweets (APDA@FIRE.Qatar-Twitter-corpus.test.csv) and 241 new headlines
(APDA@FIRE.Qatar-News-corpus.test.csv). In this case, the files contain only
the ID and the text but not the label, which has to be predicted with a statistical
model.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Deception detection algorithm</title>
      <p>In this section, we analyze the n-grams methods that can be found in the
literature in their standalone version (unigrams or bigrams) or hybrid (unigrams and
bigrams) for text processing.
3.1</p>
      <p>n-Grams
An n-gram model consists of a probabilistic method that attempts to predict
the xi item based on the items xi−(n−1), ..., xi−1, that is:</p>
      <p>Pxi = P (xi|xi−(n−1), ..., xi−1).
(1)
1 https://www.autoritas.net/APDA/corpus/
In our case, an item is a word but it could be a syllable, a set of letters, phonemes,
etc. For instance, for n = 1, the probability of the word example is written along
with the word for in English is very high. However, the probability of the word
can is written together with the word would is zero, Pcan = P (can|would) = 0.</p>
      <p>The probabilistic method proposed in this paper is based on the next steps:
1. We divide L sentences of Twitter or headlines with the following delimiters:
spaces ( ), dots (.), question marks (?), exclamation marks (!), parentheses
(()), semicolons (;) or commas (,).
2. We eliminate a set of words (stop words) that are common in the Arabic
language.
3. We remove punctuation marks, numbers, hashtags, and extra whitespaces.
4. We get all combinations of n words from all sentences. For example, in
English, the sentence In summer the temperatures are high in middle east is
divided in the following combinations: in summer, summer the, the
temperatures, temperatures are, are high, high in, in middle, middle east.
5. We calculate the Pxi probability of all combinations.
6. Given the large number of combinations, we select the T combinations with
the highest probability.
7. We build the final dataset with L rows and T columns.
8. Finally, with machine learning, we classify tweets or headlines.</p>
      <p>Note that this approach has been carried out for the training and test data.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Experiments and Results</title>
      <p>To carry out the Task 2, we use a unigram modeling to predict whether a tweet or
headline is misleading or not, and we compare its performance with the following
methods:
1. Bigrams.
2. A hybrid method that uses unigram and bigrams.</p>
      <p>We use 751 common words to remove them from the texts to analyze 2. Besides,
instead of using the probability as a metric of deception, we use the number of
times each combination appears. Finally, we have tested several machine learning
algorithms with the qdap, tm, XML, splitstackshape, caret, and RWeka libraries
of the R programming language. In particular, we have tested k-Nearest
Neighbour, Neural Networks, Random Forest, and Super Vector Machine (SVM). We
verified that SVM is the learning algorithm that ofers the best performance
with a 10‐fold cross‐validation experiment and three repetitions.</p>
      <p>
        Figure 1 shows the performance evolution of the proposed methods based
on n-grams to detect the deception in Arabic. This performance evolution is
calculated for Twitter (right figure) and new headlines (left figure). We show
the accuracy with the number of combinations T explained in Section 3.1. From
2 https://github.com/mohataher/arabic-stop-words/blob/master/list.txt
the figures, we can conclude that there is a significant performance gap between
Twitter and the headline texts. In general, as the number of combinations
increases, the accuracy increases for both texts. However, 500 combinations could
be good enough to get a good performance at a reasonable cost. We also
highlight an erratic performance behavior in the figures with few combinations (see
the hybrid method unigram and bigrams). Therefore, it is key to select correctly
which combinations can be removed. From both figures, we can conclude that
the unigram method is the best algorithm to detect deception in Arabic texts.
Table 1 shows the best accuracy values that can be achieved with the proposed
methods. Note that the Arabic language is a relatively complex language from
the NLP point of view. For instance, some prepositions are joined with some
words: ملاعلا سٔاك (the world cup) and ملاعلا سٔاكل (for the world cup). This fact causes
a performance loss in the n-grams that should be considered. This problem
can be solved by using handmade rules that separate prefixes and suffixes from
words [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>(a) Twitter
(b) New headlines</p>
      <sec id="sec-4-1">
        <title>Unigram</title>
        <p>76.19
71.45</p>
      </sec>
      <sec id="sec-4-2">
        <title>Bigram</title>
        <p>72.12
60.05</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion and Future Work</title>
      <p>In this paper, we presented our participation in Author Profiling and Deception
Detection in Arabic (APDA). We worked on a unigram approach and compared
its performance with bigrams and a hybrid method that used unigrams and
bigrams. Our proposed approach achieved good performance compared to the
other two methods. In particular, we obtained an accuracy of 76.19% for Twitter
texts and 71.45% for new headlines. As future work, we could include letters as
items instead of words to deal with the complexity of the Arabic morphology
where some prepositions are joined with words.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Statista</surname>
          </string-name>
          .
          <article-title>Social media and user-generated content</article-title>
          . https://www.statista.com/topics/2057/brands-on-social-media/,
          <year>2019</year>
          . [Online; accessed 20-aug-2019].
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Heider</given-names>
            <surname>Wahsheh</surname>
          </string-name>
          , Mohammed Al-Kabi, and
          <string-name>
            <given-names>Izzat</given-names>
            <surname>Alsmadi</surname>
          </string-name>
          .
          <article-title>Spar: A system to detect spam in arabic opinions</article-title>
          .
          <source>12</source>
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Jihen</given-names>
            <surname>Karoui</surname>
          </string-name>
          , Farah Banamara Zitoune, and
          <string-name>
            <given-names>Veronique</given-names>
            <surname>Moriceau</surname>
          </string-name>
          . Soukhria:
          <article-title>Towards an irony detection system for arabic in social media</article-title>
          .
          <source>Procedia Computer Science</source>
          ,
          <volume>117</volume>
          :
          <fpage>161</fpage>
          -
          <lpage>168</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Hossam</surname>
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Ibrahim</surname>
          </string-name>
          , Sherif Mahdy Abdou, and
          <string-name>
            <surname>Mervar</surname>
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Gheith</surname>
          </string-name>
          .
          <article-title>Sentiment analysis for modern standard arabic and colloquial</article-title>
          .
          <source>International Journal on Natural Language Computing (IJNLC)</source>
          ,
          <volume>4</volume>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Rosso</surname>
          </string-name>
          , Francisco Rangel, Irazu Hernandez Farias, Leticia Cagnina, Wajdi Zaghouani, and
          <string-name>
            <given-names>Anis</given-names>
            <surname>Charfi</surname>
          </string-name>
          .
          <article-title>A survey on author profiling, deception, and irony detection for the arabic language</article-title>
          .
          <source>Language and Linguistics Compass</source>
          ,
          <volume>12</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>20</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Francisco</given-names>
            <surname>Rangel</surname>
          </string-name>
          , Paolo Rosso, Anis Charfi, Wajdi Zaghouani, Bilal Ghanem, and
          <string-name>
            <surname>Javier</surname>
          </string-name>
          Sánchez-Junquera.
          <article-title>Overview of the track on author profiling and deception detection in arabic</article-title>
          . In: Mehta P.,
          <string-name>
            <surname>Rosso</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Majumder</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mitra</surname>
            <given-names>M</given-names>
          </string-name>
          . (Eds.)
          <article-title>Working Notes of the Forum for Information Retrieval Evaluation (FIRE 2019)</article-title>
          .
          <source>CEUR Workshop Proceedings. CEUR-WS.org, Kolkata, India, December</source>
          <volume>12</volume>
          -15.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>