<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Author Profiling f or Arabic Tweets based on n-grams</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ayoub Abbassi</string-name>
          <email>ayoub.abess@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Seifeddine Mechti</string-name>
          <email>mechtiseif@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lamia Hadrich Belguith</string-name>
          <email>l.belguith@fsegs.rnu.tn</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rim Faiz</string-name>
          <email>Rim.faiz@ihec.rnu.tn</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>ANLP Group MIRACL Laboratory</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>FSEGS, University of Sfax</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>LARODEC Laboratory, ISG of Tunis IHEC</institution>
          ,
          <addr-line>2016 Carthage</addr-line>
          ,
          <country country="TN">Tunisia</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>LARODEC Laboratory, ISG of Tunis</institution>
          ,
          <addr-line>2000 Le Bardo</addr-line>
          ,
          <country country="TN">Tunisia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents an approach for author profiling of an unknown users from their texts produced in social media. In particular, we address the identification of two profile dimensions: gender and language variety, of Arabic twitter users based on their tweets. Our approach focused on applying metaclassification technique on features extracted from tweets body. We explored two main sets of features which are character and word n-grams. The proposed approach allowed us to reach promising results for both language variety and gender identification</p>
      </abstract>
      <kwd-group>
        <kwd>Author profiling</kwd>
        <kwd>Meta-classifier</kwd>
        <kwd>N-gram features</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The rapid growth of internet and computer technology during the last two decades
makes humanity in front of an incredible increased amount of online data. According
to internet live stats1, in one second the Internet traffic is about 36,411 GB. This
impressive amount of data -mostly of text type- are shared, published, and transit in a
free (sometime in anonymous) way. In fact, an important portion of internet users are
misrepresenting themselves while surfing in the net, therefore there are a need to deal
with the data that come from unknown source.</p>
      <p>Two main sectors are interested in knowing the potential source of data. First, the
commercial sector where information such as age, gender, nationality, and native
language about customers is of higher value for marketing intelligence. Second, the
1 http://www.internetlivestats.com/
security sector that bear the burden to protect the internet from crime such as
plagiarism and identity theft, etc.</p>
      <p>Therefore, research community promotes researchers to discover and develop
effective methods and techniques in related fields such as plagiarism detection and
author profiling.</p>
      <p>This work is made in the context of the participation of the Author Profiling task in
the PAN17 shared task2. In particular, we focus on identifying the gender and
Language variety of Arabic users from their twitter tweets.</p>
      <p>2</p>
    </sec>
    <sec id="sec-2">
      <title>Dataset Description</title>
      <p>We used training dataset provided by PAN clef 2017 to train our proposed system.
We participated in the author-profiling task for the Arabic subtask. The training
dataset is composed of Twitter tweets and annotated with authors' gender and their
specific variation of their native language. A detailed statistics of the used dataset is
given in Table 1.</p>
      <p>
        As the above table shows, it is clear that the turning dataset is well distributed
across classes. However, the analysis reveals that some documents are written in
Modern Standard Arabic, not in one of the Arabic varieties [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], which can affect the
performance of our system.
      </p>
      <p>3</p>
    </sec>
    <sec id="sec-3">
      <title>System Architecture</title>
      <p>Our proposed system is divided into three steps: pre-processing, feature extraction
and Classification. Firstly, in the pre-processing step, we prepare the input data to be
used in the next step. Then, in the feature extraction phase, we extract the set of
features that seem to be useful for the task. Finally, we generate the classification
model. This model will be used to predict the class of new document.
2 http://pan.webis.de/clef17/pan17-web/</p>
    </sec>
    <sec id="sec-4">
      <title>3.1 Pre-processing</title>
      <p>As the input dataset is basically composed of Twitter tweets, these tweets have the
nature of being noisy including a lot of useless data such as links, tags, emoticons, etc.
Thus they can’t be exploited directly. The idea is to remove these noisy data.
However, in stand of looking for the variety of noisy, we simply extracted the Arabic
text. The example below shows a tweet before and after prepossessing step.
Example:
هطلاكش تلصح تحبر اناو لاؤس ع يدحت يف ناك

https://t.co/UySVCM1qwm https://t.co/wKBUpGXmZo“
Tweet after extract the Arabic text: “هطلاكش تلصح تحبر اناو لاؤس ع يدحت يف ناك
”</p>
    </sec>
    <sec id="sec-5">
      <title>3.2 Features extraction</title>
      <p>We extracted tow n-gram feature types, namely ‘character n-grams’ and ‘word
ngrams’. Accordingly, we generated two sets of features for each input document. For
each individual feature, we calculated the Inverse Document Frequency (IDF) with
which it appears. The documents are then represented as TF-IDF matrix.</p>
      <p>Given a text extracted from tweets, the set of n-grams was extracted by moving a
window of n cases across the text body. For example, based on the word as a feature,
word n-grams means all the n consecutive words in the text.</p>
      <p>For the previous tweet "هطلاكش تلصح تحبر اناو لاؤس ع يدحت يف ناك
", the word n-gram
model is illustrated in Table 2.</p>
    </sec>
    <sec id="sec-6">
      <title>3.3 Classification</title>
      <p>
        Once the documents have been transformed to their new representation, they will
be used as input to train the classifier. Training the classifier is the main key of this
work, we apply a meta-classifier technique known as 'stacking' [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] to generate the
finale module, which will be used to predict the correct class of unlabelled document.
Stacking consist in combining several base classifiers of different type, in our case,
we use the three most popular machine-learning algorithm (Support vector machines,
Decision trees and Naive Bayes) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].The principle of this technique is illustrated in
the following figure:
      </p>
      <p>We carried out several series of experiments in order to evaluate the performance
of the classifiers mentioned before individually and combined, using different sets of
features. Table 3 and Table 4 show the result of our experiments:</p>
      <p>
        For gender dimension, the best accuracy is 59.3 which is obtained using SVM, in
the case of individual classifier, and 63.2 using Stacking as classification technique.
These results are obtained by combining all features together. Such results confirm
our finding [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] of the outperformance of SVM compared with other learning
algorithms in author profiling problem.
      </p>
      <p>However, for language variety, the result obtained using word n-grams
outperformed those obtained using character n-grams or combination with 34 of
accuracy. This is obtained by combining (Stacking) the performance of classifier.
5</p>
    </sec>
    <sec id="sec-7">
      <title>Conclusion</title>
      <p>In this paper we described our approach of profiling the users of Twitter based on
meta-classifier trained on n-grams features. In particular, we focused on the
identification of gender and language variety of Arabic users. We found out that
combining the n-grams- features in a meta-classification process allowed us to
achieve higher results, on the tow tasks. The best result are obtained using word
ngrams for language variety detection and using all features combined for gender
detection.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>A.</given-names>
            <surname>FARGHALY</surname>
          </string-name>
          and
          <string-name>
            <surname>K.</surname>
          </string-name>
          <article-title>SHAALAN “Arabic natural language processing: Challenges and solutions”</article-title>
          ,
          <source>in proceedings of ACM Transactions on Asian Language Information Processing (TALIP)</source>
          , vol.
          <volume>8</volume>
          , no 4,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>S.B.</given-names>
            <surname>Kotsiantis</surname>
          </string-name>
          .,
          <string-name>
            <surname>I. Zaharakis</surname>
          </string-name>
          , and
          <string-name>
            <surname>P. Pintelas. “</surname>
          </string-name>
          <article-title>Supervised machine learning: A review of classification techniques”</article-title>
          , p.
          <fpage>3</fpage>
          -
          <lpage>24</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>K.</given-names>
            <surname>Vandana</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Namrata</surname>
          </string-name>
          , “
          <article-title>Text classification and classifiers: a survey”</article-title>
          <source>Artificial Intelligence &amp; Applications</source>
          , vol.
          <volume>3</volume>
          , n. 2,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Mechti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Abbassi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. H.</given-names>
            <surname>Belguith</surname>
          </string-name>
          ,
          <string-name>
            <surname>R.</surname>
          </string-name>
          , Faiz,,
          <string-name>
            <surname>C.</surname>
          </string-name>
          “
          <article-title>An empirical method using features combination for Arabic native language identification”</article-title>
          ,
          <source>in proceedings of the 3th International Conference of Computer Systems and Applications (AICCSA)</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>