<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>FIRE</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>BERT-Based Arabic Social Media Author Pro ling</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Chiyu Zhang</string-name>
          <email>chiyuzh@mail.ubc.ca</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Muhammad Abdul-Mageed</string-name>
          <email>muhammad.mageeed@ubc.ca</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Natural Language Processing Lab The University of British Columbia</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>12</volume>
      <fpage>12</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>We report our models for detecting age, language variety, and gender from social media data in the context of the Arabic author pro ling and deception detection shared task (APDA) [32]. We build simple models based on pre-trained bidirectional encoders from transformers (BERT). We rst ne-tune the pre-trained BERT model on each of the three datasets with shared task released data. Then we augment shared task data with in-house data for gender and dialect, showing the utility of augmenting training data. Our best models on the shared task test data are acquired with a majority voting of various BERT models trained under di erent data conditions. We acquire 54.72% accuracy for age, 93.75% for dialect, 81.67% for gender, and 40.97% joint accuracy across the three tasks.1</p>
      </abstract>
      <kwd-group>
        <kwd>author pro ling identi cation</kwd>
        <kwd>BERT</kwd>
        <kwd>Arabic</kwd>
        <kwd>social media</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The proliferation of social media has made it possible to collect user data in
unprecedented ways. These data can come in the form of usage and behavior
(e.g., who likes what on Facebook), network (e.g., who follows a given user
on Instagram), and content (e.g., what people post to Twitter). Availability
of such data have made it possible to make discoveries about individuals and
communities, mobilizing social and psychological research and employing natural
language processing methods. In this work, we focus on predicting social media
user age, dialect, and gender based on posted language. More speci cally, we use
the total of 100 tweets from each manually-labeled user to predict each of these
attributes. Our dataset comes from the Arabic author pro ling and deception
detection shared task (APDA) [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ]. We focus on building simple models using
pre-trained bidirectional encoders from transformers (BERT) [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] under various
data conditions. Our results show (1) the utility of augmenting training data,
and (2) the bene t of using majority votes from our simple classi ers.
      </p>
      <p>In the rest of the paper, we introduce the dataset, followed by our
experimental conditions and results. We then provide a literature review and conclude.</p>
    </sec>
    <sec id="sec-2">
      <title>Data</title>
      <p>For the purpose of our experiments, we use data released by the APDA shared
task organizers. The dataset is divided into train and test by organizers. The
TRAIN set is distributed with labels for the three tasks of age, dialect, and
gender. Following the standard shared tasks set up, the test set is distributed
without labels and participants were expected to submit their predictions on test.
The shared task predictions are expected by organizers at the level of users. The
distribution has 100 tweets for each user, and so each tweet is distributed with a
corresponding user id. As such, in total, the distributed training data has 2,250
users, contributing a total of 225,000 tweets. The o cial task test set contains
720,00 tweets posted by 720 users. For our experiments, we split the training
data released by organizers into 90% TRAIN set (202,500 tweets from 2,025
users) and 10% DEV set (22,500 tweets from 225 users). The age task labels
come from the tagset funder-25, between-25 and 34, above-35 g. For dialects,
the data are labeled with 15 classes, from the set fAlgeria, Egypt, Iraq, Kuwait,
Lebanon-Syria, Lybia, Morocco, Oman, Palestine-Jordan, Qatar, Saudi Arabia,
Sudan, Tunisia, UAE, Yemeng. The gender task involves binary labels from the
set fmale, femaleg.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Experiments</title>
      <p>As explained earlier, the shared task is set up at the user level where the age,
dialect, and gender of each user are the required predictions. In our experiments,
we rst model the task at the tweet level and then port these predictions at the
user level. For our core modelling, we ne-tune BERT on the shared task data.
We also introduce an additional in-house dataset labeled with dialect and gender
tags to the task as we will explain below. As a baseline, we use a small gated
recurrent units (GRU) model. We now introduce our tweet-level models.
3.1</p>
      <p>
        Tweet-Level Models
Baseline GRU. Our baseline is a GRU network for each of the three tasks.
We use the same network architecture across the 3 tasks. For each network,
the network contains a layer unidirectional GRU, with 500 units and an output
linear layer. The network is trained end-to-end. Our input embedding layer is
initialized with a standard normal distribution, with = 0, and = 1, i.e.,
W N (0; 1). We use a maximum sequence length of 50 tokens, and choose
an arbitrary vocabulary size of 100,000 types, where we use the 100,000 most
frequent words in TRAIN. To avoid over- tting, we use dropout [
        <xref ref-type="bibr" rid="ref43">43</xref>
        ] with a rate
of 0.5 on the hidden layer. For the training, we use the Adam [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] optimizer with
a xed learning rate of 1e 3. We employ batch training with a batch size of 32
for this model. We train the network for 15 epochs and save the model at the
end of each epoch, choosing the model that performs highest accuracy on DEV
as our best model. We present our best result on DEV in Table 1. We report all
our results using accuracy. Our best model obtains 42.48% for age, 37.50% for
dialect, and 57.81% for gender. All models obtain best results with 2 epochs.
BERT. For each task, we ne-tune on the BERT-Base Muultilingual Cased
model relesed by the authors [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] 2. The model was pre-trained on Wikipedia
of 104 languages (including Arabic) with 12 layer, 768 hidden units each, 12
attention heads, and has 110M parameters in entire model. The vocabulary of
the model is 119,547 shared WordPices. We ne-tune the model with maximum
sequence length of 50 tokens and a batch size of 32. We set the learning rate
to 2e 5 and train for 15 epochs. We use the same network architecture and
parameters across the 3 tasks. As Table 1 shows, comparing with GRU, BERT
is 3.16% better for age, 4.85% better for dialect, and 2.45% higher for gender.
Data Augmentation. To further improve the performance of our models, we
introduce in-house labeled data that we use to ne-tune BERT. For the gender
classi cation task, we manually label an in-house dataset of 1,100 users with
gender tags, including 550 female users, 550 male users. We obtain 162,829 tweets by
crawling the 1,100 users' timelines. We combine this new gender dataset with the
gender TRAIN data (from shared task) to obtain an extended dataset, to which
we refer as EXTENDED Gender. For the dialect identi cation task, we randomly
sample 20,000 tweets for each class from an in-house dataset gold labeled with
the same 15 classes as the shared task. In this way, we obtain 298,929 tweets
(Sudan only has 18,929 tweets). We combine this new dialect data with the shared
task dialect TRAIN data to form EXTENDED Dialect. For both the dialect and
gender tasks, we ne-tune BERT on EXTENDED Dialect and EXTENDED Gender
independently and report performance on DEV. We refer to this iteration of
experiments as BERT EXT. As Table 1 shows, BERT EXT is 2.18% better than
BERT for dialect and 0.75% better than BERT for gender. 3
Our afore-mentioned models identify user's pro ling on the tweet-level, rather
than directly detecting the labels of a user. Hence, we follow the work of Zhang
2 https://github.com/google-research/bert/blob/master/multilingual.md
3 We note that it was not possible for us to use external age-labeled data and hence
we do not report on the age task with this data augmentation setting.
&amp; Abdul-Mageed [
        <xref ref-type="bibr" rid="ref47">47</xref>
        ] to identify user-level labels. For each of the three tasks,
we use tweet-level predicted labels (and associated softmax values) as a proxy
for user-level labels. For each predicted label, we use the softmax value as a
threshold for including only highest con dently predicted tweets. Since in some
cases softmax values can be low, we try all values between 0.00 and 0.99 to take
a softmax-based majority class as the user-level predicted label, ne-tuning on
our DEV set. Using this method, we acquire the following results at the user
level: BERT models obtain an accuracy of 55.56% for age, 96.00% for dialect, and
80.00% for gender. BERT EXT models achieve 95.56% accuracy for dialect and
84.00% accuracy for gender.
3.3
First submission. For the shared task submission, we use the predictions of
BERT EXT as out rst submission for gender and dialect, but only BERT for
age (since we have no BERT EXT models for age, as explained earlier). In each
case, we acquire results at tweet-level rst, then port the labels at the
userlevel as explained in the previous section. For our second and third submitted
models, we also follow this method of going from tweet to user level. Second
submission. We combine our DEV data with our EXTENDED Dialect and
EXTENDED Gender data, for dialect and gender respectively, and train our
second submssions for the two tasks. For age second submsision, we concatenate
DEV data to TRAIN and ne-tune the BERT model. We refer to the settings
for our second submission models collectively as BERT EXT+DEV.
      </p>
      <p>
        Third submission. Finally, for our third submission, we use a majority
vote of (1) rst submission, (2) second submission, and (3) predictions from our
user-level BERT model. These majority class models (i.e., our third submission)
achieve best results on the o cial test data. We acquire 54.72% accuracy for
age, 81.67% accuracy for gender, 93.75% accuracy for dialect, and 40.97% joint
accuracy.
Arabic. Arabic is a term that refers to a collection of languages, varieties, and
dialects. The standard variety, Modern Standard Arabic (MSA), is the one
usually used in formal communication and educational settings. Arabic also has a
wide range of under-studied varieties and dialects that classically used to be
categorized in a coarse-grained fashion (e.g., Levantine, North African) [
        <xref ref-type="bibr" rid="ref1 ref13 ref17 ref2 ref46">17, 46, 1, 13,
2</xref>
        ]. More recent treatments focus on ne-grained categorizations such as country
and city levels [
        <xref ref-type="bibr" rid="ref25 ref3 ref31 ref39 ref40 ref47">25, 39, 3, 40, 31, 47</xref>
        ]. Di erences between varieties of Arabic
happen at various linguistic levels, including including phonological, morphological,
lexical, and syntactic [
        <xref ref-type="bibr" rid="ref1 ref19 ref27 ref6">19, 6, 27, 1</xref>
        ].
      </p>
      <p>
        Social Media Author Pro ling. Author pro ling is the term usually used
to refer to detecting a host of attributes of (often social media) users. This
include identifying attributes such as age, gender, educational level, economic
class, stance or ideology [
        <xref ref-type="bibr" rid="ref10 ref30">10, 30</xref>
        ], personality [
        <xref ref-type="bibr" rid="ref18 ref23 ref42 ref7">42, 23, 7, 18</xref>
        ], moral traitstraits [
        <xref ref-type="bibr" rid="ref21 ref28">21,
28</xref>
        ], and other socilogical and psychological constructs. Author pro ling based
on text [
        <xref ref-type="bibr" rid="ref15 ref5">15, 5</xref>
        ] is rooted in computational stylometry [
        <xref ref-type="bibr" rid="ref16 ref44">16, 44</xref>
        ] and has traces in the
early work of Holmes [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. The task of author pro ling has also been approached
from network perspective where cues based on friending, following, mentioning,
and commenting have been leveraged for identifying author attributes [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] for
author pro ling. In addition, the PAN author pro ling shared task [
        <xref ref-type="bibr" rid="ref33 ref34 ref35 ref36 ref37">34, 33, 37,
36, 35</xref>
        ] was established to advance related work. More information about PAN
can be found in [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ].
      </p>
      <p>
        Age and Gender. A number of studies have been conducted on
Englishbased age and gender detection, including [
        <xref ref-type="bibr" rid="ref11 ref14 ref38 ref45 ref8">38, 14, 8, 45, 11</xref>
        ]. Many of these works
use feature engineering such as text n-gram and topic models [
        <xref ref-type="bibr" rid="ref41 ref42">42, 41</xref>
        ]. In these
works, age is either cast as a multi-class classi cation task with, e.g., labels from
the set f10-19, 20-29, 30-39 g or as a regression task [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]. Other works model age
with both classi cation and regression combined [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. With rare exceptions [
        <xref ref-type="bibr" rid="ref35 ref4">35,
4</xref>
        ], we do not know of work on Arabic targeting age and gender.
5
      </p>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>
        In this work, we described our submitted models to the Arabic author pro ling
and deception detection shared task (APDA) [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ]. We focused on detecting age,
dialect, and gender using BERT models under various data conditions, showing
the utility of additional, in-house data on the task. We also showed that a
majority vote of our models trained under di erent conditions outperforms single
models on the o cial evaluation. In the future, we will investigate automatically
extending training data for these tasks as well as better representation learning
methods.
6
      </p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgement</title>
      <p>We acknowledge the support of the Natural Sciences and Engineering Research
Council of Canada (NSERC), the Social Sciences Research Council of Canada
(SSHRC), and Compute Canada (www.computecanada.ca).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Abdul-Mageed</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Subjectivity and sentiment analysis of Arabic as a morophologically-rich language</article-title>
          .
          <source>Ph.D. thesis</source>
          , Indiana University (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Abdul-Mageed</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Modeling arabic subjectivity and sentiment in lexical space</article-title>
          .
          <source>Information Processing &amp; Management</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Abdul-Mageed</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alhuzali</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elaraby</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>You tweet what you speak: A citylevel dataset of arabic dialects</article-title>
          .
          <source>In: Proceedings of the Eleventh International Conference on Language Resources</source>
          and
          <string-name>
            <surname>Evaluation (LREC-2018)</surname>
          </string-name>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Alrifai</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rebdawi</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ghneim</surname>
          </string-name>
          , N.:
          <article-title>Arabic tweeps gender and dialect prediction</article-title>
          .
          <source>In: CLEF (Working Notes)</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Argamon</surname>
            ,
            <given-names>S.E.</given-names>
          </string-name>
          :
          <article-title>Register in computational language research</article-title>
          .
          <source>Register Studies</source>
          <volume>1</volume>
          (
          <issue>1</issue>
          ),
          <volume>100</volume>
          {
          <fpage>135</fpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Bassiouney</surname>
          </string-name>
          , R.:
          <article-title>Arabic sociolinguistics</article-title>
          . Edinburgh University Press (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Bleidorn</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hopwood</surname>
            ,
            <given-names>C.J.:</given-names>
          </string-name>
          <article-title>Using machine learning to advance personality assessment and theory</article-title>
          . Personality and Social Psychology Review p.
          <volume>1088868318772990</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Burger</surname>
            ,
            <given-names>J.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Henderson</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zarrella</surname>
          </string-name>
          , G.:
          <article-title>Discriminating gender on twitter</article-title>
          .
          <source>In: Proceedings of the conference on empirical methods in natural language processing</source>
          . pp.
          <volume>1301</volume>
          {
          <fpage>1309</fpage>
          .
          <article-title>Association for Computational Linguistics (</article-title>
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Chen</surname>
            , J., Cheng,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quan</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Joint learning with both classi cation and regression models for age prediction</article-title>
          .
          <source>In: Journal of Physics: Conference Series</source>
          . vol.
          <volume>1168</volume>
          , p.
          <fpage>032016</fpage>
          . IOP Publishing (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Colleoni</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rozza</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arvidsson</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Echo chamber or public sphere? predicting political orientation and measuring political homophily in twitter using big data</article-title>
          .
          <source>Journal of communication 64(2)</source>
          ,
          <volume>317</volume>
          {
          <fpage>332</fpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Daneshvar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Inkpen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Gender identi cation in twitter using n-grams and lsa</article-title>
          .
          <source>In: Proceedings of the Ninth International Conference of the CLEF Association (CLEF</source>
          <year>2018</year>
          )
          <article-title>(</article-title>
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>M.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          . arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Elaraby</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Abdul-Mageed</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Deep models for arabic dialect identi cation on benchmarked data</article-title>
          .
          <source>In: Proceedings of the Fifth Workshop on NLP for Similar Languages, Varieties and Dialects (VarDial</source>
          <year>2018</year>
          ). pp.
          <volume>263</volume>
          {
          <issue>274</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Flekova</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Preotiuc-Pietro</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ungar</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Exploring stylistic variation with age and income on twitter</article-title>
          .
          <source>In: Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume</source>
          <volume>2</volume>
          :
          <string-name>
            <given-names>Short</given-names>
            <surname>Papers</surname>
          </string-name>
          <article-title>)</article-title>
          .
          <source>vol. 2</source>
          , pp.
          <volume>313</volume>
          {
          <issue>319</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Gamon</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Linguistic correlates of style: authorship classi cation with deep linguistic analysis features</article-title>
          .
          <source>In: Proceedings of the 20th international conference on Computational Linguistics</source>
          . p.
          <fpage>611</fpage>
          .
          <article-title>Association for Computational Linguistics (</article-title>
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Goswami</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sarkar</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rustagi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Stylometric analysis of bloggers age and gender</article-title>
          .
          <source>In: Third international AAAI conference on weblogs and social media</source>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Habash</surname>
          </string-name>
          , N.Y.:
          <article-title>Introduction to arabic natural language processing</article-title>
          .
          <source>Synthesis Lectures on Human Language Technologies</source>
          <volume>3</volume>
          (
          <issue>1</issue>
          ),
          <volume>1</volume>
          {
          <fpage>187</fpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Hinds</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joinson</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Human and computer personality prediction from digital footprints</article-title>
          . Current Directions in Psychological Science p.
          <volume>0963721419827849</volume>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Holes</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Modern Arabic: Structures, functions, and varieties</article-title>
          . Georgetown University Press (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Holmes</surname>
            ,
            <given-names>D.I.</given-names>
          </string-name>
          :
          <article-title>The evolution of stylometry in humanities scholarship</article-title>
          .
          <source>Literary and linguistic computing 13(3)</source>
          ,
          <volume>111</volume>
          {
          <fpage>117</fpage>
          (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Johnson</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goldwasser</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Classi cation of moral foundations in microblog political discourse</article-title>
          .
          <source>In: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</source>
          . pp.
          <volume>720</volume>
          {
          <issue>730</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Kingma</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ba</surname>
          </string-name>
          , J.:
          <article-title>Adam: A method for stochastic optimization</article-title>
          .
          <source>arXiv preprint arXiv:1412.6980</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Matz</surname>
            ,
            <given-names>S.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kosinski</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nave</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stillwell</surname>
            ,
            <given-names>D.J.:</given-names>
          </string-name>
          <article-title>Psychological targeting as an e ective approach to digital mass persuasion</article-title>
          .
          <source>Proceedings of the national academy of sciences 114(48)</source>
          ,
          <volume>12714</volume>
          {
          <fpage>12719</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Mitrou</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kandias</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stavrou</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gritzalis</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Social media pro ling: A panopticon or omniopticon tool? In: Proc. of the 6th Conference of the Surveillance Studies Network</article-title>
          . Barcelona,
          <string-name>
            <surname>Spain</surname>
          </string-name>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Mubarak</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Darwish</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Using twitter to collect a multi-dialectal corpus of arabic</article-title>
          .
          <source>In: Proceedings of the EMNLP 2014 Workshop on Arabic Natural Language Processing (ANLP)</source>
          . pp.
          <volume>1</volume>
          {
          <issue>7</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>N.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rose</surname>
            ,
            <given-names>C.P.</given-names>
          </string-name>
          :
          <article-title>Author age prediction from text using linear regression</article-title>
          .
          <source>In: Proceedings of the 5th ACL-HLT Workshop on Language Technology for Cultural Heritage</source>
          ,
          <source>Social Sciences, and Humanities</source>
          . pp.
          <volume>115</volume>
          {
          <fpage>123</fpage>
          .
          <article-title>Association for Computational Linguistics (</article-title>
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Palva</surname>
          </string-name>
          , H.:
          <article-title>Dialects: classi cation</article-title>
          .
          <source>Encyclopedia of Arabic Language and Linguistics</source>
          <volume>1</volume>
          ,
          <issue>604</issue>
          {
          <fpage>613</fpage>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Pang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eichstaedt</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bu</surname>
            <given-names>one</given-names>
          </string-name>
          , A.,
          <string-name>
            <surname>Sla</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ruch</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ungar</surname>
            ,
            <given-names>L.H.:</given-names>
          </string-name>
          <article-title>The language of character strengths: Predicting morally valued traits on social media</article-title>
          .
          <source>Journal of personality</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stamatatos</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>A decade of shared tasks in digital text forensics at pan</article-title>
          .
          <source>In: European Conference on Information Retrieval</source>
          . pp.
          <volume>291</volume>
          {
          <fpage>300</fpage>
          . Springer (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Preotiuc-Pietro</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , Liu,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Hopkins</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Ungar</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.</surname>
          </string-name>
          :
          <article-title>Beyond binary labels: political ideology prediction of twitter users</article-title>
          .
          <source>In: Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</source>
          . pp.
          <volume>729</volume>
          {
          <issue>740</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Qwaider</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saad</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chatzikyriakidis</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dobnik</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Shami: A corpus of levantine arabic dialects</article-title>
          .
          <source>In: Proceedings of the Eleventh International Conference on Language Resources</source>
          and
          <string-name>
            <surname>Evaluation (LREC-2018)</surname>
          </string-name>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Char</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zaghouani</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ghanem</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Snchez-Junquera</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Overview of the track on author pro ling and deception detection in arabic</article-title>
          . In: Mehta P.,
          <string-name>
            <surname>Rosso</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Majumder</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mitra</surname>
            <given-names>M</given-names>
          </string-name>
          . (Eds.)
          <article-title>Working Notes of the Forum for Information Retrieval Evaluation (FIRE 2019)</article-title>
          . CEUR Workshop Proceedings. In: CEUR-WS.org, Kolkata, India, December
          <volume>12</volume>
          -
          <fpage>15</fpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chugur</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Trenkmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verhoeven</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daelemans</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Overview of the 2nd author pro ling task at pan 2014</article-title>
          . In:
          <article-title>CLEF 2014 Evaluation Labs</article-title>
          and Workshop Working Notes Papers, She eld, UK,
          <year>2014</year>
          . pp.
          <volume>1</volume>
          {
          <issue>30</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          34.
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koppel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stamatatos</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Inches</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Overview of the author pro ling task at pan 2013</article-title>
          .
          <source>In: CLEF Conference on Multilingual and Multimodal Information Access Evaluation</source>
          . pp.
          <volume>352</volume>
          {
          <fpage>365</fpage>
          .
          <string-name>
            <surname>CELCT</surname>
          </string-name>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          35.
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Overview of the 5th author pro ling task at pan 2017: Gender and language variety identi cation in twitter</article-title>
          .
          <source>Working Notes Papers of the CLEF</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          36.
          <string-name>
            <surname>Rangel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verhoeven</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daelemans</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potthast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Overview of the 4th author pro ling task at pan 2016: cross-genre evaluations</article-title>
          .
          <source>In: Working Notes Papers of the CLEF</source>
          <year>2016</year>
          <article-title>Evaluation Labs</article-title>
          . CEUR Workshop Proceedings/Balog, Krisztian [edit.]; et al. pp.
          <volume>750</volume>
          {
          <issue>784</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          37.
          <string-name>
            <given-names>Rangel</given-names>
            <surname>Pardo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.M.</given-names>
            ,
            <surname>Celli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Daelemans</surname>
          </string-name>
          ,
          <string-name>
            <surname>W.</surname>
          </string-name>
          :
          <article-title>Overview of the 3rd author pro ling task at pan 2015</article-title>
          . In:
          <article-title>CLEF 2015 Evaluation Labs</article-title>
          and Workshop Working Notes Papers. pp.
          <volume>1</volume>
          {
          <issue>8</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          38.
          <string-name>
            <surname>Rao</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yarowsky</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shreevats</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Classifying latent user attributes in twitter</article-title>
          .
          <source>In: Proceedings of the 2nd international workshop on Search</source>
          and
          <article-title>mining user-generated contents</article-title>
          . pp.
          <volume>37</volume>
          {
          <fpage>44</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          39.
          <string-name>
            <surname>Sadat</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kazemi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Farzindar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Automatic identi cation of arabic language varieties and dialects in social media</article-title>
          .
          <source>Proceedings of SocialNLP</source>
          p.
          <volume>22</volume>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          40.
          <string-name>
            <surname>Salameh</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bouamor</surname>
          </string-name>
          , H.:
          <article-title>Fine-grained arabic dialect identi cation</article-title>
          .
          <source>In: Proceedings of the 27th International Conference on Computational Linguistics</source>
          . pp.
          <volume>1332</volume>
          {
          <issue>1344</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          41.
          <string-name>
            <surname>Sap</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Park</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eichstaedt</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kern</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stillwell</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kosinski</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ungar</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schwartz</surname>
            ,
            <given-names>H.A.</given-names>
          </string-name>
          :
          <article-title>Developing age and gender predictive lexica over social media</article-title>
          .
          <source>In: Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP)</source>
          . pp.
          <volume>1146</volume>
          {
          <issue>1151</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          42.
          <string-name>
            <surname>Schwartz</surname>
            ,
            <given-names>H.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eichstaedt</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kern</surname>
            ,
            <given-names>M.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dziurzynski</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ramones</surname>
            ,
            <given-names>S.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Agrawal</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kosinski</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stillwell</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Seligman</surname>
            ,
            <given-names>M.E.</given-names>
          </string-name>
          , et al.:
          <article-title>Personality, gender, and age in the language of social media: The open-vocabulary approach</article-title>
          .
          <source>PloS one 8</source>
          (
          <issue>9</issue>
          ),
          <year>e73791</year>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          43.
          <string-name>
            <surname>Srivastava</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hinton</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krizhevsky</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salakhutdinov</surname>
          </string-name>
          , R.:
          <article-title>Dropout: a simple way to prevent neural networks from over tting</article-title>
          .
          <source>The Journal of Machine Learning Research</source>
          <volume>15</volume>
          (
          <issue>1</issue>
          ),
          <year>1929</year>
          {
          <year>1958</year>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          44.
          <string-name>
            <surname>Verhoeven</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Two authors walk into a bar: studies in author pro ling</article-title>
          .
          <source>Ph.D. thesis</source>
          , University of Antwerp (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          45.
          <string-name>
            <surname>Verhoeven</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daelemans</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Plank</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Twisty: a multilingual twitter stylometry corpus for gender and personality pro ling</article-title>
          .
          <source>In: Proceedings of the 10th Annual Conference on Language Resources and Evaluation (LREC</source>
          <year>2016</year>
          )/Calzolari, Nicoletta [edit.]; et al. pp.
          <volume>1</volume>
          {
          <issue>6</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          46.
          <string-name>
            <surname>Versteegh</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>The arabic language</article-title>
          . Edinburgh University Press (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref47">
        <mixed-citation>
          47.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Abdul-Mageed</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>No army, no navy: Bert semi-supervised learning of arabic dialects</article-title>
          .
          <source>In: Proceedings of the Fourth Arabic Natural Language Processing Workshop</source>
          . pp.
          <volume>279</volume>
          {
          <issue>284</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>