<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Linguistic Cues of Deception in a Multilingual April Fools' Day Context</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Katerina Papantoniou</string-name>
          <email>papanton@ics.forth.gr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Panagiotis Papadakos</string-name>
          <email>papadako@ics.forth.gr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giorgos Flouris</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dimitris Plexousakis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>. Computer Science Department, University of Crete</institution>
          ,
          <country country="GR">Greece</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>. Institute of Computer Science</institution>
          ,
          <addr-line>FORTH</addr-line>
          ,
          <country country="GR">Greece</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this work we consider the collection of deceptive April Fools' Day (AFD) news articles as a useful addition in existing datasets for deception detection tasks. Such collections have an established ground truth and are relatively easy to construct across languages. As a result, we introduce a corpus that includes diachronic AFD and normal articles from Greek newspapers and news websites. On top of that, we build a rich linguistic feature set, and analyze and compare its deception cues with the only AFD collection currently available, which is in English. Following a current research thread, we also discuss the individualism/collectivism dimension in deception with respect to these two datasets. Lastly, we build classiifers by testing various monolingual and crosslingual settings. The results showcase that AFD datasets can be helpful in deception detection studies, and are in alignment with the observations of other deception detection works.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>April Fools’ Day (for short AFD) is a long
standing custom, mostly in Western societies. It is the
only day of the year when practical jokes and
deception are expected. This is the case for all social
interactions, including journalism, which is
generally considered to aim at the presentation of truth.
Every year on this day, newspapers and news
websites take part in an unofficial competition to
invent the most believable, but untrue story. In this
respect, AFD news articles fall into the deception</p>
      <p>Copyright © 2021 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
spectrum, as they satisfy widely acceptable
definitions of deception as in Masip et al. (2005).</p>
      <p>
        The massive participation of news media in this
custom establishes a rich corpus of deceptive
articles from a diversity of sources. Although AFD
articles may exploit common linguistic instruments
with satire news, like exaggeration, humour, irony
and paralogism, they are usually considered a
distinct category. This is mainly due to the fact that
they also employ other mechanisms which
characterize deception in general, like sophisms, and
changes in cognitive load and emotions
        <xref ref-type="bibr" rid="ref5">(Hauch et
al., 2015)</xref>
        to deceive their audience. AFD articles
are often believable, and there exist cases where
sophisticated AFD articles have been reproduced
by major international news agencies worldwide1.
      </p>
      <p>
        This motivated us to extend our previous work
on linguistic cues of deception and their relation
to the cultural dimension of individualism and
collectivism
        <xref ref-type="bibr" rid="ref19 ref24">(Papantoniou et al., 2021)</xref>
        , in the context
of the AFD. That work examines if differences
in the usage of linguistic cues of deception (e.g.,
pronouns) across cultures can be identified and
attributed to the individualism/collectivism divide.
      </p>
      <p>
        Specifically, the contributions of this work are:
• A new corpus that includes diachronic AFD
and normal articles from Greek newspapers
and news websites2, adding one more AFD
collection to the currently unique one in
English
        <xref ref-type="bibr" rid="ref1 ref22 ref25">(Dearden and Baron, 2019)</xref>
        .
• A study and discussion of the linguistic cues
of deception that prevail in the Greek and
English collection, along with their similarities.
• A discussion on whether the consideration
of the individualism/collectivism cultural
di1https://www.nationalgeographic.com/history/article/150331april-fools-day-hoax-prank-history-holiday
      </p>
      <p>2The collection is available in: https://gitlab.i
sl.ics.forth.gr/papanton/elaprilfoolcorp
us
mension in the context of AFD aligns with
the results of our previous work.
• An examination of the performance of
various classifiers in identifying AFD articles,
including multilanguage setups.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        The creation of reliable and realistic ground truth
datasets for the deception detection task is a
challenging task
        <xref ref-type="bibr" rid="ref3">(Fitzpatrick and Bachenko, 2012)</xref>
        .
Crowdsourcing, in the form of online campaigns
in which people express themselves in truthful
and/or deceitful manner for a small payment are
a well established way to collect deceptive data
        <xref ref-type="bibr" rid="ref18">(Ott et al., 2011)</xref>
        . Real-life situations such as
trials
        <xref ref-type="bibr" rid="ref25">(Soldner et al., 2019)</xref>
        or the use of data from
board games have also been employed
        <xref ref-type="bibr" rid="ref21">(Peskov et
al., 2020)</xref>
        . Also a popular approach is the reuse
of content from sites that debunk articles like fake
news and hoaxes
        <xref ref-type="bibr" rid="ref10 ref30">(Wang, 2017; Kochkina et al.,
2018)</xref>
        . Lastly, satire news are another way to
collect deceptive texts, but with some particularities
due to humorous deception
        <xref ref-type="bibr" rid="ref23">(Skalicky et al., 2020)</xref>
        .
      </p>
      <p>The only work that explores AFD articles is that
of Dearden et al. (2019). They collected 519 AFD
and 519 truthful stories and articles in English for
a period of 14 years. A large set of features was
exploited to identify deception cues in AFD
stories. Structural complexity and level of detail were
among the most valuable features while the
exploitation of the same feature set to a fake news
dataset resulted in similar observations.</p>
      <p>To the best of our knowledge, the only
deception related dataset for the Greek language is that
of Karidi et al. (2019). This work proposed an
automatic process for the creation of a fake news
and hoaxes articles corpus, but unfortunately the
created corpus over Greek websites is not
available. If we also consider that the creation of a
Greek dataset for deception through
crowdsourcing is a cumbersome and expensive task, that is
further hindered by the exceptionally limited
number of native Greek crowd workers, it is easy to
understand why there is a lack of datasets.</p>
      <p>
        Regarding the individualism/collectivism
cultural dimension, it constitutes a well-known
division of cultures that concerns the degree in which
members of a culture value more individual over
group goals and vice versa. In individualism, ties
between individuals are loose and individuals are
expected to take care of only themselves and their
immediate families, whereas in collectivism ties in
society are stronger. In Papantoniou et al. (2021)
there is an preliminary effort driven by prior work
in psychology discipline
        <xref ref-type="bibr" rid="ref27">(Taylor et al., 2017)</xref>
        to
examine if deception cues are altered across
cultures and if this can be attributed to this divide.
Among the conclusions were that people from
individualistic cultures employ more third and less
ifrst person pronouns to distance themselves from
the deceit when they are deceptive, whereas in the
collectivism group this trend is milder, signalling
the effort of the deceiver to distance the group
from the deceit. In addition, in individualistic
cultures positive sentiment is employed in deceptive
language, whereas in collectivists there is a
restraint of expression of sentiment both in truthful
and deceptive texts.
      </p>
      <p>
        To this end, this work explores the
deceptionrelated characteristics of a new Greek corpus
based on AFD articles from a variety of sources,
and compares them with the English ones3.
Further, since related studies
        <xref ref-type="bibr" rid="ref11 ref28 ref6">(Triandis and
Vassiliou, 1972; Hofstede, 1980; Koutsantoni, 2005)</xref>
        describe Greece as a culture with more
collectivistic characteristics (by using country as proxy from
culture), we also discuss differences in deception
cues along this cultural dimension.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Corpus Creation</title>
      <p>The AFD articles have been hand gathered
because a crawling based collection approach was
not applicable in our case. Since the news web
sites industry in Greece is not huge to establish
an acceptable number of crawled AFD articles, we
had to additionally collect articles from the press,
including articles from the pre-WWW era.
Specifically, we visited the local library that maintains
a printed archive of newspapers and searched for
disclosure articles in the issues after the 1st April,
took photos of the AFD articles, and then used
OCR and manual inspection to extract the text.
In addition we contacted national and local news
media providers to get access in their digitalized
archives. The rest were gathered from the Web.</p>
      <p>The articles were categorized thematically into
the following vfie categories: society, culture,
politics, world, and sports. If no category was
pro3We also experimented with data from the limited number
of satirical and hoaxes sources of the Greek Web. We do not
discuss them here though, since the classifiers reported
excellent accuracy showcasing the lack of diversity and the
existence of domain specific information in the collected data.
vided by the original source, we manually
annotated the articles. For each article we kept the
title, the main body, the published date, the name,
the type of the source (newspaper or news
website), and (if available) the caption, the subtitle
and the author. As preprocesing steps we
applied spellcheck and normalization. The
correction of spelling mistakes was necessary
primarily for articles extracted through OCR tools,
although spelling errors were identified in other
articles too. Normalization was performed for
homogeneity reasons in the texts retrieved from the 80’s,
since we observed language differences in some
forms (e.g., in the suffix of genitive case), which
are remains of an old form of Modern Greek4.</p>
      <p>For the truthful collection we used the same
manual procedure and we tried to have a balanced
dataset in terms of thematic categories. The
truthful collection consists of articles that have been
published in days relatively close to the 1st of
April in order to have articles that do not differ
significantly in respect to their topics, mentioned
named entities, etc.</p>
      <p>Since the AFD tradition is vivid in Greece, we
were able to locate a lot of such articles from
various newspapers and new websites for our corpus
(112 different sources). Specifically, we managed
to collect 254 truthful and 254 deceptive articles
spanning over the period 1979 - 2021. In Tables 1
to 2 some statistics of the corpus are depicted.</p>
      <sec id="sec-3-1">
        <title>Measure</title>
        <p>Num. of articles
Avg. length
Min. length
Max. length</p>
        <sec id="sec-3-1-1">
          <title>4https://en.wikipedia.org/wiki/Katharevousa</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Features Analysis</title>
      <p>For the analysis of AFD articles we adapt and
build upon the feature set used in Papantoniou et
al. (2021), but for the Greek language. The
resulting feature set consists of 64 features for the Greek
language and 75 for the English, due to the smaller
availability of linguistic resources for Greek (e.g.,
in sentiment lexicons). For the analysis we
performed the non-parametric Mann–Whitney U test
(two-tailed) with a 99% confidence interval (CI)
and α = 0.01. Table 3 depicts the results of this
analysis for elAFD and enAFD datasets5.</p>
      <p>
        In both datasets, positive sentiment is related
to the deceptive articles, while negative sentiment
with the truthful articles. The only exception
concerns the enAFD dataset, where for the NRC
lexicon the opposite holds (NRC is one of the six
sentiment lexicons used for features in English).
In addition, negative emotions like anger, fear and
sadness are related to truthful news articles in both
datasets. The use of positive emotive language
during deception may be a strategy for deceivers to
maintain social harmony as noticed also by other
studies
        <xref ref-type="bibr" rid="ref17 ref20">(Newman et al., 2003; Pérez-Rosas et al.,
2018)</xref>
        . The difference in the use of emotional
language between truthful and deceptive news is
more intense in the enAFD dataset, where vfie out
of the eight emotions in the NRC lexicon are found
statistical significant. This is in alignment with the
results in Papantoniou et al. (2021) for
individualistic and collectivistic cultures.
      </p>
      <p>
        Further, deceptive texts seem to be related with
an increased use of adverbs in both datasets. This
can be related to the less concreteness of deceptive
texts as discussed in Kleinberg et al. (2019) and
it is in line with many theories of deception like
the Reality Monitoring
        <xref ref-type="bibr" rid="ref8">(Johnson et al., 1998)</xref>
        ,
Criteria based Content Analysis
        <xref ref-type="bibr" rid="ref29">(Undeutsch, 1989)</xref>
        and Verifiability Approach
        <xref ref-type="bibr" rid="ref16">(Nahari et al., 2014)</xref>
        .
This also explains the prevalence of the number of
named entities, spatial related words, conjunctions
and WDAL imagery score in truthful texts in the
enAFD dataset and the use of more motion verbs
in deceptive texts in the elAFD dataset. According
to cognitive load theory
        <xref ref-type="bibr" rid="ref26">(Sweller, 2011)</xref>
        in
deceptive texts the language is less specific and consists
of simpler constructs. The same holds for
modality, another common feature among the datasets,
that is considered a signal of subjectivity that
pro
      </p>
      <sec id="sec-4-1">
        <title>5All the features are described in</title>
        <p>https://gitlab.isl.ics.forth.gr/papanton/elaprilfoolcorpus
vides a degree of uncertainty. In addition, hedges
in enAFD dataset, also express some feeling of
doubt or hesitancy.</p>
        <p>Lexical diversity as expressed by the token-type
ratio (TTR), that is the ratio of unique words to the
total number of tokens, is related to the deceptive
texts. This seems to contradict all the above, but
could be attributed to the fact that deceptive texts
are shorter. Although this is more evident in the
case of the enAFD dataset, it also holds for elAFD
dataset (see Table 1).</p>
        <p>Boosters, which are words that express
confidence (e.g., certainly) are quite discriminative for
deceptive texts for the enAFD dataset. Moreover
we observe the connection of the future tense with
deception and of the past with truth. The above
were also marked in Papantoniou et al. (2021) in
different domain from the news articles domain.</p>
        <p>Finally, first personal pronouns have been found
to be rather discriminative of deceptive texts in
various deception detection and cultural studies,
including Papantoniou et al. (2021). However, in
this study pronouns are statistical important only
for the enAFD dataset. This probably reflects
idiosyncrasies of the news domain, since articles
mainly present objectively facts and not opinions,
and as a result the use of first personal pronouns
is avoided. This holds for the elAFD dataset that
includes AFD articles from the news sites and the
press, and not for the enAFD dataset that consists
of various types of AFD articles and stories
collected from the web through crowdsourcing6.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Classification</title>
      <p>
        We evaluated the predictive performance of
different feature sets and approaches for AFD datasets,
including logistic regression experiments7 and
ifne-tuned monolingual BERT models for each
language8
        <xref ref-type="bibr" rid="ref12 ref2">(Devlin et al., 2019; Koutsikakis et
al., 2020)</xref>
        . We also performed cross lingual
experiments by exploiting the multilingual BERT
model (mBERT) to examine if there are
similarities among AFD datasets captured by the BERT.
      </p>
      <p>A stratified split to the datasets was used to
create training, testing, and validation subsets with
a 70-20-10 ratio. For the cross lingual
experiment we trained and validated a model over the</p>
      <sec id="sec-5-1">
        <title>6https://aprilfoolsdayontheweb.com/2004.html</title>
        <p>
          7We employ the Weka API
          <xref ref-type="bibr" rid="ref4">(Hall et al., 2009)</xref>
          8We used tensorflow 2.2.0, keras 2.3.1, and the
bert-fortf2 0.14.4 implementation of google-research/bert, over an
AMD Radeon VII card and the ROCm 3.7 platform.
Deceptive
adverbs (0.31)
adj. &amp; adv. (0.27)
TTR (0.27)
pos. sentiment (0.21)
modal verbs (0.17)
motion verbs (0.117)
boosters (0.39)
modal verbs (0.35)
TTR(0.31)
future (0.27)
adverbs (0.2)
1st pers. pp (0.2)
mpqa pos. (0.2)
nrc neg.* (-0.2)
2nd pers. pp (0.19)
1st pers. pp pl. (0.18)
sentiwordnet pos. (0.17)
demonstrative (0.17)
hedges (0.17)
adj &amp; adv (0.16)
present (0.15)
vader sentiment (0.14)
verb num. (0.14)
pers. pron. (0.12)
total pronouns (0.11)
        </p>
        <p>Truthful
elAFD
punctuation (-0.17)
nrc sadness(-0.17)
plosives (-0.16)
nrc anger (-0.15)
nrc fear (-0.14)
vowels (-0.14)
consonants (-0.14)
enAFD</p>
        <p>
          NE num. (-0.27)
spatial num. (-0.26)
conjuctions (-0.24)
nrc fear (-0.23)
past (-0.23)
nrc sadness (-0.23)
nrc anger (-0.21)
nrc trust (-0.21)
avg. word len. (-0.17)
collectivism (-0.16)
nrc pos.* (-0.16)
wdal imagery (-0.15)
mpqa neg. -0.14)
nasals (-0.14)
fbs neg. (-0.14)
consonants (-0.13)
anew arousal (-0.13)
prepositions (-0.12)
fricatives (-0.11)
3rd per. pp sg. (-0.11)
avg. preverb len. (-0.11)
nrc disgust (-0.1)
80% and 20% of a language specific dataset
respectively, and then tested the performance of
the model over the other dataset. We report
the results on test sets, while validation subsets
were used for fine-tuning the hyper-parameters of
the algorithms. For the logistic regression the
tuned through brute force parameters were: a)
Weka algorithm (SimpLog|Log: simple logistic
          <xref ref-type="bibr" rid="ref13">(Landwehr et al., 2005)</xref>
          or logistic
          <xref ref-type="bibr" rid="ref14">(Le Cessie and
Van Houwelingen, 1992)</xref>
          ) b) all n-grams of size in
[a, b], with a ≥ b and a, b ∈ [1, 3] ((a, b)), c)
stemming (stem), d) attribute selection (attrsel)
(applicable only to Log algorithm since it is the
default for SimpLog ), e) stopwords removal (stop)
and, f) lowercase conversion (lowercase). For
the BERT experiments, the hyperparameters were
tuned by random sampling 60 combinations of
values, keeping the combination that gave the
minimum validation loss. Early stopping with
patience 4 was used and the max epochs number
was set to 20. The tuned hyperparameters were:
learning rate, batch size, dropout rate, max token
length, and randomness seeds.
        </p>
        <p>
          In all cases, we report Recall (R), Precision
(P ), F-measure (F ), Accuracy (A) and AUC (A′).
Since the datasets are balanced the majority
baseline is 50%. The input for the models consists of
the concatenation of the title, the subtitle, the body
of the articles and the caption text. Since titles
are important for deception detection
          <xref ref-type="bibr" rid="ref7">(Horne and
Adali, 2017)</xref>
          and BERT processes texts of up to
512 wordpieces, we placed the title first.
5.1
        </p>
        <sec id="sec-5-1-1">
          <title>Logistic Regression Experiments</title>
          <p>The examined features sets were: a) the
features presented in section 4 (ling), b) n-grams
features i.e., phoneme-gram (ph-gram),
charactergram (char-gram), word-gram (w-gram),
POSgram (pos-gram), and syntactic-gram (sn-gram)
(the latter for the enAFD only), and c) the
linguistic+ model that represents the best model that
combines the linguistic features with any of the
n-gram features. The results are presented in
Tables 4 and 5. With * we mark the setups with a
statistically significant difference to the best setup
regarding accuracy, based on a two proposition
ztest (1-tailed) with a 99% CI. We observe that the
combination of lingustic features with
uni/bi/trigrams for the elAFD dataset and the unigrams for
the enAFD are the best setups. For the enAFD
dataset, the second best model is the
combination of linguistic features with trigrams. SimpLog
seems to perform better, while stemming,
lowercase conversion and stopwords removal are
generally beneficiary.
5.2</p>
        </sec>
        <sec id="sec-5-1-2">
          <title>BERT Experiments</title>
          <p>In these experiments, we fine-tuned BERT by
adding a task-specific linear classification layer on
top, using the sigmoid activation function. We also
combined BERT with linguistics features by
concatenating the embedding of the [CLS] token with
the linguistic features, and pass the resulting
vector to the task-specific classifier (with a slightly
modified architecture). The results of the
experiBest setup
ling.SimpLog
ph-gram(1,2),attrsel,Log*
char-gram(3,3),SimpLog*
w-gram(1,2),SimpLog
pos-gram(2,3),SimpLog*
ling.+word,(1,3),stop,
lowercase,SimpLog
ments are presented in Table 6. Although it
outperformed logistic regression experiments in both
datasets, the differences are not statistical
significant. In addition, the combination with
linguistic features is not beneficial. Multilingual BERT
models perform worse, especially for Greek. In
the cross lingual experiments the classifiers
performance is limited to about 60% accuracy in both
experiments, showcasing that the BERT layers are
not able to capture language agnostic information
from our datasets.
6</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusion and Future Work</title>
      <p>
        We introduced a new dataset with AFD news
articles in Greek and analyzed and compared its
deception cues with another English one. The results
showcased the use of emotional language,
especially of positive sentiment, for deceptive articles
which is even more prevalent in the
individualistic English dataset. Further, deceptive articles use
less concrete language, as manifested by the
increased use of adverbs, hedges, and boosters and
less usage of named entities, spatial related words
and conjunctions compared to the truthful ones.
The future and past tenses were correlated with
deceptive and truthful articles respectively. All the
above, mainly align with previous work
        <xref ref-type="bibr" rid="ref19 ref24">(Papantoniou et al., 2021)</xref>
        , except from some differences in
the usage of pronouns for the Greek dataset, which
is attributed to the idiosyncrasies of the news
domain. The accuracy of the deployed classifiers
offered adequate performance, with no statistically
significant differences between the best logistic
regression and the BERT models.
      </p>
      <p>
        In the future we aim at creating even more
crosslingual datasets for deception detection tasks
through crowdsourcing and by employing the
Chattack platform
        <xref ref-type="bibr" rid="ref24">(Smyrnakis et al., 2021)</xref>
        .
      </p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgement</title>
      <p>This work has received funding by the Hellenic
Foundation for Research and Innovation (H.F.R.I.)
under the “1st Call for H.F.R.I. Research Projects
to support Faculty Members &amp; Researchers and
the Procurement of High-and the procurement
of high-cost research equipment grant” (Project
Number:4195).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Edward</given-names>
            <surname>Dearden</surname>
          </string-name>
          and
          <string-name>
            <given-names>Alistair</given-names>
            <surname>Baron</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Fool's Errand: Looking at April Fools Hoaxes as Disinformation through the Lens of Deception and Humour</article-title>
          .
          <source>April. 20th International Conference on Computational Linguistics and Intelligent Text Processing</source>
          , CICLing 2019 ; Conference date:
          <fpage>07</fpage>
          -
          <lpage>04</lpage>
          - 2019 Through 13-
          <fpage>04</fpage>
          -
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding</article-title>
          .
          <source>In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (Long and Short Papers), pages
          <fpage>4171</fpage>
          -
          <lpage>86</lpage>
          , Minneapolis, Minnesota, June. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Eileen</given-names>
            <surname>Fitzpatrick</surname>
          </string-name>
          and
          <string-name>
            <given-names>Joan</given-names>
            <surname>Bachenko</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Building a Data Collection for Deception Research</article-title>
          .
          <source>In Proceedings of the Workshop on Computational Approaches</source>
          to Deception Detection, pages
          <fpage>31</fpage>
          -
          <lpage>8</lpage>
          , Avignon, France, April. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Mark</given-names>
            <surname>Hall</surname>
          </string-name>
          , Eibe Frank, Geoffrey Holmes, Bernhard Pfahringer,
          <string-name>
            <given-names>Peter</given-names>
            <surname>Reutemann</surname>
          </string-name>
          , and
          <string-name>
            <surname>Ian</surname>
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Witten</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>The WEKA data mining software: an update</article-title>
          .
          <source>SIGKDD Explorations</source>
          ,
          <volume>11</volume>
          (
          <issue>1</issue>
          ):
          <fpage>10</fpage>
          -
          <lpage>18</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Valerie</given-names>
            <surname>Hauch</surname>
          </string-name>
          , Iris Blandón-Gitlin,
          <string-name>
            <given-names>Jaume</given-names>
            <surname>Masip</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Siegfried L.</given-names>
            <surname>Sporer</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Are Computers Effective Lie Detectors? A Meta-Analysis of Linguistic Cues to Deception</article-title>
          . Personality and Social Psychology Review,
          <volume>19</volume>
          (
          <issue>4</issue>
          ):
          <fpage>307</fpage>
          -
          <lpage>342</lpage>
          . PMID:
          <volume>25387767</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Geert</given-names>
            <surname>Hofstede</surname>
          </string-name>
          .
          <year>1980</year>
          .
          <article-title>Culture's consequences: International differences in work-related values</article-title>
          .
          <source>Sage Publications.</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Benjamin D.</given-names>
            <surname>Horne</surname>
          </string-name>
          and
          <string-name>
            <given-names>Sibel</given-names>
            <surname>Adali</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>This Just In: Fake News Packs a Lot in Title, Uses Simpler, Repetitive Content in Text Body, More Similar to Satire than Real News</article-title>
          . ArXiv, abs/1703.09398.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Marcia K. Johnson</surname>
            , Julie G. Bush, and
            <given-names>Karen J.</given-names>
          </string-name>
          <string-name>
            <surname>Mitchell</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <article-title>Interpersonal Reality Monitoring: Judging the Sources of Other People's Memories</article-title>
          .
          <source>Social Cognition</source>
          ,
          <volume>16</volume>
          (
          <issue>2</issue>
          ):
          <fpage>199</fpage>
          -
          <lpage>224</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Bennett</given-names>
            <surname>Kleinberg</surname>
          </string-name>
          , Isabelle van der Vegt, Arnoud Arntz, and
          <string-name>
            <given-names>Bruno</given-names>
            <surname>Verschuere</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Detecting deceptive communication through linguistic concreteness</article-title>
          , Mar.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Elena</given-names>
            <surname>Kochkina</surname>
          </string-name>
          , Maria Liakata, and
          <string-name>
            <given-names>Arkaitz</given-names>
            <surname>Zubiaga</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>PHEME dataset for Rumour Detection</article-title>
          and
          <string-name>
            <given-names>Veracity</given-names>
            <surname>Classification</surname>
          </string-name>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Dimitra</given-names>
            <surname>Koutsantoni</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Greek Cultural Characteristics and Academic Writing</article-title>
          .
          <source>Journal of Modern Greek Studies</source>
          ,
          <volume>23</volume>
          :
          <fpage>97</fpage>
          -
          <lpage>138</lpage>
          ,
          <fpage>05</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>John Koutsikakis</surname>
            , Ilias Chalkidis, Prodromos Malakasiotis, and
            <given-names>Ion</given-names>
          </string-name>
          <string-name>
            <surname>Androutsopoulos</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>GREEKBERT: The Greeks Visiting Sesame Street</article-title>
          .
          <source>In 11th Hellenic Conference on Artificial Intelligence , SETN</source>
          <year>2020</year>
          , page 110-
          <fpage>117</fpage>
          , New York, NY, USA. Association for Computing Machinery.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Niels</given-names>
            <surname>Landwehr</surname>
          </string-name>
          , Mark Hall, and
          <string-name>
            <given-names>Eibe</given-names>
            <surname>Frank</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Logistic Model Trees</article-title>
          .
          <source>Machine Learning</source>
          ,
          <volume>59</volume>
          (
          <issue>1</issue>
          ):
          <fpage>161</fpage>
          -
          <lpage>205</lpage>
          , May.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>S. Le</given-names>
            <surname>Cessie</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.C. Van</given-names>
            <surname>Houwelingen</surname>
          </string-name>
          .
          <year>1992</year>
          .
          <article-title>Ridge Estimators in Logistic Regression</article-title>
          . Applied Statistics,
          <volume>41</volume>
          (
          <issue>1</issue>
          ):
          <fpage>191</fpage>
          -
          <lpage>201</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Jaume</given-names>
            <surname>Masip</surname>
          </string-name>
          , Siegfried L. Sporer, Eugenio Garrido, and
          <string-name>
            <given-names>Carmen</given-names>
            <surname>Herrero</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>The detection of deception with the reality monitoring approach: a review of the empirical evidence</article-title>
          . Psychology, Crime &amp; Law,
          <volume>11</volume>
          (
          <issue>1</issue>
          ):
          <fpage>99</fpage>
          -
          <lpage>122</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Galit</given-names>
            <surname>Nahari</surname>
          </string-name>
          , Aldert Vrij, and Ronald P. Fisher.
          <year>2014</year>
          .
          <article-title>The Verifiability Approach: Countermeasures Facilitate its Ability to Discriminate Between Truths and Lies</article-title>
          . Applied Cognitive Psychology,
          <volume>28</volume>
          (
          <issue>1</issue>
          ):
          <fpage>122</fpage>
          -
          <lpage>128</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Matthew L. Newman</surname>
          </string-name>
          , James W. Pennebaker, Diane S. Berry, and
          <string-name>
            <surname>Jane</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Richards</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Lying Words: Predicting Deception from Linguistic Styles</article-title>
          .
          <source>Personality and Social Psychology Bulletin</source>
          ,
          <volume>29</volume>
          (
          <issue>5</issue>
          ):
          <fpage>665</fpage>
          -
          <lpage>75</lpage>
          . PMID:
          <volume>15272998</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Myle</given-names>
            <surname>Ott</surname>
          </string-name>
          , Yejin Choi,
          <string-name>
            <given-names>Claire</given-names>
            <surname>Cardie</surname>
          </string-name>
          , and
          <string-name>
            <surname>Jeffrey</surname>
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Hancock</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Finding Deceptive Opinion Spam by Any Stretch of the Imagination</article-title>
          .
          <source>In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies - Volume 1, HLT '11</source>
          , pages
          <fpage>309</fpage>
          -
          <lpage>19</lpage>
          , Stroudsburg, PA, USA. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Katerina</given-names>
            <surname>Papantoniou</surname>
          </string-name>
          , Panagiotis Papadakos, Theodore Patkos, Giorgos Flouris, Ion Androutsopoulos, and
          <string-name>
            <given-names>Dimitris</given-names>
            <surname>Plexousakis</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Deception detection in text and its relation to the cultural dimension of individualism/collectivism. Natural Language Engineering</article-title>
          .
          <article-title>Also appeared as an arXiv preprint</article-title>
          arXiv:
          <volume>2105</volume>
          .
          <fpage>12530</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Verónica</given-names>
            <surname>Pérez-Rosas</surname>
          </string-name>
          , Bennett Kleinberg, Alexandra Lefevre, and
          <string-name>
            <given-names>Rada</given-names>
            <surname>Mihalcea</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Automatic Detection of Fake News</article-title>
          .
          <source>In Proceedings of the 27th International Conference on Computational Linguistics</source>
          , pages
          <fpage>3391</fpage>
          -
          <lpage>3401</lpage>
          ,
          <string-name>
            <given-names>Santa</given-names>
            <surname>Fe</surname>
          </string-name>
          , New Mexico, USA,
          <year>August</year>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Denis</given-names>
            <surname>Peskov</surname>
          </string-name>
          , Benny Cheng, Ahmed Elgohary, Joe Barrow,
          <article-title>Cristian Danescu-Niculescu-</article-title>
          <string-name>
            <surname>Mizil</surname>
          </string-name>
          , and
          <string-name>
            <surname>Jordan</surname>
          </string-name>
          Boyd-Graber.
          <year>2020</year>
          . It Takes Two to Lie:
          <article-title>One to Lie, and One to Listen</article-title>
          .
          <source>In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics</source>
          , pages
          <fpage>3811</fpage>
          -
          <lpage>3854</lpage>
          , Online, July. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Danae</given-names>
            <surname>Pla</surname>
          </string-name>
          <string-name>
            <surname>Karidi</surname>
          </string-name>
          , Harry Nakos, and
          <string-name>
            <given-names>Yannis</given-names>
            <surname>Stavrakas</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Automatic Ground Truth Dataset Creation for Fake News Detection in Social Media</article-title>
          . In Hujun Yin, David Camacho,
          <string-name>
            <given-names>Peter</given-names>
            <surname>Tino</surname>
          </string-name>
          , Antonio J.
          <string-name>
            <surname>Tallón-Ballesteros</surname>
            ,
            <given-names>Ronaldo</given-names>
          </string-name>
          <string-name>
            <surname>Menezes</surname>
          </string-name>
          , and Richard Allmendinger, editors,
          <source>Intelligent Data Engineering and Automated Learning - IDEAL 2019</source>
          , pages
          <fpage>424</fpage>
          -
          <lpage>436</lpage>
          , Cham. Springer International Publishing.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Stephen</given-names>
            <surname>Skalicky</surname>
          </string-name>
          ,
          <source>Nicholas Duran, and Scott A Crossley</source>
          .
          <year>2020</year>
          . Please, Please, Just Tell Me:
          <article-title>The Linguistic Features of Humorous Deception</article-title>
          .
          <source>Dialogue &amp; Discourse</source>
          ,
          <volume>11</volume>
          (
          <issue>2</issue>
          ):
          <fpage>128</fpage>
          -
          <lpage>149</lpage>
          , December.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Emmanouil</given-names>
            <surname>Smyrnakis</surname>
          </string-name>
          , Katerina Papantoniou, Panagiotis Papadakos, and
          <string-name>
            <given-names>Yannis</given-names>
            <surname>Tzitzikas</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Chattack: A Gamified Crowd-sourcing Platform for Tagging Deceptive &amp; Abusive Behaviour</article-title>
          .
          <source>In European Conference on Information Retrieval</source>
          , pages
          <fpage>549</fpage>
          -
          <lpage>553</lpage>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>Felix</given-names>
            <surname>Soldner</surname>
          </string-name>
          , Verónica Pérez-Rosas, and
          <string-name>
            <given-names>Rada</given-names>
            <surname>Mihalcea</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Box of Lies: Multimodal Deception Detection in Dialogues</article-title>
          .
          <source>In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (Long and Short Papers), pages
          <fpage>1768</fpage>
          -
          <lpage>1777</lpage>
          , Minneapolis, Minnesota, June. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <given-names>John</given-names>
            <surname>Sweller</surname>
          </string-name>
          .
          <year>2011</year>
          . Chapter Two - Cognitive
          <source>Load Theory</source>
          . volume
          <volume>55</volume>
          <source>of Psychology of Learning and Motivation</source>
          , pages
          <fpage>37</fpage>
          -
          <lpage>76</lpage>
          . Academic Press.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>Paul J.</given-names>
            <surname>Taylor</surname>
          </string-name>
          , Samuel Larner,
          <string-name>
            <surname>Stacey M. Conchie</surname>
            , and
            <given-names>Tarek</given-names>
          </string-name>
          <string-name>
            <surname>Menacere</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Culture moderates changes in linguistic self-presentation and detail provision when deceiving others</article-title>
          .
          <source>Royal Society Open Science</source>
          ,
          <volume>4</volume>
          (
          <issue>6</issue>
          ):
          <fpage>170128</fpage>
          , June.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <given-names>Harry C.</given-names>
            <surname>Triandis</surname>
          </string-name>
          and
          <string-name>
            <given-names>Vasso</given-names>
            <surname>Vassiliou</surname>
          </string-name>
          .
          <year>1972</year>
          .
          <article-title>Interpersonal influence and employee selection in two cultures</article-title>
          .
          <source>Journal of Applied Psychology</source>
          ,
          <volume>56</volume>
          :
          <fpage>140</fpage>
          -
          <lpage>145</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <given-names>Udo</given-names>
            <surname>Undeutsch</surname>
          </string-name>
          ,
          <year>1989</year>
          .
          <source>The Development of Statement Reality Analysis</source>
          , pages
          <fpage>101</fpage>
          -
          <lpage>19</lpage>
          . Springer Netherlands, Dordrecht.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <given-names>William</given-names>
            <surname>Yang Wang</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>"Liar, Liar Pants on Fire": A New Benchmark Dataset for Fake News Detection</article-title>
          . In Regina Barzilay and
          <string-name>
            <surname>Min-Yen</surname>
            <given-names>Kan</given-names>
          </string-name>
          , editors,
          <source>Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics</source>
          ,
          <string-name>
            <surname>ACL</surname>
          </string-name>
          <year>2017</year>
          , Vancouver, Canada,
          <source>July 30 - August 4</source>
          , Volume
          <volume>2</volume>
          :
          <string-name>
            <given-names>Short</given-names>
            <surname>Papers</surname>
          </string-name>
          , pages
          <fpage>422</fpage>
          -
          <lpage>426</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>