<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>On The Pursuit of Fake News : Graph Neural Network meets NLP</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>France zpehlivan@ina.fr</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>This paper presents the methods proposed by FakeINA team to participate The FakeNews: Corona Virus and Conspiracies Multimedia Analysis tasks. We concentrate our work on text-based misinformation and conspiracy detection. We proposed a multimodal neural network that combines a graph neural network (GNN) where a document is represented as graph and a multi-layer perceptron model where textual statistics are used as features. Experimental results show that however GNNs are able to classify the data, a multimodal performs better.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Mediaeval Fake News task[
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ] focuses on the classification of
tweet texts aiming detection of fast spreading misinformation. This
task contains three sub-tasks : Text-Based Misinformation
Detection, Text-Based Conspiracy Theories Recognition and Text-Based
Combined Misinformation and Conspiracies Detection. This work
proposed a multimodal neural network approach which is only
applied to the first two sub-tasks.
      </p>
      <p>
        Graph neural network (GNN) methods have been profoundly
useful in several domains including natural language processing[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
While it is probably most apparent to regard text as sequential
data, there are several methods to represent text as various kinds
of graphs. Dependency graph construction generates a graph by
extracting the dependency relations from the dependency parsing
tree. Constituency graph construction captures phrase-based
syntactic relations in a sentence. Another way of representing text
as a graph is to use word co-occurrence and/or document word
relations.
      </p>
      <p>
        TextGCN [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] proposes to build a single heterogeneous graph for
whole corpus and captures global word co-occurrence information.
This approache converts text classification task to node
classification task. On the other hand, TextING[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] builds a graph for each
document based on the co-occurrence of words where each node is
represented as a word embedding and sliding window is used to
capture the relation between words. It learns the fine-grained word
representation of the local structure by GNN to efectively produce
embeddings for obscure words in the new text. By representing
each document as a graph, text classification task becomes graph
classification task for GNNs.
      </p>
      <p>We present our approach in Section 2 and we discuss the results
in Section 3.</p>
    </sec>
    <sec id="sec-2">
      <title>APPROACH</title>
      <p>In this section, we present the preprocessing applied to the raw
text before graph construction. Two models proposed and their
implementation details are also presented in this section.
2.1</p>
    </sec>
    <sec id="sec-3">
      <title>Preprocessing</title>
      <p>
        We use the spaCy[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] library for Python in order to create a text file
containing the original tokens and their normalized counterparts.
For this normalization we have transformed each token to lowercase
and removed stop words. Also, considering the fact that BERT
allows a maximum of 512 tokens per sequence and the given dataset
contains sentences above that range, Bert tokenizer is used with
truncation option. However, BERT is recommended to use with
the raw text, lemmatization and stemming remain important to
generate a graph since a node (word) would be disrupted by an
irrelevant inflection like a simple plural.
2.2
      </p>
    </sec>
    <sec id="sec-4">
      <title>Graph Construction</title>
      <p>
        We use the same graph construction approach as described in
TextING[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Each document is represented as a undirected graph
where nodes are words and co-occurrences between words
represented as edges. The co-occurrences is calculated by using a sliding
window. Embedding of the nodes are initialized by extracting word
embeddings from BERT model[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
2.3
      </p>
    </sec>
    <sec id="sec-5">
      <title>Models</title>
      <p>
        After graph construction, the task converts to the graph
classification task. The main idea behind GNNs is to compute a state for
each node and update this state according to neighbouring nodes
states at each iteration. Graph Isomorphism Network (GIN)[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] was
proposed as a special case of spatial GNN suitable for graph
classification tasks to overcome the issue of distinguishing non-isomorphic
graphs. The authors argues that GIN is possibly as powerful as the
Weisfeiler-Leman test [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] test for graph classification tasks. Thus,
GIN is used in our experiments as GNN choice. Figure 1 resumes
our framework where two GIN convolutional layers are followed
by pooling layer (sum is used) and fully connected layers.
      </p>
      <p>
        We also implemented a multimodal approach by combining GIN
model with multilayer perceptron (MLP). As seen in Figure 2, we do
not only generate graphs from input texts but also extract textual
features s listed in Table 1, by using textstat[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] Python package.
These features become input layer for MLP. We extract the
embeddings from the last hidden layer and concatenat them with the graph
embeddings obtained just after pooling layer. They are sent to MLP
whose output layer will return the predictions for classification
task.
      </p>
      <p>lfesch reading ease
lfesch kincaid grade
automated readability index
dale chall readability score</p>
      <p>
        reading time
monosyllab count
syllable count
lexicon count
sentence count
char count
letter count
emoji count
All the models are implemented by using Pytorh Geometric[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. For
text-based misinformation task, we used the negative log
likelihood loss function, Adam optimizer [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], and StepLR scheduler. For
the conspiracy detection we also used Adam optimizer [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] StepLR
scheduler with binary cross entropy with logits loss function. As
the dataset is not balanced, weights are provided for both loss
functions. We implemented a grid search to find best values for sliding
window size, number of GNN layer and hidden layers. Best value
for sliding window size was 3 and number of GNN layer was 2.
3
      </p>
    </sec>
    <sec id="sec-6">
      <title>RESULTS AND DISCUSSIONS</title>
      <p>Stratified K-Fold cross validation model (with k=10) is used to
measure the performance. For each fold, dataset is split into
training(80%), validation (20%). Due to the small size of the dataset and
overfitting issues during training we did not use test split. Table
2 shows the results for K-Fold CV by using Matthews correlation
coeficient.</p>
      <sec id="sec-6-1">
        <title>Task</title>
      </sec>
      <sec id="sec-6-2">
        <title>Task 1</title>
        <p>Task 1
Task 1</p>
      </sec>
      <sec id="sec-6-3">
        <title>Task 2 Task 2 Task 2</title>
      </sec>
      <sec id="sec-6-4">
        <title>Model</title>
        <p>multimodal
multimodal</p>
        <p>GIN
multimodal
multimodal
multimodal</p>
      </sec>
      <sec id="sec-6-5">
        <title>Hidden Layers Val MCC Oficial MCC 128</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <year>2018</year>
          .
          <article-title>textstat: Textstat is an easy to use library to calculate statistics from text. It helps determine readability, complexity, and grade level</article-title>
          . (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding</article-title>
          . (
          <year>2019</year>
          ).
          <article-title>arXiv:cs</article-title>
          .CL/
          <year>1810</year>
          .04805
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Matthias</given-names>
            <surname>Fey</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jan E.</given-names>
            <surname>Lenssen</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Fast Graph Representation Learning with PyTorch Geometric</article-title>
          .
          <source>In ICLR Workshop on Representation Learning on Graphs and Manifolds.</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Matthew</given-names>
            <surname>Honnibal</surname>
          </string-name>
          and
          <string-name>
            <given-names>Ines</given-names>
            <surname>Montani</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing</article-title>
          . (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Diederick</surname>
            <given-names>P</given-names>
          </string-name>
          <string-name>
            <surname>Kingma and Jimmy Ba</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Adam: A method for stochastic optimization</article-title>
          .
          <source>In International Conference on Learning Representations (ICLR).</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Konstantin</given-names>
            <surname>Pogorelov</surname>
          </string-name>
          , Daniel Thilo Schroeder, Stefan Brenner, and
          <string-name>
            <given-names>Johannes</given-names>
            <surname>Langguth</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>FakeNews: Corona Virus and Conspiracies Multimedia Analysis Task at MediaEval 2021</article-title>
          .
          <source>Proc. of the MediaEval 2021 Workshop</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Konstantin</given-names>
            <surname>Pogorelov</surname>
          </string-name>
          , Daniel Thilo Schroeder, Petra Filkuková, Stefan Brenner, and
          <string-name>
            <given-names>Johannes</given-names>
            <surname>Langguth</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>WICO Text: A Labeled Dataset of Conspiracy Theory and 5G-Corona Misinformation Tweets</article-title>
          .
          <source>Proc. of the 2021 Workshop on Open Challenges in Online Social Networks</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Boris</given-names>
            <surname>Weisfeiler</surname>
          </string-name>
          and
          <string-name>
            <given-names>Andrei</given-names>
            <surname>Leman</surname>
          </string-name>
          .
          <year>1968</year>
          .
          <article-title>The reduction of a graph to canonical form and the algebra which appears therein</article-title>
          .
          <source>NTI, Series</source>
          <volume>2</volume>
          ,
          <issue>9</issue>
          (
          <year>1968</year>
          ),
          <fpage>12</fpage>
          -
          <lpage>16</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Lingfei</given-names>
            <surname>Wu</surname>
          </string-name>
          , Yu Chen, Kai Shen, Xiaojie Guo,
          <string-name>
            <given-names>Hanning</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Shucheng</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Jian</given-names>
            <surname>Pei</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Bo</given-names>
            <surname>Long</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Graph Neural Networks for Natural Language Processing: A Survey</article-title>
          . arXiv:
          <volume>2106</volume>
          .06090 [cs] (
          <year>June 2021</year>
          ). http://arxiv.org/abs/2106.06090 arXiv:
          <fpage>2106</fpage>
          .
          <fpage>06090</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Keyulu</surname>
            <given-names>Xu</given-names>
          </string-name>
          , Weihua Hu, Jure Leskovec, and
          <string-name>
            <given-names>Stefanie</given-names>
            <surname>Jegelka</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>How Powerful are Graph Neural Networks?</article-title>
          .
          <source>In 7th International Conference on Learning Representations, ICLR</source>
          <year>2019</year>
          ,
          <article-title>New Orleans</article-title>
          , LA, USA, May 6-
          <issue>9</issue>
          ,
          <year>2019</year>
          . OpenReview.net. https://openreview.net/forum?id=ryGs6iA5Km
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Liang</surname>
            <given-names>Yao</given-names>
          </string-name>
          , Chengsheng Mao, and
          <string-name>
            <given-names>Yuan</given-names>
            <surname>Luo</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Graph Convolutional Networks for Text Classification</article-title>
          . arXiv:
          <year>1809</year>
          .05679 [cs] (
          <year>Nov</year>
          .
          <year>2018</year>
          ). http://arxiv.org/abs/
          <year>1809</year>
          .05679 arXiv:
          <year>1809</year>
          .05679.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Yufeng</surname>
            <given-names>Zhang</given-names>
          </string-name>
          , Xueli Yu, Zeyu Cui, Shu Wu, Zhongzhen Wen, and
          <string-name>
            <given-names>Liang</given-names>
            <surname>Wang</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Every Document Owns Its Structure: Inductive Text Classification via Graph Neural Networks</article-title>
          . arXiv:
          <year>2004</year>
          .13826 [cs] (May
          <year>2020</year>
          ). http://arxiv.org/abs/
          <year>2004</year>
          .13826 arXiv:
          <year>2004</year>
          .13826.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>