<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Development of Software for Fake News Detection</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Olga N. Kaneva</string-name>
          <email>onkaneva@omgtu.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aleksandr I. Goncharenko</string-name>
          <email>kesha787898@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Omsk State Technical University</institution>
          ,
          <addr-line>Omsk, Russia, Mira Ave. 11, 644050</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>During the research process, a fake news detection approach has been developed. The proposed algorithm includes a FastText-based vectorizer, a clustering algorithm, and a set of fully connected neural networks. The algorithm is developed in Python using machine learning libraries and an API is created with the frameworkFlask. The software package is available on the Internet and is accessible to users.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Copyright ⃝c by the paper's authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).
In: Sergei S. Goncharov, Yuri G. Evtushenko (eds.): Proceedings of the Workshop on Applied Mathematics and Fundamental
Computer Science 2021 , Omsk, Russia, 24-04-2021, published at http://ceur-ws.org</p>
      <p>For tasks where the emotional component of the text is important, it is best to average vectors for each
feature. The heuristic is that when using embedding, the resulting embedding space holds information about
the emotional polarity of each word. Therefore, averaging the polarity, one gets the average emotional polarity
of the text.</p>
      <p>This method was applied to the obtained vectors, and as a result, a vector with the dimension of 300 was
obtained for each text.Further, to preserve more information from the original texts, the principal component
analysis was used [Ebi018].</p>
      <p>For each text, the method of reducing the dimension by words was applied. One principal component was
selected for each text. As a result, a vector with the dimension of 300 was obtained.</p>
      <p>Further, the vectors obtained as a result of calculating the average and were concatenated into a vector with
the dimension of 600through the principal component analysis.</p>
      <p>Then, the vectors with the dimension of 600 were employed for clustering and the use of neural networks.
3</p>
    </sec>
    <sec id="sec-2">
      <title>Clustering of texts by headlines</title>
      <p>The clustering task belongs to unsupervised learning. Clustering can be used to divide the data into several
groups, called clusters. Objects in the same cluster can be called related.Clustering can be considered as an
optimization problem where one minimizes a metric that evaluates how well a set is partitioned into subsets.</p>
      <p>Clustering is applied in this problem because of the large amount of different data. The heuristic is that it will
be easier for a neural network to separate data that has been previously partitioned by a qualitatively different
algorithm.Lloyd’s algorithm has been chosen as a clustering algorithm; it minimizes the sum of the squares of
the intra-cluster distances.</p>
      <p>The main idea is as follows. At each iteration, the centre of mass is recalculated for each cluster obtained at
the previous step. Then, the vectors are split into clusters again according to the new centre which turned out
to be closely based on the chosen metric.</p>
      <p>The algorithm ends when there is no change in the intra-cluster distance at some iteration. This happens in
a finite number of iterations since the number of possible partitions of a finite set is finite, and at each step the
total quadratic deviation decreases; therefore, looping is impossible.</p>
      <p>The hyper-parameters of the algorithm are K, the number of clusters, and p, the measure of distance.
To solve this problem it was decided to use a complex metric [Dav2018], combining
1) inertia;
2) CalinskiHarabasz metric;
3) Davies-Bouldin metric;
4) the size of the smallest cluster;
5) the range of cluster sizes.</p>
      <p>The first three metrics are responsible for making the clusters look alike, and the last two are in charge of
cluster size.</p>
      <p>All these metrics were aggregated into a single metric.</p>
      <p>During the clustering experiment, the dimension of the data was reduced through the principal component
analysis. The components giving 90% dispersion were selected.</p>
      <p>The optimization problem was solved by the direct search with the upper restriction of the hyper-parameter,
i. e. the number of clusters.</p>
      <p>Metrics were calculated for each hyper-parameter value and then aggregated into a single one. The learning
was conducted with the M iniBatchKM eans algorithm, making it possible to learn not on the entire dataset,
but mini-batches.</p>
      <p>For this task, the optimal number of clusters was 8. It was with this value of the hyper-parameter that the
final model was trained.</p>
      <p>Since a small range of cluster sizes was critical for the solution of the problem, it was decided to combine
some clusters.All small clusters were combined into a separate cluster, and objects with low confidence about
belonging to the cluster were added. Confidence was derived by predicting the probability of belonging to each
cluster.</p>
      <p>Clusters 3 and 6 and elements with less than 20 % confidence were combined into a new cluster. The
visualization of clusters using the dimension reduction algorithm can be seen in Figure 1.</p>
    </sec>
    <sec id="sec-3">
      <title>Selecting models for each cluster</title>
      <p>For each cluster, the neural network must be trained to perform the binary classification task.</p>
      <p>A neural network was chosen as a model because it allows for minimal preparation of features, can work with
feature spaces of large dimension, and can extract dependencies of a sufficiently high order. At the same time,
fully connected neural networks are quite fast classifiers [Ben2019].</p>
      <p>To solve this problem, it was decided to use the architecture of a fully connected neural network consisting
of 3 layers with a patch-normalization, a dropout and activation function P ReLU . Crossentropy was chosen as
the loss function, and the Adamoptimizer was used to train the model. Also, weights for classes were used to
solve the problem of class imbalance. F 1 score was chosen as the metric.</p>
      <p>The results of hold-out validation can be seen in Figure 2. These figures show the precision, recall and
F 1 score.
5</p>
    </sec>
    <sec id="sec-4">
      <title>Embedding in a server application</title>
      <p>The resulting models for feature extraction, clustering, and prediction were trained in jupyter notebooks, which
is convenient enough to perform experiments. However, this format of writing code does not allow it to be
integrated into the application. To solve this problem, the models were saved as files and then loaded into the
server application.</p>
      <p>The server application was used to create a simple AP I that allowed any user to utilize the trained models.
The AP I also allows other programs to connect to the models even if the third-party programs are not written in
P ython. As a result, this AP I can potentially be used by desktop applications as well as conventional W ebsites
and mobile applications.</p>
      <p>The framework F lask [Gri2014] for the P ython language was used to create the AP I. It is often referred
to as a class of micro-frameworks since it is rarely used to write large applications and is rather intended for
creating small frameworks for web applications.</p>
      <p>The algorithm reproduced in the jupyter notebooks was integrated into the application.</p>
      <p>The algorithm itself expects an input of the text and the headline of the news in English. To do this, the
resulting entities are cleaned of garbage characters using regular expressions. Next, the lines are checked for
non-emptiness.</p>
      <p>If both lines are empty, the algorithm sends an empty response, which should be processed on the user’s side.</p>
      <p>Next, the input data are converted using the written T extT oV ec class, which implements a text conversion
algorithm. Firstly, the class divides the text into tokens, each token is converted into a vector using F astT ext,
and then the average of the vectors and their first principal component are selected.</p>
      <p>Next, if there is a headline in the input data, a cluster is defined, using Clusterizer inherited from P redictor.
This class implements the definition of a cluster depending on the distance to the cluster centres. Additionally,
if there is no confidence in belonging to the right cluster, then the object is sent to the garbage cluster and
further clustering is not used. Also, since some clusters are combined into one, the dictionary determines the
true cluster.</p>
      <p>The presence of clustering allows us to add new model data. To do this, it is only needed to find new data,
determine their feature, train a neural network on them, and build this cluster into an existing clustering. In
this case, it is not necessary to retrain the entire model, which is a time-consuming process.</p>
      <p>Next, if the cluster of the data is not defined or it is defined as a garbage one, the neural network trained on
the entire dataset is activated.</p>
      <p>If there is no problem with the cluster number, then, depending on the number, a model trained on similar
data is activated. T orchP redictor class is used for this purpose. Probability is received as a result.</p>
      <p>This algorithm was successfully integrated into the Flask application.
6</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>During the research process, a fake news detection approach has been developed. To do this, data have been
collected from various sources and then used as a basis for a rather large dataset containing fake news on
various topics. A clustering algorithm has been developed, which is based on the algorithm Kmeans, which
groups vectors of texts according to the closeness of their value. Also, the sizes of the resulting clusters have
been improved. A detection algorithm has been developed for each cluster, the algorithm being a set of neural
networks trained on the source data. The developed algorithm for detecting fake news has been built into the
server application on f lask and an open AP I has been created.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [Suk2017]
          <string-name>
            <given-names>A. P.</given-names>
            <surname>Sukhodolov</surname>
          </string-name>
          .
          <article-title>The phenomenon of ”fake news” in modern media space</article-title>
          . - Eurasian
          <string-name>
            <surname>Partnership</surname>
          </string-name>
          :
          <article-title>Humanitarian Aspects</article-title>
          .
          <source>Proceedings of the International Scienti c and Practical Conference</source>
          . Irkutsk: Publishing house of Baikal State University,
          <volume>1</volume>
          :
          <fpage>93</fpage>
          -
          <lpage>112</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [Fat2018]
          <string-name>
            <given-names>I. A.</given-names>
            <surname>Fateeva</surname>
          </string-name>
          .
          <article-title>Social networks in the aspect of media safety. Media-education 2014: Proceedings of the All-Russian scienti c and practical conference (with international participation) "Media-education 2014</article-title>
          .
          <article-title>Regional aspect" 7 October -</article-title>
          9
          <source>October</source>
          <year>2014</year>
          / Edited by
          <string-name>
            <given-names>I.V.</given-names>
            <surname>Zhilavskaya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.A.</given-names>
            <surname>Karyagina</surname>
          </string-name>
          ,
          <volume>304</volume>
          -
          <fpage>310</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>[Ebi018] H. M. Ebied</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Yu</surname>
            . Savelev,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Yu</surname>
          </string-name>
          .Fink.
          <article-title>Feature extraction using PCA and Kernel-PCA for face recognition</article-title>
          .
          <source>2012 8th International Conference on Informatics and Systems (INFOS)</source>
          , pp.
          <source>MM-72- MM-77</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [Dav2018]
          <string-name>
            <given-names>D. L.</given-names>
            <surname>Davies</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. W.</given-names>
            <surname>Bouldin</surname>
          </string-name>
          .
          <article-title>A Cluster Separation Measure</article-title>
          .
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          , PAMI-
          <volume>1</volume>
          (
          <issue>2</issue>
          ):
          <fpage>224</fpage>
          -
          <lpage>227</lpage>
          , april
          <year>1979</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [Ben2019]
          <string-name>
            <given-names>B.</given-names>
            <surname>Bengfort</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bilbro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ojeda</surname>
          </string-name>
          .
          <article-title>Applied Text Analysis with Python</article-title>
          .
          <source>AIP Spb: Peter</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [Gri2014] Flask Web Development / Edited by Miguel
          <string-name>
            <surname>Grinberg. O'Reilly</surname>
          </string-name>
          /DMK Press,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>