<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>SEBD</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Deepfake Detection: Challenges and Solutions</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Davide Alessandro Coccomini</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>ISTI-CNR</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italy</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>31</volume>
      <fpage>02</fpage>
      <lpage>05</lpage>
      <abstract>
        <p>ing the concept of deepfake to such an extent that it could detect images or videos that had been manipulated even with novel techniques. In [2] and [5] we compared Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) of various kinds by putting them in a cross-forgery context revealing the superiority of the ViTs, which are less tied to the specific anomalies they see during training. After that, noting a scarcity of methods based on ViT and even more so those based on hybrid architectures, we developed our first real deepfake detector. In [1] we created a new architecture, combining an EficientNet-B0 and Cross ViT, which we have named Convolutional Cross Vision Transformer. Thanks to the local-global attention mechanism within it and the exploitation of features extracted from the CNN, the model was able to efectively detect deepfake videos, achieving SOTA results on DFDC[6] and FaceForensics++[9] dataset, all while keeping the number of parameters low. The model was also used to participate in the competition presented in [7]. In [4] we designed a new type of Convolutional TimeSformer that take into account both the spatial position of faces in the frame and their temporal position in the video. It is also capable of managing multiple identities and being robust to face-size movements thanks to the introduction of a novel attention mechanism and positional embedding. Our method surpassed the SOTA on in-dataset tests on [8] and performed robustly in real-world situations. Future work will mainly focus on improving deepfake detectors in order to make them more robust to other real-world problems. We also want to make detectors capable of combining information also of a textual nature, context, and the reputation of the account disseminating it, to understand video veracity. Also, as we started doing in [3], we will work on the more generic problem of synthetic content detection.</p>
      </abstract>
    </article-meta>
  </front>
  <body />
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Davide</given-names>
            <surname>Alessandro</surname>
          </string-name>
          Coccomini et al. “
          <article-title>Combining EficientNet and Vision Transformers for Video Deepfake Detection”</article-title>
          .
          <source>In: Image Analysis and Processing - ICIAP 2022</source>
          . Ed. by Stan Sclarof et al. Cham: Springer International Publishing,
          <year>2022</year>
          , pp.
          <fpage>219</fpage>
          -
          <lpage>229</lpage>
          . isbn:
          <fpage>978</fpage>
          -3-
          <fpage>031</fpage>
          -06433-3.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Davide</given-names>
            <surname>Alessandro</surname>
          </string-name>
          Coccomini et al. “
          <article-title>Cross-Forgery Analysis of Vision Transformers and CNNs for Deepfake Image Detection”</article-title>
          .
          <source>In: Proceedings of the 1st International Workshop on Multimedia AI against Disinformation. MAD '22</source>
          .
          <string-name>
            <surname>Newark</surname>
          </string-name>
          , NJ, USA: Association for Computing Machinery,
          <year>2022</year>
          , pp.
          <fpage>52</fpage>
          -
          <lpage>58</lpage>
          . isbn:
          <volume>9781450392426</volume>
          . doi:
          <volume>10</volume>
          .1145/3512732.3533582.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>url: https://doi.org/10.1145/3512732.3533582.</mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Davide</given-names>
            <surname>Alessandro</surname>
          </string-name>
          Coccomini et al.
          <source>Detecting Images Generated by Difusers</source>
          .
          <year>2023</year>
          . doi:
          <volume>10</volume>
          .48550/ARXIV.2303.05275. url: https://arxiv.org/abs/2303.05275.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Davide</given-names>
            <surname>Alessandro</surname>
          </string-name>
          Coccomini et al. MINTIME:
          <string-name>
            <surname>Multi-Identity Size-Invariant Video Deepfake Detection</surname>
          </string-name>
          .
          <year>2022</year>
          . doi:
          <volume>10</volume>
          .48550/ARXIV.2211.10996. url: https://arxiv.org/abs/2211.10996.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Davide</given-names>
            <surname>Alessandro</surname>
          </string-name>
          Coccomini et al. “
          <article-title>On the Generalization of Deep Learning Models in Video Deepfake Detection”</article-title>
          .
          <source>In: Journal of Imaging 9.5</source>
          (
          <year>2023</year>
          ). issn:
          <fpage>2313</fpage>
          -
          <lpage>433X</lpage>
          . url: https://www.mdpi.com/2313-433X/9/5/89.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Brian</given-names>
            <surname>Dolhansky</surname>
          </string-name>
          et al. “
          <article-title>The deepfake detection challenge (dfdc) dataset”</article-title>
          . In: arXiv preprint arXiv:
          <year>2006</year>
          .
          <volume>07397</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Luca</given-names>
            <surname>Guarnera</surname>
          </string-name>
          et al. “
          <article-title>The Face Deepfake Detection Challenge”</article-title>
          .
          <source>In: Journal of Imaging</source>
          <volume>8</volume>
          .10 (
          <year>2022</year>
          ). issn:
          <fpage>2313</fpage>
          -
          <lpage>433X</lpage>
          . doi:
          <volume>10</volume>
          .3390/jimaging8100263. url: https://www.mdpi.com/2313- 433X/8/10/263.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Yinan</given-names>
            <surname>He</surname>
          </string-name>
          et al. “
          <article-title>ForgeryNet: A Versatile Benchmark for Comprehensive Forgery Analysis”</article-title>
          .
          <source>In: 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          .
          <year>2021</year>
          , pp.
          <fpage>4358</fpage>
          -
          <lpage>4367</lpage>
          . doi:
          <volume>10</volume>
          .1109/CVPR46437.
          <year>2021</year>
          .
          <volume>00434</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Andreas</given-names>
            <surname>Rossler</surname>
          </string-name>
          et al. “
          <article-title>Faceforensics++: Learning to detect manipulated facial images”</article-title>
          .
          <source>In: Proceedings of the IEEE/CVF International Conference on Computer Vision</source>
          .
          <year>2019</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>11</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>