<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Responsible and Reliable AI at PICUS Lab</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Narendra Patwardhan</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lidia Marassi</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michela Gravina</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antonio Galli</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Monica Zuccarini</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tannistha Maiti</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tarry Singh</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefano Marrone</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carlo Sansone</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Deepkapha AI</institution>
          ,
          <addr-line>Assen</addr-line>
          ,
          <country country="NL">Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Naples, Federico II</institution>
          ,
          <addr-line>Naples</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The PICUS Lab has been conducting research activities in the field of artificial intelligence (AI) with a focus on ethics. One area of research has been the modification of transformer-based models to make them more sustainable while maintaining performance. To achieve this goal, we have explored sustainable alternatives for the internal components of these models, such as retrieval-based techniques and difusion modules for programmability. The aim is to pave the way for the development of ethical and sustainable AI systems that do not rely on massive computing and data, which can lead to high energy consumption and carbon footprint. Another area of research at PICUS Lab has been deepfake detection, a pressing concern due to the potential spread of false information and manipulation of public opinion. To address this issue, we developed FEAD-D (Face Expression Analysis for Deepfake Detection), a tool based on facial expression analysis. The system uses a bidirectional Long Short-Term Memory (BiLSTM) model and data from the DeepFake Detection Challenge (DFDC) to detect fake videos in about two minutes.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Generative Models</kwd>
        <kwd>Transformers</kwd>
        <kwd>Sustainable AI</kwd>
        <kwd>Foundation Models</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        In particular, in the first project we focused on the
development of sustainable AI systems. The increasing
Artificial intelligence (AI) has made significant progress use of AI models raises concerns regarding energy
conin recent years, yielding promising results in various sumption and carbon footprint, prompting researchers
downstream tasks. However, AI models often rely on to explore sustainable alternatives for the internal
commassive computing and data, raising concerns due to ponents of these models. Also, the increasing use of such
high energy consumption and carbon footprint. To ad- generative models for fake news creation is posing
seridress these concerns, at the PICUS Lab we have been ous ethical and sociological concerns. The researchers
conducting research activities in the field of AI with a propose modifications to transformer-based models that
focus on ethical concerns. This paper aims to present maintain performance while reducing their reliance on
two of these research projects carried out at the PICUS massive computing and data, while also opening for a
Lab: modification of transformer-based models for sus- more ethic-by-design training strategy. Retrieval-based
tainability and deepfake detection using facial expression techniques and difusion modules for programmability
analysis. The importance of these topics is supported by are discussed as sustainable alternatives. These
modifirecent studies and reports. The World Economic Forum cations can pave the way for the development of ethical
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] has highlighted the importance of ethical and sus- and sustainable AI systems.
tainable AI systems, stating that "ethical AI can drive Instead, the second project focuses on deepfake
deinnovation, and sustainability should be at the heart of tection, a pressing concern due to the potential spread
AI." Additionally, a report by the European Union on AI of false information and manipulation of public opinion.
regulation highlights the need for ethical and sustainable To address this issue, researchers at the PICUS Lab have
AI systems to ensure the long-term benefit of society. developed a tool called FEAD-D, which stands for Face
These studies and reports highlight the urgency of ad- Expression Analysis for Deepfake Detection. This tool
dressing ethical and sustainability concerns in the field of uses facial expression analysis to detect fake videos,
overAI and provide further support for the research projects coming limitations in current deepfake detection systems.
presented in this paper. The system is based on a bidirectional Long Short-Term
Ital-IA 2023: 3rd National Conference on Artificial Intelligence, orga- Memory (BiLSTM) model and data from the DeepFake
nized by CINI, May 29–31, 2023, Pisa, Italy Detection Challenge (DFDC). FEAD-D had been founded
* Corresponding author. under the CINECA ISCRA-C program (ID. HP10CMJKEO,
$ stefano.marrone@unina.it (S. Marrone) IsC93).
      </p>
      <p>
        0000-0002-4807-5664 (N. Patwardhan); 0000-0001-5033-9617 The results of these studies are important in the
con((SM. .MGarrarvoinnea));; 00000000--00000021--89197161--61955107 ((CA.. SGaanlslio)n;0e0)00-0001-6852-0377 text of the growing use of AI and its potential impact
© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License on society. The development of ethical and sustainable
CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g ACttEribUutRion W4.0oInrtekrnsahtioonpal (PCCroBYce4.0e).dings (CEUR-WS.org) AI systems is crucial for the long-term benefit of
society. This paper highlights the importance of addressing of human societies and avoid becoming outdated or less
ethical and sustainability concerns in the field of AI and reliable over time.
presents two projects that contribute to this efort. In this section, we will illustrate project Hominis,
which exploits AI techniques in the realm of generative
AI and foundation models, carried out at the
Univer2. HOMINIS: Towards Sustainable sity of Naples Federico II in collaboration with industrial
Foundation Models partners (DeepKapha). We will highlight the innovative
aspects and contributions of this project, emphasizing its
Artificial intelligence (AI) has emerged as a transforma- significance in advancing the state-of-the-art in
sustaintive force in modern society, with generative modelling able and programmable AI. By leveraging the expertise
serving as a key driver behind its rapid advancements. of both academic and industrial stakeholders, project
Foundation models [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] are large-scale machine learning Hominis aims to develop cutting-edge solutions and
apmodels, pre-trained on vast amounts of diverse data, that plications that harness the potential of AI to address
serve as a backbone for various downstream applica- real-world challenges and deliver tangible benefits across
tions through fine-tuning and adaptation. Such models, a wide range of sectors.
particularly those based on transformer architectures,
have achieved remarkable performance in domains such 2.1. Essential Additions to Foundation
as natural language processing, computer vision, and
reinforcement learning. Despite their success, the de- Models
velopment and deployment of these models have raised To optimize large-scale language models, data curation
critical concerns regarding the ethical and environmental processes have shifted from human-led to
heuristicsimplications of their high computational requirements based automated filtering. However, this automation can
and energy consumption. lead to biases, memorization of private data, and
vulner
      </p>
      <p>As foundation models grow in size and complexity, ability to adversarial attacks. Our research collaboration
the search for optimal hyperparameters becomes increas- with RealAI aims to develop sustainable foundation
modingly challenging and resource-intensive. Solely rely- els by addressing issues related to data sourcing, key
ing on scaling up can lead to overfitting and may result components, and essential additions. Project Hominis
in diminished returns on model performance improve- aims to sanitize public datasets and develop crawling
ments. It also overlooks opportunities to optimize smaller strategies for capturing diverse, multi-faceted data. This
models, which could deliver similar performance with project will also develop tools for the community to
anareduced costs and resource demands. The high parameter lyze, curate, and critique datasets while ensuring fairness,
count of foundation models presents challenges for infer- privacy, and legality.
ence on standard hardware, such as personal computers The transformer block, based on self-attention, is the
and mobile devices. This limitation restricts access to fundamental constituent of modern foundation models.
the benefits of these models for a broader audience and The attention mechanism plays a crucial role in the
Transmay contribute to a digital divide between those who former architecture, enabling the model to focus on
relcan aford to deploy and utilize advanced AI systems and evant features in the input data. In this study, we will
those who cannot. perform ablation experiments on three attention
acceler</p>
      <p>
        Sustainable AI refers to managing the life cycle of arti- ation techniques: Flash Attention [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] (hardware-aware
ifcial intelligence systems in a way that minimizes nega- acceleration), Linear Approximations [
        <xref ref-type="bibr" rid="ref4 ref5 ref6 ref7">4, 5, 6, 7</xref>
        ]
(acceltive environmental, social, and economic impacts while eration via proxying), and Synthetic Attention [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]
(acmaximizing long-term benefits for society. This holistic celeration through replacement). Our aim is to evaluate
approach emphasizes the importance of ethical consider- the efectiveness of these methods in improving model
ations, energy eficiency, and resource optimization in AI eficiency while maintaining performance levels.
design, as well as fostering inclusive collaboration and The linear layer, another key component of the
Transequitable access to AI-driven technologies. former architecture, can be optimized using routing
tech
      </p>
      <p>
        As the world evolves rapidly, AI models must be able niques (such as those popularized by Switch Transformer
to adapt and learn from new information to remain rele- [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]) to enhance its performance without adversely
afvant and useful. Finetuning large models incurs signif- fecting inference time. By employing routing strategies,
icant costs and hence programmability in AI models is we can achieve a more eficient distribution of resources
necessary so that they can be updated to accommodate within the linear layer, resulting in improved
compuemerging trends, knowledge, and societal shifts. This tational eficiency and reduced resource consumption
capability is crucial for maintaining the accuracy and without sacrificing the quality of the model’s output.
efectiveness of AI systems in real-world applications, Large AI models can benefit from multimodal data,
as it allows them to keep pace with the dynamic nature which improves performance, generalization ability, and
robustness. Networks that process independent modal- essential additions in AI systems.
ities using a common structure exhibit synergistic
effects on generalization. Additionally, diversity in data
sources, particularly in low-resource languages, can im- 3. FEAD-D
prove downstream generalization.
      </p>
      <p>
        To leverage data from diverse domains and take ad- Emotions are a valuable tool for detecting deepfakes
bevantage of this synergetic efect, we rely on two primary cause they are dificult to replicate convincingly. This
factors, tokenization and choice of architecture. This is a limitation of deepfake creation algorithms, as
emostudy proposes Universal Tokenization, an innovative tions are essential for human communication and are
method that focuses on byte-level representation aug- easily conveyed through facial expressions, voice tone,
mented by special tokens, enabling a unified encoding of and body language. Identifying the emotional
characdiverse data types. The other component we rely on is teristics of a video has become an increasingly popular
cross-attention, which allows models to decouple com- method for detecting manipulated content, as deepfakes
putational complexity from sequence length, as demon- often struggle to accurately reproduce emotions,
includstrated by the Perceiver architecture [
        <xref ref-type="bibr" rid="ref10 ref11 ref12">10, 11, 12</xref>
        ]. By using ing micro-expressions that reveal true emotions. FEAD-D
cross-attention mechanisms, these architectures can ef- (Face Expression Analysis for Deepfake Detection)
exifciently handle long sequences and large-scale inputs ploits the inconsistencies in facial expressions introduced
without incurring excessive computational costs. Our by deepfake creation artefacts. These videos often fail to
work will harness this decoupling to efectively support accurately capture the full range of emotions and
micromultimodality. expressions of the original subject. By analyzing a large
      </p>
      <p>
        Retrieval-Augmented Generation (RAG) is a technique sample of fake videos and comparing them to their
origithat combines the strengths of pre-trained language mod- nal counterparts, FEAD-D identified temporal
inconsisels with external knowledge sources, enabling hot up- tencies in the non-natural sequences of facial expressions.
dates and ofering several benefits. The RETRO paper Figure 1 illustrates the diferences in emotional patterns
[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], which popularized this approach, showcases its po- across consecutive frames of real and fake videos,
emtential for enhancing the performance and adaptability of phasizing the importance of considering the temporal
AI models. By incorporating real-time information from evolution of facial expressions in detecting deepfakes. To
external databases, RAG allows models to stay up-to- this aim, FEAD-D consists of the following modules:
date with the latest developments and trends, improving • Face detection, in which the target video is
anatheir relevance and accuracy. This dynamic integration lyzed to detect one or more faces.
of knowledge also facilitates faster updates, reducing the • Features Extraction, that analyzes the texture
need for resource-intensive retraining processes. As a re- of target images and extracts the emotion from
sult, retrieval-augmented generation contributes to more the detected face frame by frame.
eficient and versatile AI models that can efectively
address the ever-evolving needs and challenges in various • Features temporal analysis, in which all the
domains. features extracted in the previous stages are
an
      </p>
      <p>We are also exploring the incorporation of difusion as alyzed together in a cross-frame fashion to spot
the last layer of our AI model to enhance controllability. incoherent and unnatural patterns in the
emoBy adding a difusion-based mechanism, we aim to re- tional evolution of the target subject.
ifne the generated output while preserving the structure The resulting system can process a video in two minutes
and coherence of the model’s predictions. This approach and is easy to adopt with minimal technical knowledge.
allows us to exert greater control over the model’s be- It is worth noting that although emotional analysis is a
haviour and adjust its responses based on specific require- promising approach, it presents challenges related to
variments or constraints. ations across individuals, cultures, and contexts, and the</p>
      <p>
        Although our work primarily emphasizes inference possibility of creating algorithms specifically designed to
time optimizations, we are also committed to reducing mimic emotional expressions. Further research is
necesthe carbon footprint associated with training time. To sary to establish emotional analysis as a reliable method.
achieve this, we employ muTransfer [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], an innovative Future work can focus on developing more robust
deeptechnique that enables the discovery of optimal hyper- fake detection methods using advanced computer vision
parameters by training a scaled-down model, efectively techniques, training models to analyze both lip
moveeliminating the need for resource-intensive hyperparam- ments and speech signals, and investigating the
efectiveeter searches. ness of combining multiple detection methods.
      </p>
      <p>Overall, this research direction aims to create more
sustainable, adaptable, and responsible foundation
models by addressing data sourcing, key components, and</p>
    </sec>
    <sec id="sec-2">
      <title>4. Conclusions</title>
      <p>This paper presented two research projects carried out at
the PICUS Lab with a focus on sustainability and ethical
concerns in artificial intelligence (AI). The first project
explored modifications to transformer-based models to
make them more sustainable while maintaining
performance. The researchers proposed sustainable alternatives
such as retrieval-based techniques and difusion modules
for programmability. The second project focused on
deepfake detection, a pressing concern due to the potential
spread of false information and manipulation of public
opinion. The researchers developed a tool called FEAD-D,
which stands for Face Expression Analysis for Deepfake
Detection, based on facial expression analysis. The
importance of these topics is supported by recent studies
and reports that highlight the urgency of addressing
ethical and sustainability concerns in the field of AI. The
results of these studies contribute to the development of
ethical and sustainable AI systems, which are crucial for
the long-term benefit of society.</p>
      <p>In the first project, we proposed modifications to
transformer-based models that maintain performance
while reducing their reliance on massive computing and
data, thus paving the way for the development of ethical
and sustainable AI systems. In the second project, the
researchers developed a tool called FEAD-D that uses
facial expression analysis to detect fake videos,
overcoming limitations in current deepfake detection systems.
Overall, the research presented in this paper highlights
the importance of addressing ethical and sustainability
concerns in the field of AI and provides important
contributions towards this efort. Future research in this area
should continue to explore sustainable alternatives for AI
models and further improve deepfake detection systems.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>[1] Towards a more sustainable and equitable future for ai</article-title>
          , http://www3.weforum.org/docs/WEF_ Towards_a_More_Sustainable_and_Equitable_ Future_for_
          <source>AI_Report_2018</source>
          .pdf , ????
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>R.</given-names>
            <surname>Bommasani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Hudson</surname>
          </string-name>
          , E. Adeli,
          <string-name>
            <given-names>R.</given-names>
            <surname>Altman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Arora</surname>
          </string-name>
          , S. von Arx,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Bernstein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bohg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bosselut</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Brunskill</surname>
          </string-name>
          , et al.,
          <article-title>On the opportunities and risks of foundation models</article-title>
          ,
          <source>arXiv preprint arXiv:2108.07258</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>T.</given-names>
            <surname>Dao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. Y.</given-names>
            <surname>Fu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ermon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rudra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ré</surname>
          </string-name>
          ,
          <article-title>Flashattention: Fast and memory-eficient exact attention with io-awareness</article-title>
          ,
          <source>arXiv preprint arXiv:2205.14135</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Khabsa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Fang</surname>
          </string-name>
          , H. Ma, Linformer:
          <article-title>Self-attention with linear complexity</article-title>
          , arXiv preprint arXiv:
          <year>2006</year>
          .
          <volume>04768</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>N.</given-names>
            <surname>Kitaev</surname>
          </string-name>
          , Ł. Kaiser,
          <string-name>
            <given-names>A.</given-names>
            <surname>Levskaya</surname>
          </string-name>
          ,
          <article-title>Reformer: The eficient transformer</article-title>
          , arXiv preprint arXiv:
          <year>2001</year>
          .
          <volume>04451</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>I.</given-names>
            <surname>Beltagy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. E.</given-names>
            <surname>Peters</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Cohan</surname>
          </string-name>
          ,
          <article-title>Longformer: The long-document transformer</article-title>
          , arXiv preprint arXiv:
          <year>2004</year>
          .
          <volume>05150</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>B.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Dao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Winsor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rudra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ré</surname>
          </string-name>
          , Scatterbrain:
          <article-title>Unifying sparse and low-rank attention</article-title>
          ,
          <source>Advances in Neural Information Processing Systems</source>
          <volume>34</volume>
          (
          <year>2021</year>
          )
          <fpage>17413</fpage>
          -
          <lpage>17426</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Zaheer</surname>
          </string-name>
          , G. Guruganesh,
          <string-name>
            <given-names>K. A.</given-names>
            <surname>Dubey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ainslie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Alberti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ontanon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Pham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ravula</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Yang</surname>
          </string-name>
          , et al.,
          <article-title>Big bird: Transformers for longer sequences</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>33</volume>
          (
          <year>2020</year>
          )
          <fpage>17283</fpage>
          -
          <lpage>17297</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>W.</given-names>
            <surname>Fedus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zoph</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          , Switch transformers:
          <article-title>Scaling to trillion parameter models with simple and eficient sparsity</article-title>
          , ????
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Jaegle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Gimeno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Brock</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Vinyals</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zisserman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carreira</surname>
          </string-name>
          , Perceiver:
          <article-title>General perception with iterative attention</article-title>
          ,
          <source>in: International conference on machine learning, PMLR</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>4651</fpage>
          -
          <lpage>4664</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Jaegle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Borgeaud</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.-B. Alayrac</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Doersch</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Ionescu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Ding</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Koppula</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Zoran</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Brock</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Shelhamer</surname>
          </string-name>
          , et al.,
          <article-title>Perceiver io: A general architecture for structured inputs &amp; outputs</article-title>
          , arXiv preprint arXiv:
          <volume>2107</volume>
          .14795 (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>C.</given-names>
            <surname>Hawthorne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jaegle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Cangea</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Borgeaud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Nash</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Malinowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dieleman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Vinyals</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Botvinick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Simon</surname>
          </string-name>
          , et al.,
          <article-title>General-purpose, longcontext autoregressive modeling with perceiver ar</article-title>
          ,
          <source>arXiv preprint arXiv:2202.07765</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S.</given-names>
            <surname>Borgeaud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mensch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hofmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Cai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Rutherford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Millican</surname>
          </string-name>
          ,
          <string-name>
            <surname>G. B. Van Den Driessche</surname>
          </string-name>
          , J.
          <string-name>
            <surname>-B. Lespiau</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Damoc</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Clark</surname>
          </string-name>
          , et al.,
          <article-title>Improving language models by retrieving from trillions of tokens</article-title>
          , in: International conference on machine learning,
          <source>PMLR</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>2206</fpage>
          -
          <lpage>2240</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>G.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. J.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Babuschkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sidor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Farhi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ryder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pachocki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <article-title>Tensor programs v: Tuning large neural networks via zero-shot hyperparameter transfer</article-title>
          ,
          <source>arXiv preprint arXiv:2203.03466</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>