<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Epistemic Defenses against Scientific and Empirical Adversarial AI Attacks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nadisha-Marie Aliman</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Leon Kester</string-name>
          <email>leon.kester@tno.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>TNO Netherlands</institution>
          ,
          <addr-line>The Hague</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Utrecht University</institution>
          ,
          <addr-line>Utrecht</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we introduce “scientific and empirical adversarial AI attacks” (SEA AI attacks) as umbrella term for not yet prevalent but technically feasible deliberate malicious acts of specifically crafting AI-generated samples to achieve an epistemic distortion in (applied) science or engineering contexts. In view of possible socio-psychotechnological impacts, it seems responsible to ponder countermeasures from the onset on and not in hindsight. In this vein, we consider two illustrative use cases: the example of AI-produced data to mislead security engineering practices and the conceivable prospect of AI-generated contents to manipulate scientific writing processes. Firstly, we contextualize the epistemic challenges that such future SEA AI attacks could pose to society in the light of broader i.a. AI safety, AI ethics and cybersecurityrelevant efforts. Secondly, we set forth a corresponding supportive generic epistemic defense approach. Thirdly, we effect a threat modelling for the two use cases and propose tailor-made defenses based on the foregoing generic deliberations. Strikingly, our transdisciplinary analysis suggests that employing distinct explanation-anchored, trustdisentangled and adversarial strategies is one possible principled complementary epistemic defense against SEA AI attacks - albeit with caveats yielding incentives for future work.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Progress in the AI field unfolds a wide growing array of
beneficial societal effects with AI permeating more and more
crucial application domains. To forestall ethically-relevant
ramifications, research from a variety of disciplines tackling
pertinent AI safety [Amodei et al., 2016; Bostrom, 2017; Burden
and Herna´ndez-Orallo, 2020; Fickinger et al., 2020; Leike
et al., 2017], AI ethics and AI governance issues [Floridi
et al., 2018; Jobin et al., 2019; O´ hE´ igeartaigh et al., 2020;
Raji et al., 2020] gained momentum at an international
level. In addition, cybersecurity-oriented frameworks in AI
safety [Aliman et al., 2021; Brundage et al., 2018; Pistono
and Yampolskiy, 2016] stressed the necessity to not only
address unintentional errors, unforeseen repercussions and bugs
in the context of ethical AI design but also AI risks linked to
intentional malice i.e. deliberate unethical design, attacks and
sabotage by malicious actors. In parallel, the convergence of
AI with other technologies increases and diversifies the attack
surface available to malevolent actors. For instance, while
AI-enhanced cybersecurity opens up novel valuable
possibilities for defenders [Zeadally et al., 2020], AI
simultaneously provides new affordances for attackers [Ashkenazy and
Zini, 2019] from AI-aided social engineering [Seymour and
Tully, 2016] to AI-concealed malware [Kirat et al., 2018].
Next to the capacity of AI to extend classical cyberattacks
in scope, speed and scale [Kaloudi and Li, 2020], a
notable emerging threat is what we denote AI-aided epistemic
distortion. The latter represents a form of AI
weaponization and is increasingly studied in its currently most salient
form, namely AI-aided disinformation [Aliman et al., 2021;
Chesney and Citron, 2019; Kaloudi and Li, 2020; Tully and
Foster, 2020] which is especially relevant to information
warfare [Hartmann and Giles, 2020]. Recently, the
weaponization of Generative AI for information operations has been
described as “a sincere threat to democracies” [Hartmann and
Steup, 2020]. In this paper, we analyze attacks and defenses
pertaining to another not yet prevalent but technically feasible
and similarly concerning form of AI-aided epistemic
distortion with potentially profound societal implications: scientific
and empirical adversarial AI attacks (SEA AI attacks).</p>
      <p>
        With SEA AI attacks, we refer to any deliberately
malicious AI-aided epistemic distortion which predominantly and
directly targets (applied) science and technology assets (as
opposed to information operations where a wider societal
target is often selected on ideological/political grounds). In
short, the expression acts as an umbrella term for malicious
actors utilizing or attacking AI at pre- or post-deployment
stages with the deliberate adversarial aim to deceive,
sabotage, slow down or disrupt (applied) science, engineering or
related endeavors. Obviously, SEA AI attacks could be
performed in a variety of modalities
        <xref ref-type="bibr" rid="ref2 ref44 ref51 ref54 ref59 ref61">(see e.g. “deepfake
geography” [Zhao et al., 2021] related to vision)</xref>
        . However, for
illustrative purposes, we base our two exemplary use cases
on misuses of language models. The first use case treats SEA
AI attacks on security engineering via schemes in which a
malicious actor poisons training data resources [Mahlangu et
al., 2019] that are vital to data-driven defenses in the
cybersecurity ecosystem. Lately, a proof-of-concept for an AI-based
data poisoning attack has been implemented in the context
of cyber threat intelligence (CTI) [Ranade et al., 2021]. The
authors utilized a fine-tuned version of the GPT-2 language
model [Radford et al., 2019] and were able to generate fake
CTI which was indistinguishable from its legitimate
counterpart when presented to cybersecurity experts. The
second use case studies conceivable SEA AI attacks on
procedures that are essential to scientific writing. Related
examples that have been depicted in recent work encompass
plagiarism studies with transformers like BERT [Wahle et al., 2021]
and with the pre-trained GPT-3 language model [Brown et
al., 2020] that “may very well pass peer review” [Dehouche,
2021] but also AI-generated fake reviews (with a fine-tuned
version of GPT-2) apt to mislead experienced researchers in
a small user study [Tallo´n-Ballesteros, 2020]. Future
malicious actors could deliberately breed a large-scale agenda in
the spirit of “fake science news” [Ho et al., 2020] and
AIgenerated papers that would widely exceed in quality (later
withdrawn) computer-generated research papers [Van
Noorden, 2014] published at respected venues. In short,
technically already practicable SEA AI attacks could have
considerable negative effects if jointly potentiated with regard to
scale, scope and speed by malicious actors equipped with
sufficient resources. As later exemplified in Subsection 3.1,
the security engineering use case could e.g. involve dynamic
domino-effects leading to large financial losses and even risks
to human lives while the scientific writing use case seems to
moreover reveal a domain-general epistemic problem. The
mere existence of the latter also affects the former and could
engender serious pitfalls whose generically formulated
principled management is compactly treated in the next Section 2.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Theoretical Generic Epistemic Defenses</title>
      <p>As reflected in the law of requisite variety (LRV) known
from cybernetics, “only variety can destroy variety” [Ashby,
1961]. Applied to SEA AI attacks, it signifies that since
malicious adversaries are not only exploiting vulnerabilities
from a heterogeneous socio-psycho-technological landscape
but also specially vulnerabilities of epistemic nature, suitable
defense methods may profit from an epistemic stance.
Applying the cybernetic LRV offers a valuable domain-general
transdisciplinary tool able to stimulate and invigorate novel
tailored defenses in a diversity of harm-related problems
from cybersecurity [Vinnakota, 2013] to AI safety [Aliman,
2020a] over AI ethics [Ashby, 2020]. In short, utilizing
insights from epistemology as complementary basis to frame
defense methods against SEA AI attacks seems
indispensable. Past work predominantly analyzed countermeasures of
socio-psycho-technological nature to combat the spread of
(audio-)visual, audio and textual deepfakes as well as “fake
news” more broadly. For instance, the technical detection of
AI-generated content [Wahle et al., 2021] has been often
thematized and even lately applied to “fake news” in the
healthcare domain [Baris and Boukhers, 2021]. Furthermore, in the
context of counteracting risks posed by the deployment of
sophisticated online bots, it has been suggested that
“technical solutions, while important, should be complemented with
efforts involving informed policy and international norms to
accompany these technological developments” and that “it
is essential to foster increased civic literacy of the nature
of ones interactions” [Boneh et al., 2019]. Another
analysis presented a set of defense measures against the spread of
deepfakes [Chesney and Citron, 2019] which contained i.a.
legal solutions, administrative agency solutions, coercive and
covert responses as well as sanctions (when effectuated by
state actors) and speech policies for online platforms.
Concerning “fake science news” and their impacts on “credibility
and reputation of the science community” [Ho et al., 2020],
it has been even postulated by Makri that “science is losing
its relevance as a source of truth” and “the new focus on
post-truth shows there is now a tangible danger that must be
addressed” [Makri, 2017]. Following the author, scientists
could equip citizens with sense-making tools without which
“emotions and beliefs that pander to false certainties become
more credible” [Makri, 2017].</p>
      <p>
        While some of those socio-psycho-technological
countermeasures and underlying assumptions are debatable, we
complementarily zoom in different epistemic defenses against
SEA AI attacks being directed against scientific and empirical
frameworks. Amidst an information ecosystem with
quasiomnipresent terms such as “post-truth” or “fake news” and
in light of data-driven research trends embedded within
trustbased infrastructures, it seems daunting to face a threat
landscape populated by AI-generated artefacts such as: 1) “fake
data” and “fake experiments”, 2) “fake research papers”
        <xref ref-type="bibr" rid="ref25 ref32 ref43 ref50 ref60">(or
“fraudulent academic essay writing” [Brown et al., 2020])</xref>
        and 3) “fake reviews”. More broadly, it has been stated that
deepfakes “seem to undermine our confidence in the original,
genuine, authentic nature of what we see and hear” [Floridi,
2018]. Taking the perspective of an empiricism-based
epistemology grounded in justification with the aim to obtain truer
beliefs via (probabilistic) belief updates given evidence, a
recent in-depth analysis found that the existence of deepfake
videos confronts society with epistemic threats [Fallis, 2020].
Thereby, it is assumed that “deepfakes reduce the amount
of information that videos carry to viewers” [Fallis, 2020]
which analogously quantitatively affected the amount of
information in text-based news due to earlier “fake news”
phenomena. In our view, when applying this stance to
audiovisual and textual samples of scientific material but also broadly
to the context of security engineering and scientific
communication where the deployment of deepfakes for SEA AI attacks
could occur in multifarious ways, the consequences seem
disastrous. In brief, SEA AI defenses seem relevant to AI safety
since an inability to build up resiliency against those attacks
may suggest that already present-day AI could (be used to)
outmaneuver humans on a large scale – without any
“superintelligent” competency. However, empiricist
epistemology is not without any alternative. In the following, we thus
first mentally enact one alternative epistemic stance (without
claiming that it represents the only possible alternative). We
present its key generic epistemic suppositions serving as a
basis for the next Section 3 where we tailor defenses against
SEA AI attacks for the specific use cases.
      </p>
      <p>
        Firstly, it has been lately propounded that the societal
perception of a “post-truth” era is often linked to the implicit
assumption that truth can be equated with consensus which
is why it seems recommendable to consider a deflationary
account of truth [Bufacchi, 2021] – i.e. where the concept
is for instance strictly reserved to scientifically-relevant
epistemic contexts. On such a deflationary account of truth
disentangled from consensus, it has been argued that even if
consensus and trust seem eroded, we neither inhabit a post-truth
nor a science-threatening post-falsification age [Aliman and
Kester, 2020]. Secondly, we never had a direct access to
physical reality which we could have suddenly lost with the advent
of “fake news”. In fact, as stated by Karl Popper: “Once we
realize that human knowledge is fallible, we realize also that
we can never be completely certain that we have not made a
mistake” [Popper, 1996]. Thirdly, the epistemic aim in
science can neither be truth directly [Frederick, 2020] nor can it
be truer beliefs via justifications. The former is not directly
experienced and the latter has been shown to be logically
invalid by Popper [Popper, 2014]. Science is
quintessentially explanatory i.e. it is based on explanations [Deutsch,
2011] and not merely on data. While the epistemic aim
cannot be certainty or justification
        <xref ref-type="bibr" rid="ref1 ref11 ref14 ref17 ref28 ref29 ref30 ref35 ref56 ref8">(and not even “truer
explanations” [Frederick, 2020]1 for lack of direct access to truth)</xref>
        ,
a pragmatic way to view it is that our epistemic aim can be
to achieve better explanations [Frederick, 2020]. One can
collectively agree on practical updatable criteria which better
explanations should fulfill. In short, one does not assess a
scientific theory in isolation, but in comparison to rival theories
and one is thereby embedded in a context with other
scientists. Fourthly, there are distinct ways to handle falsification
and integrate empirical findings in explanation-anchored
science. One can e.g. criticize an explanation and pinpoint
inconsistencies at a theoretical level. One can attempt to make
a theory problematic via falsifying experiments whose results
are accepted to seem to conflict with the predictions that the
theory entailed [Deutsch, 2016]. Vitally, in the absence of a
better rival theory, it holds that “an explanatory theory
cannot be refuted by experiment: at most it can be made
problematic” [Deutsch, 2016].
      </p>
      <p>Against the background of this epistemic bedrock, one can
now re-assess the threat landscape of SEA AI attacks. Firstly,
one can conclude that AI-generated “fake data” and “fake
experiments” could slow down but not terminally disrupt
scientific and empirical procedures. In the case of misguiding
confirmatory data, it has no epistemic effect since as opposed to
empiricist epistemology, explanation-anchored science does
not utilize any scheme of credence updates for a theory and
it is clear that “a severely tested but unfalsified theory may
be false” [Frederick, 2020]. In the case of misleading data
that is accepted to falsify a theory T , one runs the risk to
con1That our epistemic aim can be “truer explanations” or
explanations that lead us “closer to the truth” has been sometimes
confusingly written by Deutsch and Popper respectively but this type of
account requires a semantic refinement [Frederick, 2020].
sider mistakenly that T has been made problematic.
However, since it is not permissible to drop T in the absence of a
rival theory T 0 representing a better explanation than T , the
adverarial capabilities of the SEA AI attacker are limited. In
short, theories cannot be deleted from the collective
knowledge via such SEA AI attacks without more ado. Secondly,
when contemplating the case of AI-generated “fake research
papers”, it seems that they could slow down but not disrupt
scientific methodology. Overall, one could state that the
danger lies in the uptake of deceptive theories. However,
theories are only integrated in explanatory-anchored science if
they represent better explanations in comparison to
alternatives or in the absence of alternatives if they explain novel
phenomena. In a nutshell, it takes explanations that are
simultaneously misguiding and better for such a SEA AI attack
to succeed. This is a high bar for imitative language
models if meant to be repeatedly and systematically performed2
and not merely as a unique event by chance. Further, even
in the case a deceptive theory has been integrated in a field,
that is always only provisionally such that it could be revoked
at any suitable moment e.g. once a better explanation arises
and repeated experiments falsify its claims. If in the course of
this, an actually better explanation had been mistakenly
considered as refuted, it can always be re-integrated once this is
noticed. In fact, “a falsified theory may be true” [Frederick,
2020] if the accepted observations believed to have falsified it
were wrong. Thirdly, when now considering the final case of
AI-generated “fake reviews”, it becomes clear that they could
similarly slow down but not terminally disrupt the scientific
method. At worst some existing theories could be
unnecessarily problematized and misguiding theories uptaken, but all
these epistemic procedures can be repealed retrospectively.</p>
      <p>In short, explanation-anchored science is resilient (albeit
not immune) against SEA AI attacks but one can humbly face
the idea that it is not because scientists can “tease out
falsehood from truths” [Ho et al., 2020], but because
explanationanchored science attempts to tease out better from worse
explanations while permanently requiring the creation of new
ones whereby the steps made can always be revoked, revised
and even actively adversarially counteracted. That entails a
sort of epistemic dizziness and one can never trust one’s own
observations. Also, human mental constructions are
inseparably cognitive-affective and science is not detached from
social reality [Barrett, 2017]. In our view, for a systematic
management of this epistemic dizziness, one may profit from
an adversarial approach that permanently brings to mind that
one might be wrong. Last but not least, an important feature
discussed is that the epistemic aim not being truth (which
itself is also not consensus and does not rely on trust to
exist) but instead better explanations, none of the mentioned
2That there could exist a task which imitative language
models are “theoretically incapable of handling” has been often put
into question [Sahlgren and Carlsson, 2021]. However, on
epistemic grounds elaborated in-depth previously [Aliman, 2020a;
Aliman et al., 2021] which might be amenable to experimental
falsifiability [Aliman, 2020b], we assume that the task to consciously
create and understand novel yet unknown explanatory
knowledge [Deutsch, 2011] – which humans are capable of performing
if willing to – cannot be learned by AI systems by mere imitation.
methods are dependent on trust per se – making it a
trustdisentangled view. To sum up, we identified 3 key generic
features for epistemic defenses against SEA AI attacks:</p>
      <sec id="sec-2-1">
        <title>1. Explanation-anchored instead of data-driven</title>
      </sec>
      <sec id="sec-2-2">
        <title>2. Trust-disentangled instead of trust-dependent</title>
      </sec>
      <sec id="sec-2-3">
        <title>3. Adversarial instead of (self-)compliant</title>
        <p>3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Practical Use of Theoretical Defenses</title>
      <p>In the following Subsection 3.1, we briefly perform an
exemplary threat modelling for the two specific use cases
introduced in Section 1. The threat model narratives are naturally
non-exhaustive and are selected for illustrative purposes to
display plausible downward counterfactuals projecting
capabilities to the recent counterfactual past in the spirit of
cocreation design fictions in AI safety [Aliman et al., 2021]. In
Subsection 3.2, we then derive corresponding tailor-made
defenses from the generic characteristics that have been carved
out in the last Section 2 while thematizing notable caveats.
3.1</p>
      <sec id="sec-3-1">
        <title>Threat Modelling for Use Cases</title>
        <p>Use Case Security Engineering
• Adversarial goals: As briefly mentioned in
Section 1, CTI (which is information related to
cybersecurity threats and threat actors to support analysts and
security systems in the detection and mitigation of
cyberattacks) can be polluted via misleading AI-generated
samples to fool cyber defense systems at the training
stage [Ranade et al., 2021]. Among others, CTI is
available as unstructured texts but also as knowledge
graphs taking CTI texts as input. A textual data
poisoning via AI-produced “fake CTI” represents a form
of SEA AI attack that was able to succesfully deceive
(AI-enhanced) automated cyber defense and even
cybersecurity experts which “labeled the majority of the fake
CTI samples as true despite their expertise” [Ranade et
al., 2021]. It is easily conceivable that malicious
actors could specifically tailor such SEA AI attacks in
order to subvert cyber defense in the service of subsequent
covert time-efficient, micro-targeted and large-scale
cybercrime. For 2021, cybercrime damages are estimated
to reach 6 trillion USD [Benz and Chatterjee, 2020;
Ozkan et al., 2021] making cybercrime a top
international risk with a growing set of affordances which
malicious actors do not hesitate to enact. Actors interested
in “fake CTI” attacks could be financially motivated
cybercriminals or state-related actors. Adversarial goals
could e.g. be to acquire private data, CTI poisoning in
a cybercrime-as-a-service form, gain strategical
advantages in cyber operations, conduct espionage or even
attack critical infrastructure endangering human lives.
• Adversarial knowledge: Since it is the attacker that
finetunes the language model generating the “fake CTI”
samples for the SEA AI attack, we consider a white
box setting for this system. The attacker does not
require knowledge about the internal details of the
targeted automated cyber defense allowing a black-box
setting with regard to this system at training time. In case
the attacker directly targets human security analysts by
exposing them to misleading CTI, the SEA AI attack
can be interpreted as a type of adversarial example on
human cognition in a black-box setting. However, in
such cases “open-source intelligence gathering and
social engineering are exemplary tools that the adversary
can employ to widen its knowledge of beliefs,
preferences and personal traits exhibited by the victim”
[Aliman et al., 2021]. Hence, depending on the required
sophistication, a type of grey-box setting is achievable.
• Adversarial capabilities: The use of SEA AI attacks
could have been useful at multiple stages. CTI text could
have been altered in a micro-targeted way offering
diverse capacities to a malicious actor: to distract
analysts from patching existing vulnerabilities, to gain time
for the exploitation of zero-days, to let systems
misclassify malign files as benign [Mahlangu et al., 2019] or
to covertly take over victim networks. In the light of
complex interdependencies, the malicious actor might
not even have had a full overview of all repercussions
that AI-generated “fake CTI” attacks can engender.
Poisoned knowledge graphs could have led to unforeseen
domino-effects inducing unknown second-order harm.
As long-term strategy, the malicious actor could have
harnessed SEA AI attacks on applied science writing to
automate the generation of cybersecurity reports (for it
to later serve as CTI inputs) corroborating the robustness
of actually unsafe defenses to covertly subvert those or
simply to spread confusion.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Use Case Scientific Writing</title>
        <p>•</p>
        <p>Adversarial goals: The emerging issue of (AI-aided)
information operations in social media contexts which
involves entities related to state actors has gained
momentum in the last years [Prier, 2017; Hartmann and
Giles, 2020]. A key objective of information operations
that has been repeatedly mentioned is the intention to
blur what is often termed as the line between facts and
fictions [Jakubowski, 2019]. Naturally, when logically
applying the epistemic stance introduced in the last
Section 2, it seems recommendable to avoid such
formulations for clarity since potentially confusing. Hence,
we refer to it simply as epistemic distortion. SEA AI
attacks on scientific writing being a form of AI-aided
epistemic distortion, it could represent a lucrative
opportunity for state actors or politically motivated
cybercriminals willing to ratchet up information operations. On a
smaller scale, other potential malicious goals could also
involve companies with a certain agenda for a product
that could be threatened by scientific research. Another
option could be advertisers that monetize attention via
AI-generated research papers in click-bait schemes.
• Adversarial knowledge: As in the first use case, the
language model is available in a white-box setting.
Moreover, since this SEA AI attack directly targets human
entities, one can again assume a black-box or grey-box
scenario depending on the required sophistication of the
attack. For instance, since many scientists utilize social
media platforms, open source intelligence gathering on
related sources can be utilized to tailor contents.
• Adversarial capabilities: In the domain of
adversarial machine learning, it has been stressed that for
security reasons it is important to also consider adaptive
attacks [Carlini et al., 2019], namely reactive attacks
that adapt to what the defense did. A malicious
actor aware of the discussed explanation-anchored,
trustdisentangled and adversarial epistemic defense approach
could have exploited a wide SEA AI attack surface in
case of no consensus on the utility of this defense. For
instance, a polarization between two dichotomously
opposed camps in that regard could have offered an ideal
breeding ground for divisive information warfare
endeavors. For some, the perception of increasing
disagreement tendencies may have confirmed post-truth
narratives. Not for malicious reasons, but because it
was genuinely considered. This in turn could have
cemented echo chamber effects now fuelled by a divided
set of scientists one part of which considered science
to be epistemically defeated. This combined with
posttruth narratives and the societal-level automated
disconcertion [Aliman et al., 2021] via the mere existence of
AI-generated fakery could have destabilized a fragile
society and incited violence. Massive and rapid
largescale SEA AI attacks in the form of a novel type of
scientific astroturfing could have been employed to
automatically reinforce the widespread impression of
permanently conflicting research results on-demand and
tailored to a scientific topic. The concealed or ambiguous
AI-generated samples (be it data, experiments, papers or
reviews) would not even need to be overrepresented in
respected venues but only made salient via social media
platforms being one of the main information sources for
researchers – a task which could have been automated
via social bots influencing trending and sharing patterns.
A hinted variant of such SEA AI attacks could have been
a flood of confirmatory AI-generated texts that
corroborate the robustness of defenses across a large array of
security areas in order to exploit any reduced
vulnerability awareness. Finally, hyperlinks with attention-driving
fake research contribution titles competing with science
journalism and redirecting to advertisement pages could
have polluted results displayed by search engines.
3.2</p>
      </sec>
      <sec id="sec-3-3">
        <title>Practical Defenses and Caveats</title>
        <p>As is also the case with other advanced not yet prevalent
but technically already feasible AI-aided information
operations [Hartmann and Giles, 2020] and cyberattacks targeting
AIs [Hartmann and Steup, 2020], consequences could have
ranged from severe financial losses to threats to human lives.
Multiple socio-psycho-technological solutions including the
ones reviewed in Section 1 which may be (partially) relevant
to SEA AI attack scenarios have been previously presented.
Here, we complementarily focus on the epistemic dimensions
one can add to the pool of potential solutions by applying the
3 generic features extracted in Section 2 to both use cases. We
also emphasize novel caveats. Concerning the first use case
of “fake CTI” SEA AI attacks, the straightforward thought to
restrict the use of data from open platforms is not conducive
to practicability not only due to the amount of crucial
information that a defense might miss, but also because it does
not protect from insider threats [Ranade et al., 2021].
However, common solutions such as the AI-based detection of
AI-generated outputs or trust-reliant scoring systems to flag
trusted sources do not seem sufficient either without more
ado since the former may fail in the near future if the
generator tends to win and the latter is at risk due to impersonation
possibilities that AI itself augments and due to the mentioned
insider threats. Interestingly, the issue of malicious insider
threats is also reflected in the second use case with scientific
writing being open to arbitrary participants.</p>
        <p>
          Defense for Security Engineering Use Case and Caveats
1. Explanation-anchored instead of data-driven: An
explanation-anchored solution can be formulated from
the inside out. Although AI does not understand
explanations, it is thinkable that a technically feasible future
hybrid active intelligent system3 for automated cyber
defense could use knowledge graph inconsistencies
[Heyvaert et al., 2019] as signals to calculate when it will
epistemically seek clarification from a human analyst,
when to actively query differing sources and sensors or
when to follow habitual courses of action. But the
creativity of human malicious actors cannot be predicted
and thus neither the system nor human analysts are able
to prophesy over a space of not yet created attacks. Also,
as long as the system’s sensors are learning-based AI, it
stays an Achilles heel due to the vulnerability to attacks.
2. Trust-disentangled instead of trust-dependent: Such a
procedure could seem disadvantageous given the fast
reactions required in cyber defense. However, an
adversarial explanation-anchored framework is orthogonal to
the trust policy used. Trust-disentangled does not
necessarily signify zero-trust4 at all levels if impracticable.
3. Adversarial instead of (self-)compliant: A permanently
rotating in-house adversarial team is required.
Activities can include red teaming, penetration testing and the
development of (adaptive) attacks i.a. with AI-generated
“fake CTI” text samples. A staggered approach is
cogitable in which automated defense processes that
happen at fast scales (e.g. requiring rapid access to open
source CTI) rely on interim (distributed) trust while all
others – especially those involving human deliberation
to create novel defenses and attacks – strive for
zerotrust information sharing (e.g. via a closed blockchain
with a restricted set of authorized participants having
read and write rights). In this way, one can create an
interconnected 3-layered epistemically motivated
security framework: a slow creative human-run adversarial
3Such a system could instantiate technical self-awareness
[Aliman, 2020a]
          <xref ref-type="bibr" rid="ref2 ref44 ref51 ref54 ref59 ref61">(e.g. via active inference [Smith et al., 2021])</xref>
          .
        </p>
        <p>4The zero-trust [Kindervag, 2010] paradigm advanced in
cybersecurity in the last decade which assumes “that adversaries are
already inside the system, and therefore imposes strict access and
authentication requirements” [Collier and Sarkis, 2021] seems highly
appropriate in this increasingly complex security landscape.
counterfactual layer on top of a slow creative
humanrun defensive layer steering a very fast
hybrid-active-AIaided automated cyber defense layer. Important caveats
are that such a framework: 1) can be resilient but not
immune, 2) can not and should not be entirely automated.
Defense for Science Writing Use Case and Caveats
1. Explanation-anchored instead of data-driven: A
practical challenge for SEA AI attacks may seem the need
for scientists to agree on pragmatic criteria for
“better” explanations (but widely accepted cases are e.g. the
preference for “simpler”, “more innovative” and “more
interesting” ones). Also, due to automated
disconcertion, reviewers could always suspect that a paper was
AI-generated (potentially at the detriment of human
linguistic statistical outliers). However, this is not a
sufficient argument since explanation-anchored science and
criticism focus on content and not on source or style.</p>
        <sec id="sec-3-3-1">
          <title>2. Trust-disentangled instead of trust-dependent: Via</title>
          <p>trust-disentanglement, a paper generated by a
presentday AI would not only be rejected on provenance
grounds but due to its merely imitative and
nonexplanatory content. Though, an important asset is
the review process which if infiltrated by imitative
AI-generated content could slow down
explanationanchored criticism if not thwarted fastly. A zero-trust
scheme could mitigate this risk time-efficiently (e.g. via
a consortium blockchain for review activities). Another
zero-trust method would be to taxonomically monitor
SEA AI attack events at an international level e.g. via
an AI incident base [McGregor, 2020] tailored to these
attacks and complemented by adversarial retrospective
counterfactual risk analyses [Aliman et al., 2021] and
defensive solutions. The monitoring can be AI-aided (or
in the future hybrid-active-AI-aided) but human analysts
are indispensable for a deep semantic understanding
[Aliman et al., 2021]. In short, also here, we suggest an
interconnected 3-layered epistemic framework with
adversarial, defensive and hybrid-active-AI-aided elements.
3. Adversarial instead of (self-)compliant: As advanced
adversarial strategy which would also require
responsible coordinated vulnerability disclosures [Kranenbarg
et al., 2018], one could perform red teaming,
penetration tests and (adaptive) attacks employing AI-generated
“fake data and experiments”, “fake papers” and “fake
reviews” [Tallo´n-Ballesteros, 2020]. Candidates for a blue
team are e.g. reviewers and editors. Concurrently, urgent
AI-related plagiarism issues arise [Dehouche, 2021].
4</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion and Future Work</title>
      <p>For requisite variety, we introduced a complementary generic
epistemic defense against not yet prevalent but technically
feasible SEA AI attacks. This generic approach
foregrounded explanation-anchored, trust-disentangled and
adversarial features that we instantiated within two illustrative
use cases involving language models: AI-generated samples
to fool security engineering practices and AI-crafted contents
to distort scientific writing. For both use cases, we compactly
worked out a transdisciplinary and pragmatic 3-layered
epistemically motivated security framework composed of
adversarial, defensive and hybrid-active-AI-aided elements with
two major caveats: 1) it can be resilient but not immune, 2) it
can not and should not be entirely automated. In both cases, a
proactive exposure to synthetic AI-generated material could
foster critical thinking. Vitally, the existence of truth stays
a legitimate raison d’eˆtre for science. It is only that in
effect, one is not equipped with a direct acces to truth, all
observations are theory-laden and what one think one knows is
linked to what is co-created in one’s collective enactment of
a world with other entities shaping and shaped by physical
reality. Thereby, one can craft explanations to try to improve
one’s active grip on a field of affordances but it stays an
eternal mental tightrope walking of creativity. In view of this
inescapable epistemic dizziness, the main task of
explanationanchored science is then neither to draw a line between truth
and falsity nor between the trusted and the untrusted. Instead,
it is to seek to robustly but provisionally separate better from
worse explanations. While this steadily renewed societally
relevant act does not yield immunity against AI-aided
epistemic distortion, it enables resiliency against at-present
thinkable SEA AI attacks. To sum up, the epistemic dizziness of
conjecturing that one could always be wrong could stimulate
intellectual humility, but also unbound(ed) (adversarial)
explanatory knowledge co-creation. Future work could study
how language AI – which could be exploited for future SEA
AI attacks e.g. instrumental in performing cyber(crime) and
information operations – could conversely serve as
transformative tool to augment anthropic creativity and tackle the
SEA AI threat itself. For instance, language AI could be used
to stimulate human creativity in future AI and security design
fictions for new threat models and defenses. In retrospective,
AI is already acting as a catalyst since the very defenses
humanity now crafts can broaden, deepen and refine the scope of
explanations i.a. also about better explanations – an
unceasing but also potentially strengthening safety relevant quest.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>[Aliman and Kester</source>
          , 2020]
          <article-title>Nadisha-Marie Aliman</article-title>
          and
          <string-name>
            <given-names>Leon</given-names>
            <surname>Kester</surname>
          </string-name>
          .
          <source>Facing Immersive “Post-Truth” in AIVR? Philosophies</source>
          ,
          <volume>5</volume>
          (
          <issue>4</issue>
          ):
          <fpage>45</fpage>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [Aliman et al.,
          <year>2021</year>
          ]
          <string-name>
            <surname>Nadisha-Marie</surname>
            <given-names>Aliman</given-names>
          </string-name>
          , Leon Kester, and
          <string-name>
            <given-names>Roman</given-names>
            <surname>Yampolskiy. Transdisciplinary AI Observatory-Retrospective Analyses</surname>
          </string-name>
          and
          <string-name>
            <surname>Future-Oriented Contradistinctions</surname>
          </string-name>
          . Philosophies,
          <volume>6</volume>
          (
          <issue>1</issue>
          ):
          <fpage>6</fpage>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Aliman, 2020a]
          <string-name>
            <surname>Nadisha-Marie Aliman</surname>
          </string-name>
          .
          <article-title>Hybrid CognitiveAffective Strategies for AI Safety</article-title>
          .
          <source>PhD thesis</source>
          , Utrecht University,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [Aliman, 2020b]
          <string-name>
            <surname>Nadisha-Marie Aliman</surname>
          </string-name>
          .
          <article-title>Self-Shielding Worlds</article-title>
          . https://nadishamarie.jimdo.com/clipboard/,
          <year>2020</year>
          . Online; accessed 23-November-
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [Amodei et al.,
          <year>2016</year>
          ]
          <string-name>
            <given-names>Dario</given-names>
            <surname>Amodei</surname>
          </string-name>
          , Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mane´.
          <article-title>Concrete problems in AI safety</article-title>
          .
          <source>arXiv preprint arXiv:1606.06565</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <source>[Ashby</source>
          , 1961]
          <string-name>
            <given-names>W Ross</given-names>
            <surname>Ashby</surname>
          </string-name>
          .
          <article-title>An introduction to cybernetics</article-title>
          . Chapman &amp; Hall
          <string-name>
            <surname>Ltd</surname>
          </string-name>
          ,
          <year>1961</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <source>[Ashby</source>
          , 2020]
          <string-name>
            <given-names>Mick</given-names>
            <surname>Ashby</surname>
          </string-name>
          .
          <article-title>Ethical regulators and superethical systems</article-title>
          .
          <source>Systems</source>
          ,
          <volume>8</volume>
          (
          <issue>4</issue>
          ):
          <fpage>53</fpage>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <source>[Ashkenazy and Zini</source>
          , 2019]
          <string-name>
            <given-names>Adi</given-names>
            <surname>Ashkenazy</surname>
          </string-name>
          and
          <string-name>
            <given-names>Shahar</given-names>
            <surname>Zini</surname>
          </string-name>
          .
          <source>Attacking Machine Learning - The Cylance Case Study</source>
          . https://skylightcyber.com/
          <year>2019</year>
          / 07/18/cylance-i
          <string-name>
            <surname>-</surname>
          </string-name>
          kill-you/Cylance%
          <fpage>20</fpage>
          -%
          <source>20Adversarial% 20Machine%20Learning%20Case%20Study.pdf</source>
          ,
          <year>2019</year>
          . Skylight; accessed 24-May-
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>[Baris and Boukhers</source>
          , 2021]
          <string-name>
            <given-names>Ipek</given-names>
            <surname>Baris</surname>
          </string-name>
          and
          <string-name>
            <given-names>Zeyd</given-names>
            <surname>Boukhers</surname>
          </string-name>
          . ECOL:
          <article-title>Early Detection of COVID Lies Using Content, Prior Knowledge and Source Information</article-title>
          .
          <source>arXiv preprint arXiv:2101.05499</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <source>[Barrett</source>
          , 2017]
          <article-title>Lisa Feldman Barrett</article-title>
          .
          <article-title>Functionalism cannot save the classical view of emotion</article-title>
          .
          <source>Social Cognitive and Affective Neuroscience</source>
          ,
          <volume>12</volume>
          (
          <issue>1</issue>
          ):
          <fpage>34</fpage>
          -
          <lpage>36</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <source>[Benz and Chatterjee</source>
          , 2020]
          <string-name>
            <given-names>Michael</given-names>
            <surname>Benz</surname>
          </string-name>
          and
          <string-name>
            <given-names>Dave</given-names>
            <surname>Chatterjee</surname>
          </string-name>
          .
          <article-title>Calculated risk? A cybersecurity evaluation tool for SMEs</article-title>
          .
          <source>Business Horizons</source>
          ,
          <volume>63</volume>
          (
          <issue>4</issue>
          ):
          <fpage>531</fpage>
          -
          <lpage>540</lpage>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [Boneh et al.,
          <year>2019</year>
          ]
          <string-name>
            <given-names>Dan</given-names>
            <surname>Boneh</surname>
          </string-name>
          , Andrew J Grotto,
          <string-name>
            <surname>Patrick McDaniel</surname>
            ,
            <given-names>and Nicolas</given-names>
          </string-name>
          <string-name>
            <surname>Papernot</surname>
          </string-name>
          .
          <article-title>How relevant is the Turing test in the age of sophisbots?</article-title>
          <source>IEEE Security &amp; Privacy</source>
          ,
          <volume>17</volume>
          (
          <issue>6</issue>
          ):
          <fpage>64</fpage>
          -
          <lpage>71</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <source>[Bostrom</source>
          , 2017]
          <string-name>
            <given-names>Nick</given-names>
            <surname>Bostrom</surname>
          </string-name>
          . openness in
          <source>AI development. 148</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <source>Strategic implications of Global policy</source>
          ,
          <volume>8</volume>
          (
          <issue>2</issue>
          ):
          <fpage>135</fpage>
          -
          <lpage>[</lpage>
          Brown et al.,
          <year>2020</year>
          ] Tom
          <string-name>
            <surname>B Brown</surname>
          </string-name>
          , Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry,
          <string-name>
            <given-names>Amanda</given-names>
            <surname>Askell</surname>
          </string-name>
          , et al.
          <article-title>Language models are few-shot learners</article-title>
          .
          <source>arXiv preprint arXiv:2005.14165</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [Brundage et al.,
          <year>2018</year>
          ]
          <string-name>
            <given-names>Miles</given-names>
            <surname>Brundage</surname>
          </string-name>
          , Shahar Avin, Jack Clark, Helen Toner,
          <string-name>
            <given-names>Peter</given-names>
            <surname>Eckersley</surname>
          </string-name>
          , Ben Garfinkel, Allan Dafoe, Paul Scharre, Thomas Zeitzoff,
          <string-name>
            <given-names>Bobby</given-names>
            <surname>Filar</surname>
          </string-name>
          , et al.
          <article-title>The malicious use of artificial intelligence: Forecasting, prevention, and mitigation</article-title>
          . arXiv preprint arXiv:
          <year>1802</year>
          .07228,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <source>[Bufacchi</source>
          , 2021]
          <string-name>
            <given-names>Vittorio</given-names>
            <surname>Bufacchi</surname>
          </string-name>
          .
          <article-title>Truth, lies and tweets: A consensus theory of post-truth</article-title>
          .
          <source>Philosophy &amp; Social Criticism</source>
          ,
          <volume>47</volume>
          (
          <issue>3</issue>
          ):
          <fpage>347</fpage>
          -
          <lpage>361</lpage>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <article-title>[Burden and Herna´ndez-</article-title>
          <string-name>
            <surname>Orallo</surname>
          </string-name>
          ,
          <year>2020</year>
          ]
          <string-name>
            <given-names>John</given-names>
            <surname>Burden</surname>
          </string-name>
          and Jose´ Herna´
          <fpage>ndez</fpage>
          -Orallo.
          <article-title>Exploring AI Safety in Degrees: Generality, Capability and Control</article-title>
          .
          <source>In SafeAI@ AAAI</source>
          , pages
          <fpage>36</fpage>
          -
          <lpage>40</lpage>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [Carlini et al.,
          <year>2019</year>
          ] Nicholas Carlini,
          <string-name>
            <given-names>Anish</given-names>
            <surname>Athalye</surname>
          </string-name>
          , Nicolas Papernot, Wieland Brendel, Jonas Rauber, Dimitris Tsipras, Ian Goodfellow, Aleksander Madry, and
          <string-name>
            <given-names>Alexey</given-names>
            <surname>Kurakin</surname>
          </string-name>
          .
          <article-title>On evaluating adversarial robustness</article-title>
          .
          <source>arXiv preprint arXiv:1902.06705</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <source>[Chesney and Citron</source>
          , 2019]
          <string-name>
            <given-names>Bobby</given-names>
            <surname>Chesney</surname>
          </string-name>
          and
          <string-name>
            <given-names>Danielle</given-names>
            <surname>Citron</surname>
          </string-name>
          .
          <article-title>Deep fakes: A looming challenge for privacy, democracy, and national security</article-title>
          . Calif. L. Rev.,
          <volume>107</volume>
          :
          <fpage>1753</fpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <source>[Collier and Sarkis</source>
          , 2021]
          <article-title>Zachary A Collier and Joseph Sarkis. The zero trust supply chain: Managing supply chain risk in the absence of trust</article-title>
          .
          <source>International Journal of Production Research</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>16</lpage>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <source>[Dehouche</source>
          , 2021]
          <string-name>
            <given-names>Nassim</given-names>
            <surname>Dehouche</surname>
          </string-name>
          .
          <article-title>Plagiarism in the age of massive Generative Pre-trained Transformers (GPT-3)</article-title>
          .
          <source>Ethics in Science and Environmental Politics</source>
          ,
          <volume>21</volume>
          :
          <fpage>17</fpage>
          -
          <lpage>23</lpage>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <source>[Deutsch</source>
          , 2011]
          <string-name>
            <given-names>David</given-names>
            <surname>Deutsch</surname>
          </string-name>
          .
          <article-title>The beginning of infinity: Explanations that transform the world</article-title>
          .
          <source>Penguin UK</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <source>[Deutsch</source>
          , 2016]
          <string-name>
            <given-names>David</given-names>
            <surname>Deutsch</surname>
          </string-name>
          .
          <article-title>The logic of experimental tests, particularly of Everettian quantum theory</article-title>
          .
          <source>Studies in History and Philosophy of Science Part B: Studies in History and Philosophy of Modern Physics</source>
          ,
          <volume>55</volume>
          :
          <fpage>24</fpage>
          -
          <lpage>33</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <source>[Fallis</source>
          , 2020]
          <string-name>
            <given-names>Don</given-names>
            <surname>Fallis</surname>
          </string-name>
          .
          <source>The Epistemic Threat of Deepfakes. Philosophy &amp; Technology</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>21</lpage>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [Fickinger et al.,
          <year>2020</year>
          ]
          <string-name>
            <given-names>Arnaud</given-names>
            <surname>Fickinger</surname>
          </string-name>
          , Simon Zhuang,
          <article-title>Dylan Hadfield-Menell, and Stuart Russell. Multi-principal assistance games</article-title>
          .
          <source>arXiv preprint arXiv:2007.09540</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [Floridi et al.,
          <year>2018</year>
          ]
          <string-name>
            <given-names>Luciano</given-names>
            <surname>Floridi</surname>
          </string-name>
          , Josh Cowls, Monica Beltrametti, Raja Chatila, Patrice Chazerand, Virginia Dignum, Christoph Luetge, Robert Madelin, Ugo Pagallo,
          <string-name>
            <given-names>Francesca</given-names>
            <surname>Rossi</surname>
          </string-name>
          , et al.
          <article-title>AI4People-an ethical framework for a good AI society: opportunities, risks, principles, and recommendations</article-title>
          .
          <source>Minds and Machines</source>
          ,
          <volume>28</volume>
          (
          <issue>4</issue>
          ):
          <fpage>689</fpage>
          -
          <lpage>707</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <source>[Floridi</source>
          , 2018]
          <string-name>
            <given-names>Luciano</given-names>
            <surname>Floridi</surname>
          </string-name>
          .
          <article-title>Artificial intelligence, deepfakes and a future of ectypes</article-title>
          .
          <source>Philosophy &amp; Technology</source>
          ,
          <volume>31</volume>
          (
          <issue>3</issue>
          ):
          <fpage>317</fpage>
          -
          <lpage>321</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <source>[Frederick</source>
          , 2020]
          <string-name>
            <given-names>Danny</given-names>
            <surname>Frederick</surname>
          </string-name>
          .
          <article-title>Against the Philosophical Tide: Essays in Popperian Critical Rationalism</article-title>
          .
          <source>Critias Publishing</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <source>[Hartmann and Giles</source>
          , 2020]
          <string-name>
            <given-names>Kim</given-names>
            <surname>Hartmann</surname>
          </string-name>
          and
          <string-name>
            <given-names>Keir</given-names>
            <surname>Giles</surname>
          </string-name>
          .
          <article-title>The Next Generation of Cyber-Enabled Information Warfare</article-title>
          .
          <source>In 2020 12th International Conference on Cyber Conflict (CyCon)</source>
          , volume
          <volume>1300</volume>
          , pages
          <fpage>233</fpage>
          -
          <lpage>250</lpage>
          . IEEE,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <source>[Hartmann and Steup</source>
          , 2020]
          <string-name>
            <given-names>Kim</given-names>
            <surname>Hartmann</surname>
          </string-name>
          and
          <string-name>
            <given-names>Christoph</given-names>
            <surname>Steup</surname>
          </string-name>
          .
          <article-title>Hacking the AI - the Next Generation of Hijacked Systems</article-title>
          .
          <source>In 2020 12th International Conference on Cyber Conflict (CyCon)</source>
          , volume
          <volume>1300</volume>
          , pages
          <fpage>327</fpage>
          -
          <lpage>349</lpage>
          . IEEE,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [Heyvaert et al.,
          <year>2019</year>
          ]
          <string-name>
            <given-names>Pieter</given-names>
            <surname>Heyvaert</surname>
          </string-name>
          , Ben De Meester, Anastasia Dimou, and
          <string-name>
            <given-names>Ruben</given-names>
            <surname>Verborgh</surname>
          </string-name>
          .
          <article-title>Rule-driven inconsistency resolution for knowledge graph generation rules</article-title>
          .
          <source>Semantic Web</source>
          ,
          <volume>10</volume>
          (
          <issue>6</issue>
          ):
          <fpage>1071</fpage>
          -
          <lpage>1086</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [Ho et al.,
          <year>2020</year>
          ]
          <article-title>Shirley S Ho, Tong Jee Goh, and Yan Wah Leung. Let's nab fake science news: Predicting scientists' support for interventions using the influence of presumed media influence model</article-title>
          .
          <source>Journalism, page 1464884920937488</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <source>[Jakubowski</source>
          , 2019]
          <string-name>
            <given-names>G</given-names>
            <surname>Jakubowski</surname>
          </string-name>
          .
          <article-title>What's not to like? Social media as information operations force multiplier</article-title>
          .
          <source>Joint Force Quarterly</source>
          ,
          <volume>3</volume>
          :
          <fpage>8</fpage>
          -
          <lpage>17</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [Jobin et al.,
          <year>2019</year>
          ]
          <string-name>
            <given-names>Anna</given-names>
            <surname>Jobin</surname>
          </string-name>
          , Marcello Ienca, and
          <string-name>
            <given-names>Effy</given-names>
            <surname>Vayena</surname>
          </string-name>
          .
          <article-title>The global landscape of AI ethics guidelines</article-title>
          .
          <source>Nature Machine Intelligence</source>
          ,
          <volume>1</volume>
          (
          <issue>9</issue>
          ):
          <fpage>389</fpage>
          -
          <lpage>399</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          <source>[Kaloudi and Li</source>
          , 2020]
          <string-name>
            <given-names>Nektaria</given-names>
            <surname>Kaloudi</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jingyue</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <article-title>The AI-based Cyber Threat Landscape: A Survey</article-title>
          .
          <source>ACM Computing Surveys (CSUR)</source>
          ,
          <volume>53</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>34</lpage>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          <source>[Kindervag</source>
          , 2010
          <string-name>
            <given-names>] John</given-names>
            <surname>Kindervag</surname>
          </string-name>
          .
          <article-title>Build security into your network's DNA: The zero trust network architecture</article-title>
          .
          <source>Forrester Research Inc</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>26</lpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [Kirat et al.,
          <year>2018</year>
          ]
          <string-name>
            <given-names>Dhilung</given-names>
            <surname>Kirat</surname>
          </string-name>
          , Jiyong Jang, and
          <string-name>
            <given-names>Marc</given-names>
            <surname>Stoecklin</surname>
          </string-name>
          .
          <article-title>Deeplocker-concealing targeted attacks with AI locksmithing</article-title>
          .
          <source>Blackhat USA</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [Kranenbarg et al.,
          <year>2018</year>
          ] Marleen Weulen Kranenbarg, Thomas J Holt, and Jeroen van der Ham.
          <article-title>Don't shoot the messenger! A criminological and computer science perspective on coordinated vulnerability disclosure</article-title>
          .
          <source>Crime Science</source>
          ,
          <volume>7</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [Leike et al.,
          <year>2017</year>
          ]
          <string-name>
            <given-names>Jan</given-names>
            <surname>Leike</surname>
          </string-name>
          , Miljan Martic, Victoria Krakovna,
          <source>Pedro A Ortega</source>
          , Tom Everitt,
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Lefrancq</surname>
          </string-name>
          , Laurent Orseau, and Shane Legg.
          <article-title>AI safety gridworlds</article-title>
          .
          <source>arXiv preprint arXiv:1711.09883</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [Mahlangu et al.,
          <year>2019</year>
          ]
          <string-name>
            <given-names>Thabo</given-names>
            <surname>Mahlangu</surname>
          </string-name>
          , Sinethemba January, Thulani Mashiane, Moses Dlamini, Sipho Ngobeni, Nkqubela Ruxwana, and
          <string-name>
            <given-names>Sun</given-names>
            <surname>Tzu</surname>
          </string-name>
          .
          <article-title>Data Poisoning: Achilles Heel of Cyber Threat Intelligence Systems</article-title>
          .
          <source>In Proceedings of the ICCWS 2019 14th International Conference on Cyber Warfare and Security: ICCWS</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          <source>[Makri</source>
          , 2017]
          <string-name>
            <given-names>Anita</given-names>
            <surname>Makri</surname>
          </string-name>
          .
          <article-title>Give the public the tools to trust scientists</article-title>
          .
          <source>Nature News</source>
          ,
          <volume>541</volume>
          (
          <issue>7637</issue>
          ):
          <fpage>261</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          <source>[McGregor</source>
          ,
          <year>2020</year>
          ]
          <string-name>
            <given-names>Sean</given-names>
            <surname>McGregor. Preventing Repeated Real World AI Failures by Cataloging</surname>
          </string-name>
          <article-title>Incidents: The AI Incident Database</article-title>
          . arXiv preprint arXiv:
          <year>2011</year>
          .08512,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          [ O´hE´ igeartaigh et al.,
          <year>2020</year>
          ]
          <article-title>Sea´n S O´ hE´ igeartaigh</article-title>
          , Jess Whittlestone, Yang Liu,
          <string-name>
            <surname>Yi Zeng</surname>
          </string-name>
          , and Zhe Liu.
          <article-title>Overcoming barriers to cross-cultural cooperation in AI ethics and governance</article-title>
          .
          <source>Philosophy &amp; Technology</source>
          ,
          <volume>33</volume>
          (
          <issue>4</issue>
          ):
          <fpage>571</fpage>
          -
          <lpage>593</lpage>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          [Ozkan et al.,
          <year>2021</year>
          ]
          <string-name>
            <given-names>Bilge</given-names>
            <surname>Yigit</surname>
          </string-name>
          <string-name>
            <surname>Ozkan</surname>
          </string-name>
          , Sonny van Lingen,
          <string-name>
            <given-names>and Marco</given-names>
            <surname>Spruit</surname>
          </string-name>
          .
          <article-title>The Cybersecurity Focus Area Maturity (CYSFAM) Model</article-title>
          .
          <source>Journal of Cybersecurity and Privacy</source>
          ,
          <volume>1</volume>
          (
          <issue>1</issue>
          ):
          <fpage>119</fpage>
          -
          <lpage>139</lpage>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          <source>[Pistono and Yampolskiy</source>
          , 2016]
          <article-title>Federico Pistono and Roman V Yampolskiy</article-title>
          . Unethical Research:
          <article-title>How to Create a Malevolent Artificial Intelligence</article-title>
          . arXiv e-prints, pages
          <fpage>arXiv</fpage>
          -
          <lpage>1605</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          <source>[Popper</source>
          , 1996]
          <string-name>
            <given-names>Karl</given-names>
            <surname>Popper</surname>
          </string-name>
          .
          <article-title>In search of a better world: Lectures and essays from thirty years</article-title>
          . Psychology Press,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref47">
        <mixed-citation>
          <source>[Popper</source>
          , 2014]
          <string-name>
            <given-names>Karl</given-names>
            <surname>Popper</surname>
          </string-name>
          .
          <article-title>Conjectures and refutations: The growth of scientific knowledge</article-title>
          . routledge,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref48">
        <mixed-citation>
          <source>[Prier</source>
          , 2017]
          <string-name>
            <given-names>Jarred</given-names>
            <surname>Prier</surname>
          </string-name>
          .
          <article-title>Commanding the trend: Social media as information warfare</article-title>
          .
          <source>Strategic Studies Quarterly</source>
          ,
          <volume>11</volume>
          (
          <issue>4</issue>
          ):
          <fpage>50</fpage>
          -
          <lpage>85</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref49">
        <mixed-citation>
          [Radford et al.,
          <year>2019</year>
          ]
          <string-name>
            <given-names>Alec</given-names>
            <surname>Radford</surname>
          </string-name>
          , Jeffrey Wu, Rewon Child, David Luan,
          <string-name>
            <given-names>Dario</given-names>
            <surname>Amodei</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Ilya</given-names>
            <surname>Sutskever</surname>
          </string-name>
          .
          <article-title>Language models are unsupervised multitask learners</article-title>
          .
          <source>OpenAI blog</source>
          ,
          <volume>1</volume>
          (
          <issue>8</issue>
          ):
          <fpage>9</fpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref50">
        <mixed-citation>
          [Raji et al.,
          <year>2020</year>
          ]
          <string-name>
            <given-names>Inioluwa</given-names>
            <surname>Deborah</surname>
          </string-name>
          <string-name>
            <surname>Raji</surname>
          </string-name>
          , Timnit Gebru, Margaret Mitchell, Joy Buolamwini,
          <string-name>
            <given-names>Joonseok</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Emily</given-names>
            <surname>Denton</surname>
          </string-name>
          .
          <article-title>Saving face: Investigating the ethical concerns of facial recognition auditing</article-title>
          .
          <source>In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society</source>
          , pages
          <fpage>145</fpage>
          -
          <lpage>151</lpage>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref51">
        <mixed-citation>
          [Ranade et al.,
          <year>2021</year>
          ]
          <string-name>
            <given-names>Priyanka</given-names>
            <surname>Ranade</surname>
          </string-name>
          , Aritran Piplai, Sudip Mittal, Anupam Joshi, and
          <string-name>
            <given-names>Tim</given-names>
            <surname>Finin</surname>
          </string-name>
          .
          <source>Generating Fake Cyber Threat Intelligence Using Transformer-Based Models. arXiv preprint arXiv:2102.04351</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref52">
        <mixed-citation>
          <source>[Sahlgren and Carlsson</source>
          , 2021]
          <string-name>
            <given-names>Magnus</given-names>
            <surname>Sahlgren</surname>
          </string-name>
          and
          <string-name>
            <surname>Fredrik Carlsson.</surname>
          </string-name>
          <article-title>The Singleton Fallacy: Why Current Critiques of Language Models Miss the Point</article-title>
          .
          <source>arXiv preprint arXiv:2102.04310</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref53">
        <mixed-citation>
          <source>[Seymour and Tully</source>
          , 2016
          <string-name>
            <given-names>] John</given-names>
            <surname>Seymour</surname>
          </string-name>
          and
          <string-name>
            <given-names>Philip</given-names>
            <surname>Tully</surname>
          </string-name>
          .
          <article-title>Weaponizing data science for social engineering: Automated E2E spear phishing on Twitter</article-title>
          .
          <source>Black Hat USA</source>
          ,
          <volume>37</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>39</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref54">
        <mixed-citation>
          <string-name>
            <surname>[Smith</surname>
          </string-name>
          et al.,
          <year>2021</year>
          ]
          <string-name>
            <given-names>Ryan</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Karl</given-names>
            <surname>Friston</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Christopher</given-names>
            <surname>Whyte</surname>
          </string-name>
          .
          <article-title>A Step-by-Step Tutorial on Active Inference and its Application to Empirical Data</article-title>
          .
          <source>PsyArXiv</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref55">
        <mixed-citation>
          <article-title>[Tallo´n-</article-title>
          <string-name>
            <surname>Ballesteros</surname>
          </string-name>
          ,
          <year>2020</year>
          ]
          <string-name>
            <given-names>AJ</given-names>
            <surname>Tallo</surname>
          </string-name>
          <article-title>´n-Ballesteros. Exploring the Potential of GPT-2 for Generating Fake Reviews of Research Papers</article-title>
          .
          <source>Fuzzy Systems and Data Mining VI: Proceedings of FSDM</source>
          <year>2020</year>
          ,
          <volume>331</volume>
          :
          <fpage>390</fpage>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref56">
        <mixed-citation>
          <source>[Tully and Foster</source>
          , 2020]
          <string-name>
            <given-names>Philip</given-names>
            <surname>Tully</surname>
          </string-name>
          and
          <string-name>
            <given-names>Lee</given-names>
            <surname>Foster</surname>
          </string-name>
          .
          <article-title>Repurposing Neural Networks to Generate Synthetic Media for Information Operations</article-title>
          . https://www.blackhat.com/ us-20/briefings/schedule/,
          <year>2020</year>
          . Session at blackhat USA 2020; accessed 08-August-
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref57">
        <mixed-citation>
          <string-name>
            <surname>[Van Noorden</surname>
          </string-name>
          ,
          <year>2014</year>
          ] Richard Van Noorden.
          <article-title>Publishers withdraw more than 120 gibberish papers</article-title>
          .
          <source>Nature News</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref58">
        <mixed-citation>
          <source>[Vinnakota</source>
          , 2013]
          <string-name>
            <given-names>Tirumala</given-names>
            <surname>Vinnakota</surname>
          </string-name>
          .
          <article-title>A cybernetics paradigms framework for cyberspace: Key lens to cybersecurity</article-title>
          .
          <source>In 2013 IEEE International Conference on Computational Intelligence and Cybernetics</source>
          (CYBERNETICSCOM), pages
          <fpage>85</fpage>
          -
          <lpage>91</lpage>
          . IEEE,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref59">
        <mixed-citation>
          [Wahle et al.,
          <year>2021</year>
          ]
          <string-name>
            <given-names>Jan</given-names>
            <surname>Philip</surname>
          </string-name>
          <string-name>
            <surname>Wahle</surname>
          </string-name>
          , Terry Ruas, Norman Meuschke, and
          <string-name>
            <given-names>Bela</given-names>
            <surname>Gipp</surname>
          </string-name>
          .
          <article-title>Are neural language models good plagiarists? A benchmark for neural paraphrase detection</article-title>
          .
          <source>arXiv preprint arXiv:2103.12450</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref60">
        <mixed-citation>
          [Zeadally et al.,
          <year>2020</year>
          ]
          <string-name>
            <given-names>Sherali</given-names>
            <surname>Zeadally</surname>
          </string-name>
          , Erwin Adi, Zubair Baig, and Imran A Khan.
          <article-title>Harnessing artificial intelligence capabilities to improve cybersecurity</article-title>
          .
          <source>IEEE Access</source>
          ,
          <volume>8</volume>
          :
          <fpage>23817</fpage>
          -
          <lpage>23837</lpage>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref61">
        <mixed-citation>
          <string-name>
            <surname>[Zhao</surname>
          </string-name>
          et al.,
          <year>2021</year>
          ]
          <string-name>
            <given-names>Bo</given-names>
            <surname>Zhao</surname>
          </string-name>
          , Shaozeng Zhang, Chunxue Xu,
          <string-name>
            <given-names>Yifan</given-names>
            <surname>Sun</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Chengbin</given-names>
            <surname>Deng</surname>
          </string-name>
          .
          <article-title>Deep fake geography? When geospatial data encounter Artificial Intelligence</article-title>
          .
          <source>Cartography and Geographic Information Science</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>15</lpage>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>