<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Requisite Variety in Ethical Utility Functions for AI Value Alignment</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nadisha-Marie Aliman</string-name>
          <email>nadishamarie.aliman@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Leon Kester</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>TNO Netherlands</institution>
          ,
          <addr-line>The Hague</addr-line>
          ,
          <country country="NL">Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Utrecht University</institution>
          ,
          <addr-line>Utrecht</addr-line>
          ,
          <country country="NL">Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Being a complex subject of major importance in AI Safety research, value alignment has been studied from various perspectives in the last years. However, no final consensus on the design of ethical utility functions facilitating AI value alignment has been achieved yet. Given the urgency to identify systematic solutions, we postulate that it might be useful to start with the simple fact that for the utility function of an AI not to violate human ethical intuitions, it trivially has to be a model of these intuitions and reflect their variety - whereby the most accurate models pertaining to human entities being biological organisms equipped with a brain constructing concepts like moral judgements, are scientific models. Thus, in order to better assess the variety of human morality, we perform a transdisciplinary analysis applying a security mindset to the issue and summarizing variety-relevant background knowledge from neuroscience and psychology. We complement this information by linking it to augmented utilitarianism as a suitable ethical framework. Based on that, we propose first practical guidelines for the design of approximate ethical goal functions that might better capture the variety of human moral judgements. Finally, we conclude and address future possible challenges.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        AI value alignment, the attempt to implement systems
adhering to human ethical values has been recognized as highly
relevant subtask in AI Safety at an international level and
studied by multiple AI and AI Safety researchers across
diverse research subareas [Hadfield-Menell et al., 2016; Soares
and Fallenstein, 2017; Yudkowsky, 2016]
        <xref ref-type="bibr" rid="ref2 ref22 ref35 ref38">(a review is
provided in [Taylor et al., 2016])</xref>
        . Moreover, the need to
investigate value alignment has been included in the Asilomar
AI Principles [2018] with a worldwide support of researchers
from the field. While value alignment has often been
tackled using reinforcement learning [Abel et al., 2016]
        <xref ref-type="bibr" rid="ref17 ref27 ref30 ref31 ref34 ref40">(and also
reward modeling [Leike et al., 2018])</xref>
        or inverse
reinforcement learning [Abbeel and Ng, 2004] methods, we focus on
the approach to explicitly formulate cardinal ethical utility
functions crafted by (a representation of) society and assisted
by science and technology which has been termed ethical
goal functions [Aliman and Kester, 2019b; Werkhoven et al.,
2018]. In order to be able to formulate utility functions that
do not violate the ethical intuitions of most entities in a
society, these ethical goal functions will have to be a model of
human ethical intuitions. This simple but important insight
can be derived from the good regulator theorem in
cybernetics [Conant and Ross Ashby, 1970] stating that “every good
regulator of a system must be a model of that system”. We
believe that instead of learning models of human intuitions in
their apparent complexity and ambiguity, AI Safety research
could also make use of the already available scientific
knowledge on the nature of human moral judgements and ethical
conceptions as made available e.g. by neuroscience and
psychology. The human brain did not evolve to facilitate rational
decision-making or the experience of emotions, but instead
to fulfill the core task of allostasis (anticipating the needs
of the body in an environment before they arise in order to
ensure growth, survival and reproduction) [Barrett, 2017a;
Kleckner et al., 2017]. Thereby, psychological functions
such as cognition, emotion or moral judgements are closely
linked to the predictive regulation of physiological needs of
the body [Kleckner et al., 2017] making it indispensable to
consider the embodied nature of morality when aspiring to
model it for AI value alignment.
      </p>
      <p>For the purpose of facilitating the injection of requisite
knowledge reflecting the variety of human morality in ethical
goal functions, Section 2 provides information on the
following variety-relevant aspects: 1) the essential role of affect and
emotion in moral judgements from a modern
constructionist neuroscience and cognitive science perspective followed
by 2) dyadic morality as a recent psychological theory on
the nature of cognitive templates for moral judgements. In
Section 3, we propose first guidelines on how to
approximately formulate ethical goal functions using a recently
proposed non-normative socio-technological ethical framework
grounded in science called augmented utilitarianism [Aliman
and Kester, 2019a] that might be useful to better incorporate
the requisite variety of human ethical intuitions (especially in
comparison to classical utilitarianism). Thereafter, we
propose how to possibly validate these functions within a
sociotechnological feedback-loop [Aliman and Kester, 2019b].
Finally, in Section 4, we conclude and specify open challenges</p>
    </sec>
    <sec id="sec-2">
      <title>Variety in Embodied Morality</title>
      <p>
        While value alignment is often seen as a safety problem, it
is possible to interpret and reformulate it as a related
security problem which might offer a helpful different perspective
on the subject emphasizing the need to capture the variety of
embodied morality. One possible way to look at AI value
alignment is to consider it as being an attempt to achieve
advanced AI systems exhibiting adversarial robustness against
malicious adversaries attempting to lead the system to
action(s) or output(s) that are perceived as violating human
ethical intuitions. From an abstract point of view, one could
distinguish different means by which an adversary might achieve
successful attacks: e.g. 1) by fooling the AI at the
perceptionlevel
        <xref ref-type="bibr" rid="ref17 ref19 ref20 ref24 ref30 ref31 ref34 ref4 ref5">(in analogy to classical adversarial examples
[Goodfellow, 2018], this variant has been denoted ethical
adversarial examples [Aliman and Kester, 2019a])</xref>
        which could lead
to an unethical behavior even if the utility function would
have been aligned with human ethical intuitions or 2)
simply by disclosing dangerous (certainly unintended from the
designer) unethical implications encoded in its utility
function by targeting specific mappings from perception to output
or action (this could be understood as ethical adversarial
examples on the utility function itself). While the existence of
point 1) yields one more argument for the importance of
research on adversarial robustness at the perception-level for AI
Safety reasons [Goodfellow, 2019] and a sophisticated
combination of 1) and 2) might be thinkable, our exemplification
focuses on adversarial attacks of the type 2).
      </p>
      <p>One could consider the explicitly formulated utility
function U as representing a separate model1 that given a sample,
outputs a value determining the perceived ethical
desirabil1a conceptually similar separation of objective function model
and optimizing agent has been recently performed for reward
modeling [Leike et al., 2018]
ity of that sample which should ideally be in line with the
society that crafted this utility function. The attacker which
has at his disposal the knowledge on human ethical intuitions,
can attempt targeted misclassifications at the level of a
single sample or at the level of an ordering of multiple samples
whereby the ground-truth are the ethical intuitions of most
people in a society. The Law of Requisite Variety from
cybernetics [Ashby, 1961] states that “only variety can destroy
variety”, with other words in order to cope with a certain
variety of problems or environmental variety, a system needs
to exhibit a suitable and sufficient variety of responses.
Figure 1 offers an intuitive explanation of this law. Transferring
it to the mentioned utility function U , it is for instance
conceivable that if U does not encode affective information that
might lead to a difference in ethical evaluations, an attacker
can easily craft a sample which U might misclassify as ethical
or unethical or cause U to generate a total ordering of samples
that might appear unethical from the perspective of most
people. Given that U does not have an influence on the variety of
human morality, the only way to respond to the disturbances
of the attacker and reduce the variety of possible undesirable
outcomes, is by increasing the own variety – which can be
achieved by encoding more relevant knowledge.
2.1</p>
      <sec id="sec-2-1">
        <title>Role of Emotion and Affect in Morality</title>
        <p>One fundamental and persistent misconception about human
biology (which does not only affect the understanding of the
nature of moral judgements) is the assumption that the brain
incorporates a layered architecture in which a battle between
emotion and cognition is given through the very anatomy of
the “triune brain” [MacLean, 1990] exhibiting three
hierarchical layers: a reptilian brain on top of which an emotional
animalistic paleomammalian limbic system is located and a
final rational neomammalian cognition layer implemented in
the neocortex. This flawed view is not in accordance with
neuroscientific evidence and understanding [Barrett, 2017a;
Miller and Clark, 2018]. In fact, the assumed reactive and
animalistic limbic regions in the brain are predictive (e.g. they
send top-down predictions to more granular cortical regions),
control the body as well as attention mechanisms while being
the source of the brain’s internal model of the body [Barrett
and Simmons, 2015; Barrett, 2017b].</p>
        <p>
          Emotion and cognition do not represent a dichotomy
leading to a conflict in moral judgements [Helion and Pizarro,
2015]. Instead, the distinction between the experience of an
instance of a concept as belonging to the category of emotions
versus the category of cognition is grounded in the focus of
attention of the brain [Barrett et al., 2015] whereby “the
experience of cognition occurs when the brain foregrounds
mental contents and processes” and “the experience of emotion
occurs when, in relation to the current situation, the brain
foregrounds bodily changes” [Hoemann and Barrett, 2019].
The mental phenomenon of actively dynamically simulating
different alternative scenarios (including anticipatory
emotions) has also been termed conceptual consumption [Gilbert
and Wilson, 2007] and plays a role in decision-making and
moral reasoning. While emotions are discrete constructions
of the human brain, core affect allows a low-dimensional
experience of interoceptive sensations (sensory array from
within the body) and is a continuous property of conciousness
with the dimensions of valence (pleasantness/unpleasantness)
and arousal (activation/deactivation) [Kleckner et al., 2017].
It has been argued that core affect provides a basis for
moral judgements in which different events are qualitatively
compared to each other [Cabanac, 2002]. Like other
constructed mental states, moral judgements involve
domaingeneral brain processes which simply put combine 1) the
interoceptive sensory array, 2) the exteroceptive sensory inputs
from the environment and 3) past experience/ knowledge for
a goal-oriented situated conceptualization (as tool for
allostasis) [Oosterwijk et al., 2012]. From these key constituents
of mental constructions one can extract the following:
concepts (including morality) are perceiver-dependent and
timedependent. Thereby, affect,
          <xref ref-type="bibr" rid="ref10 ref14">(but not emotion [Cameron et
al., 2015])</xref>
          is a necessary ingredient of every moral
judgement. More fundamentally, “the human brain is
anatomically structured so that no decision or action can be free of
interoception and affect” [Barrett, 2017a] – this includes any
type of thoughts that seem to correspond to the folk terms
of “rational” and “cold”. Therefore, a utility function without
affect-related parameters might not exhibit a sufficient variety
and might lead to the violation of human ethical intuitions.
        </p>
        <p>Morality cannot be separated from a model of the body,
since the brain constructs the human perception of reality
based on what seems of importance to the brain for the
purpose of allostasis which is inherently strongly linked to
interoception [Barrett, 2017a]. Interestingly, even the
imagination of future not yet experienced events is facilitated through
situated recombinations of sensory-motor and affective
nature in a similar way as the simulation of actually experienced
events [Addis, 2018]. To sum up, there is no battle between
emotion and cognition in moral judgements. Moreover, there
is also no specific moral faculty in the brain, since moral
judgements are based on domain-general processes within
which affect is always involved to a certain degree. One could
obtain insufficient variety in dealing with an adversary
crafting ethical adversarial examples on a utility model U if one
ignores affective parameters. Further crucial parameters for
ethical utility functions could be e.g. of cultural, social and
socio-geographical nature.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2 Variety through “Dyadicness”</title>
        <p>The psychological theory of dyadic morality [Schein and
Gray, 2018] posits that moral judgements are based on a fuzzy
cognitive template and related to the perception of an
intentional agent (iA) causing damage (d) to a vulnerable patient
(vP ) denoted iA !d vP . More precisely, the theory
postulates that the perceived immorality of an act is related to
the following three elements: norm violations, negative
affect and importantly perceived harm. According to a study,
the reaction times in describing an act as immoral predict the
reaction times in categorizing the same act as harmful [Schein
and Gray, 2015]. The combination of these basic constituents
is suggested to lead to the emergence of a rich diversity of
moral judgements [Gray et al., 2017]. Dyadicness is
understood as a continuum predicting the condemnation of moral
acts. The more a human entity perceives an intentional agent
inflicting damage to a vulnerable patient, the more immoral
this human perceives the act. As stated by Schein and Gray,
the dyadic harm-based cognitive template “is rooted in innate
and evolved processes of the human mind; it is also shaped
by cultural learning, therefore allowing cultural pluralism”.
Importantly, the nature of this cognitive template reveals that
moral judgements besides being perceiver-dependent, might
vary across diverse parameters such as especially e.g. in
relation to the perception of agent, act and patient in the
outcome of the action. Further, the theory also foresees a
possible time-dependency of moral judgements by introducing
the concept of a dyadic loop, a feedback cycle resulting in an
iterative polarization of moral judgements through social
discussion modulating the perception of harm as time goes by.
Overall, moral judgements are understood as constructions
in the same way visual perception, cognition or emotion are
constructed by the human mind. Similarly to the existence
of variability in visual perception, variability in morality is
the norm which often leads to moral conflicts [Schein et al.,
2016]. However, the understanding that humans share the
same harm-based cognitive template for morality has been
described as reflecting “cognitive unity in the variety of
perceived harm” [Schein and Gray, 2018].</p>
        <p>Analyzing the cognitive template of dyadic morality, one
can deduce that human moral judgements do not only
consider the outcome of an action as prioritized by
consequentialist frameworks like classical utilitarianism, nor do they
only consider the state of the agent which is in the focus
of virtue ethics. Furthermore, as opposed to deontological
ethics, the focus is not only on the nature of the performed
action. The main implications for the design of utility functions
that should ideally be aligned with human ethical values, is
that they might need to encode information on agent, action,
patient as well as on the perceivers – especially with regard
to the cultural background. This observation is fundamental
as it indicates that one might have to depart from classical
utilitarian utility functions U (s0) which are formulated as
total orders at the abstraction level of outcomes i.e. states (of
affairs) s0. In line with this insight, is the context-sensitive
and perceiver-dependent type of utility functions considering
agent, action and outcome which has been recently proposed
within a novel ethical framework denoted augmented
utilitarianism [Aliman and Kester, 2019a] (abbreviated with AU in
the following). Reconsidering the dyadic morality template
iA !d vP , it seems that in order to better capture the
variety of human morality, utility functions – now transferring it
to the perspective of AI systems – would need to be at least
formulated at the abstraction level of a perceiver-dependent
a
evaluation of a transition s ! s0 leading from a state s to a
state s0 via an action a. We encode the required novel type
of utility function with Ux(s; a; s0) with x denoting a specific
perceiver. This formulation could enable an AI system
implemented as utility maximizer to jointly consider parameters
specified by a perceiver which are related to its perception of
agent, the action and the consequences of this action on a
patient. Since the need to consider time-dependency has been
formulated, one would consequently also require to add the
time dimension to the arguments of the utility function
leading to Ux((s; a; s0); t).</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Approximating Ethical Goal Functions</title>
      <p>
        While the psychological theory of dyadic morality was useful
to estimate the abstraction level at which one would at least
have to specify utility functions, the closer analysis on the
nature of the construction of mental states performed in
Section 2, abstractly provides a superset of primitive relevant
parameters that might be critical elements of every moral
judgement (being a mental state). Given a perceiver x, the
components of this set are the following subsets: 1) parameters
encoding the interoceptive sensory array Bx (from within the
body) which are accessible to the human consciousness via
the low-dimensional core affect, 2) the exteroceptive sensory
array Ex encoding information from the environment and 3)
the prior experience Px encoding memories. Moreover, these
set of parameters obviously vary in time. However, to
simplify, it has been suggested within the mentioned AU
framework, that ethical goal functions will have to be updated
regularly
        <xref ref-type="bibr" rid="ref24 ref4 ref5">(leading to a so-called socio-technological
feedbackloop [Aliman and Kester, 2019b])</xref>
        in the same way as votes
take place at regular intervals in a democracy. One could
similarly assume that this regular update will be sufficient to
reflect a relevant change in moral opinion and perception.
3.1
      </p>
      <sec id="sec-3-1">
        <title>Injecting Requisite Variety in Utility</title>
        <p>
          For simplicity, we assume that the set of parameters Bx, Px
and Ex are invariant during the utility assignment process
in which a perceiver x has to specify the ethical
desirability of a transition s !a s0 by mapping it to a cardinal value
Ux(s; a; s0) obtained by applying a not-nearer defined type of
scientifically determined transformation vx (chosen by x) on
the mental state of x. This results in the following naive and
simplified mapping however adequately reflecting the
property of mental-state-dependency formulated in the AU
framework
          <xref ref-type="bibr" rid="ref24 ref4 ref5">(the required dependency of ethical utility functions on
parameters of the own mental state function mx in order to
avoid perverse instantiation scenarios [Aliman and Kester,
2019a])</xref>
          :
        </p>
        <p>Ux(s; a; s0) = vx(mx((s; a; s0); Bx; Px; Ex))
(1)</p>
        <p>Conversely, the utility function of classical utilitarianism
is only defined at the impersonal and context-independent
abstraction level of U (s0) which has been argued to lead to
both perverse instantiation problem but also to the repugnant
conclusion and related impossibility theorems in population
ethics for consequentialist frameworks which do not apply to
mental-state-dependent utility functions [Aliman and Kester,
2019a]. The idea to restrict human ethical utility functions
to the considerations of outcomes of actions alone –
ignoring affective parameters of the own current self – as practiced
in classical utilitarianism while later referring to the
resulting total orders with emotionally connoted adjectives such
as “repugnant” or “perverse” has been termed the
perspectival fallacy of utility assignment [Aliman and Kester, 2019b].
The use of consequentialist utility functions affected by the
impossibility theorems of Arrhenius [2000] has been
justifiably identified by Eckersley [2018] as a safety risk if used
in AI systems without more ado. It seems that the isolated
consideration of outcomes of actions (for consequentialism)
or actions (for deontological ethics) or the involved agents
(for virtue ethics) does not represent a good model of human
ethical intuitions. It is conceivable, that if a utility model U
is defined as utility function U (s0), the model cannot
possibly exhibit a sufficient variety and might more likely violate
human ethical intuitions than if it would be implemented as a
context-sensitive utility function Ux(s; a; s0). (Beyond that, it
has been argued that consequentialism implies the rejection of
“dispositions and emotions, such as regret, disappointment,
guilt and resentment” from “rational” deliberation [Verbeek,
2001] and should i.a. for this reason be disentangled from the
notion of rationality for which it cannot represent a plausible
requirement.)</p>
        <p>
          It is noteworthy that in the context of reinforcement
learning (e.g. in robotics) different types of reward functions are
usually formulated ranging from R(s0) to R(s; a; s0). For the
purpose of ethical utility functions for advanced AI systems
in critical application fields, we postulate that one does not
have the choice to specify the abstraction level of the utility
function, since for instance U (s0) might lead to safety risks.
Christiano et al. [2017] considered the elicitation of human
preferences on trajectory (state-action pairs) segments of a
reinforcement learning agent i.a. realized by human feedback
on short movies. For the purpose of utility elicitation in an
AU framework exemplarily using a naive model as specified
in equation (1), people will similarly have to assign utility to a
movie representing a transition in the future
          <xref ref-type="bibr" rid="ref24 ref4 ref5">(either in a
mental mode or augmented by technology such as VR or AR
[Aliman and Kester, 2019b])</xref>
          . However, it is obvious that this
naive utility assignment would not scale in practice.
Moreover, it has not yet been specified how to aggregate ethical
goal functions at a societal level. In the following
Subsection 3.2, we will address these issues by proposing a
practicable approximation of the utility function in (1) and a possible
societal aggregation of this approximate solution.
3.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Approximation, Aggregation and Validation</title>
        <p>
          So far, it has been stated throughout the paper that one has
to adequately increase the variety of a utility function meant
to be ethical in order to avoid violations of human ethical
intuitions and vulnerability to attackers crafting ethical
adversarial examples against the model. However, it is
important to note that despite the negatively formulated motivation
of the approach, the aim is to craft a utility model U which
represents a better model of human ethical intuitions in
general, thus ranging from samples that are perceived as highly
unethical to those that are assigned a high ethical
desirability. In order to craft practical solutions that lead to optimal
results, it might be advantageous to perform a thought
experiment imagining a utopia and from that impose practical
constraints on its viability. It might not seem realistic to
deliberate a future utopia 1 as a sustainable society which is
stable across a very large time interval in which every human
being acts according to the ethical intuitions of all humans
including the own and every artificial intelligent system
fulfills the ethical intuitions of all humans. However, it seems
more likely that within a utopia 2 being a stable society in
which every human achieves a high level of a scientific
definition of well-being
          <xref ref-type="bibr" rid="ref36">(such as e.g. PERMA [Seligman, 2012])</xref>
          with artificial agents acting as to maximize context-sensitive
utility according to which (human or artificial) agents
promoting the (measurable) well-being of human patients is
regarded as the most utile type of events, the ethical intuitions
of humans might tend to get closer to each other. The reason
being that the variety of human moral judgements might
interestingly decrease since it is conceivable that they will tend
to exhibit more similar prior experiences (all imprinted by
well-being) and have more similar environments (full of
stable people with a high level of well-being). The main factor
drawing differences could be the body – especially
biological factors. However, the parameters related to interoception
might be closer to each other, since all humans exhibit a high
level of well-being which classically includes frequent
positive affect. It is conceivable that with time, such a society
could converge towards the utopia 1.
        </p>
        <p>
          In the following, we will denote the mentioned
utopiarelated ideal cognitive template of a (human or artificial)
agent A performing an act w that contributes to the
wellcboeignngitiovfeatehmu mplaantepoaftiednytaPdicwmithorAalit!yw. P
          <xref ref-type="bibr" rid="ref36">(Thineraenbayl,oAgy t!wo thPe
is perceiver-dependent i.a. because psychological measures
of well-being include subjective and self-reported elements
such as e.g. life satisfaction or furthermore positive
emotions [Seligman, 2012].)</xref>
          Augmented utilitarianism foresees
the need to at least depict a final goal at the abstraction level
of a perceiver-dependent function on a transition as reflected
in Ux(s; a; s0). The ideal cognitive template A !w P
formulated for utopia 2 by which it has been argued that a decrease
in the variety of human morality might be achievable in the
long-term exhibits an abstraction level that is compatible with
Ux(s; a; s0).
        </p>
        <p>A thinkable strategy for the design of a utility model U that
is robust against ethical adversarial examples and a model
of human ethical intuitions is to try to adequately increase
its variety using relevant scientific knowledge and to
complementarily attempt to decrease the variety of human moral
judgements for instance by considering A !w P as high-level
final goal such that the described utopia 2 ideally becomes
a self-fulfilling prophecy. For it to be realizable in
practice, we suggest that the appropriateness of a given
aggregated societal ethical goal function could be approximately
validated against its quantifiable impact on well-being for
society across the time dimension. Since it seems however
unfeasible to directly map all important transitions of a
domain to their effect on the well-being of human entities, we
propose to consider perceiver-specific and domain-specific
utility functions indicating combined preferences that each
perceiver x considers to be relevant for well-being from the
viewpoint of x himself in that specific domain. For these
combined utility functions to be grounded in science, they
will have to be based on scientifically measurable
parameters. We postulate that a possible aggregation at a societal
level could be performed by the following steps: 1)
agreement on a common validation measure of an ethical goal
function (for instance the temporal development of societal
satisfaction with AI systems in a certain domain or with future
AGI systems, their aptitude to contribute to sustainable
wellbeing), 2) agreement on superset O of scientifically
measurable and relevant parameters (encoding e.g. affective, dyadic,
cultural, social, political, socio-geographical but importantly
also law-relevant information) that are considered as
important across the whole society, 3) specification of personal
utility functions for each member n of a society of N members
allowing personalized and tailored combinations of a
subset of O, 4) aggregation to a societal ethical goal function
UT otal(s; a; s0). Taken together, these considerations lead us
to the following possible approximation for an aggregated
societal ethical goal function given a domain:</p>
        <p>N j
UT otal(s; a; s0) = X X winffni(Ci)
with N standing for the number of participating entities in
society, Ci = (pi1; pi2; :::; pim) being a cluster of m 1
correlated parameters (whereby independent factors are assigned
an own cluster each) and f representing a set of preference
functions (form functions). For instance f = ff1; f2; :::; ff g
where f1 could be a linear transformation, f2 a concave, f3
a convex preference function and so on. Each entity n
assigns a weight win to a form function ffni applied to a
cluster of parameters Ci whereby Pij=1 win = 1. We define
O = fC1; C2; :::g as the superset of all parameters considered
in the overall aggregated utility function. Moreover, a 2 A
with A representing the foreseen discrete action space at the
disposal of the AI. (It is important to note that while the AI
could directly perform actions in the environment, it could
also be used for policy-making and provide plans for human
agents.) Further, we consider a continuous state space with
the states s and s0 2 S = RjOj. Other aspects including
e.g. legal rules and norms on the action space can be imposed
as constraints on the utility function. In a nutshell, the
utility aggregation process can be understood as a voting process
in which each participating individual n distributes his vote
across scientifically measurable clusters of parameters Ci on
which he applies a preference function ffni to which weights
win are assigned as identified as relevant by n given a to be
approximated high-level societal goal (such as A !w P ). In
short, people do not have to agree on personal preferences
and weightings, but only on a superset of acceptable
parameters, an aggregation method and an overall validation
measure. (Note that instead of involving society as a whole for
each domain, the utility elicitation procedure can as well be
approximated by a transdisciplinary set of representative
experts (e.g. from the legislative) crafting expert ethical goal
functions that attempt to ideally emulate UT otal(s; a; s0)).</p>
        <p>Finally, it is important to note that the societal ethical goal
function specified in (2) will need to be updated (and
evalutated) at regular intervals due to the mental-state-dependency
of utility entailing time-dependency [Aliman and Kester,
2019a]. This leads to the necessity of a socio-technological
feedback-loop which might concurrently offer the
possibility of a dynamical ethical enhancement [Aliman and Kester,
2019b; Werkhoven et al., 2018]. Pre-deployment, one could
in the future attempt a validation via selected preemptive
simulations [Aliman and Kester, 2019b] in which (a
representation of) society experiences simulations of future events
(s; a; s0) as movies, immersive audio-stories or later in VR
and AR environments. During these experiences, one could
approximately measure the temporal profile of the so-called
artificially simulated future instant utility [Aliman and Kester,
2019b] denoted UT otalAS being a potential constituent of
future well-being. Thereby, UT otalAS refers to the instant
utility [Kahneman et al., 1997] experienced during a
technologyaided simulation of a future event whereby instant utility
refers to the affective dimension of valence at a certain time
t. The temporal integral that a measure of UT otalAS could
approximate is specified as:</p>
        <p>UT otalAS (s; a; s0)</p>
        <p>In(t)dt</p>
        <p>(3)
XN Z T
with t0referring to the starting point of experiencing the
simulation of the event (s; a; s0) augmented by technology (movie,
audio-story, AR, VR) and T the end of this experience. In(t)
represents the valence dimension of core affect experienced
by n at time t. Finally, post-deployment, the ethical goal
function of an AI system can be validated using the
validation measure agreed upon before utility aggregation (such as
the temporal development of societal-level satisfaction with
an AI system, well-being or even the perception of
dyadicness) that has to be a priori determined.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion and Future Work</title>
      <p>In this paper, we motivated the need in AI value alignment
to attempt to model utility functions capturing the variety of
human moral judgements through the integration of relevant
scientific knowledge – especially from neuroscience and
psychology – (instead of learning) in order to avoid violations of
human ethical intuitions. We reformulated value alignment as
a security task and introduced the requirement to increase the
variety within classical utility functions positing that a
utility function which does not integrate affective and
perceiverdependent dyadic information does not exhibit sufficient
variety and might not exhibit robustness against
corresponding adversaries. Using augmented utilitarianism as a suitable
non-normative ethical framework, we proposed a
methodology to implement and possibly validate societal
perceiverdependent ethical goal functions with the goal to better
incorporate the requisite variety for AI value alignment.</p>
      <p>In future work, one could extend and refine the discussed
methodology, study a more systematic validation approach
for ethical goal functions and perform first experimental
studies. Moreover, the “security of the utility function itself is
essential, due to the possibility of its modification by malevolent
actors during the deployment phase” [Aliman and Kester,
2019a]. For this purpose, a blockchain-based solution might
be advantageous. In addition, it is important to note that
even with utility functions exhibiting a sufficient variety for
AI value alignment, it might still be possible for a malicious
attacker to craft adversarial examples against a utility
maximizer at the perception-level which might lead to
unethical behavior. Besides that, one might first need to perform
policy-by-simulation [Werkhoven et al., 2018] prior to the
deployment of advanced AI systems equipped with ethical goal
functions for safety reasons. Last but not least, the usage of
ethical goal functions might represent an interesting approach
to the AI coordination subtask in AI Safety, since an
international use of this method might contribute to reduce the AI
race to the problem-solving ability dimension [Aliman and
Kester, 2019b].</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>[Abbeel and Ng</source>
          , 2004]
          <string-name>
            <given-names>Pieter</given-names>
            <surname>Abbeel</surname>
          </string-name>
          and Andrew Y Ng.
          <article-title>Apprenticeship learning via inverse reinforcement learning</article-title>
          .
          <source>In Proceedings of the twenty-first international conference on Machine learning, page 1. ACM</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [Abel et al.,
          <year>2016</year>
          ]
          <string-name>
            <given-names>David</given-names>
            <surname>Abel</surname>
          </string-name>
          , James MacGlashan, and Michael L Littman.
          <article-title>Reinforcement learning as a framework for ethical decision making</article-title>
          .
          <source>In Workshops at the Thirtieth AAAI Conference on Artificial Intelligence</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <source>[Addis</source>
          , 2018]
          <article-title>Donna Rose Addis</article-title>
          .
          <article-title>Are episodic memories special? On the sameness of remembered and imagined event simulation</article-title>
          .
          <source>Journal of the Royal Society of New Zealand</source>
          ,
          <volume>48</volume>
          (
          <issue>2-3</issue>
          ):
          <fpage>64</fpage>
          -
          <lpage>88</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <source>[Aliman and Kester</source>
          , 2019a]
          <string-name>
            <surname>Nadisha-Marie Aliman</surname>
            and
            <given-names>Leon</given-names>
          </string-name>
          <string-name>
            <surname>Kester</surname>
          </string-name>
          .
          <article-title>Augmented Utilitarianism for AGI Safety</article-title>
          .
          <source>In International Conference on Artificial General Intelligence</source>
          , page to appear. Springer,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <source>[Aliman and Kester</source>
          , 2019b]
          <string-name>
            <surname>Nadisha-Marie Aliman</surname>
            and
            <given-names>Leon</given-names>
          </string-name>
          <string-name>
            <surname>Kester</surname>
          </string-name>
          .
          <article-title>Transformative AI Governance and AIEmpowered Ethical Enhancement Through Preemptive Simulations</article-title>
          . Delphi - Interdisciplinary
          <source>Review of Emerging Technologies</source>
          ,
          <volume>2</volume>
          (
          <issue>1</issue>
          ):
          <fpage>23</fpage>
          -
          <lpage>29</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <source>[Arrhenius</source>
          , 2000]
          <string-name>
            <given-names>Gustaf</given-names>
            <surname>Arrhenius</surname>
          </string-name>
          .
          <article-title>An impossibility theorem for welfarist axiologies</article-title>
          .
          <source>Economics &amp; Philosophy</source>
          ,
          <volume>16</volume>
          (
          <issue>2</issue>
          ):
          <fpage>247</fpage>
          -
          <lpage>266</lpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <source>[Ashby</source>
          , 1961]
          <string-name>
            <given-names>W Ross</given-names>
            <surname>Ashby</surname>
          </string-name>
          .
          <article-title>An introduction to cybernetics</article-title>
          . Chapman &amp; Hall
          <string-name>
            <surname>Ltd</surname>
          </string-name>
          ,
          <year>1961</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <source>[Asilomar</source>
          ,
          <year>2018</year>
          ]
          <string-name>
            <given-names>AI</given-names>
            <surname>Asilomar. Principles</surname>
          </string-name>
          .(
          <year>2017</year>
          ).
          <article-title>In Principles developed in conjunction with the 2017 Asilomar conference</article-title>
          [Benevolent
          <source>AI</source>
          <year>2017</year>
          ],
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>[Barrett and Simmons</source>
          , 2015]
          <string-name>
            <given-names>Lisa</given-names>
            <surname>Feldman Barrett</surname>
          </string-name>
          and
          <string-name>
            <given-names>W Kyle</given-names>
            <surname>Simmons</surname>
          </string-name>
          .
          <article-title>Interoceptive predictions in the brain</article-title>
          .
          <source>Nature Reviews Neuroscience</source>
          ,
          <volume>16</volume>
          (
          <issue>7</issue>
          ):
          <fpage>419</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [Barrett et al.,
          <year>2015</year>
          ]
          <string-name>
            <given-names>Lisa</given-names>
            <surname>Feldman</surname>
          </string-name>
          <string-name>
            <given-names>Barrett</given-names>
            ,
            <surname>Christine D</surname>
          </string-name>
          Wilson-Mendenhall, and Lawrence W Barsalou.
          <article-title>The conceptual act theory: A roadmap</article-title>
          . pages
          <fpage>83</fpage>
          -
          <lpage>110</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [Barrett, 2017a]
          <article-title>Lisa Feldman Barrett</article-title>
          .
          <article-title>How emotions are made: The secret life of the brain</article-title>
          .
          <source>Houghton Mifflin Harcourt</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [Barrett, 2017b]
          <article-title>Lisa Feldman Barrett. The theory of constructed emotion: an active inference account of interoception and categorization</article-title>
          .
          <source>Social cognitive and affective neuroscience</source>
          ,
          <volume>12</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>23</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <source>[Cabanac</source>
          , 2002]
          <string-name>
            <given-names>Michel</given-names>
            <surname>Cabanac</surname>
          </string-name>
          . What is emotion? havioural processes,
          <volume>60</volume>
          (
          <issue>2</issue>
          ):
          <fpage>69</fpage>
          -
          <lpage>83</lpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Be</surname>
          </string-name>
          [Cameron et al.,
          <year>2015</year>
          ]
          <string-name>
            <given-names>C</given-names>
            <surname>Daryl</surname>
          </string-name>
          <article-title>Cameron, Kristen A Lindquist,</article-title>
          and
          <string-name>
            <given-names>Kurt</given-names>
            <surname>Gray</surname>
          </string-name>
          .
          <article-title>A constructionist review of morality and emotions: No evidence for specific links between moral content and discrete emotions</article-title>
          .
          <source>Personality and Social Psychology Review</source>
          ,
          <volume>19</volume>
          (
          <issue>4</issue>
          ):
          <fpage>371</fpage>
          -
          <lpage>394</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [Christiano et al.,
          <year>2017</year>
          ] Paul F Christiano,
          <string-name>
            <surname>Jan</surname>
            <given-names>Leike</given-names>
          </string-name>
          , Tom Brown, Miljan Martic, Shane Legg, and
          <string-name>
            <given-names>Dario</given-names>
            <surname>Amodei</surname>
          </string-name>
          .
          <article-title>Deep reinforcement learning from human preferences</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          , pages
          <fpage>4299</fpage>
          -
          <lpage>4307</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <source>[Conant and Ross Ashby</source>
          ,
          <year>1970</year>
          ] Roger C Conant and
          <string-name>
            <given-names>W Ross</given-names>
            <surname>Ashby</surname>
          </string-name>
          .
          <article-title>Every good regulator of a system must be a model of that system</article-title>
          .
          <source>International journal of systems science</source>
          ,
          <volume>1</volume>
          (
          <issue>2</issue>
          ):
          <fpage>89</fpage>
          -
          <lpage>97</lpage>
          ,
          <year>1970</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <source>[Eckersley</source>
          , 2018]
          <string-name>
            <given-names>Peter</given-names>
            <surname>Eckersley</surname>
          </string-name>
          .
          <article-title>Impossibility and Uncertainty Theorems in AI Value Alignment (or why your AGI should not have a utility function)</article-title>
          .
          <source>arXiv preprint arXiv:1901.00064</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <source>[Gilbert and Wilson</source>
          , 2007] Daniel T Gilbert and
          <string-name>
            <surname>Timothy D Wilson</surname>
          </string-name>
          .
          <article-title>Prospection: Experiencing the future</article-title>
          .
          <source>Science</source>
          ,
          <volume>317</volume>
          (
          <issue>5843</issue>
          ):
          <fpage>1351</fpage>
          -
          <lpage>1354</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <source>[Goodfellow</source>
          , 2018]
          <string-name>
            <given-names>Ian</given-names>
            <surname>Goodfellow</surname>
          </string-name>
          .
          <article-title>Defense Against the Dark Arts: An overview of adversarial example security research and future research directions</article-title>
          . arXiv preprint arXiv:
          <year>1806</year>
          .04169,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <source>[Goodfellow</source>
          , 2019]
          <string-name>
            <given-names>Ian</given-names>
            <surname>Goodfellow</surname>
          </string-name>
          .
          <article-title>Adversarial Robustness for AI Safety</article-title>
          . https://safeai.webs.upv.es/wp-content/ uploads/2019/02/2019-01-27-goodfellow.pdf,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [Gray et al.,
          <year>2017</year>
          ]
          <string-name>
            <given-names>Kurt</given-names>
            <surname>Gray</surname>
          </string-name>
          , Chelsea Schein, and
          <string-name>
            <given-names>C</given-names>
            <surname>Daryl</surname>
          </string-name>
          <article-title>Cameron</article-title>
          .
          <article-title>How to think about emotion and morality: circles, not arrows</article-title>
          . Current opinion in psychology,
          <volume>17</volume>
          :
          <fpage>41</fpage>
          -
          <lpage>46</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [
          <string-name>
            <surname>Hadfield-Menell</surname>
          </string-name>
          et al.,
          <year>2016</year>
          ]
          <string-name>
            <given-names>Dylan</given-names>
            <surname>Hadfield-Menell</surname>
          </string-name>
          , Stuart J Russell, Pieter Abbeel, and
          <string-name>
            <given-names>Anca</given-names>
            <surname>Dragan</surname>
          </string-name>
          .
          <article-title>Cooperative inverse reinforcement learning</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          , pages
          <fpage>3909</fpage>
          -
          <lpage>3917</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <source>[Helion and Pizarro</source>
          , 2015]
          <string-name>
            <given-names>Chelsea</given-names>
            <surname>Helion</surname>
          </string-name>
          and David A Pizarro.
          <article-title>Beyond dual-processes: the interplay of reason and emotion in moral judgment</article-title>
          .
          <source>Handbook of neuroethics</source>
          , pages
          <fpage>109</fpage>
          -
          <lpage>125</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <source>[Hoemann and Barrett</source>
          , 2019]
          <article-title>Katie Hoemann and Lisa Feldman Barrett</article-title>
          .
          <article-title>Concepts dissolve artificial boundaries in the study of emotion and cognition, uniting body, brain, and mind</article-title>
          .
          <source>Cognition and Emotion</source>
          ,
          <volume>33</volume>
          (
          <issue>1</issue>
          ):
          <fpage>67</fpage>
          -
          <lpage>76</lpage>
          ,
          <year>2019</year>
          . PMID:
          <volume>30336722</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [Kahneman et al.,
          <year>1997</year>
          ]
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Kahneman</surname>
          </string-name>
          ,
          <string-name>
            <surname>Peter P Wakker</surname>
            , and
            <given-names>Rakesh</given-names>
          </string-name>
          <string-name>
            <surname>Sarin</surname>
          </string-name>
          .
          <article-title>Back to Bentham? Explorations of experienced utility</article-title>
          .
          <source>The quarterly journal of economics</source>
          ,
          <volume>112</volume>
          (
          <issue>2</issue>
          ):
          <fpage>375</fpage>
          -
          <lpage>406</lpage>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [Kleckner et al.,
          <year>2017</year>
          ] Ian R Kleckner, Jiahe Zhang, Alexandra Touroutoglou, Lorena Chanes,
          <string-name>
            <given-names>Chenjie</given-names>
            <surname>Xia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W Kyle</given-names>
            <surname>Simmons</surname>
          </string-name>
          , Karen S Quigley, Bradford C Dickerson, and Lisa Feldman Barrett.
          <article-title>Evidence for a large-scale brain system supporting allostasis and interoception in humans</article-title>
          .
          <source>Nature human behaviour</source>
          ,
          <volume>1</volume>
          (
          <issue>5</issue>
          ):
          <fpage>0069</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [Leike et al.,
          <year>2018</year>
          ]
          <string-name>
            <given-names>Jan</given-names>
            <surname>Leike</surname>
          </string-name>
          , David Krueger, Tom Everitt, Miljan Martic, Vishal Maini, and
          <string-name>
            <given-names>Shane</given-names>
            <surname>Legg</surname>
          </string-name>
          .
          <article-title>Scalable agent alignment via reward modeling: a research direction</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          arXiv preprint arXiv:
          <year>1811</year>
          .07871,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <source>[MacLean</source>
          , 1990]
          <string-name>
            <given-names>Paul D</given-names>
            <surname>MacLean</surname>
          </string-name>
          .
          <article-title>The triune brain in evolution: Role in paleocerebral functions</article-title>
          .
          <source>Springer Science &amp; Business Media</source>
          ,
          <year>1990</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <source>[Miller and Clark</source>
          , 2018]
          <string-name>
            <given-names>Mark</given-names>
            <surname>Miller</surname>
          </string-name>
          and
          <string-name>
            <given-names>Andy</given-names>
            <surname>Clark</surname>
          </string-name>
          .
          <article-title>Happily entangled: prediction, emotion, and the embodied mind</article-title>
          .
          <source>Synthese</source>
          ,
          <volume>195</volume>
          (
          <issue>6</issue>
          ):
          <fpage>2559</fpage>
          -
          <lpage>2575</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [Norman and
          <string-name>
            <surname>Bar-Yam</surname>
          </string-name>
          ,
          <year>2018</year>
          ]
          <string-name>
            <given-names>Joseph</given-names>
            <surname>Norman and Yaneer</surname>
          </string-name>
          Bar-Yam.
          <source>Special Operations Forces: A Global Immune System? In International Conference on Complex Systems</source>
          , pages
          <fpage>486</fpage>
          -
          <lpage>498</lpage>
          . Springer,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [Oosterwijk et al.,
          <year>2012</year>
          ]
          <string-name>
            <given-names>Suzanne</given-names>
            <surname>Oosterwijk</surname>
          </string-name>
          , Kristen A Lindquist,
          <string-name>
            <surname>Eric Anderson</surname>
          </string-name>
          , Rebecca Dautoff, Yoshiya Moriguchi, and Lisa Feldman Barrett.
          <article-title>States of mind: Emotions, body feelings, and thoughts share distributed neural networks</article-title>
          .
          <source>NeuroImage</source>
          ,
          <volume>62</volume>
          (
          <issue>3</issue>
          ):
          <fpage>2110</fpage>
          -
          <lpage>2128</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <source>[Schein and Gray</source>
          , 2015]
          <string-name>
            <given-names>Chelsea</given-names>
            <surname>Schein</surname>
          </string-name>
          and
          <string-name>
            <given-names>Kurt</given-names>
            <surname>Gray</surname>
          </string-name>
          .
          <article-title>The unifying moral dyad: Liberals and conservatives share the same harm-based moral template</article-title>
          .
          <source>Personality and Social Psychology Bulletin</source>
          ,
          <volume>41</volume>
          (
          <issue>8</issue>
          ):
          <fpage>1147</fpage>
          -
          <lpage>1163</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          <source>[Schein and Gray</source>
          , 2018]
          <string-name>
            <given-names>Chelsea</given-names>
            <surname>Schein</surname>
          </string-name>
          and
          <string-name>
            <given-names>Kurt</given-names>
            <surname>Gray</surname>
          </string-name>
          .
          <article-title>The theory of dyadic morality: Reinventing moral judgment by redefining harm</article-title>
          .
          <source>Personality and Social Psychology Review</source>
          ,
          <volume>22</volume>
          (
          <issue>1</issue>
          ):
          <fpage>32</fpage>
          -
          <lpage>70</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [Schein et al.,
          <year>2016</year>
          ]
          <string-name>
            <given-names>Chelsea</given-names>
            <surname>Schein</surname>
          </string-name>
          , Neil Hester, and
          <string-name>
            <given-names>Kurt</given-names>
            <surname>Gray</surname>
          </string-name>
          .
          <article-title>The visual guide to morality: Vision as an integrative analogy for moral experience, variability and mechanism</article-title>
          .
          <source>Social and Personality Psychology Compass</source>
          ,
          <volume>10</volume>
          (
          <issue>4</issue>
          ):
          <fpage>231</fpage>
          -
          <lpage>251</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          <source>[Seligman</source>
          , 2012]
          <article-title>Martin EP Seligman</article-title>
          .
          <article-title>Flourish: A visionary new understanding of happiness and well-being</article-title>
          .
          <source>Simon and Schuster</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          <source>[Soares and Fallenstein</source>
          , 2017]
          <string-name>
            <given-names>Nate</given-names>
            <surname>Soares</surname>
          </string-name>
          and
          <string-name>
            <given-names>Benya</given-names>
            <surname>Fallenstein</surname>
          </string-name>
          .
          <article-title>Agent foundations for aligning machine intelligence with human interests: a technical research agenda</article-title>
          .
          <source>In The Technological Singularity</source>
          , pages
          <fpage>103</fpage>
          -
          <lpage>125</lpage>
          . Springer,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [Taylor et al.,
          <year>2016</year>
          ]
          <string-name>
            <given-names>Jessica</given-names>
            <surname>Taylor</surname>
          </string-name>
          , Eliezer Yudkowsky, Patrick LaVictoire, and
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Critch</surname>
          </string-name>
          .
          <article-title>Alignment for advanced machine learning systems</article-title>
          .
          <source>Machine Intelligence Research Institute</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          <source>[Verbeek</source>
          , 2001]
          <string-name>
            <given-names>Bruno</given-names>
            <surname>Verbeek</surname>
          </string-name>
          .
          <article-title>Consequentialism, rationality and the relevant description of outcomes</article-title>
          .
          <source>Economics &amp; Philosophy</source>
          ,
          <volume>17</volume>
          (
          <issue>2</issue>
          ):
          <fpage>181</fpage>
          -
          <lpage>205</lpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [Werkhoven et al.,
          <year>2018</year>
          ]
          <string-name>
            <given-names>Peter</given-names>
            <surname>Werkhoven</surname>
          </string-name>
          , Leon Kester, and
          <string-name>
            <given-names>Mark</given-names>
            <surname>Neerincx</surname>
          </string-name>
          .
          <article-title>Telling autonomous systems what to do</article-title>
          .
          <source>In Proceedings of the 36th European Conference on Cognitive Ergonomics, page 2. ACM</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          <source>[Yudkowsky</source>
          , 2016]
          <string-name>
            <given-names>Eliezer</given-names>
            <surname>Yudkowsky</surname>
          </string-name>
          .
          <article-title>The AI Alignment Problem: Why it is Hard, and Where to Start</article-title>
          .
          <source>Symbolic Systems Distinguished Speaker</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>