<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Multimodal Turing Test for Realistic Humanoid Robots with Embodied Artificial Intelligence</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Carl Strathearn</string-name>
          <email>Carl.Strathearn@research.staffs.ac.uk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Minhua Ma</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Provost, Falmouth University</institution>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Computing and Digital Technologies, Staffordshire University</institution>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2010</year>
      </pub-date>
      <abstract>
        <p>Alan Turing developed the Turing Test as a method to determine whether artificial intelligence (AI) can deceive human interrogators into believing it is sentient by competently answering questions at a confidence rate of 30%+. However, the Turing Test is concerned with natural language processing (NLP) and neglects the significance of appearance, communication and movement. The theoretical proposition at the core of this paper: 'can machines emulate human beings?' is concerned with both functionality and materiality. Many scholars consider the creation of a realistic humanoid robot (RHR) that is perceptually indistinguishable from a human as the apex of humanity's technological capabilities. Nevertheless, no comprehensive development framework exists for engineers to achieve higher modes of human emulation, and no current evaluation method is nuanced enough to detect the causal effects of the Uncanny Valley (UV) effect. The Multimodal Turing Test (MTT) provides such a methodology and offers a foundation for creating higher levels of human likeness in RHRs for enhancing human-robot interaction (HRI)</p>
      </abstract>
      <kwd-group>
        <kwd>Turing Test</kwd>
        <kwd>Humanoid Robots</kwd>
        <kwd>Artificial Intelligence</kwd>
        <kwd>Embodied Artificial Intelligence</kwd>
        <kwd>HRI</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The Turing Test hypothetically evaluated
computational AI using typesetters and pre-written
scripture to emulate human thought (Turing, 1950).
However, modern conversational AI systems function
with greater accuracy at a higher rate of processing than
the analogue methods outlined in Turing’s paper.
Landgrebe &amp; Smith (
        <xref ref-type="bibr" rid="ref5">2019</xref>
        ) explain that unlike the
original Turing Test, the updated Turing Test for AI
utilises two computer interfaces to replace the
typesetter methodology. One computer system implements
a conversational AI application and the other controlled
by a human agent concealed from the view of the
human interrogator. The role of the human interrogator
is to evaluate the authenticity and accuracy of the
agent’s responses to determine which system is
artificial and which is human. There are accounts of AI
systems which claim to have passed the Turing Test.
For example, Warwick &amp; Shah (2015) and Aamoth
(2014), advocate that a chatbot program named Eugene
Goostman passed the 30% benchmark of the Turing
Test in 2014, scoring a marginal 33%, at the Royal
Society AI competition in 2014. However,
commentators such as Copeland (2014), Hern (2014)
and Robbins (2014) contest the validity of this
achievement, stating two significant flaws in the
evaluation procedure. Firstly, human interrogators had
prior knowledge that the AI system emulated a 13-year
old Ukrainian boy. This approach dissolves the
integrity of the Turing Test, which states the removal
of all identifiers is vital in maintaining impartiality
(Turing, 1950). Secondly, the creators of the Eugene
Goostman chatbot hand-selected the human
interrogators for the test, significantly increasing the
probability for participant bias. Sample &amp; Hern (2014)
argue that claiming the Eugene Goostman chatbot
passed the Turing Test is fundamentally absurd as
Turing’s prediction that in 50 years conversational AI
could pass as a human was merely hypothetical, akin to
a statistical survey or Gallup poll. Turing’s acumen is
a methodology to explain how the human mind
functions by developing a computer capable of
proximal behaviour and intelligence, which includes
verbal processing and sensorimotor/robotic dimensions
in which AI is systematically grounded (Sample &amp;
Hern, 2014).
      </p>
      <p>In consideration, Harnad (2000) argues that the Turing
Test is not a measure of how an AI system operates
over five minutes; it is the system’s ability to simulate
the human mind over a lifetime. According to Gehl
(2013), a similar text-based chatbot named Cleverbot
claimed to pass the Turing Test in 2011 at the Technie
festival in India, four years before the Eugene
Goostman chatbot. However, Cleverbot did not receive
the media coverage and scholarly attention of the
Eugene Goostman program due to numerous
irregularities in the results.</p>
      <p>
        Aron (2011), Jacquet et al. (
        <xref ref-type="bibr" rid="ref5">2019</xref>
        ) and Mann (2014)
argue that although Cleverbot claimed to exceed the
30% benchmark of the Turing Test scoring an
exceptional 59.3%, human interrogators rated human
agents as AI at an even higher rate of 63.3%. Thus,
significant discrepancies in the results indicate
fundamental flaws in the evaluation procedure and
recruitment process.
However, Landgrebe &amp; Smith (
        <xref ref-type="bibr" rid="ref5">2019</xref>
        ), Jacquet et al.
(
        <xref ref-type="bibr" rid="ref5">2019</xref>
        ) and Pereira (
        <xref ref-type="bibr" rid="ref5">2019</xref>
        ) argue that although
numerous chatbot systems claim to pass the Turing
Test. The modernised tests are weak variations of
Turing’s original proposition, which are not
representative of Turing’s hypothesis and therefore do
not qualify as certified passes. Fawaz (
        <xref ref-type="bibr" rid="ref5">2019</xref>
        ) and
Wakefield (
        <xref ref-type="bibr" rid="ref5">2019</xref>
        ) explain that creating chatbots to pass
the Turing Test is a developer’s past-time as there is no
serious scientific research in developing AI to pass the
Turing Test. In support, Sharkey (2012) suggest that as
Turing is long deceased, clarifying the terms and
conditions of passing the Turing Test is impossible.
In RHR design, Mori’s (1970) UV accounts for the
negative psychological stimulus propagated by RHRs
upon observation, as the more human-like artificial
humans appear, the greater the potential for humans to
feel repulsed by their appearance. However, per
Burleigh (2013), there are considerable arguments
against the scientific value of the UV theorem, as many
scholars regard it as purely academic. Thus, the UV
like the Turing Test remains a controversial topic in AI
and robotics.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. The Turing Test</title>
      <p>Alan Turing (1950) formulated the Turing Test to
determine if a machine agent could mislead human
interrogators into believing answers provided by a
computer are those of a human. If the machine
convinces 30%+ of human interrogators into thinking
it is sentient, the system passes the test and the higher
this percentage, the more humanistic the AI functions.
Turing argues that if a machine agent is capable of
exhibiting human behaviour indistinguishable to that of
a human, then the artificial mind functions in a manner
akin to the human mind (it can think). However, Turing
questions a machine’s ability to think as ‘thinking’ is
problematic to define and thus proposes the Turing Test
as a methodology to explore this concept.</p>
      <p>Turing supposes that if a machine agent replaced either
of the male or female agents in the imitation game and
could operate with a level of intelligence proximal to
the responses of a human, then it would replace his
original hypothesis ‘can machines think?’ (Turing,
1950).In the Turing Test, the objective of the human
interrogator is to identify which agent is AI and human
by posing a series of questions to evaluate the
authenticity of the responses to differentiate between
the AI agent and human agent. It is the agent’s role to
deceive the human interrogator into believing that they
are the opposite agent by providing type-written
answers that simulate the responses of the other.
However, Turing applies constraints to the Turing Test
to establish equilibrium between the agents. Firstly,
Turing narrows the scope of interaction between the
human interrogator and the human/machine agents to a
single topic of conversation, to prevent the human
interrogator asking questions outside of the scope of the
AI system’s capabilities which may allude to the
artificiality of the system. Similarly, Turing restricts
the human interrogator’s ability to propose
mathematical inquiries to the agents as machine’s are
capable of correctly answering complex equations
consistently, unlike humans.</p>
      <p>Secondly, Turing imposes a 15-30 second time delay
between the responses of the human interrogator as
machine agents require time to formulate and respond
to questions, unlike the human mind to which
responses are immediate. Thirdly, Turing limits the
time-scale of the evaluation to 5 minutes to prevent the
machine agent producing incorrect or repetitive
responses as the longer the interaction, the higher the
potentiality for error. However, Turing considers the
physical emulation of the human being as a distraction
from the pursuit of intelligent machine’s (Turing, 1950,
p.2). Although Turing is correct in stating that the
appearance of a machine is not indicative of its
intellectual capabilities, he neglects the capacity of the
human body in tactile learning, socialisation and
nonverbal communication which are vital processes in
social learning and communication.</p>
    </sec>
    <sec id="sec-3">
      <title>2.1 Arguments and Limitations of the Turing Test</title>
      <p>In Searl’s experiment, a human agent sat in the middle
of a room is passed a series of random Chinese symbols
from under a door. The agent uses an instruction
manual to arrange the symbols to form coherent
sentences. After a while, the agent becomes efficient in
arranging the symbols into sentences and no longer
requires the instruction manual. The instruction manual
is removed and interrogators who are fluent in written
Chinese observe the agent arrange the symbols into
sentences and state whether they think the agent is
literate in Chinese or not.</p>
      <p>
        In the experiment, the interrogators agree that the agent
is fluent in Chinese to form coherent sentences using
the symbols. However, the agent only understands the
order of the symbols and not their meaning and
therefore, lacks the vital process of comprehension.
Thus, the perception of the interrogators in Searl’s
experiment is critical in understanding how humans
interpret the appearance of intelligent behaviour as in
real-life conditions; there are no visual distinctions
between functional intelligence and comprehension,
visualised in Fig.1
Similarly, Cole (
        <xref ref-type="bibr" rid="ref5">2019</xref>
        ) and Warwick &amp; Shah (2015)
argue that the Turing Test is susceptible to human
interference by a fundamental design flaw which
inverts the human perception of the nature of
computing by remaining silent. Ghose (
        <xref ref-type="bibr" rid="ref11 ref8">2016</xref>
        ) explains
that if an AI system does not answer questions when
prompted, the human interrogators cannot distinguish
between the silence of the AI and human responses;
hence, the AI agent would pass as human by default.
Thus, it is the expectancy for a computer system to
respond to the actions of a human operator. If a
computer does not perform tasks in a manner
accustomed in HCI, this processual irregularity has the
potentiality to influence human perception of the nature
of the agent (Reynolds, 2016). In consideration, Hern
(
        <xref ref-type="bibr" rid="ref5">2019</xref>
        ) and Landgrebe &amp; Smith (
        <xref ref-type="bibr" rid="ref5">2019</xref>
        ), suggest that
silence during the Turing Test is not uncommon and
typically the result of poor programming.
      </p>
      <p>However, stricter policies regarding the time limit of
agent responses are crucial in maintaining the integrity
of the Turing Test to irradicate purposeful exploitation
of this loophole. Whitby (1996), argues that AI
developer’s and scholars have long misinterpreted the
purpose of the Turing Test as Alan Turing designed the
‘Imitation Game’ as a game and not a formal test.
Whitby argues that Alan Turing never intended the
imitation game as an evaluation of machine
intelligence, but rather as a thought experiment for
assessing a machine’s capacity to portray the
behaviours of a human authentically. Whitby suggests
that Turing’s paper is not an operational guide for AI,
but a theoretical treatise to examine the sociological
and scientific value of creating machine’s which can
mislead human beings into believing they are human.
However, Whitby explains that simulating human
personalities and emotion in AI is damaging as these
attributes tend to be misleading rather than progress the
intellectual capacity of AI. Thus, the practical value of
Turing’s hypothesis is not in creating machine’s with
intelligence proximal to humans known as artificial
general intelligence (AGI), but in emulating the
conditions of the human mind and behaviours using
computers.</p>
      <p>
        This concept is significant in HRI and HCI as it
considers how humans interface and interact with
technologies that simulate human intelligence,
personalities and behaviour. Rapaport (2000) argues
that the Turing Test is limited in its scope of evaluation
as it only considers HCI via NLP. Stock-Homburg et
al. (2020) describe the Handshake Turing Test (HTT)
and similarly, Karniel et al. (2010) the Turing
Handshake Test (THT) as tests to determine if human
interrogators can identify the differences between a
human and RHR by the act of a handshake (tactile
HRI). Moreover, this approach neglects the emulation
of appearance, communication, AI and movement by
focusing on secondary aspects such as touch and
temperature. Ishiguro (2005) developed the Total
Turing Test (TTT) for RHRs in HRI, formulated on
Harnad’s (1992) TTT for human-computer interaction
and Harnad’s (2000) Robot Turing Test (RTT) to
comprehensively evaluate the appearance, behaviour
and movement of RHRs against a human counterpart.
Ishiguro’s (2005) TTT implements point of view
(POV) cameras mounted on the heads of the human and
RHR agents. The agents conduct logistical tasks, and it
is the role of the human interrogator to discern which
agent is human and RHR from observation. Secondly,
the human interrogator observes live ‘full body’ video
streams of the agents for two seconds and decides
which agent is human and RHR. Kasaki et al. (
        <xref ref-type="bibr" rid="ref11 ref8">2016</xref>
        )
cite 70% of subjects identified the movements of RHRs
as human. Ishiguro argues that the Turing Test
evaluates the intellectual capabilities of a computer on
the assumption that the human mind is divisible from
the body.
      </p>
      <p>Thus, the TTT evaluates embodied artificial
intelligence (EAI) by combining intelligent behaviour
with a robotic body for assessing the human likeness of
robotic behaviour, appearance and movement.
However, the TTT is susceptible to design flaws;
Firstly, live video footage is inaccessible. Secondly,
Marzano &amp; Novembre (2017) argue that the 2-second
evaluation window is too limited. Thirdly, according to
Schweizer (1998) &amp; Bringsjord et al. (2000), the TTT
is not a comprehensive approach as it neglects the
evaluation of NLP to robotic mouth articulation during
HRI. Fourthly, Oppy (2003) stipulates that judging the
authenticity of intelligent behaviour by manipulating
objects is not indicative of a machine’s intellectual
capacity. In consideration, Schweizer (1998) created
the Truly Total Turing Test (TTTT) to remove
telepresence from the TTT and evaluate automated
RHR’s with EAI. However, the TTTT lacks vital
processes such as physical examination, movement,
appearance, materiality, EAI and communication when
operating as one robotic system.</p>
    </sec>
    <sec id="sec-4">
      <title>3. The Multimodal Turing Test</title>
      <p>
        Per the findings of the literature review, current
evaluation methods used to determine degrees of
human likeness in RHRs in HRI and HCI, such as The
Turing Test, TTT, TTTT, RTT, THT and HTT are too
limited in their scope of evaluation as they neglect the
significance of amalgamating; communication (speech
and gesturing), movement, vision, aesthetics and
conversational AI into a single system, which is not
representative of the human condition. In
consideration, this study lays the foundations of a
comprehensive theoretical evaluation methodology
named the Multimodal Turing Test (MTT) to
determine if RHRs can attain a level of emulation
perceptually indivisible from a human being, (Houser,
2019). As cited in a recent article in the Guardian UK,
the MTT is more holistic than the original Turing Test,
and previous evaluation methods in HRI by evaluating
an RHRs appearance, communication, movement and
AI (Mathieson, 2019), shown in Fig. 2.
The MTT incorporates the examination structure of the
1950 Turing Test by employing human interrogators to
evaluate the perceptual authenticity of RHRs.
However, unlike the binary pass / fail system of the
original Turing Test, the MTT provides engineers,
designers and programmers with a developmental
framework to benchmark progress up to and in advance
of Turing’s 30% pass rate (Strathearn, 2019). Each
stage of the MTT increases in complexity, which forms
the hierarchy of human emulation shown in Fig. 3. Like
Turing, it is not argued that an RHR metamorphosis
into an organic system by replicating the conditions of
a human being. However, if an RHR can appear and
function in a manner indistinguishable from a human
being in real-world conditions, then that RHR is
perceptually indivisible from a living human being,
The World Economic Forum (
        <xref ref-type="bibr" rid="ref5">2019</xref>
        ). Thus, equal
consideration to the appearance and functionality of
RHRs is essential to develop higher modes of human
emulation.
      </p>
      <p>However, replicating the appearance and materiality of
a human is more straightforward than simulating
human movement due to the complexity of natural
kinetic variance. Therefore, per Baudrillard’s (1994)
order of simulacra, appearance forms the bedrock of
the hierarchy of human emulation because it is the
elementary form of simulation. Aesthetical appearance
envelops a body to which movement is applied, as
natural movement is more complicated to replicate than
a still model; kinetics forms the second level of
emulation. For speech to be a useful communication
tool in RHRs, requires both an authentic appearance
and naturalistic. AI is the apex of human emulation as
the human mind is the most challenging element to
simulate authentically due to its complexity. However,
for AI to be a useful tool in RHRs, the emulated mind
requires a human-like body and a method of
communication for naturalistic HRI.</p>
      <p>
        The four evaluation categories of the hierarchy of
human emulation formulate a unified whole, which
constitutes an RHR that can emulate (to degrees of
likeness) a living human being, as reviewed in an
article by Khatib (
        <xref ref-type="bibr" rid="ref5">2019</xref>
        ) which outlines the scope of the
MTT. Furthermore, the MTT is an approach towards
humanising forms of AI as current robotic AI
predominantly focuses on logical, linguistical and
kinesthetic intelligence and neglects interpersonal and
intrapersonal intelligence to create higher modes of
EAI. Interpersonal and intrapersonal AI is synergetic,
incorporating various visual and audible stimuli such
as facial expressions, vocal tonality, gesturing, and
emotive responsivity to humanise AI interaction. This
approach enhances the capacity for natural
communication and responsivity between humans, and
RHRs founded on authentically assimilating natural
human-human interaction, (Barnfield, 2020).
Previous evaluation methods fall into the MTTs
categories of human emulation, but none are inclusive
of all four stages of development. For example, The
THT and HTT, in movement (handgrip), the TTT falls
under appearance, AI: Wizard of Oz (WOZ) method
and kinetics (robotic vision, aesthetics and movement),
the Turing Test in AI in (AI) and the TTTT in
appearance, movement and AI. However, developing
an RHR as a complete system with components across
all four categories of the hierarchy of human emulation
(without consideration of the stages) will not achieve
levels of human likeness indivisible from a human
being. For example, comparing two RHR heads to
determine which one is more visually authentic than the
other is a viable methodology for evaluating and testing
new components by increasing the realism of one
robotic head over the other.
      </p>
      <p>However, this approach is futile when comparing
RHRs against a living human being to determine
authenticity as the distinctions in form and function are
highly apparent, as exemplified in Fig. 4.
Therefore, a multimodal approach is required using a
controlled evaluation methodology by combining
features that belong to the same body (subgroup), such
as, EAI, natural speech synthesis and a robotic jaw,
tongue and lips. This evaluation procedure applies to
other subgroups such as eyes: (sclera, pupil dilation,
iris, eyelid, eyelashes, veins, eye movement, blink rate,
skin, hair, aesthetics) and so on. This approach is
similar to the functional constraints of the original
Turing Test to control the direction and flow of a
conversation by narrowing it to a specified theme or
topic of discussion. This technique permits the
refinement of smaller intricate motor functions and
aesthetics within the subgroups, indicated in Fig. 5.
The MTT is a method for overcoming many of the
design issues that are prevalent in RHRs such as
inaccurate eye emulation, poor aesthetical design and
unnatural movement. Furthermore, according to the
uncanny valley hypothesis, realistic humanoids
instigate negative perceptual feedback in humans
because they are void of variable organic nuances.
This consideration is vital in the development and
progression of modern RHRs, as traditional methods of
evaluation and design overlook the significance of
replicating nuances such as pupil dilation, gestures and
accurate lip movement. These facial expressions act as
visual cues and signifiers of sentience when discerning
the authenticity of an RHR.</p>
      <p>Thus, when evaluating an RHR, all elements are
interconnected to the perceptual whole. To achieve
this, an imitation head structure and cloaking device to
cover empty areas around the developed feature is a
practical method of resolving this issue. This approach
permits a holistic evaluation compared to analysing
individual facial features outside of the body (unified
whole).</p>
      <p>The Multimodal Turing Test: three orders of human
emulation: The three orders of human emulation are a
framework for developing RHRs that appear and
function in a manner that is indistinguishable from the
natural human being under the conditions and
limitations of the MTT evaluation procedure.
1. Fragmentary Emulation: A unified subgroup that
qualifies as perceptually indistinguishable in form / and
or function when compared to a human.
2. Synchronised Emulation: A set of two or more
subgroups that are perceptually indivisible in form /
and or function from a living human being.
3. Absolute Emulation: A fully assembled human
replicant consisting of all subgroups working as a
unified whole to emulate the human form and function.
The total length of the MTT is 20 minutes and divided
into four 5-minute evaluation sections, covering:
appearance, movement, voice and AI founded on the
five-minute evaluation rule of the original Turing test.
The MTT has broader applications outside the field of
RHRs and EAI in realistic virtual humanoids (RVHs)
with EAI for HCI. Developing higher modes of human
likeness in RVHs is significant in EAI interface design
for HCI and exploring the UV in RVHs. Therefore, it
is essential to provide evaluation conditions for
assessing the perceptual authenticity of RVHs for the
future progression of virtual humanoids towards a
simulacrum indivisible from living humans.</p>
    </sec>
    <sec id="sec-5">
      <title>4. The Multimodal Turing Test for RHRs</title>
    </sec>
    <sec id="sec-6">
      <title>4.2. Second Stage: Movement and Dexterity</title>
      <p>The MTT is more comprehensive than the Turing Test,
TTT, RTT, THT and HTT by systematically examining
appearance, functionality, AI and voice processing to
provide a universal evaluation procedure for all types
of humanoid robots with varying degrees of human
likeness. This multimodality requires several
constraints to ensure the integrity of the evaluation
procedure. In Fig. 6, the human Interrogator (A)
evaluates the authenticity of agents (B) and (C) who are
separated by a solid screen to minimise interference.
Significantly, both agents (B) and (C) inhabit the same
physical environment and visual spectrum as the
human interrogator for greater perceptual authenticity.</p>
    </sec>
    <sec id="sec-7">
      <title>4.1 First Stage: Appearance</title>
      <p>The first stage of the MTT requires human interrogator
(A) to evaluate the appearance of agents (B) and (C).
Different subgroups contain different visual elements
such as lips, hair, skin tone and wrinkles. Therefore,
imperfections in synthetic skin such as wrinkles, spots
and blemishes are essential as these defects are not
typically associated with RHRs. The first level
examines the visual authenticity of the agents, such as
an area of natural skin of Agent (B) with the
corresponding synthetic skin area from Agent (C). The
MTT is significant to the progression of RHRs as the
Turing Test does not provide a developmental
framework due to the binary pass/fail system.
Thus, allowing engineers to gauge the authenticity of
specific facial/bodily areas individually, as a group, or
as a complete form towards attaining the pass threshold
(emulation that is indivisible from a living human) is
essential. It is crucial to evaluate Agent (B) against (C)
and then Agent (C) against (B) for a detailed and
comprehensive analysis. For example, imagine Agent
(B) is a robotic mouth and (C) a human mouth, and the
human interrogator (A) identifies a visual irregularity
in the bottom lip of Agent (B) leading to the human
interrogator identifying Agent (B) as an RHR. This
process applies to every item within a subgroup to
pinpoint the precise location of the visual irregularity.
It is vital to access the aesthetical quality of the inside
of the robotic mouth during the first stage evaluation as
this area is exposed during operation.</p>
      <p>The second stage of the MTT incorporates both
movement and appearance; The human interrogator
(A) selects an expression or gesture from a list of
commands, such as smile, frown, wave, open mouth.
The Human interrogator (A) selects which agent
performs the command by addressing the agent and
saying aloud the command. As in the Turing Test, a
delay in the response time (5-10s) of the agents allows
time for NLP. Servomotor sounds must be triggered by
the human agent when performing physical movements
to reduce signifiers such as sound interference that may
allude to the mechanical nature of the RHR. It is
essential to assess tongue movement to match vowel
and consonant sound as the internal components are
exposed by the robotic mouth during verbal
communication. The accurate replication of acute
motor functions such as pupil dilation, breathing, facial
tics and blink rate must be considered in the second
stage. Furthermore, the complexity and level of
movement are variable on the style of the humanoid
robot; for instance, robotic heads do not require the
evaluation of body movement such as hand gesturing.
However, evaluating hand gestures is essential for a
‘waist up’ robot design. Comparatively, a waist up
robot does not require the evaluation of leg movement
and balance, unlike a full-body humanoid robot which
needs the robot to stand and move the lower parts of its
body. Therefore, applying constraints to control the
evaluation area for different styles of RHRs is
significant, for example; seating robotic heads and
waist-up robots and at a table during the evaluation
procedure will reduce and concentrate the evaluation
area. This method is standard in HRI to conceal an
RHRs lower body and external mechanical
components from the observer. If an RHR can pass the
first two stages of the MTT at a rate of 30%+, is the
same as saying in real-world conditions, an RHR is
visually indistinguishable from a living human being
(without speaking or AI interaction).</p>
    </sec>
    <sec id="sec-8">
      <title>4.3 Third Stage: Speech and Mouth Articulation</title>
      <p>The third stage of the MTT evaluates an RHRs speech,
lip dexterity and aesthetical appearance. It is not the
objective of the MTT to develop a more
humansounding robotic voice as this field is continually
evolving outside of RHR design. However, the MTT
examines the compatibility and accuracy of speech
synthesis with robotic mouth articulation. Speech
synthesis technologies are advancing rapidly and
continually improving in human likeness, and the use
of current and future speech synthesis technologies in
RHRs is significant towards total automation.
Using NLP in the MTT is preferable to human speech
as it protects the integrity of the test environment by
seamlessly interchanging between the previous
evaluation stages. However, as speech synthesis is yet
to replicate human speech, implementing current
speech synthesis is counterproductive when developing
RHRs that are perceptually indivisible from humans.
Therefore, it is essential to outline an alternative
methodology of natural speech processing to overcome
the current limitations of computerised speech
technologies. The WOZ approach permits a second
human agent (D) to speak in place of the robotic voice,
as demonstrated in Fig. 7. The speech of Agents (D)
and (C) are relayed to the human interrogator (A) by
headphones to minimise the sound difference between
the speaker system and natural human voice.
This approach permits the examination of human
speech using a robotic mouth system, allowing for a
greater accurate comparative evaluation than current
speech synthesis. However, real-time human speech to
lip synchronisation is less reliable than speech
synthesis due to the variability in pitch, volume,
frequency and tonality of human speech. Therefore, it
is essential to configure the robotic mouth to function
with one human voice for optimum lip-synchronisation
accuracy. Although the evaluation for natural human
speech and computerised speech is different, the
procedure is identical. The human interrogator (A)
engages in an interactive game with agents (B+D) and
(C). The objective of the game is for the human
interrogator (A) to guess what animals that agents
(B+D) and (C) are thinking of by posing questions to
each of them about the animal’s appearance, habitat,
movement and diet. The human interrogator (A) rates
and compares the authenticity of Agents (B) and (C)
voice and mouth articulation. This approach is vital for
evaluating speech, as implementing a structured
gamification methodology does not require deep
learning or machine learning methods and permits the
human interrogator to focus on speech quality rather
than correct or incorrect AI responses. Finally, time
limitations on ‘silence’ are significant to upholding the
integrity of the MTT as suggested in an article on the
MTT and the Turing Test, (Cole, 2019).</p>
      <p>Therefore, a time limitation of 10 seconds is imposed
and strictly monitored throughout the evaluation
procedure, with time added to the end of each session
if silence is excessive or exceeds the 10-second
maxima. If an RHR can pass the third stage of the
MTT, then that systems autonomous speech processing
and tonal expressions are proximal to natural human
speech and mouth/lip movement, facial expressions
and appearance. However, for an RHR to progress to
the final stage (AI) of the hierarchy of human
emulation, the system must be fully automated without
human control for the integration of speech and AI.
Therefore, implementing the alternate speech
evaluation procedure is an acceptable method for
passing the third level of the MTT but not for
progressing onto the final level.</p>
    </sec>
    <sec id="sec-9">
      <title>4.4 Final Stage: AI (Absolute Emulation)</title>
      <p>The final stage of the MTT is inclusive of all four
elements: intelligence, movement, speech and
appearance. It is vital at this stage that all human
control is removed, permitting the RHR to function
autonomously and the AI to control the operations of
movement and speech. As EAI constitutes the
‘personality’ of the RHR, developer’s need to create an
AI people personality with interests and traits that
match the appearance, speech synthesis and movement
of the RHR. Passing the final stage of the MTT would
answer the question: can machines emulate a human
being? Therefore, developing an EAI program to
control accurately trigger facial expressions, voice
tone, emotions and gestures are crucial in the final
evaluation. This method is the foundation for
developing more sophisticated modes of interpersonal
AI for robots. Like the Turing Test, the final stage
evaluation focuses on a single topic of discussion
selected by the human interrogator from a
preestablished list of subjects. The final test lasts 5 minutes
with the human interrogator (A) posing 2.5 minutes of
questioning to agents (B) and (C) on the selected topic.
As technology improves NLP, RHR and AI efficiency,
this time limit should be extended until the RHR can
deceive a human interrogator indefinitely. At the end
of the evaluation procedure, the human interrogator (A)
chooses which agent (B) or (C) is human (or unsure)
and provide a detailed account of the decision-making
process covering all evaluation categories. If 30%+ of
test subjects misidentify or are unable to discern the
difference between the RHR and the human agent, then
the RHR has succeeded in passing the final stage of the
MTT However, if an RHR does not pass all stages of
the MTT, the data gathered during the test stages will
provide engineers with information concerning specific
area/s that emit irregular feedback through the layered
evaluation process for revision or calibration.</p>
    </sec>
    <sec id="sec-10">
      <title>5. Conclusion</title>
      <p>The MTT is an essential evaluation method towards
achieving higher modes of human likeness in RHRs
and EAI as in other methods of evaluation; slight
miscalculations of an otherwise realistic-looking robot
can allude to the robot’s artificiality resulting in other
high-quality components becoming part of that failure.
The objective of the MTT is to permit engineers to
work systematically and build up areas of the face and
body to ensure all components are equal to that of a
human before expanding the fields and adding more
features towards creating a complete RHR that is
perceptually indivisible from a living human being.</p>
      <p>Retrieved:</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>AAmoth. D</surname>
          </string-name>
          (
          <year>2014</year>
          )
          <article-title>The Fake Kid Who Passed the Turing Test</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <article-title>Ret:time.com/2847900/eugene-goostman-turing-test/</article-title>
          .
          <source>Acc 5.2</source>
          .20 Aron. J (
          <year>2011</year>
          )
          <article-title>AI tricks people into thinking it is human. Retrieved: newscientist</article-title>
          .com/article/dn20865-software
          <article-title>-tricks-people-into-thinking-itishuman/#ixzz6Ez0JgQi1</article-title>
          . Acc:
          <volume>25</volume>
          .
          <fpage>02</fpage>
          .20 Barnfield. N (
          <year>2020</year>
          )
          <article-title>Face to Face With The Future of AI. Horizon Magazine. Riley Raven</article-title>
          . DOI: https://www.staffs.ac.uk/alumni/ horizonalumni-magazine pp.
          <fpage>8</fpage>
          -
          <lpage>11</lpage>
          Baudrillard. J (
          <year>1994</year>
          ).
          <article-title>Simulacra and simulation</article-title>
          . Trans: Ann Arbor : University of Michigan Press, ISBN-
          <volume>10</volume>
          :
          <fpage>0472065211</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Bringsjord</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Caporale</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Noel</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>2000</year>
          ).
          <article-title>The Total Turing Test</article-title>
          , JLLI,
          <volume>9</volume>
          (
          <issue>4</issue>
          ),
          <fpage>397</fpage>
          -
          <lpage>418</lpage>
          . DOI:www.jstor.org/stable/40180234 Burleigh. T, Schoenherr. J,
          <string-name>
            <surname>Lacroix</surname>
          </string-name>
          . G (
          <year>2013</year>
          ).
          <article-title>Does the uncanny valley exist? Computers in Human Behaviour</article-title>
          .
          <volume>29</volume>
          .
          <fpage>759</fpage>
          -
          <lpage>771</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <source>DOI:10</source>
          .1016/j.chb.
          <year>2012</year>
          .
          <volume>11</volume>
          .021.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Cole</surname>
          </string-name>
          . E (
          <year>2019</year>
          )
          <article-title>What is New in Robotics? Retrieved: blog</article-title>
          .robotiq.com/whats-new-in-robotics-
          <volume>06</volume>
          .
          <fpage>12</fpage>
          .
          <year>2019</year>
          . A:
          <volume>25</volume>
          .
          <fpage>02</fpage>
          .20 Copeland. J (
          <year>2014</year>
          )
          <article-title>Why Eugene Goostman Did Not Pass the Turing Test</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Retrieved</surname>
          </string-name>
          : https://www.huffingtonpost.co.uk/jackcopeland/turingtesteugenegoostman. Acc:
          <volume>25</volume>
          .
          <fpage>02</fpage>
          .20 Fawaz. A (
          <year>2019</year>
          )
          <article-title>A tangible Turing Test</article-title>
          . Retrieved: https://www.neowin.net/news/a
          <article-title>-tangible-turing-testthe-loebner-prizeiscoming-to-swansea-this-weekend/</article-title>
          . Acc:
          <volume>21</volume>
          .
          <fpage>02</fpage>
          .20 Gehl. R (
          <year>2014</year>
          ).
          <article-title>Teaching to the Turing Test with Cleverbot</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <source>Transformations: The Journal of Inclusive Scholarship and Pedagogy</source>
          ,
          <volume>24</volume>
          (
          <issue>1-2</issue>
          ),
          <fpage>56</fpage>
          -
          <lpage>66</lpage>
          . Retrieved February 25,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Ghose</surname>
          </string-name>
          . T (
          <year>2016</year>
          )
          <article-title>Robots Could Hack Turing Test by Keeping Silent</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <article-title>Retrieved: www.scientificamerican.com/article/robots-could-hack-turingtest-by-keeping-silent/</article-title>
          . Acc:
          <volume>25</volume>
          .
          <fpage>02</fpage>
          .20 Harnad,
          <string-name>
            <surname>S.</surname>
          </string-name>
          (
          <year>1992</year>
          )
          <article-title>The Turing Test Is Not A Trick: Turing Indistinguishability Is A Scientific Criterion</article-title>
          .
          <source>SIGART Bulletin</source>
          <volume>3</volume>
          (
          <issue>4</issue>
          ) (
          <year>October 1992</year>
          ) pp.
          <fpage>9</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Harnad</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2000</year>
          )
          <article-title>Minds Machine's and Turing</article-title>
          ,
          <source>Journal of Logic, Language and Information</source>
          , vol
          <volume>9</volume>
          : p.
          <fpage>425</fpage>
          . https://doi.or g/10.1023/A:100 8315308862
          <string-name>
            <surname>Hern</surname>
          </string-name>
          . A (
          <year>2014</year>
          )
          <article-title>What is the Turing Test? Retrieved: www</article-title>
          .theguardian.com/technology/2014/jun/09/what
          <article-title>-is-the-alan-turingtest</article-title>
          .
          <source>Acc: 25.02</source>
          .20 Houser. K (
          <year>2019</year>
          )
          <article-title>Advanced Robotics Forced Scientist To Invent A New Turing Test</article-title>
          . Retrieved: https://futurism.com/the-byte/
          <article-title>scientists-inventednew-turing-test</article-title>
          .
          <source>Acc: 18.04</source>
          .20 Ishiguro,
          <string-name>
            <surname>H.</surname>
          </string-name>
          (
          <year>2005</year>
          ).
          <article-title>Android science: Toward a new cross-interdisciplinary framework</article-title>
          .
          <source>J-Comp-Sci</source>
          ,
          <string-name>
            <surname>Corpus</surname>
            <given-names>ID</given-names>
          </string-name>
          : 6105971Jacquet,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Baratgin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            , &amp;
            <surname>Jamet</surname>
          </string-name>
          ,
          <string-name>
            <surname>F.</surname>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>Cooperation in Online Conversations</article-title>
          .
          <source>Journal of Psychology</source>
          ,
          <volume>10</volume>
          , 727. https://doi.org/10.3389/fpsyg.
          <year>2019</year>
          .00727 Kasaki,
          <string-name>
            <given-names>M.</given-names>
            &amp;
            <surname>Ishiguro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            &amp;
            <surname>Asada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            &amp;
            <surname>Osaka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            &amp;
            <surname>Fujikado</surname>
          </string-name>
          ,
          <string-name>
            <surname>T.</surname>
          </string-name>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          (
          <year>2016</year>
          ).
          <article-title>Cognitive neuroscience Robotics: Synthetic Approaches to human understanding</article-title>
          .
          <volume>10</volume>
          .1007/978-4-
          <fpage>431</fpage>
          -54595.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Khatib. H</surname>
          </string-name>
          (
          <year>2019</year>
          )
          <article-title>Just because they are robots? Retrieved: www</article-title>
          .ameinfo.com/industry/technology/robots
          <article-title>-treat-humanoids-racialgender-bias</article-title>
          <source>Acc: 15.09</source>
          .19 Landgrebe.
          <string-name>
            <given-names>J</given-names>
            <surname>Smith. B</surname>
          </string-name>
          (
          <year>2019</year>
          ) There is https://arxiv.org/abs/
          <year>1906</year>
          .05833. Acc:
          <volume>25</volume>
          .
          <fpage>02</fpage>
          .20 Mann. A (
          <year>2014</year>
          )
          <article-title>The computer actually got an F on the Turing Test</article-title>
          . Ret: wired.com/
          <year>2014</year>
          /06/turing-test
          <article-title>-not-so-fast/</article-title>
          .
          <source>Acc:9.2</source>
          .20 Marzano,
          <string-name>
            <surname>G</surname>
          </string-name>
          &amp; Novembre,
          <string-name>
            <surname>A.</surname>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>Machine's that Dream: A New Challenge in Behavioral-Basic Robotics</article-title>
          .
          <source>Procedia Computer Science</source>
          .
          <volume>104</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          146-
          <fpage>151</fpage>
          .
          <fpage>10</fpage>
          .1016/j.procs.
          <year>2017</year>
          .
          <volume>01</volume>
          .089.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Mathieson. S</surname>
          </string-name>
          (
          <year>2019</year>
          )
          <article-title>Will androids ever be able to convince people they are human? Retrieved: www</article-title>
          .researchgate.net /publication/341756234_Mr_
          <article-title>Robot_Will_androids_ever_be_able_to_con vince_people_they_are_human_</article-title>
          <source>Guardian_Online. Acc</source>
          <volume>07</volume>
          .
          <fpage>07</fpage>
          .20 Mori. M (
          <year>1970</year>
          ).
          <source>The Uncanny Valley. Energy, Issue</source>
          <volume>7</volume>
          , pp.
          <fpage>33</fpage>
          -
          <lpage>35</lpage>
          . DOI:
          <volume>10</volume>
          .1109/MRA.
          <year>2012</year>
          .2192811 Oppy,
          <string-name>
            <given-names>G. R.</given-names>
            , &amp;
            <surname>Dowe</surname>
          </string-name>
          ,
          <string-name>
            <surname>D. L.</surname>
          </string-name>
          (
          <year>2003</year>
          ).
          <article-title>The Turing Test</article-title>
          .
          <source>Stanford Encyclopedia of Philosophy</source>
          ,
          <volume>1</volume>
          (
          <issue>online</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>26</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Pereira. D</surname>
          </string-name>
          (
          <year>2019</year>
          )
          <article-title>You should fear Super Stupidity, not ASI Retrieved:towardsdatascience.com/you-should-fear-super-stupidity-notsuper-intelligence-19f93a46fa4d</article-title>
          . Acc.
          <volume>17</volume>
          .02.20 Rapaport,
          <string-name>
            <surname>W.</surname>
          </string-name>
          (
          <year>2000</year>
          ).
          <article-title>How to Pass a Turing Test</article-title>
          .
          <source>Journal of Logic, Language, and Information</source>
          ,
          <volume>9</volume>
          (
          <issue>4</issue>
          ),
          <fpage>467</fpage>
          -
          <lpage>490</lpage>
          . Retrieved: www.jstor.org/stable/40180238. Acc:
          <volume>24</volume>
          .04 2020
          <string-name>
            <surname>Reynolds</surname>
          </string-name>
          . E (
          <year>2016</year>
          )
          <article-title>Does the Fifth Amendment 'expose a serious flaw' in Turing Test? Retrieved:www</article-title>
          .wired.co.uk/article/major-flaw
          <article-title>-turing-testsilence</article-title>
          .
          <source>Acc:25.0 2</source>
          .20 Robbins. M (
          <year>2014</year>
          )
          <article-title>A Machine Did not 'Pass' the Turing Test</article-title>
          . Retrieved: https://www.vice.com/en_uk/article/gq8ddw/eugene
          <article-title>-goostman-alanturing-test-kevin-warwick</article-title>
          .
          <source>Acc: 25.02.20 Sample. I &amp; Hern. A (</source>
          <year>2014</year>
          )
          <article-title>Scientists dispute if 'Eugene Goostman' passed Turing Test</article-title>
          . Retrieved: www.theguardian.com/techno logy/2014/jun/09/scientistsdisagree-over
          <article-title>-whether-Turing-test-has-beenpassed</article-title>
          .
          <source>Acc</source>
          .
          <volume>19</volume>
          .04.20 Schweizer. P (
          <year>1998</year>
          ).
          <article-title>The Truly Total Turing Test</article-title>
          .
          <source>Minds Mach. 8</source>
          ,
          <issue>2</issue>
          (May
          <year>1998</year>
          ),
          <fpage>263</fpage>
          -
          <lpage>272</lpage>
          . DOI:
          <volume>10</volume>
          .1023/A:1008229619541 Searle,
          <string-name>
            <surname>J</surname>
          </string-name>
          (
          <year>1980</year>
          ).
          <article-title>Minds, brains, and programs</article-title>
          .
          <source>Behavioural and Brain Sciences</source>
          <volume>3</volume>
          (
          <issue>3</issue>
          ):
          <fpage>417</fpage>
          -
          <lpage>457</lpage>
          , DOI: 10.1.1.83.5248.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Sharkey. N</surname>
          </string-name>
          (
          <year>2012</year>
          )
          <article-title>Alan Turing: The experiment that shaped AI</article-title>
          .
          <article-title>Retrieved: bbc</article-title>
          .co.uk/news/technology-18475646. Acc:
          <volume>22</volume>
          .
          <fpage>02</fpage>
          .20 Stock-Homburg. R,
          <string-name>
            <surname>Peters</surname>
          </string-name>
          . J,
          <string-name>
            <surname>Schneider</surname>
          </string-name>
          . K,
          <string-name>
            <surname>Prasad</surname>
          </string-name>
          . V,
          <string-name>
            <surname>Nukovic</surname>
          </string-name>
          . L (
          <year>2020</year>
          )
          <article-title>Evaluation of the HTT for anthropomorphic Robots</article-title>
          ,
          <source>Int Conf HRI. DOI: 10.1145/3371382</source>
          .3378260 The World Economic Forum (
          <year>2019</year>
          )
          <article-title>Can machine's think? A new Turing Test may have the answer</article-title>
          . Retrieved: www.weforum.org/agenda/2019/08/our-turing
          <article-title>-test-for-androids-willjudge-how-lifelike-humanoid-robots-can-be/</article-title>
          . Acc:
          <volume>14</volume>
          .
          <fpage>03</fpage>
          .20 Turing,
          <string-name>
            <surname>A.</surname>
          </string-name>
          (
          <year>1950</year>
          ).
          <source>Computing Machinery and Intelligence</source>
          , Mind, (
          <volume>236</volume>
          ), pp.
          <fpage>433</fpage>
          -
          <lpage>460</lpage>
          . doi.org/10.1093/mind/LIX.236.433.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Wakefield. J</surname>
          </string-name>
          (
          <year>2019</year>
          )
          <article-title>The hobbyists competing to make AI human</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <article-title>Retrieved:www</article-title>
          .bbc.co.uk/news/technology-49578503
          <source>Acc: 25.02</source>
          .20 Warwick,
          <string-name>
            <given-names>K.</given-names>
            , &amp;
            <surname>Shah</surname>
          </string-name>
          ,
          <string-name>
            <surname>H.</surname>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>Passing the Turing Test Does Not Mean the End of Humanity</article-title>
          .
          <source>Cognitive Computation</source>
          ,
          <volume>8</volume>
          ,
          <fpage>409</fpage>
          -
          <lpage>419</lpage>
          . DOI:
          <volume>10</volume>
          .1007/s12559-015-9372-6
          <string-name>
            <surname>Whitby</surname>
            ,
            <given-names>B</given-names>
          </string-name>
          (
          <year>1996</year>
          )
          <article-title>Reflections on AI: the legal, moral and ethical dimensions</article-title>
          .
          <source>Intellect</source>
          , Oxford, UK. ISBN 9781871516685
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>