<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Needs-Driven Cognitive Architecture for Future 'Intelligent' Communicative Agents</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Roger K. Moore Dept. Computer Science, University of Sheffield</institution>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2016</year>
      </pub-date>
      <fpage>50</fpage>
      <lpage>51</lpage>
      <abstract>
        <p>-Recent years have seen considerable progress in the deployment of 'intelligent' communicative agents such as Apple's Siri, Google Now, Microsoft's Cortana and Amazon's Alexa. Such speech-enabled assistants are distinguished from the previous generation of voice-based systems in that they claim to offer access to services and information via conversational interaction. In reality, interaction has limited depth and, after initial enthusiasm, users revert to more traditional interface technologies. This paper argues that the standard architecture for a contemporary communicative agent fails to capture the fundamental properties of human spoken language. So an alternative needs-driven cognitive architecture is proposed which models speech-based interaction as an emergent property of coupled hierarchical feedback control processes. The implications for future spoken language systems are discussed.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        In reality, users’ experiences with contemporary spoken
language systems leaves a lot to be desired. After initial
enthusiasm, users lose interest in talking to Siri or Alexa,
and they revert to more traditional interface technologies [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
One possible explanation for this state of affairs is that, while
component technologies such as automatic speech recognition
and text-to-speech synthesis are subject to continuous ongoing
improvement, the overall architecture of a spoken language
system has been standardised for some time [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] – see Fig. 2.
Standardisation is helpful because it promotes interoperability
and expands markets. However, it can also stifle innovation
by prescribing sub-optimal solutions. So, what (if anything)
might be wrong with the architecture illustrated in Fig. 2?
      </p>
      <p>
        In the context of spoken language, the main issue with the
architecture illustrated in Fig. 2 is that it reflects a traditional
stimulus–response (‘behaviourist’) view of interaction; the
user utters a request, the system replies. This is the ‘tennis
match’ analogy for language; a stance that is now regarded
as restrictive and old-fashioned. Contemporary perspectives
regard spoken language interaction as being more like a
threelegged race than a tennis match [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]: continuous coordinated
behaviour between coupled dynamical systems.
      </p>
    </sec>
    <sec id="sec-2">
      <title>II. TOWARDS A ‘COGNITIVE’ ARCHITECTURE</title>
      <p>
        What seems to be required is an architecture that
replaces the traditional ‘open-loop’ stimulus-response
arrangement with a ‘closed-loop’ dynamical framework; a
framework in which needs/intentions lead to actions, actions lead
to consequences, and perceived consequences are compared
to intentions/needs (in a continuous cycle of synchronous
behaviours). Such an architecture has been proposed by the
author [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] – see Fig. 3.
      </p>
      <p>
        One of the key concepts embedded in the architecture
illustrated in Fig. 3 is the agent’s ability to ‘infer’ (using
search) the consequences of their actions when they cannot be
observed directly. Another is the use of a forward model of
‘self’ to model ‘other’. Both of these features align well with
the contemporary view of language as “ostensive inferential
recursive mind-reading” [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Also, the architecture makes
an analogy between the depth of each search process and
‘motivation/effort’. This is because it has been known for
some time that speakers continuously trade effort against
intelligibility [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], and this maps very nicely into a
hierarchical control-feedback process [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] which is capable of
maintaining sufficient contrast at the highest pragmatic level of
communication by means of suitable regulatory compensations
at the lower semantic, syntactic, lexical, phonemic, phonetic
and acoustic levels.
      </p>
      <p>
        As a practical example, these ideas have been used to
construct a new type of speech synthesiser (known as ‘C2H’) that
adjusts its output as a function of its inferred communicative
success [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] – it listens to itself!
      </p>
    </sec>
    <sec id="sec-3">
      <title>III. FINAL REMARKS</title>
      <p>
        Whilst the proposed cognitive architecture successfully
captures some of the key elements of language-based interaction,
it is important to note that such interaction between human
beings is founded on substantial shared priors. This means
that there may be a fundamental limit to the language-based
interaction that can take place between mismatched partners
such as a human being and an autonomous social agent [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
    </sec>
    <sec id="sec-4">
      <title>ACKNOWLEDGMENT</title>
      <p>This work was partially supported by the European
Commission [EU-FP6-507422, EU-FP6-034434, EU-FP7-231868
and EU-FP7-611971], and the UK Engineering and Physical
Sciences Research Council (EPSRC) [EP/I013512/1].</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R.</given-names>
            <surname>Pieraccini</surname>
          </string-name>
          .
          <article-title>The Voice in the Machine</article-title>
          . MIT Press, Cambridge,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>R. K.</given-names>
            <surname>Moore</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          , &amp; S.
          <string-name>
            <surname>-H. Liao</surname>
          </string-name>
          .
          <article-title>Progress and prospects for spoken language technology: what ordinary people think</article-title>
          .
          <source>In INTERSPEECH</source>
          (pp.
          <fpage>3007</fpage>
          -
          <lpage>3011</lpage>
          ). San Francisco, CA,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <article-title>[3] Introduction and Overview of W3C Speech Interface Framework</article-title>
          , http: //www.w3.org/TR/voice-intro/
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>F.</given-names>
            <surname>Cummins</surname>
          </string-name>
          .
          <article-title>Periodic and aperiodic synchronization in skilled action</article-title>
          .
          <source>Frontiers in Human Neuroscience</source>
          ,
          <volume>5</volume>
          (
          <issue>170</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>R. K.</given-names>
            <surname>Moore. PRESENCE</surname>
          </string-name>
          :
          <article-title>A human-inspired architecture for speechbased human-machine interaction</article-title>
          .
          <source>IEEE Trans. Computers</source>
          ,
          <volume>56</volume>
          (
          <issue>9</issue>
          ),
          <fpage>1176</fpage>
          -
          <lpage>1188</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>R. K.</given-names>
            <surname>Moore</surname>
          </string-name>
          .
          <article-title>Spoken language processing: time to look outside</article-title>
          ? In L. Besacier,
          <string-name>
            <given-names>A.-H.</given-names>
            <surname>Dediu</surname>
          </string-name>
          , &amp;
          <string-name>
            <surname>C.</surname>
          </string-name>
          Martn-Vide (Eds.),
          <source>2nd International Conference on Statistical Language and Speech Processing (SLSP 2014), Lecture Notes in Computer Science</source>
          (Vol.
          <volume>8791</volume>
          ). Springer,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>R. K.</given-names>
            <surname>Moore</surname>
          </string-name>
          .
          <article-title>PCT and Beyond: Towards a Computational Framework for “Intelligent” Systems</article-title>
          . In A.
          <string-name>
            <surname>McElhone</surname>
          </string-name>
          &amp; W. Mansell (Eds.),
          <article-title>Living Control Systems IV: Perceptual Control Theory and the Future of the Life and Social Sciences. Benchmark Publications Inc</article-title>
          . In Press (available at https://arxiv.org/abs/1611.05379).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>T.</given-names>
            <surname>Scott-Phillips</surname>
          </string-name>
          .
          <article-title>Speaking Our Minds: Why human communication is different, and how language evolved to make it special</article-title>
          . London, New York: Palgrave MacMillan,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>E.</given-names>
            <surname>Lombard</surname>
          </string-name>
          . Le sign de l?lvation de la voix. Ann. Maladies Oreille, Larynx, Nez, Pharynx,
          <volume>37</volume>
          ,
          <fpage>101</fpage>
          -
          <lpage>119</lpage>
          ,
          <year>1911</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>B.</given-names>
            <surname>Lindblom</surname>
          </string-name>
          .
          <article-title>Explaining phonetic variation: a sketch of the H&amp;H theory</article-title>
          . In W. J.
          <string-name>
            <surname>Hardcastle</surname>
          </string-name>
          &amp; A.
          <string-name>
            <surname>Marchal</surname>
          </string-name>
          (Eds.),
          <source>Speech Production and Speech Modelling</source>
          (pp.
          <fpage>403</fpage>
          -
          <lpage>439</lpage>
          ). Kluwer Academic Publishers,
          <year>1990</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>W. T.</given-names>
            <surname>Powers</surname>
          </string-name>
          .
          <article-title>Behavior: The Control of Perception</article-title>
          . NY: Aldine: Hawthorne,
          <year>1973</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hawkins</surname>
          </string-name>
          .
          <article-title>Roles and representations of systematic fine phonetic detail in speech understanding</article-title>
          .
          <source>Journal of Phonetics</source>
          ,
          <volume>31</volume>
          ,
          <fpage>373</fpage>
          -
          <lpage>405</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>R. K.</given-names>
            <surname>Moore &amp; M. Nicolao</surname>
          </string-name>
          .
          <article-title>Reactive speech synthesis: actively managing phonetic contrast along an H&amp;H continuum, 17th International Congress of Phonetics Sciences (ICPhS)</article-title>
          .
          <source>Hong Kong</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Nicolao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Latorre &amp; R. K.</given-names>
            <surname>Moore. C2H</surname>
          </string-name>
          :
          <article-title>A computational model of H&amp;H-based phonetic contrast in synthetic speech</article-title>
          .
          <source>INTERSPEECH. Portland, USA</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>R. K.</given-names>
            <surname>Moore</surname>
          </string-name>
          .
          <article-title>Is spoken language all-or-nothing? Implications for future speech-based human-machine interaction</article-title>
          . In K. Jokinen &amp; G. Wilcock (Eds.),
          <source>Dialogues with Social Robots - Enablements, Analyses, and Evaluation. Springer Lecture Notes in Electrical Engineering</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>