<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>A theory of causal learning in children:
Causal maps and Bayes nets. Psychological Review</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Bartlett</institution>
          ,
          <addr-line>F. (1932), Remembering</addr-line>
          ,
          <institution>London: Cambridge University Press</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>1931</year>
      </pub-date>
      <volume>111</volume>
      <issue>1</issue>
      <fpage>148</fpage>
      <lpage>160</lpage>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        We view a constructivist and model-revising epistemology
as a rapprochement between the empiricist, rationalist, and
pragmatist viewpoints. The constructivist hypothesizes that
all human understanding is the result of an interaction
between energy patterns in the world and mental categories
imposed on the world by an intelligent agent
        <xref ref-type="bibr" rid="ref1 ref4">(Piaget 1954,
1970; von Glasersfeld 1978)</xref>
        . Using Piaget’s terms, we
humans assimilate external phenomena according to our
current understanding and accommodate our understanding
to phenomena that do not meet our prior expectations.
      </p>
      <p>Constructivists use the term schemata to describe the a
priori structure used to mediate the experience of the
external world. The term schemata is taken from the
writing of the British psychologist Bartlett (1932) and its
philosophical roots go back to Kant (1781/1964). On this
viewpoint observation is not passive and neutral but active
and interpretative. There are many current psychologists
and philosophers that support and expand this pragmatic
and teleological account of human developmental activity
(Glymour 2001, Gopnik et al. 2010, Gopnik 2011a, 2011b,
Kushnir et al. 2010).</p>
      <p>Perceived information, Kant’s a posteriori knowledge,
rarely fits precisely into our preconceived and a priori
schemata. From this tension to comprehend and act as an
agent, the schema-based biases a subject uses to organize
experience are strengthened, modified, or replaced. This
accommodation in the context of unsuccessful interactions
with the environment drives a process of cognitive
equilibration. The constructivist epistemology is one of
cognitive evolution and continuous model refinement. An
important consequence of constructivism is that the
interpretation of any perception-based situation involves
the imposition of the observers (biased) concepts and
categories on what is perceived. This constitutes an
inductive bias (Luger 2009, Ch 16).</p>
      <p>
        When
        <xref ref-type="bibr" rid="ref1">Piaget (1970)</xref>
        proposed a constructivist approach to
a child’s understanding the external world, he called it a
genetic epistemology. When encountering new phenomena,
the lack of a comfortable fit of current schemata to the
world “as it is” creates a cognitive tension. This tension
drives a process of schema revision. Schema revision,
Piaget’s accommodation, is the continued evolution of the
agent’s understanding towards equilibration.
      </p>
      <p>There is a blending here of empiricist and rationalist
traditions, mediated by the pragmatist requirement of agent
survival. As embodied, agents can comprehend nothing
except that which first passes through their senses. As
accommodating, agents survive through learning the
general patterns of an external world. What is perceived is
mediated by what is expected; what is expected is
influenced by what is perceived: these two functions can
only be understood in terms of each other. A Bayesian
model-refinement representation offers an appropriate
model for critical components of this constructivist
modelrevising epistemological stance (Luger et al. 2002, Luger
2012). Interestingly enough, David Hume acknowledged
the epistemic foundation of all human activity (including,
of course, the construction of AI artifacts) in A Treatise on
Human Nature (1739/2000) when he stated “All the
sciences have a relation, greater or less, to human nature;
and ... however wide any of them may seem to run from it,
they still return back by one passage or another. Even
Mathematics, Natural Philosophy, and Natural Religion,
are in some measure dependent on the science of MAN;
since they lie under the cognizance of men, and are judged
of by their powers and faculties.”</p>
      <p>Thus, we can ask why a constructivist epistemology
might be useful in addressing the problem of building
programs that are “intelligent”. How can an agent within
an environment understand its own understanding of that
situation? We believe that constructivism also addresses
this problem of epistemological access. For more than a
century there has been a struggle in both philosophy and
psychology between two factions: the positivist, who
proposes to infer mental phenomena from observable
€
physical behavior, and a more phenomenological approach
which allows the use of first person reporting to enable
access to cognitive phenomena. This factionalism exists
because both modes of access to cognitive phenomena
require some form of model construction and inference.</p>
      <p>In comparison to physical objects like chairs and doors,
which often, naively, seem to be directly accessible, the
mental states and dispositions of an agent seem to be
particularly difficult to characterize. We contend that this
dichotomy between direct access to physical phenomena
and indirect access to mental phenomena is illusory. The
constructivist analysis suggests that no experience of the
external (or internal) world is possible without the use of
some model or schema for organizing that experience. In
scientific enquiry, as well as in our normal human
cognitive experiences, this implies that all access to
phenomena is through exploration, approximation, and
continued model refinement.</p>
      <p>Bayes theorem (1763) offers a plausible model of this
constructivist rapprochement between the philosophical
traditions we have just discussed. It is also an important
modeling tool for much of modern AI, including AI
programs for natural language understanding, robotics, and
machine learning. With a high-level discussion of Bayes’
insights, we can describe the power of this approach.</p>
      <p>Consider the general form of Bayes’ relationship used to
determine the probability of a particular hypothesis, hi,
given a set of evidence E:
p(hi | E) = n</p>
      <p>p(E | hi )p(hi )
∑p(E | hk )p(hk )
k=1
 
p(hi|E) is the probability that a particular hypothesis, hi, is
true given evidence E.
p(hi) is the probability that hi is true overall.
p(E|hi) is the probability of observing evidence E when hi
is true.
n is the number of possible hypotheses.</p>
      <p>With the general form of Bayes’ theorem we have a
functional (and computational!) description (model) for a
particular situation happening given a set of perceptual
evidence clues. Epistemologically, we have created on the
right hand size of the equation a schema describing how
prior accumulated knowledge of occurrences of
phenomena can relate to the interpretation of a new
situation, the left hand side of the equation. This
relationship can be seen as an example of Piaget’s
assimilation where encountered information fits (is
interpreted by) the patterns created from prior experiences.</p>
      <p>To describe further the pieces of Bayes formula: The
probability of an hypothesis being true, given a set of
evidence, is equal to the probability that the evidence is
true given the hypothesis times the probability that the
hypothesis occurs. This number is divided (normalized) by
the probability of the evidence itself, p(E). This probability
of evidence is represented as the sum over all hypotheses
presenting the evidence times the probability of that
hypothesis itself.</p>
      <p>There are limitations to using Bayes’ theorem as just
presented as an epistemological characterization of the
phenomenon of interpreting new (a posteriori) data in the
context of (prior) collected knowledge and experience.
First, of course, is the fact that the epistemological subject
is not a calculating machine. We simply don’t have all the
prior (numerical) values for all the hypotheses and
evidence that can fit a problem situation. In a complex
situation such as medicine where there can be hundreds of
hypothesized diseases and thousands of symptoms, this
calculation is intractable (Luger 2009, Chapter 5).</p>
      <p>A second objection is that in most situations the sets of
evidence are NOT independent, given the set of
hypotheses. This makes the calculation of p(E) in the
denominator of Bayes rule as just presented unjustified.
When this independence assumption is simply ignored, as
we see shortly, the result is called naïve Bayes. More often,
however, the rationalization of the probability of the
occurrence of evidence across all hypotheses is seen as
simply a normalizing factor, supporting the calculation of a
realistic measure for the probability of the hypothesis given
the evidence (the left side of Bayes’ equation). The same
normalizing factor is utilized in determining the actual
probability of any of the hi, given the evidence, and thus,
as in most natural language processing applications, is
usually ignored.</p>
      <p>A final objection asserts that diagnostic reasoning is not
about the calculation of probabilities; it is about
determining the most likely explanation, given the
accumulation of pieces of evidence. Humans are not doing
real-time complex mathematical processing; rather we are
looking for the most coherent explanation or possible
hypothesis, given the amassed data. Thus, a much more
intuitive form of Bayes rule ignores this p(E ) denominator
entirely (as well as the associated assumption of evidence
independence). The resulting formula determines the
likelihood of any hypothesis given the evidence, as the
product of the probability of the evidence given the
hypothesis times the probability of the hypothesis itself
p(E|hi) p(hi). In most diagnostic situations we are asked to
determine which of a set of hypotheses hi is most likely to
be supported. We refer to this as determining the argmax
across all the set of hypotheses. Thus, if we wish to
determine which of all the hi has the most support we look
for the largest p(E|hi) p(hi):
argmax(hi) p(E|hi) p(hi)</p>
      <p>In a dynamic interpretation, as sets of evidence
themselves change across time, we will call this argmax of
hypotheses given a set of evidence at a particular time the
greatest likelihood of that hypothesis at that time. We show
this relationship, an extension of the Bayesian maximum a
posteriori (or MAP) estimate, as a dynamic measure over
time t:
We&amp;next&amp;illustrate&amp;the&amp;Bayesian&amp;approach&amp;in&amp;two&amp;application&amp;dom
discrete&amp;component&amp;semiconductors&amp;(Stern&amp;et&amp;al.&amp;1997,&amp;Chakraba
expecrrte,atthineg&amp;ptrhees&amp;egnrceeateosft&amp;laikberliehaokodsu&amp;fpopro&amp;hrytspoathneusmesb&amp;aercroosfs&amp;expanding&amp;
alterFniagtuivree&amp;1h,&amp;ypproetsheensetisn. gT&amp;twheo&amp;fsaeialurcrhesf&amp;oofr&amp;dtihsceremteo&amp;scto mlikpeolnyent&amp;semicon
expl“aonpaetino”n,&amp;ofro&amp;rthae&amp;bfraeilaukre&amp;inb&amp;aro&amp;wadireen&amp;scotnhneecetviindge&amp;nccoempseoanrechn:ts&amp;to&amp;others&amp;i
Howthlea&amp;rpgreesisenthcee&amp;obfr&amp;eaa&amp;kb?reIaskt&amp;hsuerpepaonryts&amp;dai&amp;sncuomlobraetrio&amp;onf&amp;areltleatrendative&amp;hypoth
to theexpbrlaenaka?tioWne&amp;froert&amp;hae&amp;fraeilaunrye&amp;(bpreoracdepetnusa&amp;lt)hseo&amp;euvniddsenorces&amp;mseeallrsch:&amp;How&amp;larg
wherneliattheadp&amp;tpoe&amp;ntheed&amp;?brWeahka?t&amp;Wweerree&amp;tthheeres&amp;aunltyi n&amp;(gpecrocnedpittiuoanls)&amp;soofunds&amp;or&amp;sme
the croemsuplotinnegn&amp;ctsoonfdtihtieonsyss&amp;otefm&amp;th?e&amp;components&amp;of&amp;the&amp;system?&amp;</p>
    </sec>
    <sec id="sec-2">
      <title>This model is both intuitive and simple: the most likely</title>
      <p>interpretation of new data, given evidence E at time t, is a
function of whWiche&amp;ninetxetr&amp;pilrleutsattrioantei&amp;tshme&amp;oBsatyleiksiealyn&amp;atoppprroodacuhce&amp;in&amp;two&amp;application&amp;domains.&amp;In&amp;the&amp;diagnosis&amp;of&amp;failures&amp;in&amp;
that evidence discrete&amp;cotmponent&amp;semiconductors&amp;(Sttheartn&amp;et&amp;al.&amp;1997,&amp;Chakrabarti&amp;et&amp;al.&amp;2005)&amp;we&amp;have&amp;an&amp;example&amp;of&amp;
at time and the probability of
interpretation itcsreelaftoinccgu&amp;trhrien&amp;gg.reatest&amp;likelihood&amp;for&amp;hypotheses&amp;across&amp;expanding&amp;data&amp;sets.&amp;Consider&amp;the&amp;situation&amp;of&amp;</p>
      <p>By the eaFrliygur1e9&amp;190,&amp;ps,resmeunctihng&amp;otfwoc&amp;foamilupuretast&amp;ioofn&amp;d-bisacsreedte&amp;component&amp;semiconductors.&amp;The&amp;failure&amp;type&amp;is&amp;called&amp;an&amp;
language unde“rosptaennd”i,n&amp;ogr&amp;the&amp;break&amp;in&amp;a&amp;wire&amp;connecting&amp;components&amp;to&amp;others&amp;in&amp;the&amp;system.&amp;For&amp;the&amp;diagnostic&amp;expert,&amp;
and generation was stochastic,
including partshineg&amp;p, respeanrtc-eo&amp;fo-sf&amp;pae&amp;becrheak&amp;tsaugpgipnogr,ts&amp;are&amp;nfeurmenbceer&amp;of&amp;alternative&amp;hypotheses.&amp;The&amp;search&amp;for&amp;the&amp;most&amp;likely&amp;
resolution, andexdpislacnouartsioenp&amp;froorc&amp;aes&amp;fsaiinlgu,reu&amp;bsuraolalydeunssi&amp;ntghet&amp;oeovlisdence&amp;search:&amp;How&amp;large&amp;is&amp;the&amp;break?&amp;Is&amp;there&amp;any&amp;discoloration&amp;
related&amp;to&amp;the&amp;break?&amp;Were&amp;there&amp;any&amp;(perceptual)&amp;sounds&amp;or&amp;smells&amp;when&amp;it&amp;happened?&amp;What&amp;were&amp;the&amp;&amp;
l2i0k0e9)g.reOattehsetr lraiekrseeulaislhtoinoogfd&amp;caomrnteidfaiisctuiiaorelnssi&amp;no(tJfeu&amp;tlrlhaigefs&amp;ecknoycme,apnoednspeMenctaisar&amp;ltoliynf&amp;the&amp;system&amp; ?&amp; &amp; &amp; (a)a.&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;b.&amp;
wmaaycshintehelseearnuisnegs, boefcasmtoechmaosrtiec Btaeycehsnioanlo-gbyasefdo.rInpamttaenrny “Fcigounrnee&amp;c1t.i&amp;Tonw&amp;bo&amp;reoxkaemnp”&amp;lfeasil&amp;uofr&amp;ed.i&amp;screte&amp;component&amp;semiconductor
recognition were another instantiation of the constructivist &amp;
tradition, as collected sets of patterns were used to Driven&amp;by&amp;the&amp;data&amp;search&amp;supporting&amp;multiple&amp;possible&amp;hypothes
condition recognition of new patterns. notes&amp;the&amp;bambooing&amp;effect&amp;in&amp;the&amp;disconnected&amp;wire,&amp;Figure&amp;1a.&amp;T</p>
      <p>
        Judea Pearl’s (1988) proposal for use of Bayesian belief likelihood&amp;hypothesis&amp;that&amp;explains&amp;the&amp;open&amp;as&amp;a&amp;break&amp;created&amp;b
nets (BBNs) and his assumption of their links reflecting caused&amp;by&amp;a&amp;sequence&amp;of&amp;lowSfrequency&amp;highScurrent&amp;pulses.&amp;The&amp;
“causal” relationships (Pearl 2000) brought the use of the&amp;example&amp;of&amp;Figure&amp;1b,&amp;where&amp;the&amp;break&amp;is&amp;seen&amp;as&amp;balled,&amp;is&amp;me
Bayesian technology to an entirely new importance. First, these&amp;diagnostic&amp;scenarios&amp;have&amp;been&amp;implemented&amp;by&amp;an&amp;expert&amp;
the assumption of these networks being directed graphs – hypothesis&amp;space&amp;
        <xref ref-type="bibr" rid="ref2">(Stern&amp;et&amp;al.&amp;1997)</xref>
        &amp;as&amp;well&amp;as&amp;reflected&amp;in&amp;a&amp;Baye
reflecting causal relationships – and disallowing cycles – Figure&amp;2&amp;presents&amp;a&amp;Bayesian&amp;belief&amp;net&amp;(BBN)&amp;capturing&amp;these&amp;an
no entity can cause itself – brought a radical improvement
to the computational costs of reasoning with BBNs (Luger The&amp;BBN,&amp;without&amp;new&amp;data,&amp;represents&amp;the&amp;a&amp;priori&amp;state&amp;of&amp;an&amp;ex
2009, Ch 9). Second, these same two assumptions made domain.&amp;In&amp;fact,&amp;these&amp;networks&amp;of&amp;causal&amp;relationships&amp;are&amp;usuall
the BBN representation much more transparent as a &amp; working&amp;with&amp;human&amp;experts’&amp;analysis&amp;of&amp;known&amp;fa&amp;ilures.&amp;Thus,&amp;t
representationa&amp;l tool th&amp;at could&amp; capture ac.&amp;&amp;a&amp;&amp;u&amp;&amp;s&amp;&amp;a&amp;&amp;&amp;l&amp;&amp;&amp;&amp;r&amp;e&amp;&amp;&amp;l&amp;a&amp;&amp;t&amp;&amp;i&amp;o&amp;&amp;&amp;n&amp;&amp;s&amp;&amp;.&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;&amp;e&amp;&amp;&amp;r&amp;&amp;t&amp;b&amp;k.&amp;nowledge&amp;implicit&amp;in&amp;a&amp;domain&amp;of&amp;interest.&amp;When&amp;new&amp;(a&amp;
exp (b)
Finally, most all the traditional powerful stochastic e.g.,&amp;the&amp;wire&amp;is&amp;“bambooed”,&amp;the&amp;color&amp;of&amp;the&amp;copper&amp;wire&amp;is&amp;norm
      </p>
      <p>Figulraen&amp;1g.u&amp;Tawgeo&amp;ewxaomrkpleas&amp;nodf&amp;dimscarcethei&amp;nceomponent&amp;semmoicsot&amp;nlidkuecltyo&amp;resx,&amp;pelaacnha&amp;etxiohinb,i&amp;wtinitgh&amp;tihne&amp;&amp;i“tosp&amp;(ean&amp;p”&amp;roiro&amp; ri)&amp;model,&amp;given&amp;this&amp;ne
lfreeoaprrmrneisneognf,taaftoiodrnysenDxaramuimvsiecepdnleB&amp;,bai&amp;“yntych&amp;oetehsnienah&amp;endicdatdnitoeaenn&amp;tsw&amp;beMoraorrakkcrehkn(&amp;oDs”v&amp;ufBapiNmlpu)oor,dret.ec&amp;ilonugiln&amp;dmtubhleetiple&amp;posbFserisomgikubroierculneneole”&amp;enh&amp;sdhy1&amp;uyfp.ocprtoo&amp;dtTrhsow,eineosaigsc&amp;&amp;thahecxiehsxa&amp;im(heLivpbueligetsies&amp;nirtg&amp;s2&amp;tog0hfr0ee9a“,dot&amp;Ceipsshectnar&amp;lep”itkteoeerrli&amp;“9hcc)ooo.o&amp;mnAdnpn,e&amp;o&amp;ocinmttiehopnentor&amp;rrtealnatt&amp;erde&amp;shuylpt&amp;
likelifhaoilooutdrhe&amp;em.seeas&amp;stuhraets&amp;c&amp;danec&amp;erxepalsaei&amp;nw&amp;tihthei&amp;n“o&amp;tpheen&amp;B”,B&amp;tNhe.&amp;&amp;expert&amp;
readily integrated into thibsanmewboroepinrgesentational formalism.</p>
      <p>notes&amp;the&amp; &amp;effect&amp;in&amp;the&amp;disconnected&amp;wire,&amp;Figure&amp;1a.&amp;This&amp;suggests&amp;a&amp;revised&amp;greatest&amp;</p>
      <p>We next illilkusetlriahtoeodth&amp;heypBotahyeessiiasn&amp;thaatp&amp;epxropalachinsi&amp;tnhet&amp;owpoen&amp;as&amp;a&amp;brDeTraihkvi&amp;ecsnr&amp;ceubarytreetdhn&amp;ebt&amp;yed&amp;xmaataemtsapella&amp;ecr&amp;rcdyhesmtsauolplnipzsoatrrttiaiontnegs&amp;t&amp;mhhaoutwl&amp;twi&amp;ptahlese&amp;&amp;lpmikooesslsyitb&amp;&amp;lliekely&amp;current&amp;
application domains. In the diagnosis of failures in discrete</p>
      <p>caused&amp;by&amp;a&amp;sequence&amp;of&amp;lowSfrequency&amp;highScurrheynptob&amp;ptehuseltss&amp;eeesxsp.&amp;tTlhahantea&amp;cgtaironenae,t&amp;gxeipsvtlea&amp;linink&amp;aet&amp;lhpyea&amp;hr“ytoipcpouetlnha”re,&amp;stiitsmh&amp;efeo&amp;erax&amp;ntphdee&amp;ra&amp;tonpn&amp;heoyntep&amp;osoft&amp;hesis&amp;spac
component semt hiceo&amp;enxdaumctporles&amp;o(Sf&amp;tFeirgnuerte&amp;a1l.b,1&amp; w99h7e,rCe&amp;hthake&amp;rbarbeaartki&amp;is&amp;seenth&amp;aessb&amp;beaatmsll&amp;beoodfo&amp;,h&amp;iinys&amp;gpmoeetfhlfteeincstegs&amp;idn&amp;auntehd&amp;et&amp;odd&amp;aeitsxacc&amp;oaencsnrseoivcsetse&amp;&amp;cdtiumwrreier,&amp;enu,ts.F&amp;iBnigogu&amp;tthrhe&amp;eo1&amp;fm&amp;a.ost&amp;likely&amp;hy
et al. 2005) wtehheasev&amp;ediaangneoxsatmicp&amp;sleceonfarciroesa&amp;thinagvet&amp;hbeeegnr e&amp;iamtepsltementeTdh&amp;ibslyiks&amp;uaegnlig&amp;heeoxsoptsder&amp;aht&amp;yrsepyvositstehemdesSgilsrike.&amp;aet&amp;essetarlickhe&amp;ltihroooudghhy&amp;apno&amp;thesis that
likelihood forhyhpyoptohtheessise&amp;sspaaccreo&amp;(sSsteerxnp&amp;eatn&amp;adli.n&amp;1g99d7a)ta&amp;as&amp;well&amp;as&amp;refelxecptlaeidn&amp;sin&amp;at&amp;hBeayeospieann&amp;beaslief&amp;anetb&amp;(rCeahkakrcarbeartetdi&amp;et&amp;bayl.&amp;20m0e5ta).l&amp;&amp;
sets.</p>
      <p>
        Consider the siFtuigautiroen&amp;2o&amp; pfrFeisgeunrtes&amp;1a,&amp;Bpareyseesnitainng&amp;betwlieof&amp;fnaeiltu&amp;(rBesBN)&amp;capctruysrgtianll(glhi&amp;zti haEteitso)en=&amp;athanardtg&amp;owmtahasexlr(i&amp;khrei)lyaptce(aEdu&amp;td|sheiadi)gbpny(ohasitsi&amp;ce&amp;qsuiteunacteioonfsl.o&amp;
w| )
of discrete component semiconductors. The failure type is frequency high-current pulses. The greatest likely
called an “opTenh”e,&amp;BoBrN,t&amp;hweithboreuatk&amp;neiwn&amp;daatwa,i&amp;rreepcroesnennectsti&amp;nthge&amp;a&amp;priohryip&amp;sottahtees&amp;iosf&amp;faonr&amp;ethxepoerpte’ns&amp;konf othweleexdagme&amp;polfe&amp;aonf&amp;aFpigpulirceat1ibo,nw&amp;here
components todoomthaerins.&amp;iInn&amp;ftahcet,&amp;styhsetseem&amp; n.eFtwororthkes&amp;odfi&amp;acganuossatli&amp;crelationtshheipbsr&amp;eaarke&amp;uissusaeellny&amp;caasrebfaulllelyd&amp;,criasftmedel&amp;ttihnrgoudguhe&amp;mtoaneyx&amp;cheosusirvse&amp;
working&amp;with&amp;human&amp;experts’&amp;analysis&amp;of&amp;known&amp;failures.&amp;Thus,&amp;the&amp;BBN&amp;can&amp;be&amp;said&amp;to&amp;capture&amp;a&amp;priori&amp;
expert&amp;knowledge&amp;implicit&amp;in&amp;a&amp;domain&amp;of&amp;interest.&amp;When&amp;new&amp;(a&amp;posteriori)&amp;data&amp;are&amp;given&amp;to&amp;the&amp;BBN,&amp;
e.g.,&amp;the&amp;wire&amp;is&amp;“bambooed”,&amp;the&amp;color&amp;of&amp;the&amp;copper&amp;wire&amp;is&amp;normal,&amp;etc,&amp;the&amp;belief&amp;network&amp;“infers”&amp;the&amp;
most&amp;likely&amp;explanation,&amp;within&amp;its&amp;(a&amp;priori)&amp;model,&amp;given&amp;this&amp;new&amp;information.&amp;There&amp;are&amp;many&amp;inference&amp;
rules&amp;for&amp;doing&amp;this&amp;(Luger&amp;2009,&amp;Chapter&amp;9).&amp;An&amp;important&amp;result&amp;of&amp;using&amp;the&amp;BBN&amp;technology&amp;is&amp;that&amp;as&amp;
one&amp;hypothesis&amp;achieves&amp;its&amp;greatest&amp;likelihood,&amp;other&amp;related&amp;hypotheses&amp;are&amp;“explained&amp;away”,&amp;i.e.,&amp;their&amp;
likelihood&amp;measures&amp;decrease&amp;within&amp;the&amp;BBN.&amp;
current. Both of these diagnostic scenarios have been
implemented by an expert system-like search through an
hypothesis space
        <xref ref-type="bibr" rid="ref2">(Stern et al. 1997)</xref>
        as well as reflected in a
Bayesian belief net (Chakrabarti et al. 2005). Figure 2
presents a Bayesian belief net (BBN) capturing these and
other related diagnostic situations.
      </p>
      <p>The BBN, without new data, represents the a priori state
of an expert’s knowledge of an application domain. In fact,
these networks of causal relationships are usually carefully
crafted through many hours working with human experts’
analysis of known failures. Thus, the BBN can be said to
capture a priori expert knowledge implicit in a domain of
interest. When new (a posteriori) data are given to the
BBN, e.g., the wire is “bambooed”, the color of the copper
wire is normal, etc, the belief network “infers” the most
likely explanation, within its (a priori) model, given this
new information. There are many inference rules for doing
this (Luger 2009, Chapter 9). An important result of using
the BBN technology is that as one hypothesis achieves its
greatest likelihood, other related hypotheses are “explained
away”, i.e., their likelihood measures decrease within the
BBN.</p>
      <p>This current example demonstrates how the most likely
current hypothesis can be used to determine the best
explanation, given a particular time and an hypothesis
space. We next demonstrate how considering sets of
hypotheses and data across time, using the most likely
hypothesis at time t, can produce a greatest likelihood
hypothesis.
gl(hi|Et) = argmax(hi) p(Et|hi) p(hi)</p>
      <p>In this model the most likely interpretation of new data,
given evidence E at time t, is a function of which
interpretation is most likely to produce that evidence at
time t and the probability of that interpretation itself
occurring. If we want to expand this to the next time
period, t + 1, we need to describe how models can evolve
across time.</p>
      <p>As an example of argmax processing, Chakrabarti et al.
(2005, 2007) analyze a continuous data stream from a set
of distributed sensors. The running “health” of the
transmission of a Navy helicopter rotor system is
represented by a steady stream of sensor data. This data
consists of temperature, vibration, pressure, and other
measurements reflecting the state of the various
components of the running transmission system. An
example of this data can be seen in the top portion of
Figure 3, where the continuous data stream is broken into
discrete and partial time slices.</p>
      <p>A Fourier transform is then used to translate these
signals into the frequency domain, as shown on the left
side of the second row of Figure 3. These frequency
readings were compared across time periods to diagnose
the running health of the rotor system. The model used to
diagnose rotor health the auto-regressive hidden Markov
model (A-RHMM) of Figure 4. The observable states of
the system are made up of the sequences of the segmented
signals in the frequency domain while the hidden states are
the imputed health states of the helicopter rotor system
itself, as seen in the lower right of Figure 3.</p>
      <p>The hidden Markov model (HMM) technology is an
important stochastic technique that can be seen as a variant
of a dynamic BBN. In the HMM, we attribute values to
states of the network that are themselves not directly
observable. For example, the HMM technique is widely
used in the computer analysis of human speech, trying to
determine the most likely word uttered, given a stream of
acoustic signals (Jurasky and Martin 2009). In our
helicopter example, training this system on streams of
normal transmission data allowed the system to make the
correct greatest likelihood measure of failure when these
signals change to indicate a possible breakdown. The US
Navy supplied data to train the normal running system as
wIn&amp;tehilsl&amp;moddeal&amp;tthae&amp;mosset&amp;ltiksely&amp;interpretation&amp;of&amp;new&amp;data,&amp;ngisven&amp;etvhidaetnce&amp;E&amp;at&amp;time&amp;t,&amp;is&amp;a&amp;funcstieone&amp;odf&amp;ewhdich&amp;</p>
      <p>for transmissio contained
finateurplrtesta.tioTn&amp;ihs&amp;umoss,t&amp;litkheley&amp;toh&amp;prioddducee&amp;nthats&amp;etvaidteenceS&amp;at&amp;timoef&amp;t&amp;atnhd&amp;etheA&amp;pro-bRabHilitMy&amp;of&amp;Mthat&amp;irnetefrplreectattison&amp;</p>
      <p>t
tithseelf&amp;ocgcurreriangt.e&amp;Ifs&amp;wte&amp;wliaknte&amp;tol&amp;iehxpoanodd&amp;thish&amp;toy&amp;tpheo&amp;ntehxt&amp;etismies&amp;peroiofd,&amp;tt'h+'1e,&amp;wset&amp;naetede&amp;too&amp;dfesctrihbee&amp;horwo&amp;mtoodrels&amp;
scayn&amp;sevtoelvme&amp;a,crgosisv&amp;timeen.&amp; the observed evidence Ot at any time t.
&amp;</p>
      <p>&amp;&amp;
FigurFeigure2&amp;2..&amp;A&amp;BAayesiaBn&amp;bealiyef&amp;enestwiaornk&amp;reprbeseenltiinegf&amp;the&amp;cnauesatlw&amp;reloatironkshiprs&amp;eanpd&amp;draetas&amp;peoinntst&amp;iimnpglicit&amp;int&amp;he
causathle&amp;driesclreatet&amp;cioomnposnhenit&amp;psemsicaonndudctodr&amp;daomtaain.p&amp;Aso&amp;dianta&amp;tiss&amp;“diismcovepreldi”c&amp;thiet&amp;(ai&amp;pnriortih)&amp;perobdabiilsiscticr&amp; ete
comphoypnotheensets&amp;chsaengme&amp;anidc&amp;sougngedstu&amp;fucrthteor&amp;rseadrcoh&amp;fmor&amp;daatian.&amp;. As data is “discovered”
the (a priori) probabilistic hypotheses change and suggest
fAus&amp;arnt&amp;ehxaemrplsee&amp;ofa&amp;arrcgmhaxf&amp;oprrocdesasintga,&amp;.Chakrabarti&amp;et&amp;al.&amp;(2005,&amp;2007)&amp;analyze&amp;a&amp;continuous&amp;data&amp;stream&amp;from&amp;
a&amp;set&amp;of&amp;distributed&amp;sensors.&amp;The&amp;running&amp;“health”&amp;of&amp;the&amp;transmission&amp;of&amp;a&amp;Navy&amp;helicopter&amp;rotor&amp;system&amp;is&amp;
represented&amp;by&amp;a&amp;steady&amp;stream&amp;of&amp;sensor&amp;data.&amp;This&amp;data&amp;consists&amp;of&amp;temperature,&amp;vibration,&amp;pressure,&amp;and&amp;
other&amp;measurements&amp;reflecting&amp;the&amp;state&amp;of&amp;the&amp;various&amp;components&amp;of&amp;the&amp;running&amp;transmission&amp;system.&amp;An&amp;
example&amp;of&amp;this&amp;data&amp;can&amp;be&amp;seen&amp;in&amp;the&amp;top&amp;portion&amp;of&amp;Figure&amp;3,&amp;where&amp;the&amp;continuous&amp;data&amp;stream&amp;is&amp;broken&amp;
into&amp;discrete&amp;and&amp;partial&amp;time&amp;slices.&amp;</p>
      <p>&amp;
faulty} at time t.
represent&amp;the&amp;observable&amp;values&amp;at&amp;time&amp;t.&amp;The&amp;St&amp;states&amp;represent&amp;the&amp;hidden&amp;“health”
systeCm,o&amp;{snacfelu,usnisoanfe:, Afaunltye}p&amp;ati&amp;stitmeem&amp;t.&amp;ological stance
Turing’s test for intelligence was agnostic both as to what a
computer was composed of – vacuum tubes, transistors, or
3.'Conclusion:'An'Epistemological'Stance''
tinker toys - as well as to the languages used to make it
run. It simply required the responses of the machine to be
roughly equivalent to the responses of humans in the same
Turing’ssi&amp;tteusatt&amp;ifoonr&amp;si n.telligence&amp;was&amp;agnostic&amp;both&amp;as&amp;to&amp;what&amp;a&amp;computer&amp;was&amp;composed&amp;</p>
      <p>&amp;</p>
      <p>Modern AI research has proposed probabilistic
t&amp;ransistroerpsr,&amp;eosre&amp;tnintakteior&amp;ntosyasn&amp;Sd&amp;asa&amp;lwgeolrli&amp;aths&amp;mtos&amp;thfoer&amp;latnhgeuraegaels-t&amp;uimseed&amp;itnot&amp;emgarkaeti&amp;iotn&amp;run.&amp;It&amp;simply&amp;r</p>
      <p>of new (a posteriori) information into previously (a priori)
of&amp;the&amp;mleaacrhnineed&amp;tpo&amp;abtete&amp;rronusgholfy&amp;ienqfuoirvmalaetniot&amp;nto&amp;(tDhee&amp;mrepspsotenrse1s9&amp;o6f8&amp;h)u.mAamnso&amp;ning&amp;the&amp;same&amp;situa
Figure&amp;3.&amp;RealStime&amp;data&amp;from&amp;the&amp;transmission&amp;system&amp;of&amp;a&amp;helicopter’s&amp;rotor.&amp;The&amp;top&amp;component&amp;of&amp;
&amp;
&amp;
&amp;
&amp;
the&amp;figure&amp;presents&amp;the&amp;original&amp;data&amp;stream&amp;(left)M&amp;aondde&amp;ranmn&amp;A&amp;ieIg&amp;nrheltasredgaerescchr&amp;ihbaes&amp;pitr.oAposceodg&amp;pnriotibvaebisliyssttice&amp;mrwepecrrae&amp;nsleebnftte&amp;atiinonas&amp;apnrdio&amp;arligorithms&amp;for&amp;th
d&amp;time&amp;slice&amp;(right).&amp;The&amp;lo
iterating towards equilibrium, or equilibration, as Piaget
figure&amp;is&amp;the&amp;result&amp;of&amp;the&amp;Fourier&amp;transform&amp;of&amp;theo&amp;ft&amp;inmewe&amp;&amp;ks(alni&amp;opcoews&amp;dlteeadrtigaoe&amp;r(.i)t&amp;riWanfnohsrefmonartmipoerned&amp;isn)e&amp;tniont&amp;eptodr e&amp;tvhiwoeui&amp;tfshrley&amp;q(nauo&amp;pevrneicol yri&amp;)i&amp;nlefaorrnmeadt&amp;pioantterns&amp;of&amp;infor
Figure&amp;4.&amp;The&amp;data&amp;of&amp;Figure&amp;3&amp;is&amp;processed&amp;using&amp;an&amp;autoSregressive&amp;hidden&amp;Markov&amp;model.&amp;States&amp;Ot&amp;
loopy,bTelhieef,pcroogpnagitai vtieon
reprdeseontm&amp;thea&amp;obisnerv.&amp;aTbleh&amp;vaelu&amp;els&amp;oat&amp;wtimee&amp;tr.&amp;T&amp;hre&amp;iSgt&amp;shtatte&amp;s&amp;friepgreusernte&amp;th&amp;er&amp;heidpdern&amp;“ehesalethn”&amp;sttaste&amp;st&amp;ohf&amp;thee&amp;&amp;rhotoird&amp;1d9e6n8)&amp;s.&amp;tAdaamtteaosn&amp;pgoe&amp;tfrh&amp;tteuhsreeb&amp;&amp;shalegthloierciotehpqmtuesir&amp;lii&amp;srb&amp; orituomr&amp;.system.&amp; s&amp;(yPseteamrl&amp;1t9h8e8n,&amp;2000)&amp;that&amp;ca
infers, using prior and posterior components of the model,
characterizing a new diagnostic situation, this a posteriori
equilibrium with its continuing states of learned diagnostic
system,&amp;{safe, unsafe, faulty}&amp;at&amp;time&amp;t.&amp;</p>
      <p>Turing’s&amp;test&amp;for&amp;intelligence&amp;was&amp;agnostic&amp;both&amp;as&amp;to&amp;what&amp;a&amp;computer&amp;was&amp;composed&amp;of&amp;–&amp;vacuum&amp;tubes,&amp;
m&amp;(left)&amp;andsl&amp;iacne&amp;e(nrilgahrtg).edT&amp;thiemleo&amp;wsleicrel&amp;(erftigfhigt)u.r&amp;Tehies&amp;ltohwe erre&amp;sleufltt&amp; of the
s&amp;the&amp;hidde nth&amp;sethaitdedse&amp;onfs&amp;ttahtee&amp;shoeflitchoephteelirc&amp;rooptteorrr&amp;soytosrtesymst.&amp;em.</p>
      <p>particular greatest likelihood hypothesis.
plausibulen&amp;btielliietfsf&amp;icnodnsstcaonntlvye&amp;irtgeerantcinego&amp;tro weqaurdilsi&amp;berqiuuimli b,riinumth,&amp;eorf&amp;oeqrmuiliobfraation,&amp;as&amp;Piaget&amp;</p>
      <p>sufficient account of human intelligence in areas such as
cognitive&amp;sTyhsetecml a&amp;ciamn&amp;obef&amp;tinh&amp;ias,ppraiopreir,eiqsutihliabtrisutmo c&amp;whaitsht&amp;iictsm&amp;coenthtiondusinogf&amp;sfetartaes&amp;of&amp;learned&amp;di
perturbs&amp;the&amp;equilibrium.&amp;The&amp;cognitive&amp;system&amp;then&amp;infers,&amp;using&amp;prior&amp;and&amp;posterior&amp;
information and an expert’s a priori cognitive equilibrium.</p>
      <p>Further, we contend that the greatest likelihood calculation
trFanosiustorrise,&amp;orr&amp;ttinrkaenr&amp;tsoyfso&amp;S&amp;rams&amp;weoll&amp;afs&amp;ttoh&amp;thee&amp;ltanimguaegess&amp;ulsiecd&amp;eto&amp;mdaaktea&amp;it&amp;r(utnr.&amp;Iat&amp;nsimspfloy&amp;rremqueiredd)&amp;thie&amp;nretsopoWnsehs&amp;en&amp;presented&amp;with&amp;novel&amp;information&amp;characterizing&amp;a&amp;new&amp;diagnostic&amp;situation,&amp;th
diagnostic reasoning. This includes the computation of a
greatest likelihood measure of hypotheses, given new
Modern&amp;AI&amp;research&amp;has&amp;proposed&amp;probabilistic&amp;representations&amp;and&amp;algorithms&amp;for&amp;the&amp;realStime&amp;integration&amp; is cognitively plausible and offers an epistemological
model,&amp;until&amp;it&amp;finds&amp;convergence&amp;or&amp;equilibrium,&amp;in&amp;the&amp;form&amp;of&amp;a&amp;particular&amp;greatest&amp;li</p>
      <p>framework for understanding the phenomena of human
1968).&amp;Among&amp;these&amp;algorithms&amp;is&amp;loopy,belief,propagation&amp;(Pearl&amp;1988,&amp;2000)&amp;that&amp;captures&amp;a&amp;system&amp;of&amp;</p>
      <p>The claim of this paper is that stochastic methods offer a sufficient account of hu
plausible&amp;beliefs&amp;constantly&amp;iterating&amp;towards&amp;equilibrium,&amp;or&amp;equilibration,&amp;as&amp;Piaget&amp;might&amp;describe&amp;it.&amp;A&amp;
When&amp;presented&amp;with&amp;novel&amp;information&amp;characterizing&amp;a&amp;new&amp;diagnostic&amp;situation,&amp;this&amp;a&amp;posteriori&amp;data&amp;
cognitive&amp;system&amp;can&amp;be&amp;in&amp;a,priori,equilibrium&amp;with&amp;its&amp;continuing&amp;states&amp;of&amp;learned&amp;diagnostic&amp;knowledge.&amp;</p>
      <p>areas such as diagnostic reasoning. This includes the computation of a greatest lik
perturbs&amp;the&amp;equilibrium.&amp;The&amp;cognitive&amp;system&amp;then&amp;infers,&amp;using&amp;prior&amp;and&amp;posterior&amp;components&amp;of&amp;thhe&amp;ypotheses, given new information and an expert’s a priori cognitive equilibrium.
diagnostic and prognostic reasoning.</p>
      <p>&amp;
that the greatest likelihood calculation is cognitively plausible and offers an episte
The claim of this paper is that stochastic methods offer a sufficient account of human intelligence in
areas such as diagnostic reasoning. This in&amp;cludes the computation of a greatest likelihood measure foofr understanding the phenomena of human diagnostic and prognostic reasoning.
O
t</p>
      <p>Bayes, T. (1763), Essay Towards Solving a Problem in the
Doctrine of Chances, Philosophic Transactions of the Royal
Society of London, London: The Royal Society, pp 370-418.</p>
      <p>Chakrabarti, C., Rammohan, R., and Luger, G. F. (2005), ‘A
First-Order Stochastic Modeling Language for Diagnosis’
Proceedings of the 18th International Florida Artificial
Intelligence Research Society Conference, (FLAIRS-18). Palo
Alto: AAAI Press.</p>
      <p>Chakrabarti, C., Pless, D. J., Rammohan, R., and Luger, G. F.
(2007), ‘Diagnosis Using a First-Order Stochastic Language That
Learns’, Expert Systems with Applications. Amsterdam: Elsevier
Press. 32 (3).</p>
      <p>Dempster, A.P. (1968), ‘A Generalization of Bayesian Inference’,
Journal of the Royal Statistical Society, 30 (Series B): 1-38.</p>
      <p>Glymour, C. (2001). The Mind's Arrows: Bayes Nets and
Graphical Causal Models in Psychology, MIT Press, 2001</p>
      <p>Gopnik, A. (2011a). A unified account of abstract structure and
conceptual change: Probabilistic models and early learning
mechanisms. Commentary on Susan Carey "The Origin of
Concepts." Behavioral and Brain Sciences 34,3:126-129.</p>
      <p>Gopnik, A. (2011b). Probabilistic models as theories of children's
minds. Behavioral and Brain Sciences 34,4:200-201.</p>
      <p>Hume, D. (1739/2000). A Treatise of Human Nature, edited by D.</p>
      <p>F. Norton and M. J. Norton, Oxford/New York: Oxford
University Press.</p>
      <p>Jurasky, D. and Martin, J. M. (2009), Speech and Language
Processing, Upper Saddle River NJ: Pearson Education.</p>
    </sec>
    <sec id="sec-3">
      <title>Pearl, J. (2000), Causality, Cambridge University Press. UK:</title>
    </sec>
    <sec id="sec-4">
      <title>Cambridge Peirce, C.S. (1958), Collected Papers 1931 – 1958, Cambridge MA: Harvard University Press.</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Piaget</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>1970</year>
          ), Structuralism, New York: Basic Books. Simon,
          <string-name>
            <surname>H. A.</surname>
          </string-name>
          (
          <year>1981</year>
          ),
          <source>The Sciences of the Artificial (2nd ed)</source>
          , Cambridge MA: MIT Press.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Stern</surname>
            ,
            <given-names>C.R.</given-names>
          </string-name>
          and Luger,
          <string-name>
            <surname>G. F.</surname>
          </string-name>
          (
          <year>1997</year>
          ),
          <article-title>Abduction and Abstraction in Diagnosis: A Schema Based Account</article-title>
          . In Android Epistemology,
          <string-name>
            <surname>Ford</surname>
          </string-name>
          et al. eds, Cambridge MA: MIT Press.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Turing</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>1950</year>
          ), 'Computing Machinery and Intelligence', Mind,
          <volume>59</volume>
          ,
          <fpage>433</fpage>
          -
          <lpage>460</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>von Glaserfeld</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          (
          <year>1978</year>
          ),
          <article-title>'An Introduction to Radical Constructivism', The Invented Reality</article-title>
          , Watzlawick, ed., pp17-
          <fpage>40</fpage>
          , New York: Norton.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>