<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Cognitive Learning Model that Combines Feature Formation and Event Prediction</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Eman Awad</string-name>
          <email>eman.awad@ucdconnect.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fintan Costello</string-name>
          <email>fintan.costello@ucd.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Computer Science and Informatics, University College of Dublin Bel eld</institution>
          ,
          <addr-line>Dublin 4</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents a computational model that addresses two central aspects of cognitive learning: feature formation and prediction. Most current cognitive models fail to account for these two aspects of learning; this model, based on frequentist probability theory, provides a uni ed account of feature formation and temporal prediction. We discribe a computational model that learns categorical and temporal relationships between events to form features and make predictions about future events, with the aim of simulating human learning and prediction processes. With this novel learning mechanism, we aim to provide for further insight into the complex cognitive process of human reasoning.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Human learning is a complex process involving di erent cognitive sub-processes.
Some studies have demonstrated that the formation of categories based on the
probabilistic occurrence of often complex features is central to human learning
and inference [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]; humans form features from perceived events and learn
the categorical relationships between them, which are then used to make
predictions [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ], [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Other studies have demonstrated that learning
temporal relationships between events (associations between features perceived
across time) is also vital to human learning and prediction [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ],
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        With this paper we aim to present a computational model, which integrates
categorisation and temporal prediction or conditioning processes as a uni ed
account of learning. Although categorisation is key to the learning process, typical
categorisation models are unable to realistically simulate the human temporal
learning process as they only categorise static features [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ], [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ].
This is insu cient as humans continually observe a dynamic environment, thus
a realistic model should incorporate the ability to learn temporal and categorical
relationships.
      </p>
      <p>
        Category formation and event prediction are both instances of
probabilistic reasoning; there is evidence that probabilistic reasoning in humans follows
frequentist probability theory [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], a mathematical process involving the
evaluation of past events in terms of their signi cance in predicting a future event.
The model in this paper follows the approach taken by Costello and Watts [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ],
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The experimental results in these papers repeatedly showed that the biases
observed in people's estimated probability could be explained by a model where
people's reasoning follows frequentist probability theory but is subjected to
random noise. This evidence suggests that people reason in a manner consistent
with the frequentist probability theory and it is this theory that we will apply
on our model.
      </p>
      <p>The aim of this research is to create a model designed to form features,
learn the categorical and temporal relationships between them, and predict
future events following frequentist probability theory. The prediction mechanism in
our model involves two stages: forming features and then calculating predictions.
In the rst stage, relationships between events or combinations of events will be
evaluated in a process referred to as the model's statistical evaluation process.
The events that represent a reliable relationship are selected and formed into
features. The temporal and categorical relationship between constituent events
of the features is also learned. This process determines which past events are
signi cant in predicting future events. In the second stage, the model combines
predictions from relevant reliable features to calculate the overall predicted
probability value for a given event.</p>
      <p>The rst section of this paper will delineate the evaluation process used to
form reliable features; a description of the prediction mechanism will also be
provided. The second section outlines an overview of the model design. Finally,
we will test the model's performance to evaluate the prediction mechanism and
present the analysis of the results.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Reliability: Feature Formation</title>
      <p>
        Our model is underpinned by the theoretical assumption that the human learning
process follows frequentist probability theory; humans evaluate past experiences
based on the statistical reliability of previously learned features to inform their
predictions about what may occur next in the environment [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. We argue
that this evaluation process is based on statistical reasoning to form predictive
features, which will be used in the prediction calculation. We illustrate this
formation of features using an example in Fig(1), where we aim to evaluate the
reliability of the relationship between two events. This gure shows the
occurrence of events A and S across time. After the rst occurrence of A followed by
S, an event N was created to represent the possibility of a relationship between
A and S; the probability of N will change as it occurs again. For example,
consider time t=8 with ve events perceived in this period. Event A occurred three
times in this period; in two of the three times, event A was followed by event S
after one second. At t=8s, the overall probability of S occurring is two out of
eight; p(S) = 0:25. The question is whether the occurrence of S after A can be
explained by the base rate occurrence of S alone; P (S) = 0:25. If it cannot be
explained by P (S), we assume a reliable causal relationship between A and S.
Otherwise, the relationship between the two events is not reliable and A does
Fig. 1: (a) illustrates a series of events occurring across time. Time is measured in
seconds. (b) shows the frequencies of events changing across time. At time 8s, the
binomial result was higher than (0:05), indicating the unreliability of A in predicting
S. However, at time 16s, the result was less than (0:05), indicating the reliability of
feature A in predicting S. Filled circle represents reliable predictor N , indicated here
as (r).
not predict S. Based on frequentist probability theory, the binomial probability
function, is used to determine the reliability of the relationship [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. We calculate
this using the following equation:
      </p>
      <p>Bin(k; x; p) =
x
k
pk(1
p)x k
(1)</p>
      <p>Where x represents the number of times event A occurred (which may or
may not be followed by event S), k represents the number of times that the
consequent event S followed the antecedent event A, and p is the base rate
probability of the consequent event P (S). This function represents the chance of
drawing a sample of x events from a population where p = P (S), and of which
k of them are instances of S.</p>
      <p>Based on frequentist probability theory, if the binomial probability result
Bin(k; x; p) is less than or equal to the statistical signi cance level of (0:05),
it indicates that the occurrence of S after A cannot be explained by the base
rate occurrence of S, P (S) = 0:25, and we can then deduce that the relationship
between A and S is causal and reliable. However, if the binomial result is higher
than 0:05, then there is insu cient evidence to prove that the relationship is
reliable at that time.</p>
      <p>At time = 8s in Fig(1), the binomial result Bin[2; 3; 0:25] = 0:14 is higher
than the signi cance level (0:05), indicating that there is not enough evidence
to prove that there is a causal relationship between A and S. As such, event A
is not a reliable predictor of S at this time.</p>
      <p>If we consider time t = 16s, event A occurred ve times within this period; in
four of these ve times, A was followed by event S after one second. The question
again is whether this occurrence of S after A can be explained by the base
rate occurrence of S; P (S) = 0:25. Again, we calculate the binomial probability
Bin[4; 5; 0:25] = 0:01, which is lower than the signi cance level (0:05), indicating
that there is enough evidence to prove that the relationship is reliable. Thus, at
this time feature A is considered a reliable predictor of S; we designate this as (r)
in Fig(1). This example illustrates the formation of a reliable feature N , which
indicates that if A occurs, there is 54 chance of seeing S one second later. (In
this example the time interval between the events is one second. This interval
was chosen for illustrative purposes. The model uses various di erent intervals,
which can be larger than 1s.)
Complex Features: The earlier example assumed that we only have a single
event as an antecedent and a single event as a consequent. In reality, we can
encounter more complicated events that have more than a single event as an
antecedent. In the case of a complex feature, the complex event N holds a
complex antecedent that includes two sub-events or more: event A as antecedent and
another event B as a consequent. The two sub-events, A and B are the complex
antecedent of N, and the consequent of that event is S (as shown in Fig.2).</p>
      <p>To form complex features, we must rst determine if the complex event
represents a reliable relationship or not. To do this, we use Eq.(1) where x represents
the number of times the complex antecedent (A then B) occurred and k
represents the number of times that the complex antecedent was followed by the
consequent S. In determining the relationship between complex features and S,
there are three probabilities that we must test the occurrences against: (1) the
probability of S, P (S) which was exactly as de ned previously in simple events;
(2) the probability of event n1 i.e. S following the antecedent A alone; P (SjA),
and (3) the probability of event n2 i.e. S following the antecedent B alone;
P (SjB). In this case, the question is whether the occurrence of S after (A then
B) can be explained by the base rate occurrence of S, P (S) = 0:28, or by the
observed rate occurrence of the simple features n1, P (SjA), and n2, P (SjB). If it
cannot be explained by these probabilities i.e. P (S), P (SjA) and P (SjB), we
deduce that there is a reliable causal relationship between the complex antecedent
and S. Otherwise, the relationship is not reliable and the complex antecedent (A
and B) together does not speci cally predict the occurrence of S.</p>
      <p>As a working example, see Fig(2). Here we consider three di erent features
that potentially predict the occurrence of S: a simple feature n1 (which predicts
that S will occur 2 seconds after the occurrence of A), a simple feature n2 (which
predicts that S will occur 1 second after the occurrence of B), and a complex
feature N (which predicts that S will occur 1 second after the occurrence of the
complex antecedent 'A then B'). At time t = 21s, we can ask whether n1 and
n2 are reliable features (whether they represent statistically reliable
relationships between A and S or B and S) by using the binomial function as before,
taking the base-rate occurrence of P (S) = 0:28. The binomial test results are
Bin[6; 7; 0:28] = 0:001 and Bin[4; 5; 0:28] = 0:019 respectively, indicating that
both n1 and n2 are reliable features: A reliably predicts S after 2 seconds (with
probability 0:85) and, independently, B reliably predicts S after 1 second (with
probability 0:8).</p>
      <p>When testing the complex feature N , which represents the complex
antecedent (A and B) predicting the consequent S, we ask whether the observed
rate of occurrence of S after this complex antecedent can be explained by the
independent occurrence of these two predictors A and B. If the observed rate of
occurrence of S after this complex feature cannot be explained by the
independent occurrence of these two separate predictors, this indicates that there is some
further causal relationship between the speci c complex event (A and B) and the
occurrence of S, and the complex feature N will be marked as a reliable
predictor. At time t = 21s feature n1 predicts S with probability P (SjA) = 0:85, and
feature n2 predicts S with probability P (SjB) = 0:8. The complex antecedent (A
then B) has occurred ve times by time t = 21s, with four of those occurrences
followed by S. This rate of occurrence is explained both by the predictor n1
(Bin[4; 5; 0:85] = 0:39) and separately by the predictor n2 (Bin[4; 5; 0:8] = 0:4).
This indicates that there is insu cient evidence to prove that the relationship is
reliable between the speci c complex antecedent (A and B) and the occurrence
of S; the integrated event N is not a reliable predictor of S (i.e. not a reliable
feature).</p>
      <p>In other situations, the complex feature N might be reliable. For example,
consider a di erent situation, where the complex antecedent (A and B) occurred
ve times. In all of those times, S followed this complex antecedent, after one
second (i.e. the integrated event, N). However, event A and B occurred frequently
without being followed by S (A occurred ten times and B occurred fteen times).
Consider the base rate of the consequent and the probability of individual events;
P (S) = 0:16, P (SjA) = 0:5 and P (SjB) = 0:33 respectively, the binomial result
for the complex event N are Bin[5; 5; 0:16] = 0:0001, Bin[5; 5; 0:5] = 0:003, and
Bin[5; 5; 0:33] = 0:004, which are all lower than the statistical signi cance level
(0:05). In this case, A alone does not predict S and B alone does not predict S.
However, both features A and B together will be a reliable predictor of S.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Predicted Probability</title>
      <p>The set of independent predictive features identi ed by the evaluation process
that was discussed earlier would subsequently be used in the prediction
calculation. For the purpose of this discussion, we will consider the example in Fig(2)
where the simple features n1 and n2 were both reliable predictors of S, and
the integrated event N that connects these two predictors (A and B) together,
was not a reliable predictor. The aim of prediction calculation is to calculate
the overall probability of S occurring next, after accounting for all recognised
reliable predictors for this event. According to standard frequentist probability
theory, if there is a set of independent features predicting an event, the
overall predicted probability is calculated by OR-ing all the independent predictors
using the following OR expression:</p>
      <p>P r(Si;t ) = 1
(1
pi)
(2)
where pi is the probability of event Si occuring, given by each independent
predictor, at a given time (t).</p>
      <p>In this example, as previously explained, the two predictors A and B are
independent from each other because the integrated event N that connects them
together, does not reliably predict S. Since they are independent, we can
combine the probabilities of the independent predictors, (P (SjA) and P (SjB)) to
calculate the overall prediction of S, using Eq.2 as follows:</p>
      <p>P r(S) = 1
[(1
( 76 )) (1</p>
      <p>( 54 )]= 0:97</p>
      <p>This indicates that the overall probability of S occurring after the reliable
predictor (simple feature n1 and n2 independently) at a given time is high.
However, at some point in the future, event N may become reliable. In this
case, the individual predictors (A followed by S and B followed by S), are no
longer independent predictors of S, and therefore, they will not be used in the
calculation for the prediction of S. The integrated feature N will be used in the
OR expression instead of n1 and n2, to calculate the overall prediction of S.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Model Implementation</title>
      <p>We utilise frequentist probability theory to build a uni ed prediction model that
forms features; the prediction mechanism in this model uses these features to
make predictions. In our model, we divide the memory into short-term (ST M )
and long-term memory (LT M ). The LTM holds a record of all the perceived
events and the relationships between them. These events will be stored in the
form of nodes. Each node contains an antecedent (earlier sub-event being
perceived), a consequent (the later sub-event being perceived), the time interval
between them (t), the total number of times the antecedent was followed by the
consequent, with the exact time interval t (represented as counter (k)), and the
total number of times the antecedent was perceived (represented as counter (x)).
On the other hand, the STM holds the most recent events only. Considering the
limited capacity of STM in humans, we simulate this computational limitation
by designing a limited STM in the models memory. The STM is also responsible
for updating the counters (x and k), for all related nodes, when new events are
perceived.</p>
      <p>In order to calculate the prediction of any given event S, the model selects
all nodes that have S as a consequent and the antecedent of the node occurring
in the STM. The binomial function will be applied on all the selected nodes,
using Eq.2, to evaluate the relationship between all the antecedents and S; this
will enable the model to identify all the independent predictors that will reliably
predict S. At the end, all the independent predictors will be combined to calculate
the overall predicted probability of S using Eq.2.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Evaluating the Models Prediction Mechanism</title>
      <p>
        We developed a simulation tool we refer to as the Random Generator (RG)to
test this model. RG generates test data simulating real world environment by
producing a sequence of events with various randomly generated temporal
relationships linking the occurrence of events together. This test data generated by
the RG will be used to evaluate our model's performance. In the initialisation
phase of the RG tool, it assigns random probabilities to di erent temporal
relationships between a large number of events and combinations of events. The next
event is produced in a way that is dependent probabilistically on the sequence
of prior events. The RG tool embodies a probabilistic context - sensitive
grammar with long-range dependencies [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Our model will perceive the sequence
of events generated by the RG tool and learn the relationships between these
events.
      </p>
      <p>In this test, our model will be challenged to identify the various patterns
within the sequence of events generated by the RG tool. The learning mechanism
in our model aims to learn and recognise reliable relationships, to form features
and estimate the predicted probability of future occurrences. If the learning
mechanism functions e ectively, it should recognise and successfully learn the
hidden relationships between these events.</p>
      <p>We will test the model by evaluating all the event predictions. Those
predictions will be compared to the actual events that occurred succeeding the
antecedents. For example, if the model predicts that event Si will occur with a
probability of 0:2, P r(Si; t) = 0:2, in the test, all the times the model makes
a prediction of an event occurring with this probability, will be counted. If the
prediction matches the actual event i.e. Si does actually occur with a probability
of 0:2, then the model's prediction is considered to be accurate.
5.1</p>
      <sec id="sec-5-1">
        <title>Material</title>
        <p>The testing process involves two phases, the learning phase and the prediction
phase. In the rst phase, we run the model j number of times (where j = 10.000),
to observe and learn the relationships between the sequence of events generated
by the RG tool. Events with reliable relationships between them are formed into
a feature. In the second phase, we will run the model for N number of times
(where N = 50.000), to calculate and obtain the prediction P r(Si;t ) at each
time-step.</p>
        <p>To examine the e ciency of the prediction values, we calculate the proportion
of the predicted events observed in the test data. First, we examine all previously
perceived events and evaluate the prediction value given by the model for each of
those events by counting the number of times that the predicted event actually
happens. Subsequently, we compare the rate of occurrence of this real event (i.e.
the proportion of its observed occurrence) with the prediction value P r(Si;t )
calculated by the model. The accuracy of the model's prediction mechanism can
be assessed by comparing these two values.
5.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>Method</title>
        <p>After observing the sequence of events generated by the RG tool, the model's
prediction mechanism calculates the predicted probability for each event
occurring at every time-step. This is represented by an overall predicted probability
value P r(Si;t ) for each event Si occurring at time t. We examine the value
of P r(Si;t ) in di erent ranges. We separate the P r(Si;t ) values into range R
speci ed in Table (1).</p>
        <p>We then count the total number of occurrences that are predicted to occur
with probability value P r(Si;t ) falling between a speci c range R; (Rmin
P r(Si;t ) Rmax) i.e. PjN jP r(Si; t) 2 Rj. We calculate this value for all
the R (see Table 1). For each range R, we identify all these occurrences and
count the total number of times the predicted event actually does occur i.e.
PN</p>
        <p>t=j jP r(Si;t ) 2 R \ Si;t j. Subsequently, we calculate proportion of those cases
in which the predicted event actually occurs i.e. O(Si;t ; R). To calculate this,
we divide the total number of times that the event does occur in each range
PtN=j jP r(Si;t ) 2 R \ Si;t j by the total number of times the event is predicted
to occur in the range PjN jP r(Si; t) 2 Rj as shown by the following formula:
O (Si; R) = j PjN jP r(Si;t ) 2 R \ Si;t j
j PjN P r(Si;t ) 2 Rj
(3)
Where P r(Si;t ) is the predicted probability of event i at time t; j is the starting
time (measured in seconds) and N is ending time (measured in seconds).</p>
        <p>For example, for range R = 0:20 0:25 the total number of times the event
was predicted to occur with this probability ((0:20 P r(Si;t ) 0:25)) was
3016 events i.e. PjN jP r(Si; t) 2 Rj = 3016. Out of these 3016 cases, only 643
actually occurred i.e. PN</p>
        <p>t=j jP r(Si;t ) 2 R \ Si;t j = 643, therefore the proportion
of occurrences for the predicted events O (Si; R) was 0:21.</p>
        <p>We assess the accuracy of the model's prediction by comparing this
proportion of occurrences for the predicted events O (Si; R) to the range it is supposed
to fall within (R). If the model predicts that an event will occur with a
probability of 0.15, and the proportion (given by O (Si; R) for the predicted events)
is close to 0:15 (e.g. 0:18), this indicates that the model is predicting the next
event accurately. The proportion should match or be close to the range R-value
range.</p>
        <p>Since the accuracy assessment for the model's prediction requires comparison
between O(Si;t ; R) and the range R, we used the absolute di erence function as a
measure of the comparison between the two. We measure the absolute di erence
between the mid point of the range M (R) and the proportion of observations
O(Si;t ; R) i.e. jM (R) O(Si;t ; R)j.</p>
        <p>R</p>
        <p>PjN P r(Si;t ) 2 R PjN jP r(Si;t ) 2 R \ Si O(Si;t ; R) jM (R)
O(Si;t ; R)j
The third column shows PjN P r(Si;t ) 2 R \ Si;t the total number of times that the
event did occur in R. The fourth column shows the overall proportion of observations
that were predicted O(Si;t ; R). Finally, the last column is the accuracy measure for
the model's prediction as the absolute di erence value jM (R) O(Si;t ; R)j.
We found that the level of agreement between the proportion of observations
for predicted events O(Si;t ; R) and the corresponding range R was high; the
overall correlation between O(Si;t ; R) and mid-point of the range M (R) was
very high (r = 0:97; (O;M(R)) &lt; 0:0001). In general, the O(Si;t ; R) fell within
the corresponding range R or was close to R. For example, for P r(Si;t ) values
between 0:15 0:20 i.e. (0:15 P r(Si;t ) 0:20), the value of the
corresponding proportion O i.e. O(Si;t ; 0:15 0:20) = 0:19 fell within the same R range
as the P r(Si;t ). In some cases, the value of O fell just of outside the
corresponding range. For example, in the case where P r(Si;t ) value was between
0:05 0:10 (i.e. (0:05 P r(Si;t ) 0:10)), the value of the corresponding O
i.e.O(Si;t ; 0:05 0:10) = 0:14 fell slightly outside of the R range of the P r(Si;t ).
However, even though this value does not exactly fall within the range, it is
close to the Rmax (Rmax = 0:10). Overall, our results suggest that in general,
the prediction mechanism in our model is e cient i.e. the model is able to predict
e ectively.</p>
        <p>We observe that there is a relationship between the total number of
occurrences in the range (PjN P r(Si;t ) 2 R) and the proportion of observations
O (O(Si;t ; R)); the bigger the number of cases (i.e. the bigger the value of
PjN P r(Si;t ) 2 R) or the sample size, the closer the proportion of
observations values O was to the corresponding range. The measure that we have for
accuracy assessment of the model's prediction was jM (R) O(Si;t ; R)j. With
this, we can estimate how the accuracy of the model's prediction changes as the
sample size varies. For example, for range (0:50 R 0:55), the total
number of occurrences (PjN P r(Si;t ) 2 R) was only 171. In this case, the absolute
di erence (jM (R) O(Si;t ; R)j) was 0.151, indicating that the accuracy of the
prediction is low. However, for range (0:25 R 0:30), the total number of
occurrences (PjN P r(Si;t ) 2 R) was 3901. In this case, the absolute di erence
(PjN P r(Si;t ) 2 R) was 0.005, indicating that the accuracy of the prediction is
high. We can deduce from this that the di erence between O and R is smaller
when the model was provided with a bigger sample size. This suggests that the
model produces more accurate predictions when more information is available
about the likelihood of occurrences, which is not surprising considering that this
information can be used in the prediction calculation P r(Si; t). We note that
the low prediction accuracy observed in some ranges coincided with the range
R with lower number of occurrences (PjN P r(Si;t ) 2 R) i.e. the model has only
perceived those events a small number of times. This indicates that the model is
accurately predicting the probability in general and that the anomalous results
are expected due to the small sample size.</p>
        <p>In conclusion, our test demonstrated that the model has learned the
relationships between events e ciently and is capable of calculating relatively accurate
predictions based on these learned relationships.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>
        In this paper, we propose an alternative approach to cognitive modeling. We
design the learning mechanism so that it is able to identify reliable relationships
between events, form features, and then categorise them in a way that simulates
human learning as much as possible. We also incorporate in our model's
prediction mechanism the frequentist probability theory, which has been shown to be
applicable in human prediction processes [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        There are various alternative models in the literature that are successful
in categorising and making predictions over time. Arti cial Neural Networks
(ANNs), using deep learning algorithms, are currently the most commonly used
model in the eld of prediction. While the model we present here is an account
of human learning and is not intended to compete with such machine learning
algorithms, a brief comparison is worthwhile. The main di erence between our
model and ANNs is that ANNs are associative models [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] whereas ours
is representational. Unlike ANNs, our model includes explicit features that can be
extracted and easily used. Moreover, as Gallistel and Gibbon highlighted [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ],
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], associative models create associations between two events that are frequently
perceived together with a short time interval between them but these associations
do not necessarily represent reliable relationships between the events, as these
relationships are not evaluated. In contrast, representational models represent
the learned data and store it in the "if-then rule" format (if an event A occurs,
then after a speci c time interval t, event B will occur) [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. In representational
models such as ours, the model can distinguish between causal and coincidental
relationships through statistical calculation. In addition, representational models
can learn temporal relationships even with a large time interval between events,
to verify if there is a causal relationship between them.
      </p>
      <p>In summary, we argue that the prediction model presented in this paper is
a more realistic simulation of human learning than ANNs. The simulation is
more realistic as its evaluation process is designed using frequentist probability
theory that enables the model to form speci c meaningful features that are easily
accessible and can be used for prediction, unlike ANNs which lack this process.</p>
      <p>In future work, we plan to conduct further tests to improve the learning
model and incorporate additional features. The experiments and tests would
focus on the models ability to process more complex data to ensure that the
prediction results are as close to human predictions as possible. We also plan to
integrate the ability to predict actions as well as perceptions into our prediction
mechanism with the overall aim of enabling our model to make decisions and
perform goal driven actions.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Balsam</surname>
            ,
            <given-names>P.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Drew</surname>
            ,
            <given-names>M.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gallistel</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Time and associative learning</article-title>
          .
          <source>Comparative Cognition &amp; Behavior Reviews 5</source>
          ,
          <issue>1</issue>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Costello</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Watts</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Surprisingly rational: probability theory plus noise explains biases in judgment</article-title>
          .
          <source>Psychological review 121(3)</source>
          ,
          <volume>463</volume>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Costello</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Watts</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Peoples conditional probability judgments follow probability theory (plus noise)</article-title>
          .
          <source>Cognitive Psychology</source>
          <volume>89</volume>
          ,
          <issue>106</issue>
          {
          <fpage>133</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Gallistel</surname>
            ,
            <given-names>C.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gibbon</surname>
          </string-name>
          , J.:
          <article-title>Time, rate, and conditioning</article-title>
          .
          <source>Psychological review 107(2)</source>
          ,
          <volume>289</volume>
          (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Gallistel</surname>
            ,
            <given-names>C.R.:</given-names>
          </string-name>
          <article-title>The organization of learning</article-title>
          . The MIT Press (
          <year>1990</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Gallistel</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gibbon</surname>
          </string-name>
          , J.:
          <article-title>Computational versus associative models of simple conditioning</article-title>
          .
          <source>Current Directions in Psychological Science</source>
          <volume>10</volume>
          (
          <issue>4</issue>
          ),
          <volume>146</volume>
          {
          <fpage>150</fpage>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Geman</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Johnson, M.:
          <article-title>Probabilistic grammars and their applications</article-title>
          .
          <source>International Encyclopedia of the Social &amp; Behavioral Sciences</source>
          <year>2002</year>
          ,
          <volume>12075</volume>
          {
          <fpage>12082</fpage>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8. Gri ths, T.,
          <string-name>
            <surname>Yuille</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>A primer on probabilistic inference. The probabilistic mind: Prospects for Bayesian cognitive</article-title>
          science pp.
          <volume>33</volume>
          {
          <issue>57</issue>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Hampton</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          :
          <article-title>Similarity-based categorization: the development of prototype theory</article-title>
          .
          <source>Psychologica Belgica</source>
          <volume>35</volume>
          (
          <issue>2-5</issue>
          ),
          <volume>104</volume>
          {
          <fpage>125</fpage>
          (
          <issue>1995b</issue>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Haykin</surname>
            ,
            <given-names>S.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haykin</surname>
            ,
            <given-names>S.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haykin</surname>
            ,
            <given-names>S.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haykin</surname>
            ,
            <given-names>S.S.:</given-names>
          </string-name>
          <article-title>Neural networks and learning machines</article-title>
          , vol.
          <volume>3</volume>
          . Pearson Upper Saddle River, NJ, USA: (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Hodges</surname>
            <given-names>Jr</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>J.L.</given-names>
            ,
            <surname>Lehmann</surname>
          </string-name>
          ,
          <string-name>
            <surname>E.L.</surname>
          </string-name>
          :
          <article-title>Basic concepts of probability and statistics</article-title>
          .
          <source>SIAM</source>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Holdershaw</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gendall</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Understanding and predicting human behaviour</article-title>
          .
          <source>In: ANZCA08 Conference</source>
          , Power and Place.
          <source>Wellington</source>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Kruschke</surname>
            ,
            <given-names>J.K.</given-names>
          </string-name>
          :
          <article-title>Alcove: An exemplar-based connectionist model of category learning</article-title>
          .
          <source>Psychological Review</source>
          <volume>99</volume>
          (
          <issue>1</issue>
          ),
          <volume>22</volume>
          {
          <fpage>44</fpage>
          (
          <year>1992</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>La</surname>
            <given-names>erty</given-names>
          </string-name>
          , J.,
          <string-name>
            <surname>Sleator</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Temperley</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Grammatical trigrams: A probabilistic model of link grammar</article-title>
          , vol.
          <volume>56</volume>
          . School of Computer Science, Carnegie Mellon University (
          <year>1992</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Michel</surname>
            ,
            <given-names>A.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Farrell</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          :
          <article-title>Associative memories via arti cial neural networks</article-title>
          .
          <source>IEEE Control Systems Magazine</source>
          <volume>10</volume>
          (
          <issue>3</issue>
          ),
          <volume>6</volume>
          {
          <fpage>17</fpage>
          (
          <year>1990</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Nan</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A learning model that combines categorisation and conditioning (2013), www</article-title>
          .summon.com, university Collage Dublin, School of Computer Science and Informatics
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Nosofsky</surname>
            ,
            <given-names>R.M.</given-names>
          </string-name>
          :
          <article-title>Attention, similarity, and the identi cation{categorization relationship</article-title>
          .
          <source>Journal of experimental psychology: General</source>
          <volume>115</volume>
          (
          <issue>1</issue>
          ),
          <volume>39</volume>
          (
          <year>1986</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Osberg</surname>
            ,
            <given-names>T.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shrauger</surname>
            ,
            <given-names>J.S.</given-names>
          </string-name>
          :
          <article-title>Self-prediction: Exploring the parameters of accuracy</article-title>
          .
          <source>Journal of Personality and Social Psychology</source>
          <volume>51</volume>
          (
          <issue>5</issue>
          ),
          <volume>1044</volume>
          (
          <year>1986</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Pavlov</surname>
            ,
            <given-names>I.P.:</given-names>
          </string-name>
          <article-title>Lectures on conditioned re exes. vol. ii. conditioned re exes and psychiatry</article-title>
          . (
          <year>1941</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Pavlov</surname>
            ,
            <given-names>I.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Anrep</surname>
            ,
            <given-names>G.V.</given-names>
          </string-name>
          :
          <article-title>Conditioned re exes</article-title>
          .
          <source>Courier Corporation</source>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Pavlov</surname>
            ,
            <given-names>I.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gantt</surname>
            ,
            <given-names>W.A.H.:</given-names>
          </string-name>
          <article-title>Lectures on conditioned re exes</article-title>
          . Liveright, New York (
          <year>1928</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Rosch</surname>
          </string-name>
          , E.:
          <article-title>Principles of categorization</article-title>
          . Concepts: core readings pp.
          <volume>189</volume>
          {
          <issue>206</issue>
          (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Savastano</surname>
            ,
            <given-names>H.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>R.R.</given-names>
          </string-name>
          :
          <article-title>Time as content in pavlovian conditioning</article-title>
          .
          <source>Behavioural Processes</source>
          <volume>44</volume>
          (
          <issue>2</issue>
          ),
          <volume>147</volume>
          {
          <fpage>162</fpage>
          (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Skinner</surname>
            ,
            <given-names>B.F.</given-names>
          </string-name>
          :
          <article-title>The behavior of organisms: An experimental analysis</article-title>
          .
          <source>(</source>
          <year>1938</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Skinner</surname>
            ,
            <given-names>B.F.</given-names>
          </string-name>
          :
          <article-title>The behavior of organisms: An experimental analysis</article-title>
          .
          <source>BF Skinner Foundation</source>
          (
          <year>1990</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Zettlemoyer</surname>
            ,
            <given-names>L.S.</given-names>
          </string-name>
          , Collins,
          <string-name>
            <surname>M.</surname>
          </string-name>
          :
          <article-title>Learning to map sentences to logical form: Structured classi cation with probabilistic categorial grammars</article-title>
          .
          <source>arXiv preprint arXiv:1207.1420</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Zurada</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          :
          <article-title>Introduction to arti cial neural systems</article-title>
          , vol.
          <volume>8</volume>
          .
          <string-name>
            <given-names>West</given-names>
            <surname>St. Paul</surname>
          </string-name>
          (
          <year>1992</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>