<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Learning of Time Varying Functions is Based on Association Between Successive Stimuli</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Function Learning</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lee-Xieng Yang</institution>
        </aff>
      </contrib-group>
      <fpage>722</fpage>
      <lpage>727</lpage>
      <abstract>
        <p>In function learning, the to-be-learned function is normally designed as time invariant. However, when the magnitudes of variable can be defined by time points, the function varies along time. Due to this difference in essence, the learning of the time-varying functions would be different from other functions. Specifically, the correlation between successive stimuli should play an important role for learning such functions. In this study, three experiments were conducted with the correlations set as positive high, negative high, and positive low. The results show people perform well when the correlation between successive stimuli is positive high or negative high. Also, people have difficulty learning the time-varying function with a low correlation between successive stimuli. A simple two-layered neural network model is evident to be able to provide good accounts for the data of all experiments. These results suggest that learning time varying function is based on association between successive stimuli.</p>
      </abstract>
      <kwd-group>
        <kwd>Function Learning</kwd>
        <kwd>Time Varying Function</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>We are living in an orderly world, in which variables are
mostly correlated with each other. For instance, the
probability of rain might be a function of the extent to which the
sky is overcast with dark clouds, or the distance to the car in
front needed to avoid a car crash is a function of the current
car speed. The study of how people learn a function and what
people form to represent a learned function is referred to as
function learning.</p>
      <p>
        There are also two contrasting theoretical accounts in
function learning. The rule-based account posits that people
construct abstract rules to summarize the ensemble of
experienced pairs of stimuli and responses used to teach the
function. Most frequently, polynomial rules have been
proposed as the representations of the mappings between
stimulus magnitudes and response magnitudes
        <xref ref-type="bibr" rid="ref3 ref7">(see Carroll, 1963;
Koh &amp; Meyer, 1991)</xref>
        . On the contrary, the
associativebased model assumes that people form direct associations
between each stimulus and corresponding response without
abstracting any summary information
        <xref ref-type="bibr" rid="ref2 ref2 ref4 ref4">(Busemeyer, Byun,
Delosh, &amp; McDaniel, 1997; DeLosh, Busemeyer, &amp; McDaniel,
1997)</xref>
        . However, the rule-based account overestimates the
participants’ performance in the extrapolation test but the
associative-based model underestimates it. To get a better
theoretical account, a hybrid model combining these two
approaches is proposed
        <xref ref-type="bibr" rid="ref9">(McDaniel &amp; Busemeyer, 2005)</xref>
        .
      </p>
      <p>
        Although these models differ on the assumption for the
type of representation formed in function learning, it is
basically agreed that the representation is formed for the whole
function. However, contrary to this idea, it was found
that people might form different representations for
different parts of the function, such that a quadratic function
was learned as the composition of two simpler monotonic
functions, which were chosen for use at different contexts
        <xref ref-type="bibr" rid="ref8">(Lewandowsky, Kalish, &amp; Ngang, 2002)</xref>
        . The POLE model
        <xref ref-type="bibr" rid="ref6">(Kalish, Lewandowsky, &amp; Kruschke, 2004)</xref>
        accounts for this
finding well, by virtue of its architecture consisting of many
modules, each of which represents a linear function
corresponding only to a small region of the function, and a gating
mechanism which always chooses one of the modules for use
according to the stimulus value. Strictly speaking, the real
function is not learned but approximated by the composition
of many smaller linear functions.
      </p>
      <p>
        Past studies have tested different functions and shown a
number of characteristics of function learning. First, the
linear functions are easier to learn than the nonlinear ones
        <xref ref-type="bibr" rid="ref2 ref4 ref7">(see
Busemeyer et al., 1997; Koh &amp; Meyer, 1991)</xref>
        . Second, it is
found that it is more accurate to predict the response for the
stimulus whose value falls in the training range (i.e.,
interpolation) than outside the range (i.e., extrapolation)
        <xref ref-type="bibr" rid="ref2 ref4 ref9">(see
Busemeyer et al., 1997; McDaniel &amp; Busemeyer, 2005)</xref>
        . Third,
although the function of simpler forms (e.g., linear or power
function) can be learned with the variables being of
nonnumeric forms (e.g., line length),
        <xref ref-type="bibr" rid="ref5">Kalish (2013)</xref>
        reported that
the periodic functions (e.g., sine function) cannot be learned
without the employment of numeric stimuli. These
characteristics reveal the limitations of human cognition for learning
the functional relation between variables.
      </p>
    </sec>
    <sec id="sec-2">
      <title>Time-Varying Function</title>
      <p>Although many forms of functions have been tested, a
particular form of function, which maps the timing of observation
to the event at that timing seems not to have been tested yet.
We call this function as time-varying function in this article,
y = f (t). An example of this function would be the height of
water accumulated in a bucket from a constant supply source.
If the bucket is cylindrical, the height will be a linear
function of time and if the bucket is conical, the height will be
a parabolic function of time. To our knowledge, how people
learn this kind of function has never been reported in
literature. However, a relevant case in category learning has been
reported recently.</p>
      <p>
        Navarro and his colleagues tested how people could learn
the categories when the category structure varies along
training trials. In one of their experiments, the members of two
categories moved up on the stimulus dimension constantly
along with the increase of trail number and the categorization
rule was set up as ”Respond A, if xt &gt; t and B otherwise”
for any item xt on trial t. Their results showed that
participants could not only learn this category structure, but also
be able to predict the item value on the next trial
        <xref ref-type="bibr" rid="ref10 ref11 ref12">(Navarro
&amp; Perfors, 2009, 2012; Navarro, Perfors, &amp; Vong, 2013)</xref>
        . It
is implied that people are able to capture some functional
relationships between the time point (or trial number) and the
stimulus value. However, the learning of the time-varying
functions might be different from the normal functions.
      </p>
    </sec>
    <sec id="sec-3">
      <title>Comparison Between Time-Varying Function and Normal Function</title>
      <p>There are some features of the time-varying functions worth
noting. First, due to that time can never return, when
learning a time-varying function, making a prediction for response
magnitude on each trial is always extrapolating what people
have learned. However, in the case of learning the function
y = f (x), both the interpolation and extrapolation tests can be
conducted.</p>
      <p>Second, a time-varying function can be viewed as a
function defining the relationships between successive stimuli,
xt = f (xt 1). A good example is the game of throwing a
Frisbee with friends. In this case, the only observable information
is the spatial position of the Frisbee at any time point.
Therefore, the best cue for us to estimate the position of the Frisbee
at time t is its position at time t 1.</p>
      <p>Third, the learnability or complexity of function would be
defined differently for the time-varying function. For the case
of y = f (x), the linear function has less parameters to
estimate than the quadratic function, hence being easier to learn.
For the case of y = f (t), learning the functional relationship
between time point to response magnitude is equivalent to
learning to predict the next response magnitude with the
current observed response magnitude. Thus, it is hypothesized
that the time-varying function would be easy to learn, if the
correlation between successive stimuli is high. If the
correlation between successive stimuli is low, it would be hard to
learn. To verify this hypothesis, three experiments were
conducted.</p>
    </sec>
    <sec id="sec-4">
      <title>Experiment 1</title>
      <p>In this experiment, we first examined whether people can
learn a linear time-varying function. The function was written
as xt = t + et , where t was trial number from 1 to 100 and e
was randomly sampled from the uniform distribution between
-0.5 and 0.5. All stimulus values were normalized between
15 and 15 for the convenience of computer programming. It
was reasonable to expect that this function could be learned
well, for (1) it was linear as well as (2) the correlation
between successive stimuli was high.</p>
      <sec id="sec-4-1">
        <title>Method</title>
        <p>
          Participants and Appartus There were in total 22
participants recruited from National Chengchi University in
Taiwan for this experiment. Each participant was reimbursed by
NTD$ 60 (' US$ 2) for their time and traffic expense. The
whole experiment was conducted on an IBM compatible PC
in a quiet booth. The processes of stimulus displaying and
response recording were under the control of a computer script
composed by PsychoPy
          <xref ref-type="bibr" rid="ref13">(Peirce, 2007)</xref>
          .
        </p>
        <p>Session 1
Session 2
15
10
5
1 7 131925313743495561677379859197
trial
1 7 131925313743495561677379859197
trial</p>
        <p>Procedure The participants were instructed that they were
playing a shooting game. In this game, they had to guess
the position of a target on a horizontal line on the computer
screen. On each trial, they moved the mouse cursor to where
they thought the target would appear. After they pressed the
space key to complete the guessing, the target would appear
as an arrow on the correct position, together with a feedback
text of ”Hit” or ”Miss” on the screen. The participants were
told that ”Hit” meant that your guess was close enough to the
true answer and otherwise you would get ”Miss”. The whole
experiment was conducted in two sessions, each of which
consisted of 100 trials. The same100 stimuli were presented
in the two sessions. The distance between the target’s correct
position and the participants’ guess was error. The amount
of squared error and the proportion of received ”Hit” (e.g.,
accuracy) were the dependent variable in this experiment.
Visual inspection on Figure 1 shows that participants
performed quite well except for the very early trials1. For
simplifying the complexity of data analysis, we divided the 100
stimuli to 10 blocks. The squared prediction error decreases
from 40.29 to 0.03 with the mean = 4:06 through 10 blocks
across two sessions. A Block (10) Session (2)
withinsubjects ANOVA reveals a significant main effect of Block
on the squared error [F(9; 189) = 72:83, MSe = 98, p &lt; :01],
no significant main effect of Session [F(1; 21) = 2:367, MSe
= 166:30, p = :139], and a significant interaction effect
between Block and Session [F(9; 189) = 2:346, MSe = 166:3,
p &lt; :05].</p>
        <p>The participant’s accuracy is another dependent variable,
which is computed as the number of ”Hit” divided by all
trials. Due to the ”Hit” range was very small in our
experiments, the highest accuracy in a block was .63 and the
lowest was .36 across all sessions. A Block (10) Session
(2) within-subjects ANOVA shows a significant main effect
of Block on the accuracy [F(9; 189) = 8:281, MSe = 0:028,
p &lt; :01], no significant main effect of Session [F(1; 21) &lt; 1],
and a significant interaction effect between Block and Session
[F(9; 189) = 5:052, MSe = 0:027, p &lt; :01].</p>
        <p>We also check the correlation between each participant’s
predictions and the true answers. The averaged Pearson’s r
across all participants is quite high [r = :97]. Together with
the visual inspection on Figure 1, it is confirmed that people
can learn the linear time-varying function very well.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Experiment 2</title>
      <p>In this experiment, the function was set up as xt = 50 +
( 1)t p100 t, which made the target jump left and right,
gradually moving toward the central point. Obviously, this
function was far more complex than the one used in
Experiment 1 and it was nonlinear. If the learning of y = f (t)
shared the same characteristics of the learning of y = f (x),
it should be expected that this function could not be learned
well. However, if our discussion about the characteristics of
time-varying function was right, it should be expected that
this function could be learned well, due to high correlation
between successive stimuli [r = :99].</p>
      <sec id="sec-5-1">
        <title>Method</title>
      </sec>
      <sec id="sec-5-2">
        <title>Participants and Apparatus There were in total 21 par</title>
        <p>ticipants recruited from National Chengchi University in
Taiwan for this experiment. Each participant was reimbursed by
NTD$ 60 (' US$ 2) for their time and traffic expense. The
testing materials and procedure are all the same as those in
Experiment 1.</p>
        <p>1For making the figure easier to read, we plot the human
prediction by circles and the correct answers by crosses on only the
even-numbered trials in the first session. The result pattern is the
same in the second session.
15
10
5
n
o
i
its 0
o
p
−5
−10
10
5
−5
−10
1 7 131925313743495561677379859197
trial
1 7 131925313743495561677379859197
trial
See the circles and crosses in Figure 2. Apparently, the
participants could capture the moving pattern of the target, although
on the early trials, they made some larger errors. Similar to
what we found in Experiment 1, the squared prediction
error drops along blocks from 73.79 to 1.57 (mean = 15:35)
across two sessions. A Block (10) Session (2)
withinsubjects ANOVA reveals a significant main effect of Block
[F(9; 180) = 14:24, MSe = 1303, p &lt; :01], a significant main
effect of Session [F(1; 20) = 17:22, MSe = 196, p &lt; :01],
and a significant interaction effect between Block and
Session [F(9; 180) = 16:12, MSe = 177:8, p &lt; :01]. Although
the error curve goes down toward 0, the mean squared
prediction error is 15.53 far larger than that in Experiment 1, which
is 4.06. This suggests that the linear function is easier to learn
than the quadratic function.</p>
        <p>The accuracy data also suggest that this function is harder
to learn than the linear function with the mean highest
accuracy in a block across all participants and sessions as .34
and the lowest as .14. A Block (10) Session (2)
withinsubjects ANOVA reveals a significant main effect of block
[F(9; 180) = 9:747, MSe = 0:018, p &lt; :01], no significant
main effect of Session [F(1; 20) &lt; 1], and no significant
interaction effect between Block and Session [F(9; 180) &lt; 1].</p>
        <p>Although the accuracy is quite low, this does not mean that
people cannot learn this function. As shown in Figure 2, the
participants’ predictions are close to the true answers. Also,
the correlation between each participant’s predictions and the
true answers is considerably high [mean r = :92]. As
expected, the participants can learn this complex time-varying
function.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Experiment 3</title>
      <p>In this experiment, we would like to examine whether
people could predict the stimulus magnitudes, when the
correlation between successive stimuli was lower. See Figure 3
as an example, which was the real case for testing one
participant2. The dashed line showed the true moving pattern
of the stimulus, which was generated by y = g[a] + z[b + 1],
where a = b((t + 4)=5)c, b = t mod 5, g was the random
permutation of the vector [1,6,11,...,96], and for each g, z was a
new random permutation of the vector [1,2,3,4,5]. The
correlations between successive stimuli were averaged across all
participants and all sessions as r = :80, which was lower than
the correlations in the previous experiments. With no matter
which view to look at this form (i.e., number of parameters
to estimate or correlation between successive stimuli), it was
expected that this function could not be learned well.
Participants and Apparatus There were in total 18
participants recruited for this experiment from National Chengchi
University in Taiwan. Each participant was reimbursed by
NTD$ 60 (' US$ 2) for their time and traffic expense. The
testing materials and procedure are all the same as those in
Experiment 1.</p>
      <sec id="sec-6-1">
        <title>Results</title>
        <p>As shown in Figure 3, apparently, the participant could not
predict the target position. Otherwise, we will see the dashed
line (for answers) and solid line (for participant’s predictions)
superimpose on each other. However, the response pattern is
not random either. In fact, the participant’s predictions seem
always to be one step behind the true answers. Although we
do not show the predictions of the rest 17 participants, their
predictions are one step behind the true answers also. Thus,
strictly speaking, we do not think that the participants learned
this function.</p>
        <p>The squared prediction error drops from 69.69 to 42.47
along blocks in Session 1 and has no clear change from 23.12
to 24.30 in Session 2. Although the performance gets better
in Session 2, the prediction error never goes close to 0. The
mean squared error for all participants across blocks and
sessions is 30.844, which is larger than 15.53 (mean error in
Experiment 2) and 4.06 (mean error in Experiment 1). Thus, the
learning performance in this experiment is the worst among
the three experiments in this study.</p>
        <p>As done for the previous experiments, a Block (10)
Session (2) within-subjects ANOVA was conducted for the
prediction error. The results show no significant main effect of
Block [F(9; 153) = 1:53, MSe = 998:4, p = :142], a
significant main effect of Session [F(1; 17) = 14:94, MSe = 424,
p &lt; :01], and a significant interaction effect between Block
and Session [F(9; 153) = 3:206, MSe = 701:6, p &lt; :01].</p>
        <p>The mean accuracy in a block across all sessions is even
lower than that in the other two experiments. The
highest mean accuracy is about .11 and the lowest is .06. It
is clear that the participants cannot capture the moving
pattern of the stimulus. A Block (10) Session (2)
withinsubject ANOVA shows no main effect of Block on accuracy
[F(9; 153) = 1:179, MSe = 0:006, p = :312], no main effect
of Session [F(1; 17) = 3:367, MSe = 0:006, p = :08], and no
interaction effect between Block and Session [F(9; 153) &lt; 1].</p>
        <p>We also computed the Person’s r for each participant’s
prediction and the true answer. Although the mean correlation is
not low (r = :76), this finding might result from the fact that
the participants’ prediction is always one step behind the true
answer. To sum up, the linear function is the easiest to learn
and the quadratic function is the second. Basically,
participants cannot learn the complex function in Experiment 3. In
order to get a better understanding about the underly
mechanism for learning the time-varying functions, we developed a
neural network model for the learning of time-varying
functions.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Model for Learning Time Varying Function</title>
      <p>A time-varying function can be rewritten as xt = f (xt 1) and
the simplest form of it would be xt = b0 + b1xt 1. Thus,
learning a time-varying function is equivalent to estimating
the optimal parameter values, with which the model makes
the smallest error. To this end, a simple two-layered neural
network is proposed. There are two input nodes, which
respectively correspond to the position of the stimulus on the
preceding trial xt 1 and the standard moving distance which
is set as 1. There is only one output node corresponding to the
predicted position on the current trail xˆt = w1 1 + w2xt 1.
The associative weight w1 represents the size of moving
distance. The weight w2 represents how much correlated the last
position is with the current position. When the true answer xt
is provided, the error is then computed as xt xˆt .</p>
      <p>The associative weights are updated with WH algorithm3
2Different participants received different moving patterns to
learn.</p>
      <p>
        3This algorithm is a special case of backpropagation algorithm,
which is specifically used for two-layered neural network models.
        <xref ref-type="bibr" rid="ref1">(Abdi, Valentin, &amp; Edelman, 1999)</xref>
        to decrease the error
made by the model. Also, we make the updating amount for
weights decay all the way through training trials. Thus, the
updated amount for w1 on trial t is Dw1;t = hexp x(t 1)(xt
xˆt ), where h 0 is the learning rate and x 0 determines
how quickly the updated amount of weight drops. Likewise,
Dw2;t = hexp x(t 1)(xt xˆt )xt 1.
      </p>
      <p>There are some features of this model worth noting. First,
the associative weight w2 actually reflects the correlation
between successive stimuli. Second, this model only learns the
correlation between successive stimuli and contains no
summary information of the whole function. In fact, it can be
applied to account for the learning of different time-varying
functions, as no matter which form (complex or simple) the
function has, the learning of a time-varying function can
always be viewed as the learning of the association between
successive stimuli. Thus, our model should be regarded as an
associative-based model, not a rule-based model.</p>
    </sec>
    <sec id="sec-8">
      <title>Modeling</title>
      <p>The model was fit to each participant’s data in each
experiment with the stimulus positions being normalized between 0
and 1. Each participant’s first response in each session was by
default the first input for the model. The initial weights of w1
and w2 were set as 0 for all experiments except Experiment 3.
The model provided the best fit for Experiment 3 data when
w2 was initially set as 1, suggesting that participants in
Experiment 3 were more likely to repeat the observed position
of stimulus on the preceding trail as the response for current
trail. The statistics of optimally estimated parameter values
and the goodness of fit (RMSD) for all experiments are listed
in Table 1.</p>
      <p>The smaller the RMSD, the better the fit is. Apparently, the
model fit all the data very well. See the crosses in Figure 4,
Figure 5, and Figure 6 for the model prediction in Session 14,
which are quite close to the circles denoting the participants’
responses.</p>
      <p>The estimated learning rate for Experiment 1 is about 1 and
the decay rate is quite small, suggesting that decay of learning
is not fast and leaning continues through training trials. The
learned associative weights for the moving size w1 = 0:30 and
the correlation with the preceding stimulus w2 = 0:70 suggest
that the participants predict the current position of the target
4The pattern is almost the same for Session 2.
by moving it a certain distance (i.e., 0.30 times of the
standard moving size) from the place a bit behind (i.e., 70%) the
position just seen in the same direction of the last move.</p>
      <p>For Experiment 2, the mean learning rate is high and so is
the mean decay rate. This suggests that the model adjusts the
associative weights largely on the early learning trials, but
quickly halts doing so. The learned associative weights are
w1 = 1:00 and w2 = 0:94. The negative weighting for the
preceding position enables the model to make symmetrical
predictions between successive trials and jw2j 1 enables the
model to gradually converge the predicted position toward the
midpoint.</p>
      <p>For Experiment 3, the mean estimated learning rate is low
and the decay rate is high, suggesting that the model has
not updated the associative weights too much since early
trials. In fact, the learned associative weights, w1 = 0:01 and
w2 = 0:98, together suggest that the model merely repeats the
preceding target position as the current prediction. As the
model captures the participants’ response patterns very well,
it is implied that the participants did not actually learn the
function but just repeated what they saw as the prediction for
the next trial.</p>
      <p>It is revealed in Experiment 2 that the larger h or x is, the
smaller the error is (r = :51, p &lt; :05 for h and r = :57, p &lt;
:01 for x) but no significant correlations between parameters
and human performance in other experiments. This might be
because that Experiment 1 and Experiment 3 are either too
easy or too hard for the participants to learn.</p>
      <p>Exp 1</p>
    </sec>
    <sec id="sec-9">
      <title>General Discussion</title>
      <p>The main purpose of this study is to examine the
characteristics of function learning with time-varying functions. Three
experiments were conducted with different time-varying
functions: linear, quadratic, and irregular. The differences
between these functions are not only the complexity of the
function form, but also the strength of correlation between
successive stimuli. In the first two experiments, the
correlation is very high regardless of the direction, whereas in the
third experiment, the correlation is lower.</p>
      <p>10
n
iito 0
s
o
P
−10
Human
Model</p>
      <p>The behavioral data show that the learning of the linear
and quadratic functions are easier than that of the irregular
function, suggesting that the correlation between successive
stimuli is critical to function learning with time-varying
functions, not the number of parameters (or the complexity) of the
function. The success of our model supports the
associativebased account and implies that a time-varying function can
be learned as a composition of many partial representations,
not a holistic representation.</p>
      <p>One may regard the learning of time-varying functions as
operant conditioning. That may or may not be true,
depending on what we think is actually conditioned. If the
response is the target for conditioning, then the learning of
time-varying functions is not operant conditioning, as every
single response is new and it is impossible to reinforce the
likelihood for the same response to be made in the future.
However, if the moving size is the target for conditioning,
then for the case in which the target moves constantly (e.g.,
the linear function in Experiment 1), we may regard the
learning of the time-varying function as a kind of operant
conditioning. However, for the case where the target moves in a
decreasing (or increasing) speed (e.g., the quadratic function in
Experiment 2), it might not be suitable to equate the learning
of time-varying functions and operant conditioning. Future
studies including the transfer trials are needed in order to
examine whether people form any concept for the time-varying
function.
Human
Model</p>
      <p>Exp 3</p>
      <p>Human
Model
1 10 19 28 37 46Tria5l5 64 73 82 91 100</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Abdi</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Valentin</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Edelman</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>1999</year>
          ).
          <article-title>Neural networks</article-title>
          .
          <source>SAGE Publications</source>
          , Inc.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Busemeyer</surname>
            ,
            <given-names>J. R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Byun</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Delosh</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>McDaniel</surname>
            ,
            <given-names>M. A.</given-names>
          </string-name>
          (
          <year>1997</year>
          ).
          <article-title>Learning functional relations based on experience with input-output pairs by humans and artificial neural networks (K. Lamberts</article-title>
          &amp;
          <string-name>
            <surname>D. R. Shanks</surname>
          </string-name>
          , Eds.). Cambridge, MA, US: The MIT Press.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Carroll</surname>
            ,
            <given-names>J. D.</given-names>
          </string-name>
          (
          <year>1963</year>
          ).
          <article-title>Function learning: The learnig of continuous functional maps relating stimulus and response coninua</article-title>
          . Princetron, NJ: Educational Testing Service.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>DeLosh</surname>
            ,
            <given-names>E. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Busemeyer</surname>
            ,
            <given-names>J. R.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>McDaniel</surname>
            ,
            <given-names>M. A.</given-names>
          </string-name>
          (
          <year>1997</year>
          ).
          <article-title>Extrapolation: The sine qua non for abstraction in function learning</article-title>
          .
          <source>Jounral of Experimental Psychology: Learning, Memory, and Cognition</source>
          ,
          <volume>23</volume>
          ,
          <fpage>968</fpage>
          -
          <lpage>986</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Kalish</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>Learning and extraploating a periodic function</article-title>
          .
          <source>Memory &amp; Cognition</source>
          ,
          <volume>41</volume>
          ,
          <fpage>886</fpage>
          -
          <lpage>896</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Kalish</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lewandowsky</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Kruschke</surname>
            ,
            <given-names>J. K.</given-names>
          </string-name>
          (
          <year>2004</year>
          ).
          <article-title>Population of linear experts: Knowledge partitioning and funciton leanring</article-title>
          .
          <source>Psychological Review</source>
          ,
          <volume>111</volume>
          ,
          <fpage>1072</fpage>
          -
          <lpage>1099</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Koh</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , &amp; Meyer,
          <string-name>
            <surname>D. E.</surname>
          </string-name>
          (
          <year>1991</year>
          ).
          <article-title>Function learning: Induction of continuous stimulus-response relations</article-title>
          .
          <source>Journal of Experimental Psychology: Learning, Memory, and Cognition</source>
          ,
          <volume>17</volume>
          ,
          <fpage>811</fpage>
          -
          <lpage>836</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Lewandowsky</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kalish</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Ngang</surname>
            ,
            <given-names>S. K.</given-names>
          </string-name>
          (
          <year>2002</year>
          ).
          <article-title>Simplified learning in complex situations: Knowledge partitioning in function learning</article-title>
          .
          <source>Journal of Experimental Psychology: General</source>
          ,
          <volume>131</volume>
          ,
          <fpage>163</fpage>
          -
          <lpage>193</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>McDaniel</surname>
            ,
            <given-names>M. A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Busemeyer</surname>
            ,
            <given-names>J. R.</given-names>
          </string-name>
          (
          <year>2005</year>
          ).
          <article-title>The conceptual basis of function leanring and extraploation: Comparison of rule-based and associative-based models</article-title>
          .
          <source>Psychonomic Bulletin &amp; Review</source>
          ,
          <volume>12</volume>
          ,
          <fpage>24</fpage>
          -
          <lpage>42</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Navarro</surname>
            ,
            <given-names>D. J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Perfors</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>Learning time-varying categories</article-title>
          .
          <source>In Proceedings of the 31st annual conference of cognitive science society</source>
          (p.
          <fpage>414</fpage>
          -
          <lpage>424</lpage>
          ). austin,
          <source>tx: Cognitive science society.</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Navarro</surname>
            ,
            <given-names>D. J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Perfors</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>Anticipating changes: Adaption and extrapolation in category learning</article-title>
          . In N. Miyake,
          <string-name>
            <given-names>D.</given-names>
            <surname>Peebles</surname>
          </string-name>
          , &amp; R. P. Cooper (Eds.),
          <source>Building bridges across cognitive sciences around the world: Proceedings of the 34th annual conference of the cognitive science society</source>
          (p.
          <fpage>809</fpage>
          -
          <lpage>814</lpage>
          ). Austin, TX: Cognitive Science Society.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Navarro</surname>
            ,
            <given-names>D. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perfors</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Vong</surname>
            ,
            <given-names>W. K.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>Learning time-varying categories</article-title>
          .
          <source>Memory and Cognition</source>
          ,
          <volume>41</volume>
          ,
          <fpage>917</fpage>
          -
          <lpage>927</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Peirce</surname>
            ,
            <given-names>J. W.</given-names>
          </string-name>
          (
          <year>2007</year>
          ).
          <article-title>Psychopy - psychophysics software in python</article-title>
          .
          <source>Journal of Neuroscience Methods</source>
          ,
          <volume>162</volume>
          ,
          <fpage>8</fpage>
          -
          <lpage>13</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>