<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Information-Theoretic Segmentation of Natural Language</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sascha S. Gri ths</string-name>
          <email>sascha.griffiths@qmul.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mariano Mora McGinity</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jamie Forth</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matthew Purver</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Geraint A. Wiggins</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Cognitive Science Research Group, School of Electronic Engineering and Computer Science, Queen Mary University of London</institution>
          ,
          <addr-line>10 Godward Square, London E1 4FZ</addr-line>
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present computational experiments on language segmentation using a general information-theoretic cognitive model. We present a method which uses the statistical regularities of language to segment a continuous stream of symbols into \meaningful units" at a range of levels. Given a string of symbols|in the present approach, textual representations of phonemes|we attempt to nd the syllables such as grea and sy (in the word greasy); words such as in, greasy, wash, and water ; and phrases such as in greasy wash water. The approach is entirely information-theoretic, and requires no knowledge of the units themselves; it is thus assumed to require only general cognitive abilities, and has previously been applied to music. We tested our approach on two spoken language corpora, and we discuss our results in the context of learning as a statistical processes.</p>
      </abstract>
      <kwd-group>
        <kwd>Arti cial Intelligence</kwd>
        <kwd>Language Acquisition</kwd>
        <kwd>Learning</kwd>
        <kwd>Language Segmentation</kwd>
        <kwd>Information Content</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The question which we address in this paper is whether language learning can
be considered to be a statistical process. This has been an ongoing and
fundamentally dividing issue in elds which consider language learning their subject
matter.</p>
      <p>We assume that language has several layers of structure. At the bottom
we nd the smallest units; in this paper, we start with phonemes, though our
method is not restricted to this level of granularity. These smallest units build
larger items of language structure:
{ Phonemes build larger units such as syllables and morphemes.
{ Morphemes and syllables build words. Words build phrases and phrases are
the building blocks of sentences (or spoken utterances).
{ These sentences or utterances go in turn to make up larger units such as
paragraphs in text or speaker turns in speech.</p>
      <p>We must assume some smallest unit such as phonemes in speech or graphemes
in text as an entry point into the language system. The question which arises is
how one can tell where one such unit above the phoneme or grapheme ends and
another one begins. In natural language processing this task is generally called
text segmentation for written language and speech segmentation for spoken
language.</p>
      <p>In the current paper, we assume that the phonemes are presented as one
continuous stream - roughly equivalent to removing the white space from
sentences in written text - and de ne our task as determining where a word or
other linguistic unit begins or ends. This is similar to the task infants face when
learning their rst language, itself an open research question. Taking the title
as an example, we need to identify that the word segmentation is composed of
syllables, which are seg, men, ta, and tion, and morphemes, which are segment
and ation. From there larger units need to be distinguished such as the words
segmentation, of and natural. Longer utterances need to be split up into phrases,
perhaps at various levels of granularity e.g. natural language and segmentation
of natural language.</p>
      <p>
        We present a computational approach to the segmentation problem [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] in
which we rely entirely on the information content of a symbol within a language
dataset. Our prediction is that information content will rise at the beginning of
a segment and fall at the end of a segment. A similar assumption was used by
Harris [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] for nding morpheme boundaries. We assume that this assumption
should hold for segments at all levels of the linguistic hierarchy. However, for
each level the nature and extent of this fall and rise will vary; but parameters
of the model will vary predictably with levels of segmentation. In the following,
we will present computational experiments which test this prediction on two
datasets of natural language.
      </p>
      <p>Our approach to the cognitive task of language processing therefore places
emphasis on domain-independent principles, rather than taking a domain-speci c
approach as has been argued as appropriate for the case of language.</p>
      <p>We outline our information-theoretic approach in the next section; we then
present the methods used in this paper in detail. Our results and the discussion
of these for computational experiments on language segmentation are presented
in the following sections. In our conclusion we return to the question outlined
above.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Information-Theoretic Speech Segmentation</title>
      <p>
        Applying machine learning and pattern recognition methods to natural language
has become a rich source of insights into language structure and theoretical issues
of linguistics [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and the learning of language [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. While many linguists take
the view that natural language requires domain-speci c, innate structures (see
e.g.[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]), this is debated; one of the tenets of cognitive linguistics is that language
processing by humans is domain general and not domain speci c. As Geo rey
Sampson puts it, language learning depends on `general human intelligence and
abilities' [6, p. IX]. Using information theory as a framework for such an approach
has been argued to be both cognitively and biologically plausible [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ].
      </p>
      <p>
        It has often been proposed that language and music share certain
properties. One can ask the same questions regarding the structure and processing of
language as one can ask about music [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Indeed, there a number of objective
similarities and di erences between these two domains [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. There seem to be
shared resources in structural processing of language and music [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], in addition
to the conceptual similarity which is that the building blocks of the structures
are \cognitive objects" { i.e. percepts . In this paper, we assume that percepts in
general can be processed via their statistical regularities in a given corpus. The
computational model used here was created for purposes of melodic grouping
[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>
        The current research is situated within the wider context of the IDyOM and
IDyOT frameworks. IDyOM [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] stands for Information Dynamics Of Music. It
was developed on the basis of natural language processing methods and can
divide melodies into perceptually correct segments using the statistical regularities
in a corpus. Generally, however, we argue here that the framework can also still
be used to segment a corpus of natural language data into syllables, words and
phrases.
      </p>
      <p>
        In previous work [
        <xref ref-type="bibr" rid="ref13 ref14">13, 14</xref>
        ] a cognitive architecture called IDyOT is outlined
which builds on the principles of IDyOM. IDyOT [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] stands for Information
Dynamics Of Thinking. The premise of both of these di erent incarnations of
the underlying research framework is that grouping and boundary perception
are central to cognitive science [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] and that the most cognitively plausible way
of approaching this task is using Shannon's [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] information theory. Especially,
we employ information content as introduced by MacKay [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
      </p>
      <p>
        IDyOT is based on the Global Workspace Theory [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. In the long term it is
predicted that IDyOT provides the basis for modelling creativity and eventually
aspects of consciousness [
        <xref ref-type="bibr" rid="ref14 ref19">14, 19</xref>
        ]. In order of testing certain claims about the
domain generality of the information dynamics approach embodied by IDyOM
and IDyOT, we look at language segmentation to see whether the approach
shown to be useful in music segmentation can be transferred back to language.
      </p>
      <p>
        The IDyOM model [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] and corresponding software1 were developed for the
statistical modelling of music in the context of music perception and cognition
research. However, one of the central features of the model is that it can also
be used for other types of sequential data, as the principles on which it is based
are cognitively inspired and meant to be general rather than domain speci c
[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. The model presented in IDyOM relies on a pattern recognition theory of
mind [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] which suggests that languages are learned by processing the underlying
statistics of the positive data contained in stimuli. Although, at the present
moment only representations of auditory stimuli have been studied, our conjecture
is that any kind of perceptual data can be processed in this way. Wiggins [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]
1 The software can be found at
https://code.soundsoftware.ac.uk/projects/idyomproject/ les.
2.1
      </p>
      <sec id="sec-2-1">
        <title>The IDyOM Model</title>
        <p>
          (1)
(2)
(3)
IDyOM is a multidimensional variable-order Markov model. The
multidimensionality within IDyOM is formalised as a multiple viewpoints system [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ], where
viewpoints can be either given basic types, or derived and combined from existing
viewpoints to form new viewpoints revealing more abstract levels of structure.
Predictions from individual variable-order viewpoint models are combined using
an entropy-weighting strategy [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ].
        </p>
        <p>
          Two basic information-theoretic measures are central to IDyOM. Information
content is the measure of unexpectedness|or surprisal to use the terminology
of [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ]|and entropy a measure of uncertainty.
1. information content (h) is a measure of how unpredictable a [given unit] is
given its context [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ];
2. entropy (H) is the expected information content of an unseen event in a
given context.
        </p>
        <p>More formally, in IDyOM these concepts are modelled as (1) and (2) below:
gives initial indications that the model can extend to language segmentation,
and we pursue that idea in more depth here.</p>
        <p>
          As aforementioned, we take an information-theoretic approach here [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ].
Predicting the next element in a sequence given the previous element is often called
a Shannon Game [23, p. 191]. Here, we assume that both music and language can
be modelled as a sequence of elements e from an alphabet E . For each element
ei in e one can calculate its probability given the context { more speci cally the
preceding context ei1 1:
        </p>
        <p>p(eijei1 1)
There is good evidence that children use transition probabilities during language
acquisition [24, p. 33], and this probability can be calculated by approximating
on the basis of a context subsequence of nite length n, i.e. by using an n-gram
model [25, pp. 845{847].</p>
        <p>1
h(eijei1 1) = log2 p(eijei1 1)
H(ei1 1) =</p>
        <p>X p(eijei1 1)h(eijei1 1)
e2E</p>
        <p>
          Entropy-based models such as these have been used in natural language
learning in the past [28, pp. 21{37]. Given an n-gram model of p(eijei1 1) which
characterises the dataset in question, we can calculate h and H at all points in a
sequence, and thereby nd local falls and rises. Such falls and rises have
previously been shown to correlate with the ending and beginning of structural
units in music [
          <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
          ] and language [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ]. For the case of music it has also been
demonstrated that this model outperforms rule-based approaches [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ].
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Segmentation of Natural Language</title>
        <p>
          Segmentation of natural language has been a topic for computational
psycholinguistics at least since 1990 [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ]. However, it can still be regarded as a current
problem in computational approaches to language learning (see for example [31{
34]). Brent [
          <xref ref-type="bibr" rid="ref35">35</xref>
          ] classi es a number of approaches to natural language
segmentation into three types of strategies. These are the utterance-boundary strategy,
the predictability strategy and the recognition strategy. Our approach employs
elements of the predictability strategy: we attempt to detect boundaries based
on changes in the information-theoretic properties of the symbol sequences in
question. In this way it is similar to, but simpler and more general than, methods
such as that of Cohen and Adams [
          <xref ref-type="bibr" rid="ref36">36</xref>
          ], who use boundary entropy but combine it
with other frequency measures via voting experts to segment words in a range of
languages, or Sun, Shen and Tsou [
          <xref ref-type="bibr" rid="ref37">37</xref>
          ], who use mutual information but combine
it with other statistical measures to segment Chinese characters into words.
        </p>
        <p>
          This contrasts with approaches in which one tries to build grammars (or
probabilistic models) of likely segment sequences (the predictability strategy), (e.g.
for Finnish morphemes [
          <xref ref-type="bibr" rid="ref38">38</xref>
          ]), and with those in which one matches patterns of
known words against the stream (the recognition strategy); in those approaches
one needs to build up a lexicon rst, either from external knowledge (e.g. [
          <xref ref-type="bibr" rid="ref39">39</xref>
          ]) or
from incremental clustering (e.g. [
          <xref ref-type="bibr" rid="ref40">40</xref>
          ]). Our boundary detection strategy needs
no knowledge of the lexicon or even of the fact that there are such concepts as
syllables, words or phrases.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Methodology</title>
      <p>Our experimental method requires two steps: rstly, building a statistical n-gram
(IDyOM) model on the basis of which to calculate information content (entropy
is left for future work); secondly, hypothesising boundaries based on local drops
and rises in information content.
3.1</p>
      <sec id="sec-3-1">
        <title>Calculating the Information Content</title>
        <p>
          IDyOM has a range of model con gurations intended to simulate di erent
aspects of musical listening behaviour. The basic distinction concerns the data
used to train an n-gram model: a model can be trained from a large dataset,
modelling the learned experience of a listener and termed the Long Term model
(LTM) in IDyOM's terminology; or from only the current sequence under
consideration, trained incrementally for each utterance being predicted [
          <xref ref-type="bibr" rid="ref20 ref26 ref41">26, 20, 41</xref>
          ],
modelling a listening experience in a speci c context, and termed the Short Term
model (STM). However, variations are possible: the LTM approach can be made
dynamic by adjusting its probabilities based on the current sequence as it is
observed (termed LTM+); and the STM and LTM models can be combined. This
results in a total of ve models:
STM model trained on stimuli only in a local context (i.e. notes of the melody
or phonemes in the utterance currently being predicted);
LTM model trained on a large training corpus;
LTM+ as LTM, but model also learns from the current example;
Both combination of STM and LTM;
Both+ combination of STM and LTM+.
3.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Segmentation</title>
        <p>
          Our overall approach is to look for characteristic local contours in information
content [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ]|what Pearce et al. [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] call `peak picking'. Rises in information
content are signals of unexpectedness, and Wiggins [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] hypothesises that these
should correlate with the beginnings of new segments; conversely, falls in
information content are signals of predictability, which we expect to correlate with
the endings of segments.
        </p>
        <p>
          Our current method is extremely simple, checking only for a simple rise
between successive data-points: the value at the current symbol ei must exceed
that at its immediate predecessor ei 1 by some speci ed amount. This amount
is our only parameter, d; thus, a new segment begins if h(ei) h(ei 1) &gt; d. We
evaluate performance using the statistic [
          <xref ref-type="bibr" rid="ref42 ref43">42, 43</xref>
          ], and set d to give the maximal
value for (for a speci c segment type) by testing all d over the interval [0; 10].
3.3
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>Data</title>
        <p>
          We test this method on two language corpora. The rst dataset is a derived
corpus2 of the CHILDES corpus of child-directed adult English speech [
          <xref ref-type="bibr" rid="ref44">44</xref>
          ], collated
and transcribed at the phoneme level for word segmentation experiments [
          <xref ref-type="bibr" rid="ref45">45</xref>
          ]. It
contains 93,555 phoneme tokens which make up 33,377 words and 9,790
utterances; average utterance length is 3.4 words. A single viewpoint with phonemes
as observed variables, denoted fphonemesg, is used as the basic IDyOM
representation.
        </p>
        <p>
          The second corpus is the TIMIT transcriptions [
          <xref ref-type="bibr" rid="ref46">46</xref>
          ], a dataset of spoken
English sentences obtained for the purpose of automatic speech recognition model
training, and transcribed at the level of sentences, words and phonemes. It
contains 81,533 phoneme tokens which make up 20,756 words and 2,342 utterances;
average utterance length is therefore 8.9 words. Again we use a simple phoneme
viewpoint; as TIMIT also contains stress annotations (represented as primary,
secondary, and no stress), this also allows us to construct a linked viewpoint
formed of the cross-product of phonemes and stress fphonemes stressg, and
a two-viewpoint system combining both viewpoints fphonemes, phonemes
stressg.
        </p>
        <p>
          To evaluate phrase-level segmentation, we used the Pattern parser [
          <xref ref-type="bibr" rid="ref47">47</xref>
          ] {
neither TIMIT nor CHILDES contains phrase structure information. Automatic
parses are noisy: we excluded cases where Pattern produced a parse which could
not be mapped back onto the phonetic form of the utterance. Thus, our
analysis on the phrase level only considers approximately half of the data for both
corpora.
2 http://www.ling.ohio-state.edu/~melsner/resources/acl12data.tgz
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>We evaluate our segmentation model in terms of accuracy of boundary
placement against the ground truth for each level|syllables, words and phrases|with
accuracy assessed via both Kappa values ( ) and the F1-score (harmonic mean
of precision and recall). Both and F1 are calculated on the basis of
individual phoneme tokens, with the gold-standard annotation classifying only the rst
token in each segment as a boundary. We also examine the mean information
content (h), and optimum value of our segmentation parameter (d). h is the
same in across all, as it is a property of the (phoneme-based) corpus and model
and not of the evaluation.
4.1</p>
      <sec id="sec-4-1">
        <title>CHILDES</title>
        <p>The results for the CHILDES corpus segmentation into words and phrases is
summarised in Table 1. Lower h values mean better predictability, as high h
signi es more \surprisal" by a new element.</p>
        <p>Performance is reasonable at word level, with F1 around 0.7 and
approaching 0.6. The performance of the STM is considerably lower than other models, as
might be expected; we note that h is considerably higher for the STM, indicating
worse t. The d parameter is therefore also correspondingly higher { and takes
longer to nd { for the STM. The lowest d for words is found in the Both model
and LTM and LTM+ for the phrase segmentation.</p>
        <p>In terms of both F1 and , the LTM and LTM+ are the best models for
the word discovery task. In the phrase segmentation task, we nd that the Both
and Both+ models do marginally better than in the word segmentation task
with respect to , but with worse performance with respect to F1. The apparent
improvement of results for the STM may be due to the lower number of segments
which need to be predicted. The improvement in performance by the short term
model also leads to improvements in the Both and Both+ con guration as these
are combinations of STM and LTM.</p>
        <p>
          In comparison to previous work on the same dataset, our best con guration
(LTM) still performs slightly worse with respect to F1-scores in the word
segmentation task. While Elsner et al. [
          <xref ref-type="bibr" rid="ref45">45</xref>
          ] obtained an F1-score of 0.8, our best
F1 score was 0.71. We also checked the baselines with respect to a random
segmentation, a segmentation which assumes every symbol is a boundary and a
segmentation which assumes no boundaries. In each case, the will be 0 as
expected with low F1-scores.
4.2
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>TIMIT</title>
        <p>
          Table 2 shows the results for the TIMIT corpus. The results for the syllable
segmentation task are very comparable for all measures with those reported by
Wiggins [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ]. As with the CHILDES dataset, the STM shows higher values for
h, with the LTM and Both models showing better performance. and F1 are
almost the same for the STM for all con gurations. Therefore, the STM, for
which the h is determined based on isolated utterances, seems not to be a good
model for this task.
        </p>
        <p>Generally, one can see a trend that the LTM, LTM+, Both and Both+ models
perform better if they receive more information, i.e. in the fphonemes stressg
and fphonemes, phonemes stressg show marginally better performance. The
smallest values for d are found in the viewpoints fphonemes stressg.</p>
        <p>There seems no improvement in performance when moving from the LTM to
LTM+ or the Both model variants (in all con gurations except the fphonemesg
condition, where the LTM+ shows slightly higher F1-score). Overall, the best
con guration for the word segmentation task is the LTM in the fphonemes,
phonemes stressg condition.</p>
        <p>Segmentation at Di erent Levels For the word segmentation task, the optimal
setting of d is higher than that for the syllable segmentation task, for all
congurations. This corresponds with intuitive expectation, as one needs to predict
fewer segments. In all con gurations, the LTM and LTM+ still show better
performance than the STM, Both and Both+; overall accuracy is slightly
improved over the syllable segmentation task, with F1 scores over 0.7. The LTM+
is the best con guration overall in the fphonemes, phonemes stressg condition.
The STM perhaps also shows some improvement here, with values marginally
higher.</p>
        <p>In the phrase segmentation task, again, the optimum d increases relative to
word and syllable tasks, as even fewer segments need to be predicted.
Performance in terms of and F1-scores is, however, much lower for phrase discovery
than for syllables and words. Thus, with regard to this measure the performance
on the TIMIT data is less e ective.</p>
        <p>The values and F1 scores for the STM, however, are considerably higher for
this task. The STM does, however, not bene t from the additional information
which it gets in the fphonemes stressg and fphonemes, phonemes stressg
conditions.</p>
        <p>In all con gurations, the LTM, LTM+, Both and Both+ models show worse
performance in the phrase segmentation task with respect to and F1. Also,
there is little di erence in the performance of these four models. The Both is the
best con guration overall in the fphonemes, phonemes stressg condition with
respect to the value.</p>
        <p>
          As expected our results are very similar to those reported in Wiggins [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] for
syllable segmentation. We also checked the baselines with respect to a random
segmentation, a segmentation which assumes every symbol is a boundary and
a segmentation which assumes no boundaries. In each case, the will be 0 as
expected with low F1-scores.
4.3
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>Overall Segmentation Performance</title>
        <p>
          vs parameter d for words and phrases.
optimum d), after which performance drops o again. This optimum value of d
is higher as we move to longer, higher-level segments (from syllables to words,
and from words to phrases): larger changes in information content correspond to
segment boundaries at di erent levels. However, the phrase segmentation curve
shows less of a peak: performance reaches a level at which it stays. This suggests
that as long as d is large enough one nds segments which have a high probability
of coinciding with phrase boundaries. The plateau in the curve after the peak
can be explained as an e ect of our segmentation method coding the beginning
of an utterance as a given start symbol. This is similar to the approach of Elsner
et al [
          <xref ref-type="bibr" rid="ref45">45</xref>
          ]. Thus, once the method stops oversegmenting at low d it nds the
optimum d and afterwards continues to agree on those given symbols at higher
d values.
        </p>
        <p>Fig. 2. TIMIT corpus</p>
        <p>vs parameter d for syllables, words and phrases.</p>
        <p>The STM shows particularly bad performance initially but then the plots
show a sudden leap in performance. This is true for all con gurations on both
corpora. Thus, short term segmentation seems to require a certain threshold to
show any noticeable segmentation performance. Exposure to isolated utterances
is insu cient to learn the distributional regularities of language.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Discussion &amp; Conclusion</title>
      <p>
        Landis and Koch [
        <xref ref-type="bibr" rid="ref48">48</xref>
        ] characterise a 2 [0:4; 0:6] as \moderate". Thus, most of
the results reported here show a moderate success. The results reported for the
CHILDES corpus with respect to phrases is slightly higher and thus falls into the
\substantial" category. However, one has to again note, that results regarding the
syntactic units are to be taken with caution. There is less to predict and less to
agree on. Therefore, one would expect a higher agreement between ground-truth
and segmentation.
      </p>
      <p>The long term model shows better performance than the short term model.
In e ect, these two model a listeners knowledge of language (LTM) and a current
listening experience (STM). It is to be expected that there is little result to be
expected from learning from a single listening experience. Thus, the results with
respect to the di erences in LTM and STM show that a long term learning from
raw stimulus is possible.</p>
      <p>The TIMIT data also indicates that learning is improved if stress information
can be included. Though, the di erences are small, the inclusion of stress in the
viewpoints selected for predicting the next phoneme do improve the results. The
di erences reported here are minor, though.</p>
      <p>The present contribution took a strong view of statistical language learning.
We claimed that it would be possible to predict syllable, word and phrase
boundaries from a raw stimulus without having explicit information about these units
encoded in the method. We succeeded in the sense that our results indicate that
this is indeed possible. In future work, we plan to explore further in what way
the inclusion of di erent viewpoints improves the results.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>The research reported in this is supported by ConCreTe: the project ConCreTe
acknowledges the nancial support of the Future and Emerging Technologies
(FET) programme within the Seventh Framework Programme for Research of
the European Commission, under FET grant number 611733.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Elman</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          :
          <article-title>Computational approaches to language acquisition</article-title>
          . In Brown, K., ed.:
          <source>Encyclopedia of Language and Linguistics</source>
          . Volume
          <volume>2</volume>
          . Second edn.
          <source>Elsevier</source>
          , Oxford (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Harris</surname>
            ,
            <given-names>Z.S.:</given-names>
          </string-name>
          <article-title>From phoneme to morpheme</article-title>
          .
          <source>Language</source>
          <volume>31</volume>
          (
          <issue>2</issue>
          ) (
          <year>1955</year>
          ) pp.
          <volume>190</volume>
          {
          <fpage>222</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Lappin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shieber</surname>
            ,
            <given-names>S.M.:</given-names>
          </string-name>
          <article-title>Machine learning theory and practice as a source of insight into universal grammar</article-title>
          .
          <source>Journal of Linguistics</source>
          <volume>43</volume>
          (
          <issue>2</issue>
          ) (
          <year>2007</year>
          )
          <volume>393</volume>
          {
          <fpage>427</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lappin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Linguistic Nativism and the Poverty of the Stimulus</article-title>
          .
          <source>WileyBlackwell</source>
          , Oxford (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Chomsky</surname>
          </string-name>
          , N.:
          <article-title>New Horizons in the Study of Language and Mind</article-title>
          . Cambridge University Press, Cambridge (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Sampson</surname>
            ,
            <given-names>G.R.</given-names>
          </string-name>
          :
          <source>The Language Instinct Debate. Continuum</source>
          , London (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Wallace</surname>
          </string-name>
          , R.:
          <article-title>Cognition and biology: perspectives from information theory</article-title>
          .
          <source>Cognitive Processing</source>
          <volume>15</volume>
          (
          <issue>1</issue>
          ) (
          <year>February 2014</year>
          )
          <volume>1</volume>
          {
          <fpage>12</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Rohrmeier</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuidema</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wiggins</surname>
            ,
            <given-names>G.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schar</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Principles of structure building in music, language and animal song</article-title>
          .
          <source>Philosophical Transactions of the Royal Society of London B: Biological Sciences</source>
          <volume>370</volume>
          (
          <issue>1664</issue>
          ) (
          <year>2015</year>
          )
          <fpage>20140097</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Jackendo</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lerdahl</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>The capacity for music: What is it, and what's special about it?</article-title>
          <source>Cognition</source>
          <volume>100</volume>
          (
          <issue>1</issue>
          ) (
          <year>2006</year>
          )
          <volume>33</volume>
          {
          <fpage>72</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Patel</surname>
            ,
            <given-names>A.D.</given-names>
          </string-name>
          :
          <article-title>Music, Language, and the Brain</article-title>
          . Oxford University Press, Oxford (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Pearce</surname>
            ,
            <given-names>M.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wiggins</surname>
            ,
            <given-names>G.A.</given-names>
          </string-name>
          :
          <article-title>The information dynamics of melodic boundary detection</article-title>
          .
          <source>In: Proceedings of the Ninth International Conference on Music Perception and Cognition</source>
          , Bologna (
          <year>2006</year>
          )
          <volume>860</volume>
          {
          <fpage>867</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Pearce</surname>
            ,
            <given-names>M.T.</given-names>
          </string-name>
          , Mullensiefen,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Wiggins</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.A.</surname>
          </string-name>
          :
          <article-title>Melodic grouping in music information retrieval: New methods and applications</article-title>
          .
          <source>In: Advances in music information retrieval</source>
          . Springer, Berlin (
          <year>2010</year>
          )
          <volume>364</volume>
          {
          <fpage>388</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Wiggins</surname>
            ,
            <given-names>G.A.</given-names>
          </string-name>
          :
          <article-title>The mind's chorus: creativity before consciousness</article-title>
          .
          <source>Cognitive Computation</source>
          <volume>4</volume>
          (
          <issue>3</issue>
          ) (
          <year>2012</year>
          )
          <volume>306</volume>
          {
          <fpage>319</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Wiggins</surname>
            ,
            <given-names>G.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Forth</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>IDyOT: A computational theory of creativity as everyday reasoning from learned information</article-title>
          . In Besold, T.R.,
          <string-name>
            <surname>Schorlemmer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smaill</surname>
          </string-name>
          , A., eds.: Computational Creativity Research:
          <article-title>Towards Creative Machines</article-title>
          . Volume
          <volume>7</volume>
          of Atlantis Thinking Machines. Atlantis Press (
          <year>2015</year>
          )
          <volume>127</volume>
          {
          <fpage>148</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Pearce</surname>
            ,
            <given-names>M.T.</given-names>
          </string-name>
          , Mullensiefen,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Wiggins</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.A.</surname>
          </string-name>
          :
          <article-title>The role of expectation and probabilistic learning in auditory boundary perception: A model comparison</article-title>
          .
          <source>Perception</source>
          <volume>39</volume>
          (
          <issue>10</issue>
          ) (
          <year>2010</year>
          )
          <volume>1365</volume>
          {
          <fpage>1389</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Shannon</surname>
            ,
            <given-names>C.E.</given-names>
          </string-name>
          :
          <article-title>A mathematical theory of communication</article-title>
          .
          <source>ACM SIGMOBILE Mobile Computing and Communications Review</source>
          <volume>5</volume>
          (
          <issue>1</issue>
          ) (
          <year>1948</year>
          )
          <volume>3</volume>
          {
          <fpage>55</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>MacKay</surname>
            ,
            <given-names>D.J.C.</given-names>
          </string-name>
          :
          <string-name>
            <surname>Information</surname>
            <given-names>Theory</given-names>
          </string-name>
          , Inference, and
          <string-name>
            <given-names>Learning</given-names>
            <surname>Algorithms</surname>
          </string-name>
          . Cambridge University Press, Cambridge (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Baars</surname>
            ,
            <given-names>B.J.:</given-names>
          </string-name>
          <article-title>A cognitive theory of consciousness</article-title>
          . Cambridge University Press, Cambridge (
          <year>1993</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Wiggins</surname>
            ,
            <given-names>G.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tyack</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schar</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rohrmeier</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>The evolutionary roots of creativity: Mechanisms and motivations</article-title>
          .
          <source>Philosophical Transactions of the Royal Society of London B: Biological Sciences</source>
          <volume>370</volume>
          (
          <issue>1664</issue>
          ) (
          <year>2015</year>
          )
          <fpage>20140099</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Pearce</surname>
          </string-name>
          , M.T.:
          <article-title>The construction and evaluation of statistical models of melodic structure in music perception and composition</article-title>
          .
          <source>PhD thesis</source>
          , City University London (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Kurzweil</surname>
          </string-name>
          , R.:
          <article-title>How to create a mind: The secret of human thought revealed</article-title>
          .
          <source>Penguin</source>
          , London (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Wiggins</surname>
            ,
            <given-names>G.A.</given-names>
          </string-name>
          :
          <article-title>\I let the music speak": Cross-domain application of a cognitive model of musical learning</article-title>
          . In Rebuschat,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Williams</surname>
          </string-name>
          , J., eds.:
          <article-title>Statistical Learning and Language Acquisition</article-title>
          . Mouton de Gruyter, Amsterdam, NL (
          <year>2012</year>
          )
          <volume>463</volume>
          {
          <fpage>494</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Manning</surname>
            ,
            <given-names>C.D.</given-names>
          </string-name>
          , Schutze, H.:
          <article-title>Foundations of statistical natural language processing</article-title>
          . MIT Press, Cambridge, MA (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Ambridge</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lieven</surname>
            ,
            <given-names>E.V.M.</given-names>
          </string-name>
          :
          <article-title>Child Language Acquisition: Contrasting Theoretical Approaches</article-title>
          . Cambridge University Press, Cambridge (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Russell</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Norvig</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Arti cial Intelligence: A Modern Approach</article-title>
          . Third edn. Prentice Hall International (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Conklin</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Witten</surname>
            ,
            <given-names>I.H.</given-names>
          </string-name>
          :
          <article-title>Multiple viewpoint systems for music prediction</article-title>
          .
          <source>Journal of New Music Research</source>
          <volume>24</volume>
          (
          <issue>1</issue>
          ) (
          <year>March 1995</year>
          )
          <volume>51</volume>
          {
          <fpage>73</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Mahowald</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fedorenko</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piantadosi</surname>
            ,
            <given-names>S.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gibson</surname>
          </string-name>
          , E.: Info/information theory:
          <article-title>Speakers choose shorter words in predictive contexts</article-title>
          .
          <source>Cognition</source>
          <volume>126</volume>
          (
          <issue>2</issue>
          ) (
          <year>2013</year>
          )
          <volume>313</volume>
          {
          <fpage>318</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Charniak</surname>
          </string-name>
          , E.:
          <article-title>Statistical Language Learning</article-title>
          . MIT Press, Cambridge, MA (
          <year>1996</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Pearce</surname>
            ,
            <given-names>M.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wiggins</surname>
            ,
            <given-names>G.A.</given-names>
          </string-name>
          :
          <article-title>Expectation in Melody: The In uence of Context and Learning</article-title>
          .
          <source>Music Perception</source>
          <volume>23</volume>
          (
          <issue>5</issue>
          ) (
          <year>2006</year>
          )
          <volume>377</volume>
          {
          <fpage>405</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Elman</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          :
          <article-title>Finding structure in time</article-title>
          .
          <source>Cognitive Science</source>
          <volume>14</volume>
          (
          <issue>2</issue>
          ) (
          <year>June 1990</year>
          )
          <volume>179</volume>
          {
          <fpage>211</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Elsner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goldwater</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Feldman</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wood</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>A joint learning model of word segmentation, lexical acquisition, and phonetic variability</article-title>
          .
          <source>In: Proceedings of the Conference on Empirical Methods in Natural Language Processing</source>
          . (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <string-name>
            <surname>Fourtassi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Borschinger,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Johnson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Dupoux</surname>
          </string-name>
          , E.:
          <article-title>Why is English so easy to segment?</article-title>
          <source>In: Proceedings of the Fourth Annual Workshop on Cognitive Modeling and Computational Linguistics (CMCL)</source>
          ,
          <article-title>So a, Bulgaria, Association for Computational Linguistics</article-title>
          (
          <year>August 2013</year>
          )
          <volume>1</volume>
          {
          <fpage>10</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <string-name>
            <surname>Pate</surname>
            ,
            <given-names>J.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Johnson</surname>
          </string-name>
          , M.:
          <article-title>Syllable weight encodes mostly the same information for english word segmentation as dictionary stress</article-title>
          . In: EMNLP, Doha, Qatar (
          <year>2010</year>
          )
          <volume>844</volume>
          {
          <fpage>853</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          34. Coltekin, C.:
          <article-title>Units in segmentation: A computational investigation</article-title>
          .
          <source>In: Proceedings of the Sixth Workshop on Cognitive Aspects of Computational Language Learning</source>
          , Lisbon, Portugal, Association for Computational Linguistics (
          <year>September 2015</year>
          )
          <volume>55</volume>
          {
          <fpage>64</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          35.
          <string-name>
            <surname>Brent</surname>
            ,
            <given-names>M.R.</given-names>
          </string-name>
          :
          <article-title>Speech segmentation and word discovery: a computational perspective</article-title>
          .
          <source>Trends in Cognitive Sciences</source>
          <volume>3</volume>
          (
          <issue>8</issue>
          ) (
          <year>August 1999</year>
          )
          <volume>294</volume>
          {
          <fpage>301</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          36.
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Adams</surname>
            ,
            <given-names>N.:</given-names>
          </string-name>
          <article-title>An algorithm for segmenting categorical time series into meaningful episodes</article-title>
          .
          <source>In: Advances in Intelligent Data Analysis, 4th International Conference</source>
          , Cascais, Portugal (
          <year>2001</year>
          )
          <volume>198</volume>
          {
          <fpage>207</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          37.
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsou</surname>
            ,
            <given-names>B.K.</given-names>
          </string-name>
          :
          <article-title>Chinese word segmentation without using lexicon and hand-crafted training data</article-title>
          .
          <source>In: Proceedings of the 36th Annual Meeting of the Association for Computational Linguistics</source>
          . (
          <year>1998</year>
          )
          <volume>1265</volume>
          {
          <fpage>1271</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          38.
          <string-name>
            <surname>Virpioja</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Turunen</surname>
            ,
            <given-names>V.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spiegler</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kohonen</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kurimo</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Empirical comparison of evaluation methods for unsupervised learning of morphology</article-title>
          .
          <source>Traitement Automatique des Langues</source>
          <volume>52</volume>
          (
          <issue>2</issue>
          ) (
          <year>2011</year>
          )
          <volume>45</volume>
          {
          <fpage>90</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          39.
          <string-name>
            <surname>Sproat</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shih</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gale</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>N.:</given-names>
          </string-name>
          <article-title>A stochastic nite-state wordsegmentation algorithm for Chinese</article-title>
          .
          <source>In: Proceedings of the 32nd Annual Meeting of the Association for Computational Linguistics</source>
          . (
          <year>1994</year>
          )
          <volume>66</volume>
          {
          <fpage>73</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          40.
          <string-name>
            <surname>Gold</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scassellati</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Audio speech segmentation without language-speci c knowledge</article-title>
          .
          <source>In: Proceedings of the 28th Annual Conference of the Cognitive Science Society</source>
          , Vancouver (
          <year>2006</year>
          )
          <volume>1370</volume>
          {
          <fpage>1375</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          41.
          <string-name>
            <surname>Wiggins</surname>
            ,
            <given-names>G.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pearce</surname>
            ,
            <given-names>M.T.</given-names>
          </string-name>
          , Mullensiefen, D.:
          <article-title>Computational modelling of music cognition and musical creativity</article-title>
          . In
          <string-name>
            <surname>Dean</surname>
          </string-name>
          , R., ed.:
          <source>The Oxford Handbook of Computer Music</source>
          . Oxford University Press (
          <year>2009</year>
          )
          <volume>383</volume>
          {
          <fpage>420</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          42.
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A Coe cient of Agreement for Nominal Scales</article-title>
          .
          <source>Educational and Psychological Measurement</source>
          <volume>20</volume>
          (
          <issue>1</issue>
          ) (
          <year>April 1960</year>
          )
          <volume>37</volume>
          {
          <fpage>46</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          43.
          <string-name>
            <surname>Carletta</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Assessing agreement on classi cation tasks: the kappa statistic</article-title>
          .
          <source>Computational Linguistics</source>
          <volume>22</volume>
          (
          <issue>2</issue>
          ) (
          <year>1996</year>
          )
          <volume>249</volume>
          {
          <fpage>254</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          44.
          <string-name>
            <surname>MacWhinney</surname>
          </string-name>
          , B.: CHILDES Project:
          <article-title>Tools for analyzing talk</article-title>
          .
          <source>3rd Edition</source>
          . Vol.
          <volume>2</volume>
          :
          <string-name>
            <given-names>The</given-names>
            <surname>Database</surname>
          </string-name>
          . Lawrence Erlbaum Associates, Mahwah, NJ (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          45.
          <string-name>
            <surname>Elsner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goldwater</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eisenstein</surname>
          </string-name>
          , J.:
          <article-title>Bootstrapping a uni ed model of lexical and phonetic acquisition</article-title>
          .
          <source>In: Proceedings of the 50th Annual Meeting of the Association of Computational Linguistics</source>
          . (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          46.
          <string-name>
            <surname>Zue</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sene</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Glass</surname>
          </string-name>
          , J.:
          <article-title>Speech database development at MIT: TIMIT and beyond</article-title>
          .
          <source>Speech Communication</source>
          <volume>9</volume>
          (
          <year>1990</year>
          )
          <volume>351</volume>
          {
          <fpage>356</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref47">
        <mixed-citation>
          47.
          <string-name>
            <surname>De Smedt</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daelemans</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Pattern for Python</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          <volume>13</volume>
          (
          <issue>1</issue>
          ) (
          <year>2012</year>
          )
          <year>2063</year>
          {
          <fpage>2067</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref48">
        <mixed-citation>
          48.
          <string-name>
            <surname>Landis</surname>
            ,
            <given-names>J.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koch</surname>
            ,
            <given-names>G.G.</given-names>
          </string-name>
          :
          <article-title>The measurement of observer agreement for categorical data</article-title>
          .
          <source>Biometrics</source>
          <volume>33</volume>
          (
          <issue>1</issue>
          ) (
          <year>1977</year>
          )
          <volume>159</volume>
          {
          <fpage>174</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>