<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Using Conditional Probabilities and Vowel Collocations of a Corpus of Orthographic Representations to Evaluate Nonce Words</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>O¨zkan Kılı¸c</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Cognitive Science, Informatics Institute, Middle East Technical University</institution>
          ,
          <addr-line>Ankara</addr-line>
          ,
          <country country="TR">Turkey</country>
        </aff>
      </contrib-group>
      <fpage>65</fpage>
      <lpage>72</lpage>
      <abstract>
        <p>Nonce words are widely used in linguistic research to evaluate areas such as the acquisition of vowel harmony and consonant voicing, naturalness judgment of loanwords, and children's acquisition of morphemes. Researchers usually create lists of nonce words intuitively by considering the phonotactic features of the target languages. In this study, a corpus of Turkish orthographic representations is used to propose a measure for the nonce word appropriateness for linearly concatenative languages. The conditional probabilities of orthographic co-occurrences and pairwise vowel collocations within the same word boundaries are used to evaluate a list of nonce words in terms of whether they would be rejected, moderately accepted or fully accepted. A group of 50 Turkish native speakers were asked to evaluate the same list of nonce words. Both the method and the participants displayed similar results.</p>
      </abstract>
      <kwd-group>
        <kwd>Nonce words</kwd>
        <kwd>Orthographic representations</kwd>
        <kwd>Conditional probabilities</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Nonce words are frequently employed in linguistic studies to evaluate areas such as
well-formedness [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], morphological productivity [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and development [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], judgment of
semantic similarity [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], and vowel harmony [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Nonce words are also used to
understand the process of adopting loan words. The majority of loaned words undergo certain
phonetic changes to more resemble the lexical entries of the language into which they
will be adopted [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. For example, television in Turkish becomes televizyon /televızjon/
because /jon/ is more frequent than /Zın/ in Turkish1. Similarly, train is adopted
as tren /tren/ because, similar to diphthongs, vowel-to-vowel co-occurrences are not
usually allowed in Turkish non-compound words. This phenomenon shows that the
speakers of a language are aware of the possible sound frequencies and collocations
of their native languages, and they can make judgements on the naturalness of loan
1In the METU-Turkish Corpus, there are 181 occurrences with the segment /Zın/
of which only 30 are at the terminating word boundaries. On the other hand, there are
5,945 occurrences with the segment /jon/ of which 3,190 are at the terminating word
boundaries, excluding the word televizyon.
words, recently invented words and nonce words by using their knowledge of the
existing Turkish lexis. Thus, the acceptability of nonce words is a logical decision based on
known-word statistics.
      </p>
      <p>
        The acceptability of nonce words can be investigated by experimental investigations
through phonotactic properties or factor-based analysis [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. In the experimental
investigations, it is observed that the participants accepted or rejected nonce words according
to probable combinations of sounds [
        <xref ref-type="bibr" rid="ref1 ref8">1, 8</xref>
        ]. In factor-based analysis, the acceptability of
nonce words is evaluated through the co-occurrences of syllables or consonant clusters
locally [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] or non-locally [
        <xref ref-type="bibr" rid="ref10 ref11 ref12">10–12</xref>
        ] or through nucleus-coda combination probabilities [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>
        In this study, the acceptability of nonce words was assessed using the conditional
probabilities of the bigram co-occurrences of the orthographic representations locally
and the pairwise collocations of the vowels within the same word boundaries. Similar
methods within the context of phonotactic modeling had been used for Finnish vowel
harmony [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Yet in this study, the local bigram phonotactic modeling was used to
evaluate Turkish nonce words. Two threshold values were set for the decision to reject,
moderately accept and fully accept. The threshold values were computed according to
the length of each input string. For the evaluation of the conditional and collocation
probabilities, the METU-Turkish Corpus containing about two million words was
employed [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. The list of nonce words was created intuitively. The same list of nonce
words evaluated by the method was also given to 50 Turkish native speakers to judge
the level of acceptability of each word. The 25 male and 25 female Turkish native
speakers, had an average age is 31.26 (s = 4.11).The results from the native speakers
were very similar to the results provided by the statistical method. In this paper, brief
information about Turkish language and plausibility of conditional probabilities will
be given then details of the method and the results will be presented.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Turkish Language and Conditional Probability</title>
      <p>
        Turkish has 8 vowels and 21 consonants, and it is agglutinative with a considerably
complex morphology [
        <xref ref-type="bibr" rid="ref16 ref17">16, 17</xref>
        ]. While communicating, the word internal structure in
Turkish is required to be segmented because Turkish morphosyntax plays a central
role in semantic analysis. For example, although Turkish is considered as an SOV
language, the sentences are usually in a free order. Thus, the subject and object of a verb
can only be determined by the morphological markers as in (1) rather than the word
order.
      </p>
      <p>(1) Ko¨pek adam-ı ısırdı.</p>
      <p>Dog man-Acc bit
The dog bit the man.</p>
      <p>Ko¨pe˘g-i adam ısırdı.</p>
      <p>Dog-ACC man bit</p>
      <p>The man bit the dog.</p>
      <p>The description of Turkish word structure depends heavily on
morphophonological constraints and morphotactics. In Turkish morphotactics, the continuation of a
morpheme is determined by the preceding morpheme or by the stem as in (2).
(2) ev-de-ki
house-Loc-Rel
The one in the house</p>
      <p>*ev-ki-de</p>
      <p>
        These morphotactic constraints in Turkish are captured by statistical models based
on conditional probabilities [
        <xref ref-type="bibr" rid="ref18 ref19">18, 19</xref>
        ]. In addition to morphotactics, the
morphophonology of Turkish needs a brief explanation because nonce words have to mimic this
morphophonology.
      </p>
      <p>Vowel harmony is dominantly effective in Turkish morphophonology in order to
preserve the roundedness and the frontness of vowels within the same word
boundaries. While a morpheme with a vowel is concatenated to a string, its vowel is modified
with respect to the roundedness and frontness properties of the most recent vowel in
the string as in (3).</p>
      <p>(3) ev-ler
house - Plu
houses
oda-lar
room - Plu
rooms
bil-di
know - Past
knew
duy-du
hear - Past
heard</p>
      <p>Another important phenomenon in Turkish morphophonology is voicing. If some
of the strings terminating with the voiceless consonant, ‘p, t, k, ¸c’, are followed by the
suffixes starting with vowels, then the consonants are voiced as ‘b, d, ˘g, c’ as in (4).
(4) sonu¸c
result
sonuc-um
result -1S.Poss
my result</p>
      <p>Consonant assimilation is also important in Turkish morphophonology. The initial
consonants of some morphemes undergo an assimilation operation if they are attached
to the strings terminating in the voiceless consonants, ‘p, t, k, ¸c, f, s, ¸s, h, g’, as in the
surface forms of the Turkish past tense -DI in (5).</p>
      <p>The final Turkish morphophonological phenomena that need to be briefly
mentioned are deletion and epenthesis occurring as in (6).</p>
      <p>(6) hak
right
hakk-ım
right - 1S.Poss
my right
isim
name
ism-im
name - 1S.Poss
my name</p>
      <p>The Turkish morphophonological phenomena described above occur in the
cooccurrences of the orthographic representations in the concatenating positions except
in vowel harmony and the deletion. This results in high conditional probabilities
evaluated using the frequencies of the pairs of consecutive orthographic representations.
Since the vowel harmony and deletion take place after or before the concatenation
positions, their pairwise collocations within the same word boundaries are also required
to be utilized in the statistical model.</p>
      <p>The transition probability between A and B is simply based on the conditional
probability statistics as in (7).
(7) P (B|A) = (frequency of AB) / (frequency of A)</p>
      <p>
        Infants are reported to successfully discriminate speech segments using transitional
probabilities of syllable pairs [
        <xref ref-type="bibr" rid="ref20 ref21">20, 21</xref>
        ]. Adults also make use of transitional probabilities
between word classes to acquire syntactic rules [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. Similarly, transition probabilities
are dominantly used in unsupervised morphological segmentation and disambiguation
[
        <xref ref-type="bibr" rid="ref18 ref19">18, 19</xref>
        ], [
        <xref ref-type="bibr" rid="ref23 ref24 ref25">23–25</xref>
        ].
      </p>
      <p>Statistical approaches to linguistics support the empiricist view; and they provide
an explanatory account of linguistic phenomena such as the decrease in performance
errors and language variations. Considering the properties of the Turkish language,
using the conditional probabilities of orthographic representations and the collocations
of vowels within the same word boundaries is a plausible method to decide whether
nonce words or loan words will be rejected, moderately accepted or accepted
3</p>
    </sec>
    <sec id="sec-3">
      <title>The Method</title>
      <p>Let s be a string such that s = u1u2. . . un, where ui is a letter in the Turkish alphabet.
The string s is unified with the empty strings σ and ε such that s = σu1u2. . . unε,
where σ denotes the initial word boundary and ε denotes the terminal word boundary.
The overall transition probability of the string s is evaluated from the METU-Turkish
Corpus using Formula 1.</p>
      <p>n+1
Pt(s) = Y P (ui|ui−1) (1)</p>
      <p>1</p>
      <p>For example, using the Formula 1, P (a|σ) gives the probability of the strings
starting with the letter a, and P (b|a) estimates the probability of the substring ab in the
corpus. Now let v be a subset of the string s such that v = ui,1uj,2 . . . uk,m where uk,m
is the mth vowel in the kth location of the string s. The overall vowel collocations of
the string s are estimated from the substring of vowels v using Formula 2.</p>
      <p>Pc(v) = Ym g(vi−1vi)</p>
      <p>f (vi−1)
2</p>
      <p>if |v| &gt; 1
Pc(v) = f (vi) if |v| = 1 (2)</p>
      <sec id="sec-3-1">
        <title>CorpusSize</title>
        <p>In the Formula 2, the function f (vi) gives the frequency of the words that contain
the vowel vi as a substring in the corpus. The function g(vi−1vi) gives the frequency of
words in which the vowels vi−1 and vi are collocating not necessarily in immediately
consecutive positions but within the same word boundaries. The acceptability
probability of the string s is calculated by Pa(s) = Pt(s)Pc(v). The acceptability decision of
the string s in the method is made by using the Formula 3.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Accept</title>
      </sec>
      <sec id="sec-3-3">
        <title>Reject if if if</title>
        <p>Pa(s) ≥ 10−(t+v)
10−(t+v+1) &gt; Pa(s)</p>
      </sec>
      <sec id="sec-3-4">
        <title>M oderately accept</title>
        <p>10−(t+v+1) ≤ Pa(s) &lt; 10−(t+v)
(3)
where t is the number of transitions (which is the length of the string + 1) and v
is the number of the vowel collocations (which is the number of the vowels - 1) in the
string. If the string s has only one vowel, then v = 1.</p>
        <p>The method was applied to the list of nonce words given in the following section.
The same list was also given to the 50 Turkish native speakers to evaluate the
acceptability of each item. The comparison of the results from the method and the native
speakers is given below.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>The nonce word talar is evaluated as in (8)
(8)</p>
      <p>Pa(talar) = Pt(σtalarε)xPc(aa)
= P (t|σ)P (a|t)P (l|a)P (a|l)P (r|a)P (ε|r)xPc(aa)
= 7.66e − 06xPc(aa) = 7.66e − 06 ∗ 4.75e − 01 = 3.63e − 06</p>
      <p>Since Pa(talar) ≥ 10−(6+1), in which 6 conditional probability estimations and 1
vowel collocation are evaluated, the nonce word talar is accepted. The word list was
evaluated by the 50 selected Turkish speakers. The distribution of the native speaker
responses and the results of the method are given in Table 1.</p>
      <p>For 82% of the words the Turkish native speaker’s responses are in agreement
with the results from the method. The method failed to simulate the responses from
the participants in 18% of the results.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Discussions and Conclusion</title>
      <p>
        The acceptability of loan words and nonce words is mainly determined by the
phonological properties of the target language and the current approaches are
syllablebased [
        <xref ref-type="bibr" rid="ref10 ref11 ref12 ref13 ref7 ref8 ref9">7–13</xref>
        ]. Since there are no lexical entries for nonce words, the method in this study
tries to estimate the acceptability of the words using the bigram conditional
probabilities and collocations of the orthographic representations within the word boundaries,
which is a simplified way of inducing Turkish morphophonology.
      </p>
      <p>The nonce word u¨lu¨ was rejected by the method but accepted by the participants. A
possible reason might be that the nonce word u¨lu¨ sounds similar to an existing Turkish
word o¨lu¨ ’death’. Similarly, the responses for the nonce word nort were in disagreement.
This nonce word has a similar pronunciation to an English word north and the most of
the participants also knew English as a foreign language. Therefore, the participants
might also make use of their foreign language knowledge to evaluate nonce words.</p>
      <p>Although the method does not assume to utilize any property of Turkish phonology
and it does not implement any phonologic filtering mechanism, it is able to mimic, in
a remarkable way, a large number of the responses from the participants. Indeed, this
study does not propose that acceptability is based on raw orthographical
representations rather than syllables and phonemes. Instead, it underlines that simple pairwise
conditional properties and vowel collocations from a corpus can give an estimation of
the acceptability of a list of nonce words. This can be used by researchers that need
an evaluation for the nonce words for their studies when no phonologically annotated
corpus with syllables exists.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Limitations and development</title>
      <p>The method needs to be tested with larger word lists. The method is successful because
there is a close correspondence between phonotactics and orthotactics in Turkish. It
requires improvements in terms of the morphophonological properties of target languages.
The method uses exact orthographic representations. Thus, it requires an additional
phonological similarity measure for the representations to increase the success rate.</p>
      <p>The threshold values for the acceptability decisions depend on word lengths. They
also need to be improved with respect to the target languages. The method also needs
to be tested and adapted for the languages with ablaut or umlaut phenomena such as
English and German, and the templatic languages such as Arabic and Hebrew.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Hammond</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Gradience, phonotactics, and the lexicon in English phonology</article-title>
          .
          <source>Int. J. of English Studies</source>
          <volume>4</volume>
          (
          <year>2004</year>
          )
          <fpage>1</fpage>
          -
          <lpage>24</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Anshen</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aronoff</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Producing morphologically complex words</article-title>
          .
          <source>Linguistics</source>
          <volume>26</volume>
          (
          <year>1988</year>
          )
          <fpage>641</fpage>
          -
          <lpage>655</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Dabrowska</surname>
          </string-name>
          , E.:
          <article-title>Low-level schemas or general rules? The role of diminutives in the acquisition of Polish case inflections</article-title>
          .
          <source>Language Sciences</source>
          <volume>28</volume>
          (
          <year>2006</year>
          )
          <fpage>120</fpage>
          -
          <lpage>135</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>MacDonald</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ramscar</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Testing the distributional hypothesis: The influence of context on judgements of semantic similarity</article-title>
          .
          <source>Proc. of the 23rd Annual Conference of the Cognitive Science Society</source>
          , University of Edinburgh (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Pycha</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Novak</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shosted</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shin</surname>
          </string-name>
          , E.:
          <article-title>Phonological rule-learning and its implications for a theory of vowel harmony</article-title>
          .
          <source>Proc. of WCCFL 22</source>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Garding and M. Tsujimura</surname>
          </string-name>
          (Eds.) (
          <year>2003</year>
          )
          <fpage>423</fpage>
          -
          <lpage>435</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Kawahara</surname>
            ,
            <given-names>S:</given-names>
          </string-name>
          <article-title>OCP is active in loanwords and nonce words: Evidence from naturalness judgment studies</article-title>
          .
          <source>Lingua</source>
          (to appear)
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Albright</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>From clusters to words: Grammatical models of nonce word acceptability</article-title>
          .
          <source>Handout of talk presented at 82nd LSA, Chicago, January</source>
          <volume>3</volume>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Shademan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>From clusters to words: Grammatical models of nonce word acceptability. Grammar and Analogy in Phonotactic Well-formedness Judgments</article-title>
          .
          <source>Ph. D. thesis</source>
          , University of California, Los Angeles (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Hay</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pierrehumbert</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beckman</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Speech perception, well-formedness and the statistics of the lexicon</article-title>
          . In: J.
          <string-name>
            <surname>Local</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Ogden</surname>
          </string-name>
          , and R. Temple (Eds.), Phonetic Interpretation: Papersbin Laboratory Phonology VI. Cambridge: Cambridge University Press (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Frisch</surname>
            ,
            <given-names>S. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zawaydeh</surname>
            ,
            <given-names>B. A.</given-names>
          </string-name>
          :
          <article-title>The psychological reality of OCP-Place in Arabic</article-title>
          .
          <source>Language</source>
          <volume>77</volume>
          (
          <year>2001</year>
          )
          <fpage>91</fpage>
          -
          <lpage>106</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Koo</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Callahan</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Tier-adjacency is not a necessary condition for learning phonotactic dependencies</article-title>
          .
          <source>Language and Cognitive Processes</source>
          <volume>77</volume>
          (
          <year>2011</year>
          )
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Finley</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Testing the limits of long-distance learning: learning beyond a threesegment window</article-title>
          .
          <source>Cognitive Science</source>
          <volume>36</volume>
          (
          <year>2012</year>
          )
          <fpage>740</fpage>
          -
          <lpage>756</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Treiman</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kessler</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Knewasser</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tincoff</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Bowman</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>English speakers' sensitivity to phonotactic patterns</article-title>
          .
          <source>In: M. B. Broe and J. Pierrehumbert (eds.)</source>
          , Papers in Laboratory Phonology V:
          <article-title>Acquisition and the Lexicon</article-title>
          . Cambridge: Cambridge University Press (
          <year>2000</year>
          )
          <fpage>269</fpage>
          -
          <lpage>282</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Goldsmith</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Riggle</surname>
          </string-name>
          , J.:
          <article-title>Information theoretic approaches to phonological structure: the case of Finnish vowel harmony</article-title>
          .
          <source>Natural Language &amp; Linguistic Theory</source>
          (to appear)
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Say</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zeyrek</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oflazer</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , O¨zge, U.:
          <article-title>Development of a corpus and a treebank for present-day written Turkish</article-title>
          .
          <source>Proc. of the Eleventh International Conference of Turkish Linguistics</source>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16. G¨oksel,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Kerslake</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            :
            <surname>Turkish</surname>
          </string-name>
          :
          <string-name>
            <given-names>A Comprehensive</given-names>
            <surname>Grammar</surname>
          </string-name>
          . Routledge: London and New York (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Lewis</surname>
            ,
            <given-names>G</given-names>
          </string-name>
          : Turkish Grammar,
          <article-title>Second edition</article-title>
          . Oxford: University Press (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18. Kılı¸c, O¨., Boz¸sahin, C.:
          <article-title>Semi-supervised morpheme segmentation without morphological analysis</article-title>
          .
          <source>Pro. of the LREC 2012 Workshop on Language Resources</source>
          and
          <article-title>Technologies for Turkic Languages, I˙stanbul</article-title>
          ,
          <string-name>
            <surname>Turkey</surname>
          </string-name>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Yatbaz</surname>
            ,
            <given-names>M. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yuret</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Unsupervised morphological disambiguation using statistical language models</article-title>
          .
          <source>Pro. of the NIPS 2009 Workshop on Grammar Induction, Representation of Language and Language Learning</source>
          , Whistler, Canada (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Aslin</surname>
            ,
            <given-names>R.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saffran</surname>
            ,
            <given-names>J.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Newport</surname>
            ,
            <given-names>E.L.</given-names>
          </string-name>
          :
          <article-title>Computation of conditional probability statistics by human infants</article-title>
          .
          <source>Psychological Science</source>
          <volume>9</volume>
          (
          <year>1998</year>
          )
          <fpage>321</fpage>
          -
          <lpage>324</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Gomez</surname>
            ,
            <given-names>R. L.</given-names>
          </string-name>
          :
          <article-title>Variability and detection of invariant structure</article-title>
          .
          <source>Psychological Science</source>
          <volume>13</volume>
          (
          <year>2002</year>
          )
          <fpage>431</fpage>
          -
          <lpage>436</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Kaschak</surname>
            ,
            <given-names>M. P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saffran</surname>
            ,
            <given-names>J. R.</given-names>
          </string-name>
          :
          <article-title>Idiomatic syntactic constructions and language learning</article-title>
          .
          <source>Cognitive Science</source>
          <volume>30</volume>
          (
          <year>2006</year>
          )
          <fpage>43</fpage>
          -
          <lpage>63</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Creutz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lagus</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Unsupervised models for morpheme segmentation and morphology learning</article-title>
          .
          <source>ACM Tran. on Speech and Language Processing</source>
          <volume>4</volume>
          (
          <issue>1</issue>
          ) (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Bernhard</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Unsupervised morphological segmentation based on segment predictability and word segments alignment</article-title>
          .
          <source>Proc. of 2nd Pascal Challenges Workshop</source>
          (
          <year>2006</year>
          )
          <fpage>19</fpage>
          -
          <lpage>24</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Demberg</surname>
          </string-name>
          , V.:
          <article-title>A language-independent unsupervised model for morphological segmentation</article-title>
          .
          <source>Ann. Meet. of Assoc. for Computational Linguistics</source>
          <volume>45</volume>
          (
          <issue>1</issue>
          ) (
          <year>2007</year>
          )
          <fpage>920</fpage>
          -
          <lpage>927</lpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>