<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>KPCA Embeddings: an Unsupervised Approach to Learn Vector Representations of Finite Domain Sequences</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Eduardo Brito</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rafet Sifa</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christian Bauckhage</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Eduardo.Alfredo.Brito.Chacon,Rafet.Sifa,Christian.Bauckhage} @iais.fraunhofer.de https://multimedia-pattern-recognition.info</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Fraunhofer IAIS Schloss Birlinghoven</institution>
          ,
          <addr-line>Sankt Augustin</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <abstract>
        <p>Most of the well-known word embeddings from the last few years rely on a prede ned vocabulary so that out-of-vocabulary words are usually skipped when they need to be processed. This may cause a signi cant quality drop in document representations that are built upon them. Additionally, most of these models do not incorporate information about the morphology of the words within the word vectors or if they do, they require labeled data. We propose an unsupervised method to generate continuous vector representations that can be applied to any sequence of nite domain (such as text or DNA sequences) by means of kernel principal component analysis (KPCA). We also show that, apart from their potential value as a preprocessing step within a more complex natural language processing system, our KPCA embeddings also can capture valuable linguistic information without any supervision, in particular word morphology of German verbs. When they are applied to DNA sequences, they also encode enough information to detect splice junctions.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Machine learning approaches for natural language processing (NLP) generally
demand a numeric vector representation for words. We can distinguish any two
di erent words from a xed vocabulary by assigning them a one-hot vector,
where all entries of the vector are zero-valued but in a single position that
identi es the word. This is a very sparse representation that encodes no information
about the words but their position in the vocabulary.</p>
      <p>A more information-rich alternative to one-hot vectors are the so-called word
embeddings. They are distributed vector representations, which are dense,
lowdimensional, real-valued and can capture latent features of the word [13]. Based
on the distributional hypothesis [5] (words that appear in similar contexts have
similar meanings), they exploit word co-occurrence so that similar words are
mapped close to each other in the word vector space. Part of the success of the
word embeddings is due to their e cient shallow neural network architectures
such as the continuous skip-gram model and the continuous bag of words model
[9], widely popularized after the release of word2vec1. Word embeddings also
inspired research in other areas di erent from NLP to learn vector representations
such as node representations from a graph [4].</p>
      <p>Although syntax and semantics can be encoded with word2vec embeddings,
they do not incorporate morphological information about the word. As a
consequence, morphologically similar words may not be nearby in the word vector
space. Some approaches make use of existing linguistic resources so that the
word embeddings capture not only contextual information but also
morphological information[3, 7]. Due to the their dependence on language-speci c resources,
they will not work for languages whose available linguistic resources are scarce.</p>
      <p>The subword-based models such as FastText [1, 6] learn indirectly
morphology by learning not only word vectors but also n-gram vectors. This enables not
only complete unsupervised learning but also the possibility of inferring
out-ofvocabulary (OOV) words, which is an important issue for all other approaches
mentioned, notably when noisy informal text needs to be processed. Despite
these advantages of subword-based models, some morphologically rich languages
(where these models are supposed to perform specially well) may contain very
long words. This leads to a dramatic increase of the number of necessary n-gram
representations, increasing time and space complexity as well.</p>
      <p>We propose an alternative unsupervised method by means of kernel
principal component analysis (KPCA) that encodes morphology while learning word
representations and that generate new vectors for OOV words after training. In
addition, our approach is general enough to learn vector representations for any
sequence whose elements belong to a xed prede ned nite set. In particular, we
test our KPCA embeddings in two di erent tasks: classifying the verb category
for German verbs and detecting splice junctions in DNA sequences.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Approach</title>
      <p>
        KPCA indirectly maps vectors to a feature space (of higher dimension) in order
to obtain the principal components from that space [12]. No explicit calculation
in the feature space is required since we only need to be able to compute the
inner product in the feature space and this can be achieved by using kernel
functions. We can exploit the freedom to select any inner product of our choice
so that we can also perform KPCA to non-numeric entities. Formally, given a
zero-mean column data matrix X = [x1; x2; : : : ; xn] containing n m-dimensional
data points, principal component analysis (PCA) deals with representing the
data through principal components maximizing the variance in X by solving
Cv =
v;
(
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
1 https://code.google.com/archive/p/word2vec
where C = n1 XXT is the covariance matrix, an arbitrary eigenvalue and v its
corresponding eigenvector. Considering (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), we can represent every eigenvector
as a linear combination of the data points the data matrix X as
Replacing the occurrences of the gram matrix XTX by a selected kernel matrix
K as
1
n
1
n
where i determines the weight of the S rensen-Dice coe cient term for each
n-gram length. We compute a similarity matrix S by applying this similarity
function s to all word pairs of our vocabulary V :
      </p>
      <sec id="sec-2-1">
        <title>XTXXTX</title>
        <p>=</p>
        <p>XTX :
1
n</p>
        <p>KK
=</p>
        <p>K :
and eliminating K as
1</p>
        <p>
          K = ; (
          <xref ref-type="bibr" rid="ref6">6</xref>
          )
n
we result in a kernelized representation of the conventional PCA.
        </p>
        <p>For the particular case of words and DNA sequences, the inner product can
be any string similarity. For our experiments, we adapt the the S rensen-Dice
coe cient by considering not only bigrams, but n-grams of any length in general.
Let V a vocabulary of words, Gn(w) the n-grams of a word w 2 V. We de ne
our similarity function s of two words x; y 2 V as follows:
s(x; y) = X
n2N+
n
2jGn(x) \ Gn(y))j ;
jGn(x) [ Gn(y)j</p>
        <p>X
n
n = 1
where</p>
        <p>
          2 Rm. Substituting (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) in (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) we obtain
which upon data projection can be represented as
        </p>
        <p>XXTv = X</p>
        <p>= v;</p>
      </sec>
      <sec id="sec-2-2">
        <title>XXTX</title>
        <p>=</p>
        <p>X ;
Sij = s(wi; wj )
8wi; wj 2 V:
Then, we calculate the kernel matrix K by applying a non-linear kernel
function (for instance RBF kernel or polynomial kernel) to the similarity matrix S.
After computing the eigenvectors and eigenvalues of the resulting matrix K, we
construct our projection matrix P by selecting d eigenvectors v1; : : : ; vd and
dividing them by their respective eigenvalues 1; : : : ; d:</p>
        <p>P = [ v1 ; : : : ; vd ]:</p>
        <p>1 d</p>
        <p>
          We can now generate a KPCA embedding for any word wt. We only need to
compute the similarity function of the word against all the words processed from
(
          <xref ref-type="bibr" rid="ref2">2</xref>
          )
(
          <xref ref-type="bibr" rid="ref3">3</xref>
          )
(
          <xref ref-type="bibr" rid="ref4">4</xref>
          )
(
          <xref ref-type="bibr" rid="ref5">5</xref>
          )
(
          <xref ref-type="bibr" rid="ref7">7</xref>
          )
(
          <xref ref-type="bibr" rid="ref8">8</xref>
          )
(
          <xref ref-type="bibr" rid="ref9">9</xref>
          )
0.2
0.1
0.0
0.1
0.2
steigen
umsteigen
einsteigen
        </p>
        <p>austeigen
steige
einusmtesitgeeige
austeige
gestiegen
umgestiegen
eingestaieugsegneusmtieggeezongen
eingezogaeunsgezogen
ein</p>
        <p>um
aus
geboren</p>
        <p>gelesen
gebore</p>
        <p>geliebtlesen
gegdeamcahctht
denken lieben
denke
lesemachen
liebe</p>
        <p>gesehen
mache</p>
        <p>
          umziehen
einziehenausziehen
sehen
sehe
einziehe auszieuhmeziehe
0.20
0.15
0.10
0.05
0.00
0.05
the vocabulary V and apply the kernel function k to the resulting vector. This
results in a kernelized distance vector rt. The product of rt with the projection
matrix P constitutes the d-dimensional KPCA embedding ut of the word wt.
rt = k(s(wt; V))
and
ut = P&gt;rt:
(
          <xref ref-type="bibr" rid="ref10">10</xref>
          )
It is important to note that, this approach can be generalized to encode any other
non-numeric entity as long as we can de ne an equivalent similarity function
between each pair of entities.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Experiments</title>
      <p>We show how our KPCA embeddings can be used for di erent kinds of sequential
data. In particular, we apply our approach to represent words to classify di erent
types of verb forms and to represent DNA sequences to recognize splice junctions.
3.Sg.Pres.Ind</p>
      <p>Psp</p>
      <p>Inf
3.Sg.Past.Ind
3.Pl.Pres.Ind
3.Sg.Pres.Subj</p>
      <p>3.Pl.Past.Ind
3.Pl.Past.Subj
3.Sg.Past.Subj</p>
      <p>Infzu
1.Pl.Pres.Ind
3.Pl.Pres.Subj
1.Sg.Pres.Ind
1.Sg.Past.Ind</p>
      <p>2.Sg.Imp
1.Pl.Past.Ind
2.Sg.Pres.Ind
1.Pl.Past.Subj</p>
      <p>2.Pl.Imp
2.Pl.Pres.Ind
1.Sg.Past.Subj</p>
      <p>Pos
2.Sg.Pres.Subj
1.Sg.Pres.Subj</p>
      <p>Pl.3.Pres.Ind
2.Sg.Past.Ind</p>
      <p>Prp
2.Pl.Past.Ind
1.Pl.Pres.Subj
2.Sg.Past.Subj</p>
      <p>Pl.1.Pres.Ind
0.00
0.05
0.10</p>
      <p>0.15
Ratio
0.20
0.25
We test how our KPCA embeddings can encode morphology with a ne-grain
POS tagging task. We restrict our vocabulary to the tokens tagged as verbs
from the TIGER treebank [2]. This simpli es the problem to classify the correct
morphological tag (consisting of grammatical person, number, tense and mode
when they apply) of a German verb only from its KPCA embedding.</p>
      <p>First, we extract all tokens tagged as verb (corresponding to the TIGER tags
VVFIN, VAFIN, VMFIN, VVIMP, VAIMP, VVINF, VVIZU, VAINF, VMINF,
VVPP, VMPP, VAPP) and remove all duplicates. This leads to 13370 unique
verbs with 31 di erent morphological tags, whose distribution is showed in gure
2. Then, we apply the approach described in section 2. We build a training set
consisting of 80% of the verbs and a test set with the remaining 20%. We consider
only bigrams and trigrams to compute the similarity function for each pair of
verbs by selecting ve di erent weight distributions (di erent values for 2 and
3 in equation 7). We also incorporate an additional character at the beginning
and at the end of each verb when producing the n-grams. After running KPCA
on the training set, we infer vector representations of the verbs from the test
set. By using the KPCA embeddings as a features and the morphological tag as
label, we train k-nearest neighbors classi ers to predict the morphological tag
of a word from only the KPCA embedding. As baseline representations, we also
learn a word2vec [10, 9] model for each di erent vector size. These word vectors
were learned applying the default hyperparameter values.</p>
      <p>
        From Table 1 we can observe that a mean accuracy above 77% can be
achieved by classi ers taking only the nearest neighbor (k = 1). This can be
interpreted as a high accuracy considering the extremely unbalanced label
distribution (see gure 2). Among the trained classi ers, we can also nd some
improvement when the trigram similarity weight (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) is at least as high as
the bigram similarity (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ). In addition, any KPCA embedding model beats all
word2vec models for this task. For the sake of a fair comparison, the displayed
word2vec results from Table 1 correspond to models where their vector size
matches with a tested number of principal components d of KPCA embeddings
and tested with the same k values for k-nearest neighbors. However, we also
tested additional word2vec models with vector sizes up to 100 and up to 100
neighbors. None of these larger models reached a mean accuracy above 27%.
3.2
      </p>
      <p>Splice junction recognition on DNA sequences
We encode DNA sequences from the dataset \Molecular Biology (Splice-junction
Gene Sequences) Data Set" from UCI Machine Learning Repository [8]2. This
dataset consists of DNA subsequences represented as 30 characters out of the
four nucleobases (A, T, C, G) plus other four characters (D, N, S, R) which
mark ambiguity. The sequences may contain a spline junction between the 30
rst and the 30 last characters. They are thus labeled with three di erent
categories depending if they contain exon/intron boundary (EI class), intron/exon
boundary (IE class) or neither (N). The distribution of the classes is displayed
in Table 2.</p>
      <p>
        We compute the similarity function considering all n-grams, giving all terms
from (
        <xref ref-type="bibr" rid="ref7">7</xref>
        ) the same weight:
i =
(1=58; i 2 f2;
0;
i 2= f2;
; 59g
; 59g
(
        <xref ref-type="bibr" rid="ref11">11</xref>
        )
Using the learned representations, we train k-nearest neighbors classi ers to
predict to which of the three classes each DNA sequence belongs to. Analyzing
the prediction performance, as we can see from Table 3, we achieve a mean
accuracy 94.67% with two di erent kernels. This result beats all baseline systems
that are presented with the dataset, including knowledge-based arti cial neural
networks (KBANN) [11].
2 https://archive.ics.uci.edu/ml/datasets/Molecular+Biology+
(Splice-junction+Gene+Sequences)
d
3 k = 1 k = 2 k = 3 k = 5 k = 10
d
3 k = 1 k = 2 k = 3 k = 5 k = 10
      </p>
      <p>Nr. sequences Ratio</p>
    </sec>
    <sec id="sec-4">
      <title>Discussion and future work</title>
      <p>We showed that our KPCA embedding approach to learn vector word
representations can encode the morphology of the words in an unsupervised fashion, at
least for the particular case of German verbs. The learned KPCA embeddings
could beat by far any word2vec model in the task of predicting the
grammatical tag. The highest accuracy was achieved by taking the nearest neighbor to
predict the verb category. Due to this fact, we suspect that our proposed word
representations tend to form clusters according to their word form, from which
predicting a grammatical tag is a feasible task with simple classi ers such as
k-nearest neighbors classi ers.</p>
      <p>Since many NLP applications require also syntactic and semantic information
about the words, good word embeddings should also incorporate information
not only from the form of the represented word but also about the context in
which they appear. In this direction, we will enhance our approach by adapting
our similarity function so that it also considers the frequency of each evaluated
word pair appearing in the same context or, alternatively, by taking our KPCA
representations as input representation of a neural network architecture. For the
latter, our KPCA would \just" substitute the one-hot encoding of most of neural
EI
IE</p>
      <p>N
language models. Additionally, we will also extend our research evaluating the
same approach on other morphologically rich in ected languages (like any of the
Romance languages) or agglutinative languages (such as Turkish). We assume
they may pro t the most from our approach since their word forms reveal more
grammatical information than word forms from more analytic languages such as
English. To this extent, KPCA embeddings may also help to overcome the lack
of linguistic resources of some of these non-English languages.</p>
      <p>Furthermore, we presented how we can learn KPCA embeddings for DNA
sequences. These representations proved to be useful in the task of predicting
splice junctions. It would be also interesting to generalize our method to encode
other types of discrete sequential data where also n-grams could be extracted,
for instance text paragraphs (word n-grams), protein sequences (amino acid
ngrams) or even sheet music (note n-grams).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>P.</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Grave</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Joulin</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          .
          <article-title>Enriching word vectors with subword information</article-title>
          .
          <source>arXiv preprint arXiv:1607.04606</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>S.</given-names>
            <surname>Brants</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dipper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Eisenberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hansen-Schirra</surname>
          </string-name>
          , E. Konig, W. Lezius,
          <string-name>
            <given-names>C.</given-names>
            <surname>Rohrer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and H.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          . Tiger:
          <article-title>Linguistic interpretation of a german corpus</article-title>
          .
          <source>Research on Language and Computation</source>
          ,
          <volume>2</volume>
          (
          <issue>4</issue>
          ):
          <volume>597</volume>
          {
          <fpage>620</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>R.</given-names>
            <surname>Cotterell</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Schu</surname>
          </string-name>
          <article-title>tze. Morphological word-embeddings</article-title>
          .
          <source>In HLT-NAACL</source>
          , pages
          <volume>1287</volume>
          {
          <fpage>1292</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>A.</given-names>
            <surname>Grover</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Leskovec</surname>
          </string-name>
          . node2vec:
          <article-title>Scalable feature learning for networks</article-title>
          .
          <source>In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source>
          , pages
          <volume>855</volume>
          {
          <fpage>864</fpage>
          . ACM,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Z. S.</given-names>
            <surname>Harris</surname>
          </string-name>
          .
          <article-title>Distributional structure</article-title>
          .
          <source>Word</source>
          ,
          <volume>10</volume>
          (
          <issue>2-3</issue>
          ):
          <volume>146</volume>
          {
          <fpage>162</fpage>
          ,
          <year>1954</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>A.</given-names>
            <surname>Joulin</surname>
          </string-name>
          , E. Grave,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          .
          <article-title>Bag of tricks for e cient text classi cation</article-title>
          .
          <source>arXiv preprint arXiv:1607.01759</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. G. Jurdzinski et al.
          <article-title>Word embeddings for morphologically complex languages</article-title>
          .
          <source>Schedae Informaticae</source>
          ,
          <volume>25</volume>
          :
          <fpage>127</fpage>
          {
          <fpage>138</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>M. Lichman.</surname>
          </string-name>
          <article-title>UCI machine learning repository</article-title>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          , G. Corrado, and
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <article-title>E cient estimation of word representations in vector space</article-title>
          .
          <source>In ICLR Workshop</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , I. Sutskever,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. S.</given-names>
            <surname>Corrado</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          , pages
          <volume>3111</volume>
          {
          <fpage>3119</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>M. O. Noordewier</surname>
            ,
            <given-names>G. G.</given-names>
          </string-name>
          <string-name>
            <surname>Towell</surname>
            , and
            <given-names>J. W.</given-names>
          </string-name>
          <string-name>
            <surname>Shavlik</surname>
          </string-name>
          .
          <article-title>Training knowledge-based neural networks to recognize genes in dna sequences</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          , pages
          <fpage>530</fpage>
          {
          <fpage>536</fpage>
          ,
          <year>1991</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12. B.
          <article-title>Scholkopf, A. Smola, and</article-title>
          K.
          <string-name>
            <surname>-R. Mu</surname>
          </string-name>
          <article-title>ller. Kernel principal component analysis</article-title>
          .
          <source>In International Conference on Arti cial Neural Networks</source>
          , pages
          <volume>583</volume>
          {
          <fpage>588</fpage>
          . Springer,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>J. Turian</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Ratinov</surname>
            , and
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Bengio</surname>
          </string-name>
          .
          <article-title>Word representations: A simple and general method for semi-supervised learning</article-title>
          .
          <source>In A. for Computational Linguistics, editor, 48th Annual Meeting of the Association for Computational Linguistics</source>
          , pages
          <volume>384</volume>
          {
          <fpage>394</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>