<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Free or Fixed Word Order: What Can Treebanks Reveal?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vladislav Kubonˇ</string-name>
          <email>vk@ufal.mff.cuni.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Markéta Lopatková</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Charles University in Prague, Faculty of Mathematics and Physics, Institute of Formal and Applied Linguistics Malostranské nám.</institution>
          <addr-line>25, Prague 1, 118 00</addr-line>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <fpage>23</fpage>
      <lpage>29</lpage>
      <abstract>
        <p>The paper describes an ongoing experiment consisting in the attempt to quantify word-order properties of three Indo-European languages (Czech, English and German). The statistics are collected from the syntactically annotated treebanks available for all three languages. The treebanks are searched by means of a universal query tool PML-TQ. The search concentrates on the mutual order of a verb and its complements (subject, object(s)) and the statistics are calculated for all permutations of the three elements. The results for all three languages are compared and a measure expressing the degree of word order freedom is suggested in the final section of the paper. This study constitutes a motivation for formal modeling of natural language processing methods.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        General linguistics, see esp. [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ] studies natural
languages from the point of view of similarities and
differences in their syntactic structure, their development and
historical changes, as well as from the point of view of
language functions. It studies mutual influence of
particular groups of features and, on the basis of similarities of
language phenomena it introduces the so called language
typology [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ]. The freedom or, on the other hand,
strictness of the word order definitely belongs among the most
important phenomena. General linguistics, for example,
studies whether and how a particular language handles the
order of words in sentences – whether the word is
determined primarily by syntactic categories (e.g., a noun or a
pronoun, without any additional morphological signs,
located on the first sentential position represents a subject
in English), or whether syntactic categories are primarily
determined by other means than by the word order (for
example, in Slavic languages, the subject tends to be a noun
in the nominative case, regardless of its position in the
sentence).
      </p>
      <p>
        Particular natural languages cannot be, of course,
strictly characterized by a single feature (for example word
order), they are typically categorized into individual
language types by a mixture of characteristic features. If we
concentrate on word order, we study the prevalent order of
the verb and its main complements – indo-european
languages are thus characterized as SVO (SVO reflecting the
order Subject, Verb, Object) languages. English and other
languages with a fixed word order typically follow this
order of words in declarative sentences; although Czech,
Russian and other Slavic languages are the so-called
languages with a high degree of word order freedom, they still
stick to the same order of word in a typical (unmarked)
sentence. As for the VSO-type languages, their
representatives can be found among semitic (Arabic, classical
Hebrew) or Celtic languages, while (some) Amazonian
languages belong to the OSV type. These characteristics,
which are traditionally mentioned in classical textbooks
of general linguistics [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], have been specified on the basis
of excerptions and careful examination by many linguists.
      </p>
      <p>Today, when we have at our disposal a wide range of
linguistic data resources for tens of languages, we can
easily confirm (or enhance by quantitative clues) their
conclusions. This paper represents one of the steps in this
direction.</p>
      <p>
        The Institute of Formal and Applied Linguistics at the
Charles University in Prague, has established a repository
for linguistic data and resources LINDAT/CLARIN1. This
repository enables experiments with syntactically
annotated corpora, so called treebanks, for several tens of
languages. Wherever it is possible due to the license
agreements, the corpora are trasformed into a common format,
which enables – after a very short period of getting
acquainted with each particular treebank – a comfortable
search and analysis of the data from a particular language.
The HamleDT2 (HArmonized Multi-LanguagE
Dependency Treebank) project has already managed to transform
more than 30 treebanks from all over the world [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] into
a common format.
      </p>
      <p>
        In this pilot study we concentrate on three
Indoeuropean languages which substantially differ by the
degree of word freedom – Czech, German and English. We
investigate their typological properties on the basis of the
Prague Dependency Treebank [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], the English part of the
Prague Czech-English Dependency Treebank[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and the
German treebank TIGER [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] by means of the interface of
PML-TQ Tree Query [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], which enables the access to the
treebanks from the HamleDT.3
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Setup of the Experiment</title>
      <p>
        The analysis of syntactic properties of natural languages
constitutes one of our long term goals. The phenomenon
of word order has been in a center of our investigations
1https://lindat.mff.cuni.cz/cs/
2http://ufal.mff.cuni.cz/hamledt
3https://lindat.mff.cuni.cz/services/pmltq/
for a very long time. Our previous investigations
concentrated both on studying individual properties of languages
with higher degree of word-order freedom (as, e.g.,
nonprojective constructions (long-distance dependencies) [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]
as well as on the endeavor to find some general measures
enabling to more precisely characterize concrete natural
languages with regard to the degree of their word-order
freedom (see, e.g. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]).
      </p>
      <p>The experiment presented in this paper continues in the
same direction. It is driven by the endeavor to find an
objective way how to compare natural languages from the
point of view of the degree of their word-order freedom.
While the previous experiments concentrated on more
formal approach, this one builds upon a thorough analysis of
available data resources. Let us briefly introduce them in
the subsequent subsections.</p>
      <p>When investigating syntactic properties of natural
languages, it is very often the case that the discussion
concentrates on individual phenomena, their properties and their
influence on the order of words. The mere presence of
some phenomenon (or its more detailed properties) is, of
course, important and definitely influences the degree of
word-order freedom but this kind of investigation cannot
be complete without stating also the quantitative
properties of the given phenomenon. A linguistically interesting,
but marginal phenomenon does not tell us so much as a
basic phenomenon occurring relatively frequently. This
observation constitutes the basis of our current experiment.
In order to capture the quantitative characteristic of a
natural language, let us take a representative sample of its
syntactically annotated data and let us calculate the
distribution of individual types of word order for the three main
syntactic components – subject, predicate and object. It is
obvious that the more free is the word order of a given
language, the more equally they are going to be distributed.
2.1</p>
      <sec id="sec-2-1">
        <title>Available Treebanks</title>
        <p>
          The extensive quantitative analysis of the same linguistic
phenomenon for different languages would not be
feasible without a common platform which makes it possible
to compare various data resources from the same point of
view. Thanks to the initiative HamleDT4 (HArmonized
Multi-LanguagE Dependency Treebank) it is now
possible to compare the data from more than 30 languages in
a uniform way [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
        <p>
          The HamleDT family of treebanks is based on the
dependency framework and technology developed for the
Prague Dependency Treebank (PDT)5 [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], i.e., large
syntactically annotated corpus for the Czech Language. Here
we focus on the so-called analytical layer, i.e., the layer
describing surface sentence structure (relevant for
studying word order properties). The framework and its
language independence was verified within (the English
4http://ufal.mff.cuni.cz/hamledt
5http://ufal.mff.cuni.cz/pdt3.0
part of) the Prague Czech English Dependency Treebank
(PCEDT)6 [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] – within this project, syntactically
annotated Penn Treebank7 [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] was automatically transformed
from the original phrase-structure trees into the
dependency annotation.8 Based on this experience, the
HamleDT initiative goes further, syntactically annotated
corpora for different languages are collected and transferred
into the common format. Here we make use of the TIGER
corpus9 for the German language [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], the corpus with
native phrase-structure annotation enriched with the
information about the head for each phrase (and thus bearing
also information on dependencies). Figures 2, 6 and 7
show sample trees for Czech, English and German,
respectively, and Table 1 summarizes the size of these corpora.
corpus
PDT
PCEDT
TIGER
# preds
79,283
51,048
36,326
lang
Czech
English
German
type
manual
automatic
automatic
genre
news
economy
news
For searching the data, we exploit a PML-TQ search
tool,10 which has been primarily designed for processing
the PDT data. PML-TQ is a query language and search
engine designed for querying annotated linguistic data [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]
– it allows users to formulate complex queries on richly
annotated linguistic data.
        </p>
        <p>Having the treebanks in the common data format, the
PML-TQ framework makes it possible to analyse the data
in a uniform way – the following sample query gives
us trees with an intransitive predicate verb (in a main
clause), i.e. Pred node with Sb node and no Obj nodes
among its dependent nodes, where Sb follows the Pred;
the filter on the last line (&gt;&gt; for $n0.lemma give $1,
count() ) outputs a table listing verb lemmas with this
marked word order position and number of their
occurrences in the corpus, see also Figure 1.</p>
        <p>a-node $n0 :=
[ afun = "Pred",
child a-node $n1 :=
[ afun = "Sb", $n1.ord &gt; $n0.ord ],
0x child a-node
[ afun = "Obj"]]
&gt;&gt; for $n0.lemma give $1, count()
6http://ufal.mff.cuni.cz/pcedt2.0/cs/index.html
7https://www.cis.upenn.edu/ treebank
8This dependency-based surface annotation then served as a basis for
deep syntactic dependency-based annotation of English; however, as for
Czech, only surface structure is interesting for the studied phenomenon
of word order.</p>
        <p>9http://www.ims.uni-stuttgart.de/forschung/
ressourcen/korpora/tiger.html
10https://lindat.mff.cuni.cz/services/pmltq/
Let us now look at the syntactic typology of natural
languages under investigation. We are going to take into
account especially the mutual position of subject,
predicate and direct object. After a thorough investigation of
the ways how indirect objects are annotated in all three
corpora, we have decided to limit ourselves – at least in
this stage of our research – to basic structures and to
extract and analyse only sentences without too complicated
or mutually interlocked phenomena. Namely we focus on
sentences with the following properties:
• A predicate under scrutiny belongs to the main clause
(as e.g. in the sentence JsouPred vám nejasná neˇkterá
ustanovení danˇ ových zákon ˚u? ‘ArePred certain
provisions of the tax laws unclear to you?’, see the
dependency tree in Fig. 2); i.e., we do not analyse word
order of dependent clauses;
• We analyse only non-prepositional subjects and
objects (compare e.g. with the sentence V 2180 meˇstech
a obcích žije na 2.6 milionu obyvatelSb; ‘There are
(about 2.6 milion of inhabitants)Sb living in 2 180
towns and villages;’, see Fig. 3);
• Sentences may contain coordinated predicates (as,
e.g., predicates následoval and opakovalo in the
corpus sentence Vzápeˇtí následovalPred další regulacˇní
stupenˇ a vše se opakovaloPred . ‘The next level of
regulation immediately followedPred and everything
repeatedPred again.’, see Fig. 4);
However, sentences with common subjects (or
objects) are not taken into account (thus sentences as,
e.g., KoupelnaSb nebo teplá vodaSb nejsou trvale k
dispozici. ‘A bathroomSb or hot water supplySb are
not at the permanent disposal.’, see Fig. 5 are not
counted in the tables).11
11Including coordination phenomena in all their complexity would
require much robust queries in any dependency framework; thus we have
decided to disregard this type of sentences at all.
The highest quality syntactically annotated Czech data can
be found in the Prague Dependency Treebank; in fact, it
is the only corpus we work with that has been manually
annotated and thoroughly tested for the annotation
consistency. The texts of PDT belong mostly to the
journalism genre, it consists of newspaper texts and (in a limited
scale) of texts from a popularizing scientific journal.</p>
        <p>The following Table 2 summarizes the number of
sentences with intransitive verbs in main clauses in PDT with
respect to the word order positions of Sb and Pred – we
can see that the marked word order (verb preceding its
subject) is quite common in Czech.12</p>
        <p>The second table displays the distribution of individual
combinations of a subject, predicate and a single object.</p>
        <p>12In our settings, we do not checked the part of speech of the
predicate; however, out of the 79,283 sentences conforming to the properties
mentioned above, only 329 have other than verbal predicate.
It is not surprising that the unmarked – intuitively "most
natural" – word order type, SVO, accounts for only slightly
more than half of cases. The relatively high degree of word
order freedom is thus supported also quantitatively.</p>
        <p>Even more interesting (and also supporting the claim
that the word order freedom of Czech is relatively high)
are the results for sentences with at least two objects. They
are summarized in Table 4. The distribution is even flatter
than in Table 3 with all types being represented (even those
starting with two objects, see the following example) and
none of them exceeding 30%.</p>
        <p>Plán mu v úterý pˇredložil velvyslanec USA v
Chorvatsku Peter Galbraith.</p>
        <sec id="sec-2-1-1">
          <title>Word order type SVOO SOVO SOOV</title>
          <p>VSOO
VOSO
VOOS
OSVO
OSOV
OOSV
OOVS
OVSO
OVOS
Total
The statistics concerning the distribution of word-order
types for English have been calculated on the English
part of the Prague Czech English Dependency Treebank
(PCEDT). This corpus actually contains the same set of
sentences as the Wall Street Journal section of Penn
Treebank,13 (see above for references) but unlike its
predecessor, its syntactic structure has been annotated using
dependency trees. As was mentioned above, the transformation
on the surface syntactic layer was fully automatic, which
has of course affected the quality of annotation.</p>
          <p>The statistics of different types of word order have been
collected in the same manner as in the previous
subsection. We have also applied identical filters as for Czech
sentences from PDT. Table 5 contains data for sentences
with intransitive verbs. Only as few as 40 sentences have
other than verbal predicate.</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>Word order type SV VS Total</title>
          <p>Number
28,236</p>
          <p>900
29,136
As we can see, the strict word order of English
sentences manifests itself in a vast majority of sentences
having the prototypical word order of the subject being
followed by a predicate. The examples of the opposite word
order include sentences containing direct speech with the
following pattern:
"It’s just a matter of time before the tide turns," says one
Midwestern lobbyist.</p>
          <p>Out of the 900 sentences with the reversed word order,
as many as 630 contained the predicate to say, 121 to
be. Each of all other verbs involved in these constructions
13The Czech part had been created as translation of original English
sentences.
were represented less than 10 times. In total, 23 verbs
appear in these sentences at least twice, out of them 16 can
be classified as verbs of communication (verba dicendi)
(in total, it means 678 occurrences out of 822, i.e., 82,5 %
of all occurrences with at least two hits in the corpus).</p>
          <p>The results for sentences containing one object also
strongly confirm the fact that the order Subject - Predicate
- Object (SVO) is practically the only acceptable order in
standard sentences. The remaining types of word order
(representing only 1.06% sentences in the corpus)
mentioned in Table 6 actually represented annotation errors in
a vast majority of cases (esp. auxiliary verbs which have
been quite often incorrectly annotated as Objects).</p>
        </sec>
        <sec id="sec-2-1-3">
          <title>Word order type SVO SOV VSO</title>
          <p>VOS
OVS
OSV
Total
It turns out that for English, it does not make sense to
construct a similar table as Table 4 sentences with more
than one object. The automatic annotation of PCEDT is,
unfortunately, biased in what should be considered an
Object (in the original Penn Treeank annotation, the verbal
complements are labeled just as noun (or prepositional)
phrases (NPs and PPs), no distinction between Objects and
Adverbials.) As a consequence, adverbial constructions
are very often incorrectly annotated as Objects and thus it
is impossible to rely on this distinction (and the analysis
shows that the numbers would be highly misleading).
3.3</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>German</title>
        <p>German has more constraints on word order than Czech
and less than English, therefore it constitutes a very
natural candidate for our experiment. On top of that, there
are also numerous high quality resources which can be
exploited. We have used the German treebank conforming
to the HamleDT initiative, which is located in the Lindat
repository.14</p>
        <p>The statistics for German were collected in the same
way and with the same constraints as Czech and English
ones. The statistics for German sentences with intransitive
predicates are presented in Table 7.</p>
        <p>The almost equal number of sentences with SV and
VS word order types is quite surprising. The fact that
SV represents the typical word order in declarative
sentences, while VS in interrogative ones provides an
obvious explanation. Unfortunately, this explanation does not
14https://lindat.mff.cuni.cz/services/pmltq/hamledt_dt_de/
cover all occurrences because the analyzed corpus
(consisting mostly of newspaper texts) contains only a very
small proportion of interrogative sentences. We have not
investigated the reason for the surprisingly high number of
VS sentences, but it definitely constitutes a very
interesting topic for further research. The same is valid also for
the results contained in Table 8, where we have found
relatively high number of sentences having the word order of
an interrogative sentence, too.
Neither for German we have investigated the sentences
with two or more objects due to annotation
inconsistencies.
The statistics presented in the previous section actually
confirm the well known fact that Czech has the highest
degree of word order freedom from all three languages
investigated in our experiment. This fact is also reflected in
the chart 8 comparing the results for sentences with one
object for all three languages.</p>
        <p>Let us now try to suggest a formula which might allow
to express the degree of word order freedom in a more
precise way. Intuitively, the more free is the word order,
the more equally distributed should be the results of all
six word order types. The more strict the word order, the
more distant are the values from the ideal (equal
distribution). This leads directly to the application of a least
squares method:</p>
        <p>v
M = 16 uut ∑6 (Vi − Av)2,
i=1
(1)
where M is the proposed measure, Vi the percentual
value of the i-th word order type and Av is the average
percentage for each word type (i.e., 100/6). For the three
languages in our experiment we then get the following
values:
• Czech: 6.82
• German: 19.20
• English: 36.79</p>
        <p>These values seem to correspond to the intuitive
feeling that the word order order of English is really strongly
fixed, while German and Czech have more free word order
with Czech having the highest degree of word order
freedom. If we express the results in the form of percentages
of the absolutely fixed word order (i.e., one of the word
order types accounts for 100% and all others do not appear
at all), we’ll get the following results:</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Conclusions</title>
      <p>The experiment described in this paper brought several
interesting results which may be taken as a basis for further
experiments. First of all, it shows that the endeavor to
unify the annotation schemes used for various treebanks in
the HamleDT project provides new opportunities for
linguistic research. The treebank data can now be studied in
a relation to other treebanks using the common search tool
and obtaining results which are not dependent on
peculiarities of individual annotation schemes.</p>
      <p>These new opportunities have been demonstrated on a
small-scale experiment involving three languages (Czech,
German and English). We have managed to extract
quantitative clues confirming the linguistic hypothesis about the
degree of word order freedom of all three languages
under consideration. The main advantage of our approach
is the fact that our research is based on a large number of
sentences of each language and thus it provides a
representative sample of the actual language usage in a given
genre. Contrary to theoretical linguistic research, our
approach does not concentrate upon marginal (but definitely
linguistically interesting) phenomena, but it is based upon
the real language captured in the treebanks.</p>
      <p>In the future we would like to continue the research in
two directions. One will be the obvious endeavor to collect
the statistics for more languages, the second one will be a
more subtle treatment of linguistic phenomena appearing
in treebanks, as, e.g. the investigation including also
subordinated clauses or interrogative sentences.</p>
      <sec id="sec-3-1">
        <title>Grant support</title>
        <p>This paper exploits language data developed and/or
distributed in the frame of the project MŠMT Cˇ R
LINDAT/CLARIN (project LM2010013).</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Saussure</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Course in general linguistics</article-title>
          . Open Court, La Salle,
          <string-name>
            <surname>Illinois</surname>
          </string-name>
          (
          <year>1983</year>
          )
          <article-title>(prepared by C. Bally and A</article-title>
          . Sechehaye, translated by R. Harris)
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Saussure</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Kurs obecné lingvistiky</article-title>
          . Academia,
          <string-name>
            <surname>Praha</surname>
          </string-name>
          (
          <year>1989</year>
          )
          <article-title>(translated by F</article-title>
          .
          <source>Cˇ ermák)</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Sapir</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <string-name>
            <surname>Language</surname>
          </string-name>
          .
          <article-title>An introduction to the study of speech</article-title>
          . Harcourt, Brace and Company, New York (
          <year>1921</year>
          )
          <article-title>(http://www</article-title>
          .gutenberg.org/files/12629/12629-h/ 12629-h.htm).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Skalicˇka</surname>
          </string-name>
          , V.:
          <article-title>Vývoj jazyka</article-title>
          .
          <source>Soubor statí. Státní pedagogické nakladatelství</source>
          ,
          <source>Praha</source>
          (
          <year>1960</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Cˇ ermák</surname>
          </string-name>
          , F.:
          <article-title>Jazyk a jazykoveˇda. Pražská imaginace</article-title>
          ,
          <source>Praha</source>
          (
          <year>1994</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Zeman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dušek</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marecˇek</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Popel</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ramasamy</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Šteˇpánek</surname>
          </string-name>
          , J.,
          <string-name>
            <surname>Žabokrtský</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hajicˇ</surname>
          </string-name>
          , J.: HamleDT:
          <article-title>Harmonized multi-language dependency treebank</article-title>
          .
          <source>Language Resources and Evaluation</source>
          <volume>48</volume>
          (
          <year>2014</year>
          ),
          <fpage>601</fpage>
          -
          <lpage>637</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Hajicˇ</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Panevová</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hajicˇová</surname>
          </string-name>
          , E.,
          <string-name>
            <surname>Sgall</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pajas</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Šteˇpánek</surname>
          </string-name>
          , J.,
          <string-name>
            <surname>Havelka</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mikulová</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Žabokrtský</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ševcˇíková-Razímová</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Prague Dependency Treebank 2.0</article-title>
          . LDC, Philadelphia, PA, USA (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Hajicˇ</surname>
          </string-name>
          , J.,
          <string-name>
            <surname>Hajicˇová</surname>
          </string-name>
          , E.,
          <string-name>
            <surname>Panevová</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sgall</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bojar</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cinková</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fucˇíková</surname>
          </string-name>
          , E.,
          <string-name>
            <surname>Mikulová</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pajas</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Popelka</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Semecký</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Šindlerová</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Šteˇpánek, J.,
          <string-name>
            <surname>Toman</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Urešová</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Žabokrtský</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          : Announcing Prague Czech-English
          <source>Dependency Treebank 2.0. In: Proceedings of the 8th International Conference on Language Resources and Evaluation (LREC</source>
          <year>2012</year>
          ), Istanbul, Turkey,
          <string-name>
            <surname>ELRA</surname>
          </string-name>
          , European Language Resources Association (
          <year>2012</year>
          ),
          <fpage>3153</fpage>
          -
          <lpage>3160</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Brants</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dipper</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eisenberg</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hansen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>König</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lezius</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rohrer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uszkoreit</surname>
          </string-name>
          , H.:
          <article-title>TIGER: Linguistic Interpretation of a German Corpus</article-title>
          .
          <source>Journal of Language and Computation</source>
          (
          <year>2004</year>
          ),
          <fpage>597</fpage>
          -
          <lpage>620</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Pajas</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Šteˇpánek</surname>
          </string-name>
          , J.:
          <article-title>System for querying syntactically annotated corpora</article-title>
          .
          <source>In: Proceedings of the ACL-IJCNLP 2009 Software Demonstrations</source>
          , Suntec, Singapore, Association for Computational Linguistics (
          <year>2009</year>
          ),
          <fpage>33</fpage>
          -
          <lpage>36</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Holan</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kubonˇ</surname>
          </string-name>
          , V.,
          <string-name>
            <surname>Oliva</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Plátek</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>On complexity of word order. Les grammaires de dépendance - Traitement automatique des langues (TAL) 41 (</article-title>
          <year>2000</year>
          )
          <fpage>273</fpage>
          -
          <lpage>300</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Kubonˇ</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lopatková</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Plátek</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>On formalization of word order properties</article-title>
          . In: Gelbukh,
          <string-name>
            <surname>A.</surname>
          </string-name>
          , (ed.), Theoretical Computer Science and General Issues,
          <source>Computational Linguistics and Intelligent Text Processing, CICLing</source>
          <year>2012</year>
          , volume
          <volume>7181</volume>
          of LNCS., Berlin / Heidelberg, SpringerVerlag (
          <year>2012</year>
          )
          <fpage>130</fpage>
          -
          <lpage>141</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Mitchell</surname>
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Marcus</surname>
            , Mary Ann Marcinkiewicz,
            <given-names>B.S.</given-names>
          </string-name>
          :
          <article-title>Building a large annotated corpus of English: the Penn Treebank</article-title>
          .
          <source>Computational Linguistics</source>
          <volume>19</volume>
          (
          <year>1993</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>