<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Linguistic Typology Meets Universal Dependencies</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>William Croft, Dawn Nordquist, Katherine Looney, Michael Regan Department of Linguistics University of New Mexico</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <fpage>63</fpage>
      <lpage>75</lpage>
      <abstract>
        <p>Current work on universal dependency schemes in NLP does not make reference to the extensive typological research on language universals, but could benefit since many principles are shared between the two enterprises. We propose a revision of the syntactic dependencies in the Universal Dependencies scheme (Nivre et al. [16, 17]) based on four principles derived from contemporary typological theory: dependencies should be based primarily on universal construction types over language-specific strategies; syntactic dependency labels should match lexical feature names for the same function; dependencies should be based on the information packaging function of constructions, not lexical semantic types; and dependencies should keep distinct the “ranks” of the functional dependency tree.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <sec id="sec-1-1">
        <title>A number of parsing and tagging schemes have recently been proposed that aim</title>
        <p>
          to be universal across languages, including the Universal Stanford Dependencies
(USD; de Marneffe et al. [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]) and the Google Universal Part-of-Speech Tagset
(Petrov et al. [
          <xref ref-type="bibr" rid="ref19">20</xref>
          ]). More recently, the Universal Dependencies (UD) initiative
(Nivre et al. [
          <xref ref-type="bibr" rid="ref15 ref16">16, 17</xref>
          ]) has brought together a slightly altered form of the
Universal Stanford Dependencies, an extended version of the Google Universal
Part-of
        </p>
      </sec>
      <sec id="sec-1-2">
        <title>Speech Tagset, and a revised subset of the Interset morphological features (Zeman</title>
        <p>
          [
          <xref ref-type="bibr" rid="ref30">31</xref>
          ]).
        </p>
        <p>
          Nivre [
          <xref ref-type="bibr" rid="ref15">16</xref>
          ] writes, ‘Traditionally, research in our [NLP] community has not
paid much attention to language typology or linguistic universals’. Unfortunately,
none of these proposals make explicit reference to the extensive typological
literature on universals based on large-scale, balanced crosslinguistic samples
(although this does not mean that typology has not influenced these proposals). As
a consequence, these proposed universal schemes offer a mixture of genuinely
universal annotations from a typological perspective, and some annotations that
are language-specific, even if they are specific to multiple languages. Universals
projects in NLP would benefit from taking a more typological perspective (cf.
Bender [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]). First, a universal scheme that better reflects typological patterns across all
of the world’s languages should be more robust and extendable to other languages
for which digital resources may be developed. Second, the Universal
Dependencies initiative is actually quite harmonious with many elements of recent
typological theory. In fact, many changes in UD version 2 independently converge with
proposals in this paper. 1
        </p>
        <p>In the following sections, we present the four most important principles behind
a typological approach to syntactic analysis. From these principles, we develop an
annotation scheme for dependencies that reflects a fuller range of crosslinguistic
variation and the universals governing the variation, but—we hope—without losing
the practical values embodied in the UD initiative. This scheme is currently used
in teaching typological syntax to undergraduates. The pedagogical and theoretical
goals can be compared to the principles guiding the UD project:2
1. UD needs to be satisfactory on linguistic analysis grounds for individual
languages.
2. UD needs to be good for linguistic typology, i.e., providing a suitable basis
for bringing out cross-linguistic parallelism across languages and language
families.
3. UD must be suitable for rapid, consistent annotation by a human annotator.
4. UD must be suitable for computer parsing with high accuracy.</p>
      </sec>
      <sec id="sec-1-3">
        <title>5. UD must be easily comprehended and used by a non-linguist, whether a language learner or an engineer with prosaic needs for language processing. We refer to this as seeking a habitable design, and it leads us to favor traditional grammar notions and terminology.</title>
      </sec>
      <sec id="sec-1-4">
        <title>6. UD must support well downstream language understanding tasks (relation extraction, reading comprehension, machine translation, . . . ).</title>
      </sec>
      <sec id="sec-1-5">
        <title>Theoretical and pedagogical goals match with all the UD goals, except 4 and</title>
      </sec>
      <sec id="sec-1-6">
        <title>6, which do not apply. We recognize that some revisions to UD in this paper</title>
        <p>proposed for theoretical and pedagogical goals may conflict with UD goals 4 and</p>
      </sec>
      <sec id="sec-1-7">
        <title>6. In this paper we restrict our attention to typologically motivated revisions to the</title>
      </sec>
      <sec id="sec-1-8">
        <title>UD syntactic dependencies; we do not address multiword expressions, fixes to deal with headless constructions in dependency syntax, or lexical tagging.</title>
        <p>2</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Constructions and strategies</title>
      <p>
        One needs to develop valid comparative concepts in crosslinguistic comparison
(Haspelmath [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]). Haspelmath is following a long tradition in typology
(Greenberg [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], Keenan and Comrie [13], Stassen [
        <xref ref-type="bibr" rid="ref24">25</xref>
        ], Croft [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]). These authors all argue
1The first author thanks Joakim Nivre, Chris Manning and Marie-Catherine de Marneffe for the
opportunity to discuss UD v2 before its release.
      </p>
      <p>2http://www.universaldependencies.org/introduction
in support of defining comparative concepts in semantic/functional terms.
Haspelmath, however, also argues for the need for “hybrid” comparative concepts defined
in both semantic and formal (morphosyntactic) terms, as long as the formal criteria
are crosslinguistically valid.</p>
      <p>
        Two types of “hybrid” comparative concepts are necessary for crosslinguistic
comparison, and in fact have been used in typological research since its
inception (Croft [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ]). The first type of comparative concept, constructions, are any
morphosyntactic structure in any language used to express a particular meaning or
function. For example, a predicate nominal construction is any construction used
to predicate the object category of a referent, as in Ivan is the best dancer. The
second type of comparative concept, strategies, are constructions that use specific
morphosyntactic devices, that is, specific formal structures, where those formal
structures are defined in a crosslinguistically valid way. For example, the English
predicate nominal construction uses a copula strategy, where ‘copula’ is defined
in crosslinguistically valid terms as a (usually free) morpheme, distinct from the
object concept word, that is part of the predication. Constructions are universal, in
fact universal by definition, since human languages are general-purpose
communication systems. Strategies are language-specific, or more precisely, not necessarily
universal; English uses a copula strategy in its predicate nominal construction, but
many languages do not. This leads to the first principle for designing a
typologically universal annotation scheme: a universal annotation scheme should have
a classification of constructions as its universal foundational layer; it should
avoid wherever possible annotations dependent on strategies which are not
universal.
      </p>
      <p>
        UD does the right thing typologically in its aim to ‘push all the way the
design principle of having direct links between content words’ (de Marneffe et al.
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]). This is because the use of an independent syntactic unit such as a copula or
adposition is not a strategy found in every language. For this reason, the
topology of the dependency trees in our proposal largely matches that of UD; the
primary differences are in the classification of the dependencies (see Figure 1). Of
course, elements in a universal scheme that represent strategies, at least the most
commonly occurring strategies, will also be needed. The most common strategies
found as function words include case (as in UD), classifiers (clf, added to UD
v2), independent indexation/agreement elements (idx) and linkers (lnk); these are
given crosslinguistically valid definitions in Croft [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]). Coordinating and
subordinating conjunctions (cc and mark, following UD) are also strategies, as opposed to
asyndetic coordination and deranking (a morphological strategy; see Stassen [
        <xref ref-type="bibr" rid="ref24">25</xref>
        ]).
3
      </p>
      <p>Adpositions and case features are different strategies for the same
construction (in the typological sense), namely the construction that relates an argument
dependent to its head. A universal scheme should capture the fact that adpositions
and case morphology are two strategies for the same construction. This leads to a
3Not all words categorized as function words are strategies, e.g. pronouns and determiners.
reference
object the sharp thorns
property sharpness
modification
the bush’s thorns
the sharp thorns
action</p>
      <sec id="sec-2-1">
        <title>I said that the thorns the thorns that scratched</title>
        <p>scratched me. me
the scratching of the the thorns scratching me
thorns
predication
It’s a thorn</p>
      </sec>
      <sec id="sec-2-2">
        <title>Those thorns are</title>
        <p>sharp</p>
      </sec>
      <sec id="sec-2-3">
        <title>The sharp thorns</title>
        <p>scratched me
second design principle for typologically universal annotation: use the same term
for morphological and syntactic strategies for the same function, across
dependency labeling and lexical tagging (morphological features and POS). UD uses the
same label, case, for the syntactic strategy (dependency of a function word) and
the morphological strategy (feature of a word form), which conforms to the second
principle. But more generally, UD, like traditional grammar, uses different terms
for the independent function word dependency and the corresponding
morphological category, and sometimes yet another term for the POS tag for the function word.</p>
        <sec id="sec-2-3-1">
          <title>Here UD principles 2 and 5 are in conflict; we opt for 2 while UD opts for 5.</title>
          <p>3</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Semantics and information packaging</title>
      <p>
        The crosslinguistic analysis of parts of speech is a long-vexed issue in typology
as well as other grammatical theories. The solution is to recognize that parts of
speech are best defined by the intersection of lexical semantic categories—object,
action, property—and information packaging functions—in the case of major parts
of speech, the propositional act functions of reference, predication, and
modification (Croft [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ]). The primary argument for this analysis is that, in principle at
least, any semantic class may be expressed through any propositional act, albeit
often with a distinct morphosyntactic strategy. This fact is illustrated with the English
examples in Table 1.
      </p>
      <p>
        We argue that this split-level analysis of meaning into semantic content and
information packaging applies to all levels of sentence meaning. Predicates have
multiple arguments. They are generally packaged into a ranking that represents
discourse salience, essentially subject &lt; object &lt; oblique. The ranking has a
preferred realization of semantic roles, with agents as subjects, patients or themes as
objects, and other roles as obliques. However, almost any semantic role can be
expressed as either subject, object or oblique, illustrated with English examples in
Table 2. The less common realizations of semantic roles may be realized
grammatically by voice, including verb alternations in English and applicative constructions
(Peterson [
        <xref ref-type="bibr" rid="ref18">19</xref>
        ]) in other languages. Although agent-like participants do not receive
object-like marking in English, they do in the inverse voice construction in
Alagent
theme
goal
subject
The protesters
sprayed green paint
on the sidewalk.
      </p>
      <p>Green paint was
sprayed on the
sidewalk.</p>
      <p>The sidewalk was
sprayed with green
paint.</p>
      <p>core
object
—</p>
      <sec id="sec-3-1">
        <title>The protesters sprayed</title>
        <p>green paint on the
sidewalk.</p>
      </sec>
      <sec id="sec-3-2">
        <title>The protesters sprayed</title>
        <p>the sidewalk with
green paint.
oblique</p>
      </sec>
      <sec id="sec-3-3">
        <title>The sidewalk was</title>
        <p>sprayed with green
paint by the protesters.</p>
      </sec>
      <sec id="sec-3-4">
        <title>The protesters sprayed</title>
        <p>the sidewalk with green
paint.</p>
      </sec>
      <sec id="sec-3-5">
        <title>The protesters sprayed</title>
        <p>green paint on the
sidewalk.
gonkian and other language families, and in the voice systems of the extensive</p>
        <sec id="sec-3-5-1">
          <title>Austronesian family, e.g. Tagalog.</title>
        </sec>
        <sec id="sec-3-5-2">
          <title>Coordination and adverbial subordination represent two different ways of pack</title>
          <p>aging temporal, causal, conditional, concessive and other semantic relations
between two events. Any of these semantic relations can be expressed in English by
either adverbial subordination or by coordination; see Table 3. Expressing a
relation between events in terms of coordination packages them as a single complex</p>
        </sec>
        <sec id="sec-3-5-3">
          <title>Gestalt, where the two events are construed in a symmetrical fashion (Wierzbicka</title>
          <p>
            [
            <xref ref-type="bibr" rid="ref29">30</xref>
            ]). Expressing the same relation between events in terms of adverbial
subordination packages them in an asymmetrical relation, usually described as figure
(main clause) vs. ground (adverbial clause; Talmy [
            <xref ref-type="bibr" rid="ref25">26</xref>
            ], Reinhart [
            <xref ref-type="bibr" rid="ref20">21</xref>
            ]).
          </p>
          <p>
            What does this mean in practical terms? The information packaging
functions are generally much more isomorphic to syntactic structures than lexical
semantic classes: The information packaging functions are also less variable across
languages—i.e. more universal—than lexical semantics. The dissociation of
lexical semantics and information packaging leads to a third design principle: the
dependencies should primarily express information packaging structure;
semantic content should be primarily annotated via lexical tagsets. This design principle
will lead mostly to a simplification of the UD dependency set. First, det, nummod
and amod would be grouped into mod, with their differences captured by lexical
semantic tags. UD v2 does not merge these, because some users use only the
dependencies while others use only the lexical tags; hence retaining redundancy is
preferred. Second, dobj and iobj would be grouped into obj. UD v2 keeps the two
separate, though dobj is renamed obj. Double object constructions generally have
the two objects in strict order; although some argue that only one of the two objects
is a “true” object, in fact there is no crosslinguistic consistency in the syntactic
behavior of the two objects (Peterson [
            <xref ref-type="bibr" rid="ref18">19</xref>
            ]). Finally, nsubj and nsubjpass would be
grouped into sbj, with semantic roles identified by voice or other semantic tags on
the verb. If special dependencies are created for passive voice, then one would have
to devise special dependencies for applicative, applicative+passive, and other
combinations of valency-changing operations. For this reason, UD v2 has eliminated
Semantic relation
          </p>
        </sec>
        <sec id="sec-3-5-4">
          <title>Anterior</title>
        </sec>
        <sec id="sec-3-5-5">
          <title>Posterior</title>
        </sec>
        <sec id="sec-3-5-6">
          <title>Overlap</title>
        </sec>
        <sec id="sec-3-5-7">
          <title>Cause</title>
        </sec>
        <sec id="sec-3-5-8">
          <title>Purpose</title>
        </sec>
        <sec id="sec-3-5-9">
          <title>Apprehensional</title>
        </sec>
        <sec id="sec-3-5-10">
          <title>Means/Positive Circumstantial</title>
        </sec>
        <sec id="sec-3-5-11">
          <title>Negative Circumstantial</title>
        </sec>
        <sec id="sec-3-5-12">
          <title>Additive</title>
        </sec>
        <sec id="sec-3-5-13">
          <title>Substitutive</title>
        </sec>
        <sec id="sec-3-5-14">
          <title>Subtractive</title>
        </sec>
        <sec id="sec-3-5-15">
          <title>Conditional</title>
        </sec>
        <sec id="sec-3-5-16">
          <title>Concessive</title>
          <p>Subordination construction</p>
        </sec>
      </sec>
      <sec id="sec-3-6">
        <title>He washed the car before driv</title>
        <p>ing to the party.</p>
      </sec>
      <sec id="sec-3-7">
        <title>He drove to the party after</title>
        <p>washing the car.</p>
      </sec>
      <sec id="sec-3-8">
        <title>He washed the car while the</title>
        <p>sun was still shining.</p>
      </sec>
      <sec id="sec-3-9">
        <title>She went to bed because she</title>
        <p>was exhausted.</p>
      </sec>
      <sec id="sec-3-10">
        <title>I will grab a stick to defend my</title>
        <p>self.</p>
      </sec>
      <sec id="sec-3-11">
        <title>I grabbed a stick lest he attack</title>
        <p>me.</p>
      </sec>
      <sec id="sec-3-12">
        <title>He got into the army by lying</title>
        <p>about his age.</p>
      </sec>
      <sec id="sec-3-13">
        <title>She carried the punch into the</title>
        <p>living room without spilling a
drop.</p>
        <p>In addition to having your
hand stamped, you must show
your ticket stub.</p>
      </sec>
      <sec id="sec-3-14">
        <title>We barbecued chicken at home</title>
        <p>instead of going out to eat.</p>
      </sec>
      <sec id="sec-3-15">
        <title>He did all the problems cor</title>
        <p>rectly except he missed the
proof on the last one.</p>
      </sec>
      <sec id="sec-3-16">
        <title>If you do that, the terrorists</title>
        <p>have won.</p>
      </sec>
      <sec id="sec-3-17">
        <title>Although John had no money,</title>
        <p>he went into this expensive
restaurant.</p>
        <p>Coordination construction</p>
      </sec>
      <sec id="sec-3-18">
        <title>He washed the car and</title>
        <p>drove to the party.</p>
      </sec>
      <sec id="sec-3-19">
        <title>He washed the car and</title>
        <p>drove to the party.</p>
      </sec>
      <sec id="sec-3-20">
        <title>The sun was shining and he</title>
        <p>was washing the car.</p>
      </sec>
      <sec id="sec-3-21">
        <title>She was exhausted and (so)</title>
        <p>went to bed.</p>
      </sec>
      <sec id="sec-3-22">
        <title>I will grab a stick and de</title>
        <p>fend myself.</p>
      </sec>
      <sec id="sec-3-23">
        <title>Grab a stick or he will attack</title>
        <p>you.</p>
      </sec>
      <sec id="sec-3-24">
        <title>He lied about his age and</title>
        <p>got into the army.</p>
      </sec>
      <sec id="sec-3-25">
        <title>She carried the punch into</title>
        <p>the living room, and she
didn’t spill a drop.</p>
      </sec>
      <sec id="sec-3-26">
        <title>You have to have your hand</title>
        <p>stamped and show your
ticket stub.</p>
      </sec>
      <sec id="sec-3-27">
        <title>We didn’t go out to eat,</title>
        <p>and/but barbecued chicken
at home.</p>
      </sec>
      <sec id="sec-3-28">
        <title>He did all the problems cor</title>
        <p>rectly but he missed the
proof on the last one.</p>
      </sec>
      <sec id="sec-3-29">
        <title>Murphy, you do that and the</title>
        <p>terrorists have won, ...</p>
      </sec>
      <sec id="sec-3-30">
        <title>John had no money, but</title>
        <p>he went into this expensive
restaurant (anyway).
Info packaging rank</p>
        <sec id="sec-3-30-1">
          <title>Complex sentences</title>
        </sec>
        <sec id="sec-3-30-2">
          <title>Predicate complex</title>
        </sec>
        <sec id="sec-3-30-3">
          <title>Arguments</title>
        </sec>
        <sec id="sec-3-30-4">
          <title>Modifiers</title>
        </sec>
        <sec id="sec-3-30-5">
          <title>Admodifiers</title>
        </sec>
        <sec id="sec-3-30-6">
          <title>Common strategies</title>
          <p>Proposed dependencies
conj, advcl
cxp, sec [[PROPERTY]], aux [[TAMP]]
sbj, obj, obl, comp [[EVENT]]
mod, nmod [[ENTITY]], acl [[EVENT]], appos
qlfy
cc, mark, case, clf, idx, lnk
the special dependencies for passive voice, retaining the label nsubj.
4</p>
          <p>“Ranks” of syntactic structure</p>
        </sec>
        <sec id="sec-3-30-7">
          <title>Finally, we propose a fourth design principle for the categorization or labeling of</title>
          <p>dependencies: dependencies should be unique to the “rank” of dependency
structure. The “ranks” of dependency structure are: predicates, which are chained
together in complex sentences (conj, advcl, as in UD) but also form complex
predicates (cxp); arguments, divided into subject, object and oblique (sbj, obj,
obl); modifiers; and admodifiers (modifiers of modifiers; qlfy); see Table 4.</p>
          <p>
            One of these “ranks”, complex predicates, is not found in Western traditional
grammar. However, complex predicates occur widely in English and other
languages, and are diverse in form and function. Certain elements of complex
predicates are analyzed as compounds in UD v2, including serial verbs, verb-particle
constructions, and more grammaticalized light verb constructions. However, these
predicate complexes are syntactically flexible, unlike nominal compounds (door
handle; sky blue) which are semi-fixed in the terminology of Sag et al. [
            <xref ref-type="bibr" rid="ref21">22</xref>
            ]. We
distinguish complex predicate structure with a dependency distinct from nominal
compounds, cxp. UD v2 retains a distinct cop dependency for copulas, albeit more
narrowly defined than in v1. We follow FrameNet in analyzing copulas as instances
of support verbs (light verbs) and therefore eliminate cop.
          </p>
          <p>
            We distinguish two special classes of complex predicate elements. The first are
property concepts expressed as secondary predication (secondary predicates
(depictives and resultatives) and manner “adverbs”. UD treats all three differently:
depictives as adverbial clauses (acl), resultatives as controlled complements (xcomp),
and manner adverbs as adverbs (advmod). Although English and other Western
European languages use a different strategy for manner “adverbs” than for depictives
and resultatives, typologically the three property concept constructions share the
same range of strategies and indeed can share the same strategy within a language
(Loeb-Diehl [
            <xref ref-type="bibr" rid="ref13">14</xref>
            ], van der Auwera and Malchukov [
            <xref ref-type="bibr" rid="ref26">27</xref>
            ], Verkerk [
            <xref ref-type="bibr" rid="ref27 ref28">28, 29</xref>
            ]). We
distinguish this element of a predicate complex as sec. Finally, forms expressing
tense, aspect, mood, modality, evidentiality and polarity are called aux. UD v1
used the aux dependency for such elements when they are verblike, which is only
one strategy for expressing these categories, and distinguished negative polarity
(neg). UD v2 extends aux to uninflected elements expressing such meanings and
eliminates the neg dependency, but uses advmod for aspectual, modal etc. adverbs.
          </p>
        </sec>
        <sec id="sec-3-30-8">
          <title>Admodifiers are “adverbs” that qualify modifiers, such as degree and hedging.</title>
          <p>In our proposed revisions, there is no advmod dependency. Manner adverbs are
subsumed under secondary predicates (sec); adverbs expressing aspect, modality
and so on are subsumed under aux; adverbs expressing spatial or temporal location
(here, yesterday) are analyzed as pronominal oblique dependents (obl); and
adverbs functioning as admodifiers, which generally express degree or hedging, are
assigned to the qlfy dependency. All of these types of “adverbs” are diverse both
in terms of syntactic distribution, morphological form and semantics. A single
“adverb” dependency is therefore not a coherent category in grammatical or semantic
terms.</p>
        </sec>
        <sec id="sec-3-30-9">
          <title>From a pure dependency syntax perspective, the “ranks” of complex sentences,</title>
          <p>complex predicates and arguments are not differentiated because all of them are
dependents of the predicate. Linguistically, however, they are very different:
adverbial clauses are an asymmetric packaging of information otherwise expressed
symmetrically in coordination; complex predicates are complex expressions
describing events and their semantic properties; arguments are participants in events.</p>
        </sec>
        <sec id="sec-3-30-10">
          <title>Although we retain the dependency structure for these three “ranks” (all dependent</title>
          <p>on the predicate), we capture these distinctions by using distinct dependencies, as
seen in Table 4.</p>
          <p>The fourth design principle makes dependency structure most closely match
information packaging structure. However, it collides with another fact—the fact
that motivates the separation of lexical semantics from information packaging in
the first place: in principle, any semantic category can be packaged in any way. To
illustrate the practical problem: an event may function as an argument of a
predicate (1); but then the event’s own participant entities (John, the request) may be
expressed as “arguments” of the event “predicate”. Of course, languages allow
these packagings of content, but they also use a range of strategies to express them.
In the example of complements, strategies range from predicate-like
tense-aspectmood inflection and argument-like subject/object/oblique coding (1), to
argumentlike nominalized event forms and possessive modifier-like coding of the participant
entities (2), to—even more problematic from a practical point of view—the
mixing in a single construction of strategies from prototypical predicate-argument and
argument-modifier constructions (for example, English gerunds; 3).
(1) She believes that John approved the request.
(2) She regrets John’s approval of the request.
(3) She regrets John’s approving the request.</p>
          <p>The solution adopted here for this practical problem is twofold. First, isolate
the most common cases of this complexity and give them their own distinct
dependency label, based on their semantic categories: events as arguments are
complements (comp), events as modifiers are adjectival (relative) clauses (acl), entities
as modifiers (possessives, etc.) are nominal modifiers (nmod), as in UD. 4
Second, allow for recursion of dependencies with these distinct dependencies. Hence
complements and relative clauses may have arguments that are dependents of the
subordinate clause predicate, and noun modifiers may themselves have modifiers
that are dependents of the noun modifier. These exceptions are justified because
they are the cases that most frequently—but by no means always—have complex
syntax “imported” from the “native” position of the semantic category in the
dependency tree. This is basically a standard analysis of recursive structures, but limited
to just these dependencies, and with the parallels to other strategies highlighted
(we would analyze 3 and even 2 as comp).</p>
        </sec>
        <sec id="sec-3-30-11">
          <title>The final scheme for syntactic dependencies that conform to the four typo</title>
          <p>
            logical principles is given in Table 4. UD dependencies left unchanged are: the
root (root); multiword expressions (fixed, flat, compound in UD v2; cf. Sag et
al. [
            <xref ref-type="bibr" rid="ref21">22</xref>
            ]); discourse (vocative, discourse), loose joining (dislocated, list, parataxis,
reparandum); elliptical elements (orphan); unspecified (dep); punctuation (punct)
and other special cases (reparandum, goeswith). We are unsure about dividing
          </p>
        </sec>
        <sec id="sec-3-30-12">
          <title>Sag et al.’s [22] semi-fixed expressions into exocentric (flat) and endocentric (compound) types, and we exclude complex predicate elements from compound. The discourse and loose joining dependencies may be revised, but at present insufficient typological research has been done on them.</title>
          <p>
            These four principles guided the proposed revisions to UD in this paper, but
there are other ways in which typology can provide input to a universal syntactic
annotation scheme. For example, sentence coordination is unheaded, and various
analyses have been proposed (Zeman et al. [
            <xref ref-type="bibr" rid="ref31">32</xref>
            ]). UD v1 attached following
conjuncts to the first conjunct. However, two typological phenomena suggest that a
chaining analysis is preferable. The first is the class of switch-reference
constructions (Haiman and Munro [
            <xref ref-type="bibr" rid="ref11">11</xref>
            ]), in which verb forms are selected based on the
coreference or lack thereof between the conjunct subject and the subject of the
preceding or following conjunct. The second is the universal phenomenon of tense
iconicity (Haiman [
            <xref ref-type="bibr" rid="ref10">10</xref>
            ]): the sequence of conjoined clauses mirrors the sequence
of events. UD v2 adopts the chaining analysis.
5
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Applications: examples and annotation task using the proposed scheme</title>
      <p>4UD divides complements into subject (scomp), obligatory control (xcomp) and nonobligatory
control (ccomp) complement types. We do not distinguish complements by grammatical role, in
order to avoid proliferation of complement types. Control properties are lexically or pragmatically
determined; if lexical, then by the third principle they would be lexically tagged. We recognize that
retaining the distinction between ccomp and xcomp may be useful for UD goal 6.
root</p>
      <p>compound
nsubj
advmod
obj
det
det
root</p>
      <p>tamp
sbj
sec
obj
mod
mod
te1 ¼jaN1 thO1 di:u2 za:k7 na:i6 taN1 te1 ¼jaN1 thO1 di:u2 za:k7 na:i6 taN1</p>
      <sec id="sec-4-1">
        <title>Maonan: UD v2</title>
      </sec>
      <sec id="sec-4-2">
        <title>Maonan: Typological UD</title>
      </sec>
      <sec id="sec-4-3">
        <title>We performed an annotation task using the proposed scheme. Two annotators,</title>
        <p>
          one an instructor of the syntax class, the other an undergraduate who had not used
the annotation scheme previously, annotated passages from six languages (Nguna
[
          <xref ref-type="bibr" rid="ref22">23</xref>
          ], Chantyal [
          <xref ref-type="bibr" rid="ref17">18</xref>
          ], Arapaho [text formerly available at the University of Colorado
        </p>
      </sec>
      <sec id="sec-4-4">
        <title>Arapaho website], Maonan [15], Supyire [2] and Mapuche [24]), already divided</title>
        <p>into 10-20 single sentences. The annotators used a hybrid annotation scheme which
includes constituents for modifier phrases, argument phrases, complex predidates
(which may be discontinuous) and clauses, as well as labeling of heads and
dependent types. The hybrid scheme has proven to be an effective pedagogical tool.
The hybrid annotation can be translated into a pure dependency tree (we have not
yet automated the translation). However, the hybrid annotation also allows for
the annotation of headless and exocentric constructions, where a pure dependency
analysis requires conventional assignment of an arbitrary head, and rather complex
rules for doing so, as seen in the UD guidelines. Scoring of interannotator
agreement is done by an alignment algorithm; this scoring provides an intuitively natural
weighting of the different types of inconsistencies or errors.</p>
      </sec>
      <sec id="sec-4-5">
        <title>The results of the annotation task are given in Table 5. Interannotator agreement is measured using Cohen’s kappa. Unsurprisingly, higher agreement is found for languages typologically similar to English; Arapaho is high because some sentences are single words in a polysynthetic language.</title>
        <p>[13] Edward L. Keenan and Bernard Comrie. Noun phrase accessibility and
universal grammar. Linguistic Inquiry, 8:63–99, 1977.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Emily</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Bender</surname>
          </string-name>
          .
          <article-title>Linguistically naïve != language independent: why NLP needs linguistic typology</article-title>
          .
          <source>In Proceedings of the EACL 2009 Workshop on the Interaction between Linguistics and Computational Linguistics</source>
          , pages
          <fpage>26</fpage>
          -
          <lpage>32</lpage>
          . Association for Computational Linguistics,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Robert</given-names>
            <surname>Joel Carlson</surname>
          </string-name>
          .
          <source>A Grammar of Supyire</source>
          . Berlin: Mouton de Gruyter,
          <year>1994</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>William</given-names>
            <surname>Croft</surname>
          </string-name>
          .
          <source>Syntactic Categories and Grammatical Relations: The Cognitive Organization of Information</source>
          . Chicago: University of Chicago Press,
          <year>1991</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>William</given-names>
            <surname>Croft</surname>
          </string-name>
          .
          <source>Radical Construction Grammar: Syntactic Theory in Typological Perspective</source>
          . Oxford: Oxford University Press,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>William</given-names>
            <surname>Croft</surname>
          </string-name>
          .
          <source>Typology and Universals</source>
          . Cambridge: Cambridge University Press, 2nd edition,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>William</given-names>
            <surname>Croft</surname>
          </string-name>
          .
          <article-title>Comparing categories and constructions crosslinguistically (again): the diversity of ditransitives (review article on Studies in Ditransitive Constructions: A Comparative Handbook, edited by Andrej Malchukov, Martin Haspelmath and Bernard Comrie)</article-title>
          .
          <source>Linguistic Typology</source>
          ,
          <volume>18</volume>
          :
          <fpage>533</fpage>
          -
          <lpage>51</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>William</given-names>
            <surname>Croft</surname>
          </string-name>
          .
          <article-title>Comparative concepts and language-specific categories: theory and practice</article-title>
          .
          <source>Linguistic Typology</source>
          ,
          <volume>20</volume>
          :
          <fpage>377</fpage>
          -
          <lpage>393</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Marie-Catherine de Marneffe</surname>
            , Timothy Dozat, Natalia Silveira, Katri Haverinen, Filip Ginter, Joakim Nivre, and
            <given-names>Christopher D.</given-names>
          </string-name>
          <string-name>
            <surname>Manning</surname>
          </string-name>
          . Universal Stanford Dependencies:
          <article-title>a cross-linguistic typology</article-title>
          .
          <source>In Proceedings of LREC 2014</source>
          , pages
          <fpage>4585</fpage>
          -
          <lpage>4592</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Joseph</surname>
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Greenberg</surname>
          </string-name>
          .
          <article-title>Some universals of grammar with particular reference to the order of meaningful elements</article-title>
          . In Joseph H. Greenberg, editor,
          <source>Universals of grammar</source>
          , pages
          <fpage>73</fpage>
          -
          <lpage>113</lpage>
          . Cambridge, Mass.: MIT Press,
          <source>2nd edition</source>
          ,
          <year>1966</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>John</given-names>
            <surname>Haiman</surname>
          </string-name>
          .
          <source>Natural Syntax: Iconicity and Erosion</source>
          . Cambridge: Cambridge University Press,
          <year>1985</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>John</given-names>
            <surname>Haiman</surname>
          </string-name>
          and Pamela Munro, editors.
          <source>Switch-Reference and Universal Grammar</source>
          . Amsterdam: John Benjamins,
          <year>1983</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Martin</given-names>
            <surname>Haspelmath</surname>
          </string-name>
          .
          <article-title>Comparative concepts and descriptive categories in crosslinguistic studies</article-title>
          .
          <source>Language</source>
          ,
          <volume>86</volume>
          :
          <fpage>663</fpage>
          -
          <lpage>87</lpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Flora</given-names>
            <surname>Loeb-Diehl</surname>
          </string-name>
          .
          <article-title>The typology of manner expressions</article-title>
          .
          <source>PhD thesis</source>
          , Radboud Universiteit Nijmegen,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Tian</surname>
            <given-names>Qiao</given-names>
          </string-name>
          <string-name>
            <surname>Lu</surname>
          </string-name>
          .
          <article-title>A Grammar of Maonan</article-title>
          . Boca Raton, Fl.: Universal Publishers,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Joakim</given-names>
            <surname>Nivre</surname>
          </string-name>
          .
          <article-title>Towards a universal grammar for natural language processing</article-title>
          . In Alexander Gelbukh, editor,
          <source>Computational Linguistics and Intelligent Text Processing</source>
          , pages
          <fpage>3</fpage>
          -
          <lpage>16</lpage>
          . New York: Springer,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Joakim</surname>
            <given-names>Nivre</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marie-Catherine de Marneffe</surname>
          </string-name>
          , Filip Ginter, Yoav Goldberg,
          <string-name>
            <surname>Jan</surname>
            <given-names>Hajicˇ</given-names>
          </string-name>
          , Christopher D. Manning,
          <string-name>
            <surname>Ryan</surname>
            <given-names>McDonald</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Slav</given-names>
            <surname>Petrov</surname>
          </string-name>
          , Sampo Pyysalo, Natalia Silveira, Reut Tsarfaty, and Daniel Zeman.
          <article-title>Universal Dependencies v1: a multilingual treebank collection</article-title>
          .
          <source>In Proceedings of the 10th International Conference on Language Resources and Evaluation</source>
          , pages
          <fpage>1659</fpage>
          -
          <lpage>1666</lpage>
          . European Language Resources Association,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Michael</given-names>
            <surname>Noonan</surname>
          </string-name>
          . Chantyal texts.
          <source>Himalayan Linguistics Archive</source>
          ,
          <volume>2</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>254</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [19]
          <string-name>
            <surname>David</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Peterson</surname>
          </string-name>
          .
          <source>Applicative Constructions</source>
          . Oxford: Oxford University Press,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Slav</surname>
            <given-names>Petrov</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dipanjan Das</surname>
          </string-name>
          , and
          <string-name>
            <surname>Ryan McDonald</surname>
          </string-name>
          .
          <article-title>A universal part-of-speech tagset</article-title>
          .
          <source>In Proceedings of LREC 2012</source>
          , pages
          <fpage>2089</fpage>
          -
          <lpage>2096</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>Tanya</given-names>
            <surname>Reinhart</surname>
          </string-name>
          .
          <article-title>Principles of gestalt perception in the temporal organization of narrative texts</article-title>
          . Linguistics,
          <volume>22</volume>
          :
          <fpage>779</fpage>
          -
          <lpage>809</lpage>
          ,
          <year>1984</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Ivan</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Sag</surname>
            , Timothy Baldwin, Francis Bond, Ann Copestake, and
            <given-names>Dan</given-names>
          </string-name>
          <string-name>
            <surname>Flickinger</surname>
          </string-name>
          .
          <article-title>Multiword expressions: a pain in the neck for NLP</article-title>
          . In Alexander Gelbukh, editor,
          <source>Computational Linguistics and Intelligent Text Processing: Third International Conference, CICLing</source>
          <year>2002</year>
          ,
          <string-name>
            <given-names>Mexico</given-names>
            <surname>City</surname>
          </string-name>
          , Mexico,
          <source>February 17-23</source>
          ,
          <year>2002</year>
          : Proceedings, pages
          <fpage>1</fpage>
          -
          <lpage>15</lpage>
          . New York: Springer,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>Alfred</given-names>
            <surname>Schütz</surname>
          </string-name>
          . Nguna Texts. Honolulu: University of Hawaii Press,
          <year>1969</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>Ineke</given-names>
            <surname>Smeets</surname>
          </string-name>
          .
          <source>A Grammar of Mapuche</source>
          . Berlin: Mouton de Gruyter,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>Leon</given-names>
            <surname>Stassen</surname>
          </string-name>
          .
          <source>Comparison and Universal Grammar. Oxford: Basil Blackwell</source>
          ,
          <year>1985</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>Leonard</given-names>
            <surname>Talmy</surname>
          </string-name>
          .
          <article-title>Figure and ground in complex sentences</article-title>
          . In Joseph H.
          <string-name>
            <surname>Greenberg</surname>
          </string-name>
          et al., editor,
          <source>Universals of Human Language</source>
          ,
          <volume>4</volume>
          : Syntax, pages
          <fpage>625</fpage>
          -
          <lpage>652</lpage>
          . Stanford: Stanford University Press,
          <year>1978</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Johan</surname>
            <given-names>van der Auwera and Andrej L.</given-names>
          </string-name>
          <string-name>
            <surname>Malchukov</surname>
          </string-name>
          .
          <article-title>A semantic map for depictive adjectivals</article-title>
          . In Nikolaus P. Himmelmann and Eva Schultze-Berndt, editors,
          <source>Secondary Predication and Adverbial Modification: The Typology of Depictives</source>
          , pages
          <fpage>393</fpage>
          -
          <lpage>421</lpage>
          . Oxford: Oxford University Press,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>Annemarie</given-names>
            <surname>Verkerk</surname>
          </string-name>
          .
          <article-title>Secondary predication in a typological context: the encoding of resultatives, depictives and manner predications and their placement within a conceptual space</article-title>
          .
          <source>Master's thesis</source>
          , Radboud Universiteit Nijmegen,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>Annemarie</given-names>
            <surname>Verkerk</surname>
          </string-name>
          .
          <article-title>A semantic map of secondary predication</article-title>
          .
          <source>In Linguistics in the Netherlands</source>
          <year>2009</year>
          , pages
          <fpage>115</fpage>
          -
          <lpage>26</lpage>
          . Amsterdam: John Benjamins,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>Anna</given-names>
            <surname>Wierzbicka</surname>
          </string-name>
          .
          <article-title>Coordination: the semantics of syntactic constructions</article-title>
          .
          <source>In Lingua Mentalis: The Semantics of Natural Language</source>
          , pages
          <fpage>223</fpage>
          -
          <lpage>85</lpage>
          . New York: Academic Press,
          <year>1980</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Zeman</surname>
          </string-name>
          .
          <article-title>Reusable tagset conversion using tagset drivers</article-title>
          .
          <source>In Proceedings of the 6th International Conference on Language Resources and Evaluation (LREC)</source>
          , pages
          <fpage>213</fpage>
          -
          <lpage>218</lpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [32]
          <string-name>
            <surname>Daniel</surname>
            <given-names>Zeman</given-names>
          </string-name>
          , Ondrˇej Dušek, Martin Popel Marecˇek, Loganathan Ramasamy, Jan Šteˇpánek, Zdeneˇk Žabortský, and Jan Hajicˇ. HamleDT:
          <article-title>Harmonized multi-language dependency treebank</article-title>
          .
          <source>Language Resources &amp; Evaluation</source>
          ,
          <volume>48</volume>
          :
          <fpage>601</fpage>
          -
          <lpage>637</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>