<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Deep Dependency Graph Conversion in English</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jinho D. Choi</string-name>
          <email>jinho.choi@emory.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Mathematics and Computer Science Emory University</institution>
        </aff>
      </contrib-group>
      <fpage>35</fpage>
      <lpage>62</lpage>
      <abstract>
        <p>This paper presents a method for the automatic conversion of constituency trees into deep dependency graphs consisting of primary, secondary, and semantic relations. Our work is distinguished from previous work concerning the generation of shallow dependency trees such that it generates dependency graphs incorporating deep structures in which relations stay consistent regardless of their surface positions, and derives relations between out-of-domain arguments, caused by syntactic variations such as open clause, relative clause, or coordination, and their predicates so the complete argument structures are represented for both verbal and non-verbal predicates. Our deep dependency graph conversion recovers important argument relations that would be missed by dependency tree conversion, and merges syntactic and semantic relations into one unified representation, which can reduce the bundle of developing another layer of annotation dedicated for predicate argument structures. Our graph conversion method is applied to six corpora in English and generated over 4.6M dependency graphs covering 20 different genres.1</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Several approaches have been proposed for the automatic conversion of constituency
trees into dependency trees in English [
        <xref ref-type="bibr" rid="ref12 ref13 ref8 ref9">8, 9, 12, 13, 22</xref>
        ]. Multiple benefits are found
by this type of conversion. First, there exists a large amount of corpora annotated
with constituency trees in English such that by converting them into dependency
trees, large data can be obtained for building robust dependency parsing models with
a minimum of manual effort. Second, long-distance dependencies are represented
by non-projective dependencies in dependency trees, which can be reliably found by
the current state-of-the-art dependency parsers [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], whereas they are represented by
empty categories in constituency trees and little to no constituency parsers produce
them well [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Third, dependency trees are more suitable for representing flexible
word order languages as well as colloquial writings such that they are often preferred
to represent universal structures over constituency trees.
1All our resources are publicly available: https://github.com/emorynlp/ddr
Most of the previous work focuses on the generation of shallow dependency trees,
which do not necessarily carry on the same dependency structures given different
syntactic alternations even when they comprise similar semantics. For the following
sentences, shallow dependency trees give different structures although the
underlying semantics of these sentence pairs are the same:
      </p>
    </sec>
    <sec id="sec-2">
      <title>John called Mary</title>
      <p>John gave Mary a book
vs.
vs.</p>
    </sec>
    <sec id="sec-3">
      <title>John made a call to Mary A book was given by John to Mary</title>
      <p>
        Furthermore, since dependency trees are bounded by tree properties, single-root,
connected, single-head, and acyclic, they cannot represent any argument structure
that would break these properties [25]. Such argument structures occur often, where
an argument is shared by multiple predicates (e.g., open clauses, coordination) or it
becomes the head of its predicate by the syntax (e.g., relative clauses). Preserving the
tree properties allows the development of efficient parsing models [
        <xref ref-type="bibr" rid="ref19 ref21 ref27">20, 30, 32, 41</xref>
        ];
however, this ends up requiring the development of another model for finding the
missing arguments (e.g., semantic role labeling), which can be more cumbersome
than developing one graph parsing model that generates deep dependency graphs.
      </p>
      <p>
        This paper presents a method that converts the Penn Treebank style constituency
trees [27] into deep dependency graphs. Our dependency graphs are motivated by
deep structures [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], where arguments take the same semantic roles regardless of
their surface positions, and give complete predicate argument structures by utilizing
function tags, empty categories, and unexplored features in coordination provided
by the constituency trees. We believe that this work will be beneficial for those who
need a large amount of dependency graphs with rich predicate argument structures,
where predicates are abstracted away from their syntactic variations.
2
      </p>
      <sec id="sec-3-1">
        <title>Related Work</title>
        <p>
          Nivre [
          <xref ref-type="bibr" rid="ref20">31</xref>
          ] proposed a deterministic conversion method using head-finding and
labeling rules for the conversion of constituency trees into dependency trees. Johansson
and Nugues [22] improved this method by adding non-projective dependencies and
semantic relations using empty categories and function tags; their representation
had been used for the CoNLL’08-09 shared tasks [
          <xref ref-type="bibr" rid="ref17 ref25">17, 37</xref>
          ]. Choi and Palmer [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]
extended this work by updating the head-finding rules for the recent Penn Treebank
format and handling several complex structures such as small clauses or gapping
relations. de Marneffe and Manning [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] suggested a separate conversion method
that gives rich dependency labels, well-known as the Stanford typed dependencies.
Choi and Palmer [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] improved this work by adding non-projective dependencies
and secondary dependencies. de Marneffe et al. [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] introduced another conversion
method aiming towards the Universal Dependencies [
          <xref ref-type="bibr" rid="ref22">33</xref>
          ], a project that attempts
to develop an universal representation for multiple languages. Our work is
distinguished from the previous work because they mostly target on the generation of tree
structures whereas our main focus is on the generation of graph structures.
Our work was highly inspired by previous frameworks on lexicalized tree adjoining
grammars (LTAG), combinatory categorial grammars (CCG), lexical functional
grammars (LFG), and head-driven phrase structure grammars (HPSG). Xia [39]
extracted LTAG from constituency trees by automatically deriving elementary trees
with linguistic knowledge. Hockenmaier and Steedman [18] converted constituency
trees into a corpus of CCG derivations by making several systematic changes in the
constituency trees, knowns as CCGbank [19]. Cahill et al. [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] extracted LFG
subcategorization frames and paths linking long distance dependencies from f-structures
converted from constituency trees. Miyao et al. [
          <xref ref-type="bibr" rid="ref18">29</xref>
          ] extracted HPSG by deriving
fine-grained lexical entries from constituency trees with heuristic annotations.
Numerous statistical parsers have been developed from the corpora generated by these
approaches where the generated structures can be viewed as direct acyclic graphs.
All of the above approaches were based on the old bracketing guidelines from the
Penn Treebank [26], whereas we followed the latest guidelines that made several
structural as well as tagging changes. Our work is similar to Schuster and Manning
[36] in a way that we both try to find the complete predicate argument structures
by adding secondary dependencies to shallow dependency trees, but distinguished
because their dependency relations are still sensitive to the surface positions whereas
such syntactic alternations are abstracted away from our representation.
        </p>
        <p>
          There exist several corpora consisting of deep dependency graphs. Kromann
[24] introduced the Danish Dependency Treebank containing dependency graphs
with long-distance dependencies, gapping relations, and anaphoric reference links.
Al-Raheb et al. [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] created the DCU 250 Arabic Dependency Bank including manual
annotation based on the theoretical framework of LFG. Yu et al. [
          <xref ref-type="bibr" rid="ref28">42</xref>
          ] generated the
Enju Chinese Treebank (ECT) from the Penn Chinese Treebank [
          <xref ref-type="bibr" rid="ref26">40</xref>
          ] by developing
a large-scale grammar based on HPSG. Flickinger et al. [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] introduced
DeepBank derived from parsing results using linguistically precise HPSG and manual
disambiguation. Hajicˇ et al. [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] created the Prague Czech-English Dependency
Treebank (PDT) consisting of parallel dependency graphs over the constituency
trees in the Penn Treebank and their Czech translations. ECT, DeepBank, and PDT
were used for the SemEval 2015 Task 18: Broad-Coverage Semantic Dependency
Parsing. Candito et al. [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] introduced the Sequoia French Treebank that added a
deep syntactic representation to the existing Sequia corpus [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
        <p>
          Although not directly related, it is worth mentioning the existing corpora
consisting of predicate argument structures. Baker et al. [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] introduced FrameNet based
on frame semantics that gave manual annotation of lexical units and their semantic
frames. Palmer et al. [
          <xref ref-type="bibr" rid="ref23">34</xref>
          ] created PropBank where each predicate was annotated
with a sense and each sense came with its own argument structure. Meyers et al. [28]
created NomBank providing annotation of nominal arguments in the Penn Treebank
by fine-tuning the lexical entries. The original PropBank included only verbal
predicates; Hwang et al. [21] extended PropBank with light verb constructions where
eventive nouns associated with light verbs were also considered. Banarescu et al.
[
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] introduced Abstract Meaning Representation which was motivated by PropBank
but richer in representation and more abstracting away from syntax.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Deep Dependency Graph</title>
        <p>
          Our deep dependency graphs (DDG) preserve only two out of the four tree properties:
single-root and connected. Two types of dependencies are used to represent DDG.
The primary dependencies, represented by the top arcs in figures, form dependency
trees similar to the ones introduced by the Universal Dependencies (UD) [
          <xref ref-type="bibr" rid="ref22">33</xref>
          ]. The
secondary dependencies, represented by the bottom arcs in figures, form dependency
graphs allowing multiple heads and cyclic relations. Separating these two types of
dependencies enables to develop either tree or graph parsing models. Additionally,
semantic roles extracted from function tags are annotated on the head nodes.2
3.1
        </p>
        <sec id="sec-3-2-1">
          <title>Non-verbal Predicates</title>
          <p>Copula Non-verbal predicates are mostly constructed by copulas. DDG considers
both the prototypical copula (e.g., be) as well as semi-copulas (e.g., become, remain).
Non-verbal predicates with copulas can be easily identified by checking the function
tag PRD (secondary predicate) in constituency trees (Figure 1a). Unlike UD, the
preposition becomes the head of a preposition phrase when it is a predicate in DDG
(Figure 1b). This is to avoid multiple subjects per predicate, which would be caused
by making a clause as the head of a prepositional phrase (Figure 1c).
Light verb construction Non-verbal predicates can also be constructed by
lightverbs, which are not annotated in constituency trees but they are in PropBank [21].
A set of light verbs L = {make, take, have, do, give, keep}, a set of 2,474 eventive
nouns N = {call, development, violation, : : :}, and a map M 2 jLj jNj ! jPj =
{(give, call) ! to, (make, development) ! of, . . . } of prepositions indicating the
objects of the nominal predicates are collected from PropBank. Given a verb v 2 L
with the direct object n 2 N, v is considered a light verb and the preposition phrase
that immediately follows n and contains the preposition p M(v; n) is considered
the object of n in DDG (Figure 2b). This lexicon-based approach yields about 2.5
times more light verb constructions than PropBank annotation; further assessment
of this pseudo annotation should be performed, which we will explore in the future.
3.2</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>Deep Arguments</title>
          <p>Dative Indirect objects as well as preposition phrases whose semantic roles are the
same as the indirect objects are considered datives. A nominal phrase is identified
as an indirect object if it is followed by another nominal phrase representing the
direct object (Figure 3a). A preposition phrase is considered a dative if it has either
the function tag DTV (dative; Figure 3b) or BNF (benefactive; Figure 3c). Whether or
not all benefactives should be considered datives is opened to a discussion; we plan
to analyze this by using large unstructured data such as Wikipedia to measure the
likelihood of dative constructions for each verb.
2All figures are provided together at the end of this paper.
Expletive Both the existential there and the extrapositional it in the subject
position are considered expletives. The existential there can be identified by checking
the part-of-speech tag EX. The extrapositional it is indicated by the empty category
*EXP*-d in constituency trees, where d is the index to the referent clause (Figure 4c).
When there exists an expletive, DDG labels the referent as the subject of the main
predicate (Figures 4a and 4d) such that it is consistently represented regardless of
the syntactic alternations, whereas it is not the case in UD (Figures 4b and 4e).
Passive construction Arguments in passive constructions are recognized as they
would be in active constructions. The NP-movement for a passive construction is
indicated by the empty category *-d in the constituency tree, where d is the index
to the antecedent (Figures 5a and 5b). However, the NP-movement for a reduced
passive construction is indicated by the empty category * with no index provided for
the antecedent (Figure 5c). To find the antecedents in reduced passive constructions,
we use the heuristic provided by NLP4J, an open source NLP toolkit, which gives
over 99% agreement to the manual annotation of this kind in PropBank.3 In Figure 5,
John, Mary, and book, are the subject (nsbj), the dative (dat), and the object (obj)
of the predicate give, regardless of their syntactic variations in the active, passive,
and reduced passive constructions, which can be achieved by deriving dependency
relations from the empty categories. Note that the object relation in Figure 5c would
cause a cyclic relation among primary dependencies such that it is represented by
the secondary dependency in DDG.</p>
          <p>Small clause A small clause is a declarative clause that consists of a subject and a
secondary predicate, identified by the function tags SBJ and PRD, respectively. There
are two kinds of small clauses found in constituency trees, one with an internal
subject and the other with an external subject. Figure 6 shows examples of small
clauses with internal subjects. In this case, John is consistently recognized as the
subject of the adjectival predicate smart in the declarative clause (Figure 6a), the
small clause (Figure 6b), and the small clause in the passive construction (Figure 6c).
The subject relation in Figure 6c causes the non-projective dependency, which adds
another complexity to DDG; nonetheless, making John as the subject of consider
instead of smart as in UD would yield different relations between active (Figure 7a)
and passive (Figure 7b) constructions, which is against the main objective of DDG.</p>
          <p>Unlike the case of a small clause with the internal subject, a small clause with
the external subject contains the empty category *PRO*-d where d is the index to
the external subject. In this case, the external subject takes two separate semantic
roles, one from its matrix verb and the other from the secondary predicate in the
small clause. In Figure 8, John is consistently recognized as the object of the verbal
predicate call and the subject of the nominal predicate baptist for both the active
(Figure 8a) and the passive (Figure 8b) constructions in DDG, whereas it is not
the case in UD. The subject relation between John and baptist is preserved by the
secondary dependency to avoid multiple heads among the primary dependencies.
3This heuristic is currently used to pseudo annotate these links in PropBank, labeled as LINK-PSV.
Open clause An open clause is a clause with the external subject indicated by the
empty category *PRO*-d (see the description above). Figure 9 shows examples of
open clauses. The external subjects are represented by the secondary dependencies
to avoid multiple heads. Notice that the head of the open clause, teach, in Figure 9b
is assigned with the semantic role prp (purpose) extracted from the function tag
PRP, which gives a more fine-grained relation to this type (Section 3.4).
Relative clause The NP-movement for the relativizer in a relative clause is noted
by the empty category *T*-d in the constituency tree. Each relativizer is assigned
with the dependency relation before its NP-movement and labeled as r-*, indicating
that there exists a referent to this relativizer that should be assigned with the same
relation. In Figure 10a, the relativizer who becomes the subject of the predicate smart
so it is labeled as r-nsbj, implying that there exists the referent John that should
be considered the real subject of smart. Similarly in Figure 10b, the relativizer who
becomes the dative of the predicate buy so labeled as r-dat, implying that there
exists John who is the real dative of buy. These referent relations are represented by
the secondary dependencies to avoid cyclic relations. The constituency trees do not
provide such referent information; we again use the heuristic provided by NLP4J,4
which has been used to pseudo generate such annotation in PropBank, LINK-SLC.
Coordination Arguments in coordination structures are shared across predicates.
These arguments can be identified in constituency trees; they are either the siblings
of the coordinated verbs (e.g., the book and last year in Figure 11a) or the siblings
of the verb phrases that are the ancestors of these verbs (e.g., John in Figure 11a).
When the coordination is not on the same level, right node raising is used, which can
be identified by the empty category *RNR*-d. In Figure 11b, John is coordinated
across the verb phrase including value and the preposition phrase including for.
Unlike the coordinated verbs in Figure 11a that are siblings, these are not siblings
so need to be coordinated through right node raising. The coordinated arguments
are represented by the secondary dependencies to avoid multiple heads.
3.3</p>
        </sec>
        <sec id="sec-3-2-3">
          <title>Auxiliaries</title>
          <p>Modal adjective Modal adjectives are connected with the class of modal verbs
such as can, may, or should that are used with non-modal verbs to express possibility,
permission, intention, etc:
able
likely
willing
unable
An adjective am is considered a modal if 1)it is a non-verbal predicate (i.g., if it
belong to an adjective phrase with the function tag PRD), 2)it is followed by a clause
whose subject is an empty category e, and 3)the antecedent of e is the subject of
am. In Figure 12, able and about are considered modal adjectives because they are
followed by the clauses whose subjects are linked to the subjects of the adjectival
predicates, John. Modal adjectives together with modal verbs give another level of
abstraction in DDG.</p>
          <p>Raising verb Distinguished from most of the previous work, raising verbs modify
the “raised” verbs in DDG. A verb is considered a raising verb if 1)it is followed
by a clause whose subject is the empty category *-d, and 2)the antecedent of the
empty category is the subject of the raise verb. In Figure 13, the raising verbs go,
have, and keep are followed by the clauses whose subjects are the empty categories
*-1, *-2, and *-3, which all link to the same subject as the raised verb, study.
have
go
continue
need
As shown in Figure 9a, semantic roles are extracted from certain function tags and
added to the terminal heads of the phrases that include such function tags. The
function tags used to extract semantic roles are: DIR: directional, EXT: extent, LOC:
locative, MNR: manner, PRP: purpose, and TMP: temporal.
4
4.1</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>Analysis</title>
        <sec id="sec-3-3-1">
          <title>Corpora</title>
          <p>
            Six corpora that consist of the Penn Treebank style constituency trees are used to
generate deep dependency graphs: OntoNotes (Weischedel et al. [38]), the English
Web Treebank (Web; Petrov and McDonald [
            <xref ref-type="bibr" rid="ref24">35</xref>
            ]), QuestionBank (Judge et al. [23]),
and the MiPACQjSharpjThyme corpora (Albright et al. [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ]). All together, these
corpora cover 20 different genres including formal, colloquial, conversational, and
clinical documents, providing enough diversities to our dependency representation.
          </p>
          <p>SC
WC</p>
          <p>OntoNotes</p>
          <p>138,566
2,620,495
Table 4 shows the distributions of the primary and secondary dependencies generated
by our deep dependency graph conversion. At a glimpse, the portion of the secondary
dependencies over the entire primary dependencies seems rather small (about 2.3%).
However, when only the core arguments (*sbj, obj, dat, comp) and the adverbials
(adv*, neg, ppmod) are considered, where the secondary dependencies are mostly
focused on, the portion increases to 8.4%, which is more significant. Few of the
secondary dependencies are generated for unexpected relations such as acl, appo,
and attr; from our analysis, we found that those were mostly caused by annotation
errors in constituency trees.
We present a conversion method that automatically transforms constituency trees
into deep dependency graphs. Our graphs consist of three types of relations,
primary dependencies, secondary dependencies, and semantic roles, which can be
processed separately or together to produce one unified dependency representation.
The primary dependencies form dependency trees that can be generated by any
non-projective dependency parser. The secondary dependencies together with the
primary dependencies form deep dependency graphs. The semantic roles together
with the syntactic dependencies form rich predicate argument structures. Our
conversion method is applied to large corpora (over 4.6 times larger than the original
Penn Treebank), which provides big data with much diversities. We plan to further
extend this approach to more semantically-oriented dependency graphs by utilizing
existing lexicons such as PropBank and VerbNet.</p>
        </sec>
      </sec>
      <sec id="sec-3-4">
        <title>Acknowledgments</title>
        <p>We gratefully acknowledge the support of the Kindi research grant. A special thank
is due to professor Martha Palmer at the University of Colorado Boulder, who had
encouraged the author to develop this representation during his Ph.D. program.
Type</p>
        <p>Label</p>
        <p>Description
csbj
expl
nsbj
comp
dat
obj
acl
appo
attr
num
relcl
adv
advcl
advnp
neg
ppmod
case
conj
com
dep
disc
meta
prn
root
r-adv
r-advcl
r-advnp
r-comp
r-nsbj
r-ppmod</p>
        <p>Total
clr
[18] Julia Hockenmaier and Mark Steedman. Acquiring Compact Lexicalized
Grammars from a Cleaner Treebank. In Proceedings of the 3rd International
Conference on Language Resources and Evaluation, LREC’02, 2002.
[19] Julia Hockenmaier and Mark Steedman. CCGbank: A Corpus of CCG
Derivations and Dependency Structures Extracted from the Penn Treebank.
Computational Linguistics, 33(3):355–396, 2007.
[20] Liang Huang and Kenji Sagae. Dynamic Programming for Linear-Time
Incremental Parsing. In Proceedings of the 48th Annual Meeting of the Association
for Computational Linguistics, ACL’10, 2010.
[21] Jena D. Hwang, Archna Bhatia, Clare Bonial, Aous Mansouri, Ashwini Vaidya,
Nianwen Xue, and Martha Palmer. PropBank Annotation of Multilingual Light
Verb Constructions. In Proceedings of ACL workshop on Linguistic Annotation,
LAW’10, pages 82–90, 2010.
[22] Richard Johansson and Pierre Nugues. Extended Constituent-to-dependency
Conversion for English. In Proceedings of the 16th Nordic Conference of
Computational Linguistics, NODALIDA’07, 2007.
[23] John Judge, Aoife Cahill, and Josef van Genabith. QuestionBank: Creating
a Corpus of Parse-Annotated Questions. In Proceedings of the 44th Annual
Meeting of the Association for Computational Linguistics, ACL’06, pages
497–504, 2006.
[24] Matthias T. Kromann. The Danish Dependency Treebank and the underlying
linguistic theory. In Proceedings of the Second Workshop on Treebanks and
Linguistic Theories, TLT’03, 2003.
[25] Sandra Kübler, Ryan T. McDonald, and Joakim Nivre. Dependency Parsing.</p>
        <p>Synthesis Lectures on Human Language Technologies. Morgan &amp; Claypool
Publishers, 2009.
[26] Mitchell Marcus, Grace Kim, Mary Ann Marcinkiewicz, Robert Macintyre,
Ann Bies, Mark Ferguson, Karen Katz, and Britta Schasberger. The Penn
Treebank: Annotating Predicate Argument Structure. In ARPA Human Language
Technology Workshop, pages 114–119, 1994.
[27] Mitchell P. Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini.
Building a Large Annotated Corpus of English: The Penn Treebank. Computational
Linguistics, 19(2):313–330, 1993.
[28] Adam Meyers, Ruth Reeves, Catherine Macleod, Rachel Szekely, Veronika
Zielinska, Brian Young, and Ralph Grishman. Annotating Noun Argument
Structure for NomBank. In Proceedings of the 4th International Conference
on Language Resources and Evaluation, LREC’04, 2004.
[36] Sebastian Schuster and Christopher D. Manning. Enhanced English Universal
Dependencies: An Improved Representation for Natural Language
Understanding Tasks. In Proceedings of the Tenth International Conference on
Language Resources and Evaluation, LREC’16, 2016.
[38] Ralph Weischedel, Eduard Hovy, Martha Palmer, Mitch Marcus, Robert
Belvin, Sameer Pradhan, Lance Ramshaw, and Nianwen Xue. OntoNotes:
A Large Training Corpus for Enhanced Processing. In Joseph Olive, Caitlin
Christianson, and John McCary, editors, Handbook of Natural Language
Processing and Machine Translation. Springer, 2011.
[39] Fei Xia. Extracting Tree Adjoining Grammars from Bracketed Corpora. In In
Proceedings of the Fifth Natural Language Processing Pacific Rim Symposium,
1999.
(a) Penn Treebank (PTB).
(b) Deep Dependency Graph (DDG).</p>
        <p>(c) Universal Dependencies (UD).
(b) Deep Dependency Graph (DDG).</p>
        <p>(c) Universal Dependencies (UD).
(b) Penn Treebank (PTB; top) and Deep Dependency Graph (DDG; bottom).</p>
        <p>(c) Penn Treebank (PTB; top) and Deep Dependency Graph (DDG; bottom).
(b) Universal Dependencies (UD).</p>
        <p>(c) Penn Treebank (PTB).
(d) Deep Dependency Graph (DDG).</p>
        <p>(e) Universal Dependencies (UD).
(b) Penn Treebank (PTB; top) and Deep Dependency Graph (DDG; bottom).</p>
        <p>(c) Penn Treebank (PTB; top) and Deep Dependency Graph (DDG; bottom).
(b) Penn Treebank (PTB; top) and Deep Dependency Graph (DDG; bottom).</p>
        <p>(c) Penn Treebank (PTB; top) and Deep Dependency Graph (DDG; bottom).
(a) Universal Dependencies (UD).</p>
        <p>(b) Universal Dependencies (UD).
(a) Penn Treebank (top), Deep Dependency Graph (middle), and Universal Dependencies (bottom).
(b) Penn Treebank (top), Deep Dependency Graph (middle), and Universal Dependencies (bottom).</p>
        <p>(b) Penn Treebank (PTB; top) and Deep Dependency Graph (DDG; bottom).</p>
        <p>(b) Penn Treebank (PTB; top) and Deep Dependency Graph (DDG; bottom).</p>
        <p>(b) Penn Treebank (PTB; top) and Deep Dependency Graph (DDG; bottom).</p>
        <p>(b) Penn Treebank (PTB; top) and Deep Dependency Graph (DDG; bottom).</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Yafa</given-names>
            <surname>Al-Raheb</surname>
          </string-name>
          , Amine Akrout, Josef van Genabith,
          <string-name>
            <given-names>and Joseph</given-names>
            <surname>Dichy</surname>
          </string-name>
          .
          <article-title>DCU 250 Arabic Dependency Bank: An LFG Gold Standard Resource for the Arabic Penn Treebank</article-title>
          .
          <source>In The Challenge of Arabic for NLP/MT at the British Computer Society</source>
          , pages
          <fpage>105</fpage>
          -
          <lpage>116</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Albright</surname>
          </string-name>
          , Arrick Lanfranchi, Anwen Fredriksen, William F. Styler, Colin Warner,
          <string-name>
            <given-names>Jena D.</given-names>
            <surname>Hwang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Jinho D.</given-names>
            <surname>Choi</surname>
          </string-name>
          , Dmitriy Dligach,
          <string-name>
            <given-names>Rodney D.</given-names>
            <surname>Nielsen</surname>
          </string-name>
          , James Martin,
          <string-name>
            <given-names>Wayne Ward</given-names>
            , Martha Palmer, and
            <surname>Guergana</surname>
          </string-name>
          <string-name>
            <given-names>K.</given-names>
            <surname>Savova</surname>
          </string-name>
          .
          <article-title>Towards comprehensive syntactic and semantic annotations of the clinical narrative</article-title>
          .
          <source>Journal of the American Medical Informatics Association</source>
          ,
          <volume>20</volume>
          (
          <issue>5</issue>
          ):
          <fpage>922</fpage>
          -
          <lpage>930</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Collin</surname>
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Baker</surname>
            ,
            <given-names>Charles J.</given-names>
          </string-name>
          <string-name>
            <surname>Fillmore</surname>
            ,
            <given-names>and John B. Lowe.</given-names>
          </string-name>
          <article-title>The Berkeley FrameNet Project</article-title>
          .
          <source>In Proceedings of the 36th Annual Meeting of the Association for Computational Linguistics and the 17th International Conference on Computational Linguistics</source>
          ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Laura</given-names>
            <surname>Banarescu</surname>
          </string-name>
          , Claire Bonial, Shu Cai, Madalina Georgescu, Kira Griffitt, Ulf Hermjakob, Kevin Knight, Philipp Koehn, Martha Palmer, and
          <string-name>
            <given-names>Nathan</given-names>
            <surname>Schneider</surname>
          </string-name>
          .
          <article-title>Abstract Meaning Representation for Sembanking</article-title>
          .
          <source>In Proceedings of the 7th Linguistic Annotation Workshop and Interoperability with Discourse</source>
          ,
          <source>LAW-ID'13</source>
          , pages
          <fpage>178</fpage>
          -
          <lpage>186</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Aoife</given-names>
            <surname>Cahill</surname>
          </string-name>
          , Michael Burke,
          <string-name>
            <surname>Ruth O'Donovan</surname>
          </string-name>
          , Josef van Genabith,
          <article-title>and Andy Way. Long-distance Dependency Resolution in Automatically Acquired Widecoverage PCFG-based LFG Approximations</article-title>
          .
          <source>In Proceedings of the 42nd Annual Meeting on Association for Computational Linguistics, ACL'04</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Marie</given-names>
            <surname>Candito</surname>
          </string-name>
          and
          <string-name>
            <given-names>Djamé</given-names>
            <surname>Seddah</surname>
          </string-name>
          .
          <article-title>The sequoia corpus : Syntactic annotation and use for a parser lexical domain adaptation method in french</article-title>
          .
          <source>In Proceedings of the Joint Conference JEP-TALN-RECITAL</source>
          , pages
          <fpage>321</fpage>
          -
          <lpage>334</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Marie</given-names>
            <surname>Candito</surname>
          </string-name>
          , Guy Perrier, Bruno Guillaume, Corentin Ribeyre, Karën Fort, Djamé Seddah, and
          <string-name>
            <surname>Eric De La Clergerie</surname>
          </string-name>
          .
          <article-title>Deep Syntax Annotation of the Sequoia French Treebank</article-title>
          .
          <source>In Proceedings of the Ninth International Conference on Language Resources and Evaluation</source>
          ,
          <source>LREC'14</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Jinho</surname>
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Choi</surname>
            and
            <given-names>Martha</given-names>
          </string-name>
          <string-name>
            <surname>Palmer</surname>
          </string-name>
          .
          <article-title>Robust Constituent-to-Dependency Conversion for Multiple Corpora in English</article-title>
          .
          <source>In Proceedings of the 9th International Workshop on Treebanks and Linguistic Theories</source>
          ,
          <source>TLT'10</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Jinho</surname>
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Choi</surname>
            and
            <given-names>Martha</given-names>
          </string-name>
          <string-name>
            <surname>Palmer</surname>
          </string-name>
          .
          <article-title>Guidelines for the Clear Style Constituent to Dependency Conversion</article-title>
          .
          <source>Technical Report 01-12</source>
          , University of Colorado Boulder,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Jinho</surname>
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Choi</surname>
            , Amanda Stent, and
            <given-names>Joel</given-names>
          </string-name>
          <string-name>
            <surname>Tetreault</surname>
          </string-name>
          .
          <article-title>It depends: Dependency parser comparison using a web-based evaluation tool</article-title>
          .
          <source>In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics, ACL'15</source>
          , pages
          <fpage>387</fpage>
          -
          <lpage>396</lpage>
          , Beijing, China,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Noam</given-names>
            <surname>Chomsky</surname>
          </string-name>
          . Lectures in Government and Binding. Dordrecht, Foris,
          <year>1981</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Marie-Catherine de Marneffe and Christopher D. Manning</surname>
          </string-name>
          .
          <article-title>The Stanford typed dependencies representation</article-title>
          .
          <source>In Proceedings of the COLING workshop on Cross-Framework and Cross-Domain Parser Evaluation</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Marie-Catherine de Marneffe</surname>
            , Timothy Dozat, Natalia Silveira, Katri Haverinen, Filip Ginter, Joakim Nivre, and
            <given-names>Christopher D.</given-names>
          </string-name>
          <string-name>
            <surname>Manning</surname>
          </string-name>
          .
          <article-title>Universal Stanford dependencies: A cross-linguistic typology</article-title>
          .
          <source>In Proceedings of the 9th International Conference on Language Resources and Evaluation</source>
          , LREC'
          <volume>14</volume>
          , pages
          <fpage>4585</fpage>
          -
          <lpage>4592</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Flickinger</surname>
          </string-name>
          , Yi Zhang, and Valia Kordoni.
          <article-title>DeepBank: A Dynamically Annotated Treebank of the Wall Street Journal</article-title>
          .
          <source>In Proceedings of the 11th International Workshop on Treebanks and Linguistic Theories. International Workshop on Treebanks and Linguistic Theories</source>
          ,
          <source>TLT'12</source>
          , pages
          <fpage>85</fpage>
          -
          <lpage>96</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Ryan</given-names>
            <surname>Gabbard</surname>
          </string-name>
          , Mitchell Marcus, and
          <string-name>
            <given-names>Seth</given-names>
            <surname>Kulick</surname>
          </string-name>
          .
          <article-title>Fully parsing the penn treebank</article-title>
          .
          <source>In Proceedings of the Conference on Human Language</source>
          Technology -
          <article-title>North American chapter of the Association for Computational Linguistics</article-title>
          ,
          <source>NAACL'06</source>
          , pages
          <fpage>184</fpage>
          -
          <lpage>191</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Jan</surname>
            <given-names>Hajicˇ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eva</surname>
            <given-names>Hajicˇová</given-names>
          </string-name>
          , Jarmila Panevová, Petr Sgall, Ondrˇej Bojar, Silvie Cinková,
          <string-name>
            <surname>Eva</surname>
            <given-names>Fucˇíková</given-names>
          </string-name>
          , Marie Mikulová, Petr Pajas, Jan Popelka, Jirˇí Semecký, Jana Šindlerová,
          <string-name>
            <surname>Jan</surname>
            <given-names>Šteˇpánek</given-names>
          </string-name>
          , Josef Toman, Zdenˇka Urešová, and
          <string-name>
            <given-names>Zdeneˇk</given-names>
            <surname>Žabokrtský. Announcing Prague</surname>
          </string-name>
          Czech-English
          <source>Dependency Treebank</source>
          <volume>2</volume>
          .0. In Nicoletta Calzolari (Conference Chair), Khalid Choukri, Thierry Declerck,
          <article-title>Mehmet Ug˘ur Dog˘an, Bente Maegaard</article-title>
          , Joseph Mariani, Asuncion Moreno, Jan Odijk, and Stelios Piperidis, editors,
          <source>Proceedings of the Eight International Conference on Language Resources and Evaluation</source>
          ,
          <source>LREC'12</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Jan</surname>
            <given-names>Hajicˇ</given-names>
          </string-name>
          , Massimiliano Ciaramita, Richard Johansson, Daisuke Kawahara, Maria Antònia Martí, Lluís Màrquez, Adam Meyers, Joakim Nivre,
          <string-name>
            <given-names>Sebastian</given-names>
            <surname>Padó</surname>
          </string-name>
          ,
          <string-name>
            <surname>Jan</surname>
            <given-names>Šteˇpánek</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pavel</surname>
            <given-names>Stranˇák</given-names>
          </string-name>
          , Mihai Surdeanu, Nianwen Xue,
          <string-name>
            <given-names>and Yi</given-names>
            <surname>Zhang. The CoNLL-2009 Shared Task</surname>
          </string-name>
          <article-title>: Syntactic and Semantic Dependencies in Multiple Languages</article-title>
          .
          <source>In Proceedings of the 13th Conference on Computational Natural Language Learning: Shared Task</source>
          ,
          <source>CoNLL'09</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>18</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [29]
          <string-name>
            <surname>Yusuke</surname>
            <given-names>Miyao</given-names>
          </string-name>
          , Takashi Ninomiya, and
          <article-title>Jun'ichi Tsujii. Corpus-Oriented Grammar Development for Acquiring a Head-driven Phrase Structure Grammar from the Penn Treebank</article-title>
          .
          <source>In Proceedings of the 1st International Joint Conference on Natural Language Processing, IJCNLP'04</source>
          , pages
          <fpage>684</fpage>
          -
          <lpage>693</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>Joakim</given-names>
            <surname>Nivre</surname>
          </string-name>
          .
          <article-title>An Efficient Algorithm for Projective Dependency Parsing</article-title>
          .
          <source>In Proceedings of the 8th International Workshop on Parsing Technologies, IWPT'03</source>
          , pages
          <fpage>149</fpage>
          -
          <lpage>160</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>Joakim</given-names>
            <surname>Nivre</surname>
          </string-name>
          .
          <source>Inductive Dependency Parsing</source>
          . Springer,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [32]
          <string-name>
            <surname>Joakim</surname>
            <given-names>Nivre</given-names>
          </string-name>
          , Johan Hall, Jens Nilsson, Güls¸en Eryigˇit, and Svetoslav Marinov.
          <article-title>Labeled Pseudo-Projective Dependency Parsing with Support Vector Machines</article-title>
          .
          <source>In Proceedings of the 10th Conference on Computational Natural Language Learning</source>
          ,
          <source>CoNLL'06</source>
          , pages
          <fpage>221</fpage>
          -
          <lpage>225</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [33]
          <string-name>
            <surname>Joakim</surname>
            <given-names>Nivre</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marie-Catherine de Marneffe</surname>
            , Filip Ginter, Yoav Goldberg, Jan Hajic,
            <given-names>Christopher D.</given-names>
          </string-name>
          <string-name>
            <surname>Manning</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ryan</surname>
            <given-names>McDonald</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Slav</given-names>
            <surname>Petrov</surname>
          </string-name>
          , Sampo Pyysalo, Natalia Silveira, Reut Tsarfaty, and Daniel Zeman.
          <article-title>Universal Dependencies v1: A Multilingual Treebank Collection</article-title>
          .
          <source>In Proceedings of the 10th International Conference on Language Resources and Evaluation</source>
          , LREC'
          <volume>16</volume>
          , pages
          <fpage>23</fpage>
          -
          <lpage>28</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [34]
          <string-name>
            <surname>Martha</surname>
            <given-names>Palmer</given-names>
          </string-name>
          , Daniel Gildea, and
          <string-name>
            <given-names>Paul</given-names>
            <surname>Kingsbury</surname>
          </string-name>
          .
          <source>The Proposition Bank: An Annotated Corpus of Semantic Roles. Computational Linguistics</source>
          ,
          <volume>31</volume>
          (
          <issue>1</issue>
          ):
          <fpage>71</fpage>
          -
          <lpage>106</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>Slav</given-names>
            <surname>Petrov</surname>
          </string-name>
          and
          <article-title>Ryan McDonald. Overview of the 2012 Shared Task on Parsing the Web</article-title>
          .
          <source>In Proceedings of the 1st Workshop on Syntactic Analysis of Non-Canonical Language, SANCL</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [37]
          <string-name>
            <surname>Mihai</surname>
            <given-names>Surdeanu</given-names>
          </string-name>
          , Richard Johansson, Adam Meyers, Lluís Màrquez, and
          <string-name>
            <given-names>Joakim</given-names>
            <surname>Nivre</surname>
          </string-name>
          .
          <article-title>The CoNLL-2008 Shared Task on Joint Parsing of Syntactic and Semantic Dependencies</article-title>
          .
          <source>In Proceedings of the 12th Conference on Computational Natural Language Learning: Shared Task</source>
          ,
          <source>CoNLL'08</source>
          , pages
          <fpage>59</fpage>
          -
          <lpage>177</lpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [40]
          <string-name>
            <surname>Nianwen</surname>
            <given-names>Xue</given-names>
          </string-name>
          , Fei Xia,
          <string-name>
            <surname>Fu-Dong Chiou</surname>
            , and
            <given-names>Martha</given-names>
          </string-name>
          <string-name>
            <surname>Palmer</surname>
          </string-name>
          .
          <article-title>The penn chinese treebank: Phrase structure annotation of a large corpus</article-title>
          .
          <source>Natural Language Engineering</source>
          ,
          <volume>11</volume>
          (
          <issue>2</issue>
          ):
          <fpage>207</fpage>
          -
          <lpage>238</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>Hiroyasu</given-names>
            <surname>Yamada</surname>
          </string-name>
          and
          <string-name>
            <given-names>Yuji</given-names>
            <surname>Matsumoto</surname>
          </string-name>
          .
          <article-title>Statistical dependency analysis with support vector machine</article-title>
          .
          <source>In Proceedings of the 8th International Workshop on Parsing Technologies, IWPT'03</source>
          , pages
          <fpage>195</fpage>
          -
          <lpage>206</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [42]
          <string-name>
            <surname>Kun</surname>
            <given-names>Yu</given-names>
          </string-name>
          , Miyao Yusuke, Xiangli Wang,
          <string-name>
            <surname>Takuya Matsuzaki</surname>
            , and
            <given-names>Junichi</given-names>
          </string-name>
          <string-name>
            <surname>Tsujii</surname>
          </string-name>
          .
          <article-title>Semi-automatically Developing Chinese HPSG Grammar from the Penn Chinese Treebank for Deep Parsing</article-title>
          .
          <source>In Proceedings of the 23rd International Conference on Computational Linguistics, COLING'10</source>
          , pages
          <fpage>1417</fpage>
          -
          <lpage>1425</lpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>