<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Simulating Dependencies to Improve Parse Error Detection</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Markus Dickinson</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Amber Smith</string-name>
          <email>smithamj@indiana.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Linguistics Indiana University</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <fpage>76</fpage>
      <lpage>88</lpage>
      <abstract>
        <p>We improve parse error detection, weighting dependency information on the basis of simulated parses. Such simulations extend the training grammar, and, although the simulations are not wholly correct or incorrect-as observed from the results with different weightings for small treebanks-they help to determine whether a new parse fits the training grammar.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction and motivation</title>
      <p>Obtaining high-quality annotation is a bottleneck for corpus and parser
development, leading to methods of leveraging existing resources to improve the
parameters of a lesser-resourced language—through the use of annotated corpora of
higher-resourced languages, parallel language data, universal treebanks, and so on
[5, 7, 11, 13, 18, 19, 22]. Building on work that envisions bootstrapping corpora
via parse error detection and human correction [3, 24]—from a larger paradigm
of exploiting a small amount of annotated data in a language [6, 8, 14]—we focus
on extracting more resources from just the treebank itself. Specifically, we explore
the use of parse simulations within a (small) treebank to supplement the parameters
learned from the treebank itself, in order to improve parse error detection.</p>
      <p>This bears much in common with contrastive estimation [e.g., 25, 26] and
negative sampling [e.g., 9, 20], where learning is the process of distinguishing a positive
example from similar negative examples in the data. With small data and the task of
error detection, however, it is not clear whether the simulated information is truly
incorrect or not: for small data sets, one may uncover structures which, although
incorrect in a particular sentence, can be correct. Thus, instead of maximizing the
differences between actually-occurring and simulated data, we employ “negative”
information heuristically, to modify the scores of syntactic dependency structures.</p>
      <p>Employing simulations has the effect of generalizing the grammar implicit in
a treebank. This allows one to explore a large amount of information in the
annotation space (e.g., 100x more structures). Our specific contributions are to:
Posit annotation simulations during training;
Replace ad hoc weighting by a weighting scheme using the simulations,
leading to improved error detection precision; and
Experiment with different weighting schemes, testing the relationship
between simulations and (un)seen structures.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Parse error detection</title>
      <p>We start with the DAPS (Detecting Anomalous Parse Structures) method [3, 4],1
effective across a variety of settings. N -gram sequences of dependency structures
are used to identify anomalous (likely erroneous) dependencies. Referring to the
suggested simulation (dotted line) in Figure 1 as an example, the process is:
1. Extract rules from dependency trees (rule = a head (H) + its dependents). e.g.,
one rule (headed by board) is S T A R T det:DT NN-H prep:IN E N D.
2. Extract rule n-grams, e.g., det:DT NN-H prep:IN or prep:IN E N D.
3. Score an element by adding up relevant training n-gram counts. e.g., to score
prep:IN here, one adds up training counts of the n-grams with prep:IN.2
Using information from complete trees, the DAPS method serves as a sanity
check on more complicated parse models. While such checks may be possible
inside a parser, this method optimizes for unlikelihood, whereas parsers more directly
maximize likelihood. In the context of low-resource languages, DAPS has the
benefit of using very general representations—as opposed to integrating information
from, e.g., word clusters—that do not require a lot of (un)annotated data.
Insights Our previous DAPS work [3, 24] has noted that not all n-grams are
equal: a parsed rule lacking any training bigrams indicates an error, and yet,
without more context, bigrams often upweight erroneous rules because they generally
occur in training. Previous variations ignored bigrams (high method: n 3) or
downweighted them (weighted all-grams [wall]: n 2, bigram counts multiplied
by 0:01). Problematic is setting the weight arbitrarily and equally across bigrams.</p>
      <p>A second insight comes from a stage of revision checking [3], where revisions
(alternative attachments and labelings) are posited for a parse. Errors are flagged
when the parser did not choose better-scoring possible revisions. The crucial
insight is that dependency trees provide information not only about what they are,
but what they could have been, an insight we will apply during training
1http://cl.indiana.edu/~md7/papers/dickinson-smith11.html
2See [3] for more details.
nsubj
aux</p>
      <p>prep
dobj
det
pobj</p>
      <p>det
prep
amod
Pierre Vinken ...</p>
      <p>NNP NNP ...</p>
      <p>will join the board as
MD VB DT NN IN
a
DT
nonexecutive director ...</p>
      <p>JJ NN ...</p>
    </sec>
    <sec id="sec-3">
      <title>Annotation simulations</title>
      <sec id="sec-3-1">
        <title>Positing simulations</title>
        <p>We focus on alternate plausible structures within the training data and use these
syntactic parse simulations to gauge the utility of each n-gram. Consider, for
example, the attachment of the preposition (IN) as in Figure 1. It may help here to
see that in the gold tree as doesn’t attach to the immediately preceding noun (NN)
board—a potential attachment, as indicated by the gray dotted line.</p>
        <p>We view the process of generating alternative structures, i.e., simulations, as
involving the reattachment of each node (i.e., word/POS) in the dependency tree.
This means that: a) a full set of alternatives is generated for every node in the tree
(e.g., for the attachment of Pierre, of Vinken, etc.), and b) no more than one node
is reattached for a given simulation (e.g., if as is re-attached, then no other word
is). Each simulation is thus only “one off” from the original tree, preventing highly
implausible generalizations from the source trees. While each simulation results in
a specific incorrect tree (token), to be useful for error detection simulations should
ideally correspond to valid structures (types) used elsewhere in the corpus, i.e.,
plausible structures a parser would likely consider for the sentence at hand.</p>
        <p>We use the reattachment procedure as implemented for revision checking
(section 2).3 From this, we keep a frequency database of n-grams that actually occurred
(A) and a separate one for simulated n-grams (S); we call each database a
grammar. The n-grams are a sequence of pairs of POS and dependency labels.</p>
        <p>To build plausible simulations, we currently stipulate [cf., e.g., 1, 10] that: a)
a posited reattachment must maintain projectivity; and b) its posited dependency
label must have been seen with its POS in training.4 We only consider
reattachments, and not simple relabelings, because: a) we already posit multiple labels for
each reattached dependency, and b) the grammar size is more manageable. One
could explore creating new POS, label pairs or filtering implausible simulations;
we leave these for future work, but will see the effects of not filtering in section 4.</p>
        <p>3The interpretation is different for revision checking, as one hopes to go from an incorrect parse
to a correct one, and here the situation is reversed, but the process is the same.</p>
        <p>4We do not enforce acyclicity, as in [3], as this unduly restricts the search space for revisions [12].
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Weighting simulations</title>
        <p>With a definition of simulations, the next step is to integrate such information with
actually-occurring rules. As mentioned, we do not simply view every simulation
as incorrect, so our methods reflect a gradable nature of the simulations.</p>
        <p>During training, we use the actual count of an n-gram (a) and its simulated
count (s) to calculate an adjusted score (a ). Previous (0) weighting used an ad
hoc weight (wall) on bigrams or no weight (all/high: a = a).</p>
        <p>The formulas in Figure 2 attempt to downweight n-grams which are too
confusable with other structures to be good indicators of quality (i.e., occur a great
deal in simulated fashion [s &gt; a]) and to upweight helpful n-grams (i.e., ones
where actual occurrences outweigh simulated ones [a &gt; s]), as can be seen in the
ratio as++11 . Starting with formula B, the training grammar is expanded by giving
non-zero scores to what we term SimOnly n-grams, where a = 0 and s &gt; 0.</p>
        <p>Formulas B–E differ in how they handle the interplay between these SimOnly
n-grams, the actually-occurring n-grams (a &gt; 0, Actual), and the never-occurring
n-grams (a = 0, s = 0), i.e., the ones Novel to both grammars. The general
principles are characterized at the bottom of each section of Figure 2: less important
than the specifics is what these indicate about how to use simulations for new data.
3.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Combining simulations and revisions</title>
        <p>As an additional evaluation, we run revision checking during the error detection
stage, in order to see the effect of combining parse simulations with parse
revisions. Note what this means: parse simulations posit additional structures to
expand grammars during training, while revisions are additional structures posited
after parsing. (For revision checking we try both reattachments and relabelings
as revisions.) These posited revisions are scored with the same training grammar
as the original parse is scored with—i.e., the grammar including simulations—and
thus no scoring proceeds in the same way for the parse and any suggested revision.
4
4.1</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Evaluation</title>
      <sec id="sec-4-1">
        <title>Experimental conditions</title>
        <p>We use the Wall Street Journal (WSJ) corpus from the Penn Treebank [15],
converted to Stanford dependencies [2], to be comparable to previous work [24].
Based on the recommendations we outline in [24], we consider two
complementary parsers (graph-based MSTParser [17]; transition-based MaltParser [21]) for
the various methods, as well as a more recent one (TurboParser [16]). We use a
small training corpus (section 02 [48k tokens]) [see 24], as this matches our focus,
and test on a slightly larger corpus (sections 00+01+22 [134k tokens]).</p>
        <p>C =1:2
C =2
-2
-2
A
B</p>
        <p>-2
a =</p>
        <p>-2
a =
D</p>
        <p>-2
E</p>
        <p>Novel &lt; SimOnly &lt; Actual
Grammar sizes Before evaluating the effectiveness on error detection, we can
see the impact of parse simulations by examining training grammar sizes, as in
Table 1. In moving from Actual (A) to Simulated (S) grammars, we see a 88x
increase in type grammar size (94k to 8.3m) and a 104x increase in token size.5 We
also see a gain in type coverage and a drop in token coverage, indicating
simulations of infrequent n-grams; given the difficulty of predicting rare events, positing
such rules may benefit many applications.6</p>
        <p>Grammar
(All)
02A
02S
02A[S
00+01+22A</p>
        <p>When the entire Actual and Simulated grammars are combined (A [ S), the
overlap with the testing data increases by an absolute 13% (types) or 6% (tokens).
Creating simulations thus shows promise for expanding the grammar in ways that
extend to new data—with a caveat of a huge leap in grammar size for a smaller
gain in grammar coverage, thus indicating a need for some grammar filtering.
Evaluation metric For evaluating error detection, we report precision (P) of
detected errors.7 We rely on equivalently-sized segments across different error
detection methods, accounting for recall. Using segment size—e.g., 5% of corpus
tokens—approximates the idea of annotators having a set amount of time to
correct parser errors [see 24, for discussion]. When it comes to segment size, we
assume that correction works from the lowest scores up to the highest.</p>
        <p>For methods employing revision checking (sections 2 and 3.3), the
methodology is slightly different: first, all flagged positions are examined, from lowest to
highest scores, and then all unflagged positions. We focus on flagged positions.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Results</title>
        <p>We report results for MST and Malt in Table 2 and for Turbo in Table 3, with
the best formula results in bold, and the best overall results underlined. To
summarize: 1) Weighting seems to work, as formula 0 is generally among the worst
5There are 32,199 types (188,698 Actual tokens) in common between the A and S grammars.
6We omit numbers broken down by n-gram size: there are many more types for higher n values.
7In [24], we also recommend reporting a revised labeled attachement score (LASr), to see the
impact of correcting errors on resulting corpus quality. As we are more interested in method
development and as the LASr trends precisely mirror the precision ones, we do not report LASr here.
formulas; 2) The best overall settings seem to be a mixture of hand- (wall) and
auto-weighting (A, B, E); and 3) A better parser (Turbo) mitigates the
improvements to some extent, but basically follows the same trends.</p>
        <p>While the wall method is generally best, for MST and Malt the all method
curiously outperforms it for formulas B–E at the 5% size, i.e., bigram weighting
seems unhelpful. For Turbo, the all method is almost always lowest, with high
and wall being fairly comparable across the different segments; unweighted
bigrams (all) are apparently more misleading with a better parser. Additionally,
formula E (Novel &lt; SimOnly &lt; Actual) is most often the best setting, but A
(ignoring the Novel/SimOnly distinction) can be better. Formulas employing a
Novel/SimOnly distinction are perhaps more unstable than formula A.</p>
        <p>If forced to pick a best setting, wall.E is consistently strong, aside from its
dip at 5% for MST and Malt. Simulations seem to be both positive and negative:
formula E equals or outperforms B (SimOnly Actual), while also
outperforming D (SimOnly &lt; Novel) and often A (SimOnly Novel). This
pattern, along with the 5% dip, may in part be due to not having sorted the simulations:
some SimOnly cases behave like Novel and some like Actual n-grams.
Model
Positions
score
0
Employing revision checking presents a slightly different—and more consistent—
picture, shown in Tables 4 and 5, where we focus on evaluation cutoffs that
highlight the effect of revisions.8 Most noticeably, formula E almost always has the
highest precision for both all the flagged positions (AF) and the flagged positions
with the lowest scores (0F), and the wall.E method is clearly best. Although
some of the differences are small (especially for the MST results), we see 4–10%
increases in precision for Malt for the 5% segment, comparing E to formulas 0/A.</p>
        <p>We suspect that this improvement comes because simulations and revisions rely
upon the same procedure of positing alternate structures, and formula E provides
better scoring for the unseen structures involved in the revisions. Because it assigns
positive scores to simulated training n-grams (SimOnly), formula E more often
positively scores revisions. Thus, it is easier to find a revision that scores better
than the original parse (compare, e.g., 86% [= 11;;740600 ] of zero-scoring dependencies
being flagged [0F] for Malt.all with formula E vs. 59% [= 23;;162094 ] with formula
A). The extended grammar thus helps identify when there is a potentially better
8We also ignore the C and D models here, as they continue to be the worst-performing ones.
Model
0F</p>
        <p>
          Model
0F
all.0 86.2% (
          <xref ref-type="bibr" rid="ref1">1,598</xref>
          ) 50.2% (
          <xref ref-type="bibr" rid="ref14">14,424</xref>
          ) 70.2%
all.A 86.0% (
          <xref ref-type="bibr" rid="ref1">1,603</xref>
          ) 54.3% (
          <xref ref-type="bibr" rid="ref13">13,430</xref>
          ) 73.5%
all.B 86.4% (397) 55.1% (
          <xref ref-type="bibr" rid="ref13">13,415</xref>
          ) 74.5%
all.E 86.4% (397) 55.7% (
          <xref ref-type="bibr" rid="ref13">13,484</xref>
          ) 75.7%
high.0 82.6% (
          <xref ref-type="bibr" rid="ref4">4,718</xref>
          ) 50.1% (
          <xref ref-type="bibr" rid="ref14">14,676</xref>
          ) 71.4%
high.A 82.6% (
          <xref ref-type="bibr" rid="ref4">4,729</xref>
          ) 54.2% (
          <xref ref-type="bibr" rid="ref13">13,524</xref>
          ) 73.3%
high.B 77.2% (
          <xref ref-type="bibr" rid="ref1">1,665</xref>
          ) 55.4% (
          <xref ref-type="bibr" rid="ref14">14,936</xref>
          ) 72.6%
high.E 77.2% (
          <xref ref-type="bibr" rid="ref1">1,663</xref>
          ) 55.9% (
          <xref ref-type="bibr" rid="ref15">15,030</xref>
          ) 74.4%
wall.0 86.2% (
          <xref ref-type="bibr" rid="ref1">1,603</xref>
          ) 52.5% (
          <xref ref-type="bibr" rid="ref15">15,595</xref>
          ) 74.7%
wall.A 86.0% (
          <xref ref-type="bibr" rid="ref1">1,608</xref>
          ) 56.3% (
          <xref ref-type="bibr" rid="ref14">14,169</xref>
          ) 76.1%
wall.B 86.4% (396) 56.5% (
          <xref ref-type="bibr" rid="ref14">14,612</xref>
          ) 74.6%
wall.E 86.4% (396) 56.8% (
          <xref ref-type="bibr" rid="ref14">14,764</xref>
          ) 76.2%
all.0 91.3% (
          <xref ref-type="bibr" rid="ref2">2,127</xref>
          ) 47.5% (
          <xref ref-type="bibr" rid="ref13">13,460</xref>
          ) 70.8%
all.A 91.4% (
          <xref ref-type="bibr" rid="ref2">2,129</xref>
          ) 51.3% (
          <xref ref-type="bibr" rid="ref12">12,218</xref>
          ) 71.8%
all.B 96.6% (
          <xref ref-type="bibr" rid="ref1">1,460</xref>
          ) 53.9% (
          <xref ref-type="bibr" rid="ref12">12,671</xref>
          ) 76.2%
all.E 96.6% (
          <xref ref-type="bibr" rid="ref1">1,460</xref>
          ) 54.3% (
          <xref ref-type="bibr" rid="ref12">12,662</xref>
          ) 77.1%
high.0 78.6% (
          <xref ref-type="bibr" rid="ref3">3,369</xref>
          ) 45.5% (
          <xref ref-type="bibr" rid="ref13">13,050</xref>
          ) 63.9%
high.A 78.5% (
          <xref ref-type="bibr" rid="ref3">3,363</xref>
          ) 49.5% (
          <xref ref-type="bibr" rid="ref11">11,632</xref>
          ) 67.0%
high.B 84.8% (
          <xref ref-type="bibr" rid="ref2">2,273</xref>
          ) 52.9% (
          <xref ref-type="bibr" rid="ref13">13,367</xref>
          ) 73.2%
high.E 84.8% (
          <xref ref-type="bibr" rid="ref2">2,271</xref>
          ) 53.2% (
          <xref ref-type="bibr" rid="ref13">13,377</xref>
          ) 74.2%
wall.0 91.4% (
          <xref ref-type="bibr" rid="ref2">2,129</xref>
          ) 49.4% (
          <xref ref-type="bibr" rid="ref14">14,179</xref>
          ) 72.4%
wall.A 91.4% (
          <xref ref-type="bibr" rid="ref2">2,130</xref>
          ) 53.1% (
          <xref ref-type="bibr" rid="ref12">12,566</xref>
          ) 74.4%
wall.B 96.6% (
          <xref ref-type="bibr" rid="ref1">1,461</xref>
          ) 55.2% (
          <xref ref-type="bibr" rid="ref13">13,535</xref>
          ) 77.1%
wall.E 96.6% (
          <xref ref-type="bibr" rid="ref1">1,461</xref>
          ) 55.5% (
          <xref ref-type="bibr" rid="ref13">13,636</xref>
          ) 78.0%
revision that the parser ignored. This seems to be an encouraging direction to
pursue if one wishes to extend error detection into automatic correction.
Discussion The best formula, E, cleanly separates Novel, SimOnly, and Actual
cases. What does this mean, then, in terms of whether simulated information is
incorrect or not, i.e., whether the grammar is appropriately being generalized? To
address this question, we first note that formula E equals or outperforms B,9
indicating that the SimOnly cases are less helpful than the Actual cases. There is
thus an indication that simulated n-grams often contain negative information.
        </p>
        <p>But the picture is more complicated: we also need to note that E almost always
has higher precision than formula D, the formula that cleanly ranks SimOnly
lower than Novel cases. Additionally, formula E generally—though, not always—
9B and E are equivalent when treating zero cases the same, i.e., as Novels only; see Figure 2.
outperforms formula A, the formula where Novel and SimOnly are grouped
together as equally bad. Taken together, the performances indicate that simulations
as a whole are better than cases which have never been seen.</p>
        <p>We conjecture that this behavior is due, at least in part, to not having sorted out
the simulations: some of the SimOnly cases seem to behave like Novel cases
and some more like Actual n-grams. As mentioned, a crucial next step is thus to
sort the SimOnly cases, so that only the valid cases receive higher scores.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Summary and outlook</title>
      <p>We have improved a method for parse error detection, weighting dependency
ngrams on the basis of parse simulations in training. As important as the
improvements in error detection is the idea of extending the training grammar for small
treebanks by such simulations. As we observed from the results with different
formulas, the simulations cannot be treated as wholly correct or incorrect—especially
given the small size of the training data—but are nonetheless useful for helping
determine whether a new parse fits into the (extended) grammar.</p>
      <p>There are several next steps for this work. There is a need to try different
(sub)sets of simulations, to see which constraints or filters can be best employed,
in particular to sort out the good and bad SimOnly cases. Adding filters will
increase the amount of time spent in building a grammar during training, but result
in smaller—as well as more accurate—grammars, thus improving the speed for
post-parsing analysis. We have experimented only with a small training corpus,
resulting in a huge extension of n-grams; one could investigate larger and more
diverse training corpora, to observe both the effectiveness of the techniques and
the efficiency with such large grammars.</p>
      <p>One could also test the methods on out-of-domain data, to see how well the
grammar extensions cover different kinds of data, as well as comparing to
alternative methods of tree exploration, such as parse reranking models employing
ngrams or subtree features over a set of k-best parses [e.g., 23]. In that light, one
also needs to explore more extrinsic evaluation, to see whether the methods are
actually identifying errors that truly impact parsing models (i.e., have a real-world
impact) vs. ones which are more ad hoc.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>We wish to thank the three anonymous reviewers for very helpful feedback, as well
as the CL discussion group at Indiana University.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Giuseppe</given-names>
            <surname>Attardi</surname>
          </string-name>
          and
          <article-title>Felice Dell'Orletta. Reverse revision and linear tree combination for dependency parsing</article-title>
          .
          <source>In Proceedings of HLT-NAACL-09</source>
          , Short Papers, pages
          <fpage>261</fpage>
          -
          <lpage>264</lpage>
          , Boulder, CO,
          <year>June 2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Marie-Catherine de Marneffe</surname>
          </string-name>
          , Bill MacCartney, and
          <string-name>
            <surname>Christopher</surname>
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Manning</surname>
          </string-name>
          .
          <article-title>Generating typed dependency parses from phrase structure parses</article-title>
          .
          <source>In LREC</source>
          <year>2006</year>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Markus</given-names>
            <surname>Dickinson</surname>
          </string-name>
          and
          <string-name>
            <given-names>Amber</given-names>
            <surname>Smith</surname>
          </string-name>
          .
          <article-title>Detecting dependency parse errors with minimal resources</article-title>
          .
          <source>In Proceedings of IWPT-11</source>
          , pages
          <fpage>241</fpage>
          -
          <lpage>252</lpage>
          , Dublin,
          <year>October 2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Markus</given-names>
            <surname>Dickinson</surname>
          </string-name>
          and
          <string-name>
            <given-names>Amber</given-names>
            <surname>Smith</surname>
          </string-name>
          .
          <article-title>Finding parse errors in the midst of parse errors</article-title>
          .
          <source>In Proceedings of the 13th International Workshop on Treebanks and Linguistic Theories (TLT13)</source>
          , Tübingen, Germany,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Long</given-names>
            <surname>Duong</surname>
          </string-name>
          , Trevor Cohn, Steven Bird, and Paul Cook.
          <article-title>Low resource dependency parsing: Cross-lingual parameter sharing in a neural network parser</article-title>
          .
          <source>In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 2: Short Papers)</source>
          , pages
          <fpage>845</fpage>
          -
          <lpage>850</lpage>
          , Beijing, China,
          <year>July 2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Long</given-names>
            <surname>Duong</surname>
          </string-name>
          , Trevor Cohn, Steven Bird, and
          <string-name>
            <given-names>Paul</given-names>
            <surname>Cook</surname>
          </string-name>
          .
          <article-title>A neural network model for low-resource universal dependency parsing</article-title>
          .
          <source>In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing</source>
          , pages
          <fpage>339</fpage>
          -
          <lpage>348</lpage>
          , Lisbon, Portugal,
          <year>September 2015</year>
          .
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Anna</given-names>
            <surname>Feldman</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jirka</given-names>
            <surname>Hana</surname>
          </string-name>
          .
          <article-title>A resource-light approach to morphosyntactic tagging</article-title>
          .
          <source>Rodopi</source>
          , Amsterdam/New York, NY,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Dan</given-names>
            <surname>Garrette</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jason</given-names>
            <surname>Baldridge</surname>
          </string-name>
          .
          <article-title>Learning a part-of-speech tagger from two hours of annotation</article-title>
          .
          <source>In Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , pages
          <fpage>138</fpage>
          -
          <lpage>147</lpage>
          , Atlanta, Georgia,
          <year>June 2013</year>
          .
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Yoav</given-names>
            <surname>Goldberg</surname>
          </string-name>
          and
          <article-title>Omer Levy. word2vec explained: Deriving Mikolov et al.'s negative-sampling word-embedding method</article-title>
          .
          <source>Technical report</source>
          , BenGurion University of the Negev,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Enrique</given-names>
            <surname>Henestroza</surname>
          </string-name>
          Anguiano and
          <string-name>
            <given-names>Marie</given-names>
            <surname>Candito</surname>
          </string-name>
          .
          <article-title>Parse correction with specialized models for difficult attachment types</article-title>
          .
          <source>In Proceedings of EMNLP-11</source>
          , pages
          <fpage>1222</fpage>
          -
          <lpage>1233</lpage>
          , Edinburgh,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Rebecca</surname>
            <given-names>Hwa</given-names>
          </string-name>
          , Philip Resnik, Amy Weinberg, Clara Cabezas, and
          <string-name>
            <given-names>Okan</given-names>
            <surname>Kolak</surname>
          </string-name>
          .
          <article-title>Bootstrapping parsers via syntactic projection across parallel texts</article-title>
          .
          <source>Natural Language Engineering</source>
          ,
          <volume>11</volume>
          :
          <fpage>311</fpage>
          -
          <lpage>325</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Mohammad</surname>
            <given-names>Khan</given-names>
          </string-name>
          , Markus Dickinson, and
          <string-name>
            <given-names>Sandra</given-names>
            <surname>Kübler</surname>
          </string-name>
          .
          <article-title>Does size matter? text and grammar revision for parsing social media data</article-title>
          .
          <source>In Proceedings of the Workshop on Language Analysis in Social Media</source>
          , Atlanta, GA USA,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Teresa</surname>
            <given-names>Lynn</given-names>
          </string-name>
          , Jennifer Foster,
          <string-name>
            <given-names>Mark</given-names>
            <surname>Dras</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Lamia</given-names>
            <surname>Tounsi</surname>
          </string-name>
          .
          <article-title>Cross-lingual transfer parsing for low-resourced languages: An Irish case study</article-title>
          .
          <source>In Proceedings of CLTW</source>
          <year>2014</year>
          , Dublin,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Teresa</surname>
            <given-names>Lynn</given-names>
          </string-name>
          , Jennifer Foster,
          <string-name>
            <given-names>Mark</given-names>
            <surname>Dras</surname>
          </string-name>
          , and Josef van Genabith.
          <article-title>Working with a small dataset - semi-supervised dependency parsing for Irish</article-title>
          .
          <source>In Proceedings of SPMRL</source>
          <year>2013</year>
          , Seattle,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Mitch</surname>
            <given-names>Marcus</given-names>
          </string-name>
          , Beatrice Santorini, and Mary Ann Marcinkiewicz.
          <article-title>Building a large annotated corpus of English: The Penn Treebank</article-title>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>19</volume>
          (
          <issue>2</issue>
          ):
          <fpage>313</fpage>
          -
          <lpage>330</lpage>
          ,
          <year>1993</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Andre</surname>
            <given-names>Martins</given-names>
          </string-name>
          , Miguel Almeida, and
          <string-name>
            <surname>Noah</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Smith.</surname>
          </string-name>
          <article-title>Turning on the turbo: Fast third-order non-projective turbo parsers</article-title>
          .
          <source>In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)</source>
          , pages
          <fpage>617</fpage>
          -
          <lpage>622</lpage>
          , Sofia, Bulgaria,
          <year>August 2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Ryan</surname>
            <given-names>McDonald</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kevin</given-names>
            <surname>Lerman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Fernando</given-names>
            <surname>Pereira</surname>
          </string-name>
          .
          <article-title>Multilingual dependency analysis with a two-stage discriminative parser</article-title>
          .
          <source>In Proceedings of CoNLL-X</source>
          , pages
          <fpage>216</fpage>
          -
          <lpage>220</lpage>
          , New York City,
          <year>June 2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Ryan</surname>
            <given-names>McDonald</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Joakim</given-names>
            <surname>Nivre</surname>
          </string-name>
          , Yvonne Quirmbach-Brundage, Yoav Goldberg,
          <string-name>
            <surname>Dipanjan Das</surname>
          </string-name>
          ,
          <string-name>
            <surname>Kuzman Ganchev</surname>
            , Keith Hall, Slav Petrov, Hao Zhang, Oscar Täckström, Claudia Bedini, Núria Bertomeu Castelló, and
            <given-names>Jungmee</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
          </string-name>
          .
          <article-title>Universal dependency annotation for multilingual parsing</article-title>
          .
          <source>In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)</source>
          , pages
          <fpage>92</fpage>
          -
          <lpage>97</lpage>
          , Sofia, Bulgaria,
          <year>August 2013</year>
          .
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Ryan</surname>
            <given-names>McDonald</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Slav</given-names>
            <surname>Petrov</surname>
          </string-name>
          , and
          <article-title>Keith Hall. Multi-source transfer of delexicalized dependency parsers</article-title>
          .
          <source>In Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing</source>
          , pages
          <fpage>62</fpage>
          -
          <lpage>72</lpage>
          , Edinburgh, Scotland,
          <year>July 2011</year>
          .
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Tomas</surname>
            <given-names>Mikolov</given-names>
          </string-name>
          , Ilya Sutskever, Kai Chen, Gregory S. Corrado, and
          <string-name>
            <given-names>Jeffrey</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          <volume>26</volume>
          , pages
          <fpage>3111</fpage>
          -
          <lpage>3119</lpage>
          ,
          <string-name>
            <surname>Lake</surname>
            <given-names>Tahoe</given-names>
          </string-name>
          ,
          <string-name>
            <surname>NV</surname>
          </string-name>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Joakim</surname>
            <given-names>Nivre</given-names>
          </string-name>
          , Johan Hall, Jens Nilsson, Atanas Chanev, Gulsen Eryigit, Sandra Kübler, Svetoslav Marinov, and Erwin Marsi.
          <article-title>MaltParser: A languageindependent system for data-driven dependency parsing</article-title>
          .
          <source>Natural Language Engineering</source>
          ,
          <volume>13</volume>
          (
          <issue>2</issue>
          ):
          <fpage>95</fpage>
          -
          <lpage>135</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>Mohammad</given-names>
            <surname>Sadegh Rasooli</surname>
          </string-name>
          and Michael Collins.
          <article-title>Density-driven crosslingual transfer of dependency parsers</article-title>
          .
          <source>In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing</source>
          , pages
          <fpage>328</fpage>
          -
          <lpage>338</lpage>
          , Lisbon, Portugal,
          <year>September 2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Mo</surname>
            <given-names>Shen</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Daisuke</given-names>
            <surname>Kawahara</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Sadao</given-names>
            <surname>Kurohashi</surname>
          </string-name>
          .
          <article-title>Dependency parse reranking with rich subtree features</article-title>
          .
          <source>IEEE/ACM Transactions on Audio, Speech, and Language Processing</source>
          ,
          <volume>22</volume>
          (
          <issue>7</issue>
          ):
          <fpage>1208</fpage>
          -
          <lpage>1218</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Noah</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Smith</surname>
            and
            <given-names>Jason</given-names>
          </string-name>
          <string-name>
            <surname>Eisner</surname>
          </string-name>
          .
          <article-title>Contrastive estimation: Training log-linear models on unlabeled data</article-title>
          .
          <source>In Proceedings of the 43rd Annual Meeting of the Association for Computational Linguistics (ACL'05)</source>
          , pages
          <fpage>354</fpage>
          -
          <lpage>362</lpage>
          , Ann Arbor, MI,
          <year>June 2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Noah</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Smith</surname>
            and
            <given-names>Jason</given-names>
          </string-name>
          <string-name>
            <surname>Eisner</surname>
          </string-name>
          .
          <article-title>Guiding unsupervised grammar induction using contrastive estimation</article-title>
          .
          <source>In International Joint Conference on Artificial Intelligence (IJCAI) Workshop on Grammatical Inference Applications</source>
          , pages
          <fpage>73</fpage>
          -
          <lpage>82</lpage>
          , Edinburgh,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>