<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Search Your Own Treebank</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alexandr Chernov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Erhard Hinrichs</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marie Hinrichs</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Linguistics University of Tübingen</institution>
        </aff>
      </contrib-group>
      <fpage>25</fpage>
      <lpage>34</lpage>
      <abstract>
        <p>This paper reports on the use of the treebank search and visualization tool TüNDRA for the purposes of inspecting linguistically annotated data that are generated by the web-based annotation tool WebLicht. The motivation for enriching WebLicht by the functionalities offered by TüNDRA is twofold: (i) it allows on-the-fly searches for specific linguistic phenomena at the word and sentence level, (ii) it provides on-the-fly visualizations of such phenomena.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>TCF (Text Corpus Format) [4] was developed for use as an internal data exchange
format. Annotation tools are wrapped as webservices that receive and return TCF.
WebLicht is a distributed system, where the annotation tool webservices are hosted
at CLARIN1 centers and are invoked via HTTP requests. Currently, WebLicht can
be used to process 11 different languages with over 100 annotation tools hosted at
9 CLARIN centers.</p>
      <p>WebLicht has been generally well accepted and has gained an increasing user
base over time, with approximately 700,000 invocations of its webservices in the
past calendar year. However, it still lacks the important feature of providing
straightforward search functionality on the resulting annotations. Such functionality
allows users to better explore their annotation results, enabling them to ask questions
about their data, such as:
• How was a particular word form annotated for part of speech?
• How many occurrences of a proper name were successfully annotated as a
named entity?
• How many occurrences of a particular syntactic construction were annotated
in the data set?</p>
      <p>Since TüNDRA provides precisely the type of querying functionality
illustrated by the above examples, and WebLicht provides the tools to easily create
custom, on-the-fly treebanks, the two applications have been more tightly coupled.
The remainder of this paper is structured as follows: Sections 2 and 3 describe
the state of visualization and exploration in WebLicht before and after TüNDRA
integration, respectively. Section 4 describes enhancements to TüNDRA,
including those required for WebLicht integration. Sections 5 and 6 describe related and
future work, respectively.
2</p>
    </sec>
    <sec id="sec-2">
      <title>WebLicht and TüNDRA Before Integration</title>
      <p>In order to better recognize the motivation and impact of the work described here, it
is necessary to understand the prior states of both WebLicht and TüNDRA in terms
of visualization and search functionality. This section gives some background
information about the two applications in isolation.
2.1</p>
      <sec id="sec-2-1">
        <title>WebLicht Before Integration</title>
        <p>
          Prior to the integration of TüNDRA into WebLicht, visualization of annotation
results were provided in ways that are appropriate for the individual annotation
layers. A table view was used for tokens, lemmas, part-of-speech tags, and
morphology annotations. Named entities were highlighted within the text using
colorcoding to distinguish between different types (person, location, organization, etc).
A graphical view is used for constituency parse trees, and dependency parse trees
were displayed using embedded brat [
          <xref ref-type="bibr" rid="ref4">14</xref>
          ] visualizations. Figure 1 shows examples
of these WebLicht visualizations.
        </p>
        <p>Before the integration with TüNDRA, WebLicht provided no direct query
functionality for annotation results. Although querying of annotation results produced
by WebLicht was in some cases possible, it was cumbersome and did not provide
a good user experience. Consider the procedure for performing a search on
annotations contained in the table view. First the table needed to be downloaded and
opened in external spreadsheet software, followed by use of the generic and rather
rudimentary search functionalities of the spreadsheet software.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>TüNDRA Before Integration</title>
        <p>Since TüNDRA was specifically designed for processing treebanks, much care was
taken to provide visualization and search support for any annotations that a
treebank may contain. TüNDRA’s flexibility enables it to support nearly any type of
treebank. No assumptions are made about what a treebank node represents or the
number or type of features it contains. Exactly this flexibility makes it easy for
TüNDRA to work with the dynamic data produced by WebLicht, which may or
may not have structural information.</p>
        <p>TüNDRA uses the query language Tiger [7, 9], which supports querying of
both constituent-based and dependency-based treebanks. The Tiger language
supports the querying of nodes and of edge labels in syntax graphs. Individual nodes
can be identified by hash-tag variables and further specified by Boolean
expressions of feature-value pairs. Tiger queries can make reference to the two primitive
node relations of precedence (.) and (labelled) dominance (&gt;). The dominance
relations can be further specified by particular edge labels. Node identifiers in a
query can, inter alia, be used to collect statistics on matches. Although some
understanding of the structure and of the features of a treebank are required in order
to form search queries, the required information can usually be gained by browsing
through the treebank itself or by the stylebook of a treebank, if such off-line
documentation is available. See section 3 for examples of using TüNDRA to query and
gather statistics on dynamic WebLicht data.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3 Integration of TüNDRA into WebLicht</title>
      <p>Using the newly integrated TüNDRA, it is possible to conveniently execute queries
and to gather statistics on annotations directly within the WebLicht application.
Queries are not restricted to parsed data, but can be executed on all levels of
annotation. However, structural queries can only be successful if the text has been
parsed, which is not always the case for WebLicht data. This section presents two
examples of using TüNDRA to query and gather statistics on WebLicht-generated
datasets.</p>
      <p>g
e n
r i
u d
g u
i l
F c
n
I
3.1</p>
      <sec id="sec-3-1">
        <title>Example 1: Searching for Named Entities in a German Corpus</title>
        <p>The following example demonstrates gathering statistics about named entities on
non-parsed data. The WebLicht data format includes an attribute for named
entity annotations for further categorizing them (e.g. person, location, etc). The
new TüNDRA implementation makes it possible to search for named entities of a
particular type. The query in (1), executed with the "statistics" option, finds the
frequency and percentage of geo-political named entities:
( 1 ) # ne : [ _ne ="GPE " ]</p>
        <p>This query was executed on a text that was automatically annotated with named
entities in WebLicht. The text was about the physicist Hans Geiger and the results
of the statistical query can be seen in the table in Figure 2.</p>
        <p>In addition to the benefit of enabling exploration of WebLicht data, the
integration of TüNDRA into WebLicht leads to a unified presentation of data in the
applications. Since TüNDRA was designed to process parsed text, it is particularly
advantageous to have its constituency-based and dependency-based visualizations
in WebLicht. Figure 2 shows parse tree visualizations as they appear in both
WebLicht and TüNDRA.</p>
        <p>VROOT</p>
        <p>PX</p>
        <p>NX
NX</p>
        <p>NX
VF
NX</p>
        <p>SIMPX</p>
        <p>LK</p>
        <p>VXFIN
1904 verbrachte
pos: CARD pos: VVFIN</p>
        <p>ROOT</p>
        <p>NX
ADV</p>
        <p>MF</p>
        <p>ADVX
OBJA</p>
        <p>NX</p>
        <p>PP
er auch ein Semester
pos: PPER pos: ADV pos: ART pos: NN</p>
        <p>an der
pos: APPR pos: ART</p>
        <p>Ludwig-Maximilians-Universität München
pos: NN pos: NE</p>
        <p>.
pos: $.</p>
        <p>ZEIT</p>
        <p>SUBJ</p>
        <p>DET</p>
        <p>PN</p>
        <p>DET</p>
        <p>APP
-PUNCT1904 verbrachte er auch ein Semester an der
pos: CARD pos: VVFIN pos: PPER pos: ADV pos: ART pos: NN pos: APPR pos: ART
Ludwig-Maximilians-Universität München
pos: NN pos: NE</p>
        <p>.
pos: $.
Since the WebLicht suite of services is not limited German, but also includes tools
for English, it is also possible to annotate and query English corpus data. The
query used in section 3.1 is simple in the sense that it only searches for a single
class of lexical tokens. However, TüNDRA supports the full expressiveness of the
Tiger query language, thus allowing also queries that involve multiple constituents
and/or lexical tokens. For example, lexicographers and lexical semanticists may be
interested in the set of prepositions that co-occur with a particular verb in a given
corpus. The query in (2) exemplifies such a query for the English verb agree in a
constituency-based treebank.
( 2 ) # vp : [ c a t ="VP " ] &gt; [ lemma =" a g r e e " ]</p>
        <p>&amp; # vp &gt; [ c a t ="PP " ] &gt; #p : [ pos =" IN " ]
This query searches for VP nodes (#vp) which dominate a lexical node with lemma
agree and a PP node, which in turn dominates a lexical node with part-of-speech
IN, the label used for prepositions in the Penn treebank tagset [8]. Since this lexical
node is identified by the hashtag variable #p, statistics on the word forms for this
lexical node can be gathered.</p>
        <p>The same type of query can be created for execution on a dependency treebank.
The following query finds lexical nodes with lemma agree and with an edge,
labelled with relation name VMOD and pointing to a lexical node with part-of-speech
label IN.
( 3 ) [ lemma =" a g r e e " ] &gt; VMOD #p : [ pos =" IN " ]
In order to complete the integration of TüNDRA into WebLicht, it was necessary to
(i) accommodate the WebLicht data format and (ii) refactor the code into front-end
and back-end components.</p>
        <p>Before WebLicht data can be used in TüNDRA, it must first be converted into
the TüNDRA internal format. This is a straightforward process for most
annotations at the token level, with the exception of named entities, which can span more
than one token. They are handled by adding a special "_ne" attribute. At the
structural level, the case where parse annotations are not present in the data must be</p>
        <p>the
pos: DT
lemma: the
text: the</p>
        <p>court .
pltoeesmx:mt:aN:Ncocourutrt ptloeesxm:tm:a.:..</p>
        <p>NP</p>
        <p>I
pos: PRP
lemma: I
text: I</p>
        <p>I
pos: PRP
lemma: I
text: I
vp
VP</p>
        <p>PP
p
agree with
pltoeesmx:mt:aV:BaPgargereee ptloeesxm:tm:aI:Nwiwtihth</p>
        <p>the
pos: DT
lemma: the
text: the</p>
        <sec id="sec-3-1-1">
          <title>ROOT SUB VMOD</title>
          <p>NP</p>
          <p>NP
decision
pos: NN
lemma: deci
sion
text: decis
ion</p>
          <p>P</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>NMOD</title>
        </sec>
        <sec id="sec-3-1-3">
          <title>PMOD</title>
        </sec>
        <sec id="sec-3-1-4">
          <title>NMOD</title>
        </sec>
        <sec id="sec-3-1-5">
          <title>PMOD</title>
          <p>NMOD
agree witph
tpleoexsmt:m:aV:BaPgargereee tlpeeoxmstm::aI:Nwiwtihth</p>
          <p>the
pos: DT
lemma: the
text: the
decision
pos: NN
lemma: deci
sion
text: decis
ion</p>
          <p>of
pos: IN
lemma: of
text: of</p>
          <p>the
pos: DT
lemma: the
text: the</p>
          <p>court .
ltpeeomxsmt::aN:Ncocourutrt tlpeeoxmstm::a.:..
handled. This is done by creating "fake" tree structures with all tokens attached
directly to the root, allowing them to be processed like all other trees in TüNDRA. In
addition, a table view, similar to that which was previously available in WebLicht,
but not in TüNDRA, has been incorporated into TüNDRA.</p>
          <p>It was necessary to refactor the TüNDRA code into well-defined front- and
back-ends. The front-end is used by both applications. This is done in WebLicht
by simply replacing the prior visualization component with TüNDRA’s front-end.
The front-end in turn communicates with the back-end which does any necessary
data conversion and performs queries. This clean division of labor into the user
interface (front-end) and the query-processing engine (back-end) made it possible
for both applications to share the same visualization and search component.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5 Related Work</title>
      <p>
        There are many well-known treebank search and visualization tools, such as INESS
[11], PML-TQ [
        <xref ref-type="bibr" rid="ref3">13</xref>
        ], GrETEL [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], and ICARUS [2]. Although some of them
provide upload or import options for processing personal treebanks (PML-TQ,
ICARUS), to our knowledge none of them are tightly integrated into a workflow
engine that allows on-the-fly annotation using custom-built pipelines. INESS has
rich support for search and visualization of hosted treebanks, including comparison
of parallel corpora, but only authorized users can run annotation pipelines.
GrETEL is a treebank search engine which is very easy to use due to its novel way of
guiding the query building process, but only supports hosted treebanks.
6
      </p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion and Future Work</title>
      <p>In this paper we have shown how the treebank visualization and search application
TüNDRA has been integrated into the annotation workflow engine WebLicht,
making it possible to apply TüNDRA’s full query/visualization/statistics capabilities to
on-the-fly treebanks created in WebLicht. It was necessary to refactor large parts
of the TüNDRA code to make its front-end available in WebLicht. At the time of
writing, the new versions of both applications are in beta phase.</p>
      <p>The main focus of future work in TüNDRA are in the areas of query building
and statistics views. Query building needs to be simplified for users who are
unfamiliar with the query language. This can be done by offering graphical guidance
and limiting elements of the query to valid values where possible. It is also planned
to provide more detailed statistics views, including more visualization options and
allowing more in-depth exploration of the statistics. A version of TüNDRA that
can be run and administered locally is also planned, enabling the local use of
treebanks that, for example, cannot be hosted by the public version of TüNDRA for
legal reasons. In the near future, the public treebanks hosted by TüNDRA will be
made available without the need for logging in.</p>
      <p>Future work on WebLicht includes providing a batch mode for more
convenient processing of very large texts. This goes beyond what WaaS can already do
(executing chains from the command line or programming code) by splitting up
very large texts into smaller chunks for processing if necessary, invoking the tool
chain in parallel on the smaller chunks, piecing it all back together, and storing the
finished result for later download. Users will be able to monitor the progress of
their jobs and will be notified when it is finished.
[2] Markus Gärtner, Gregor Thiele, Wolfgang Seeker, Anders Björkelund and
Jonas Kuhn (2013). ICARUS – An Extensible Graphical Search Tool for
Dependency Treebanks. In: Proceedings of the 51st Annual Meeting of the
Association for Computational Linguistics: System Demonstrations. Sofia,
Bulgaria.
[3] Dirk Goldhahn, Thomas Eckart und Uwe Quasthoff (2012). Building Large
Monolingual Dictionaries at the Leipzig Corpora Collection: From 100 to
200 Languages. In: Proceedings of the 8th International Conference on
Language Resources and Evaluation (LREC-2012). Istanbul, Turkey. European
Language Resources Association (ELRA).
[4] Ulrich Heid, Helmut Schmid, Kerstin Eckart, and Erhard Hinrichs (2010).</p>
      <p>A corpus representation format for linguistic web services: The d-spin text
corpus format and its relationship with iso standards. In: Proceedings of the
7th International Conference on Language Resources and Evaluation (LREC
2010). Valletta, Malta. European Language Resources Association (ELRA).
[5] Erhard Hinrichs, Marie Hinrichs, and Thomas Zastrow (2010). WebLicht:
Web-Based LRT Services for German. In: Proceedings of the Systems
Demonstrations at the 48th Annual Meeting of the Association for
Computational Linguistics (ACL-2010). pp. 25-29. Uppsala, Schweden.
[6] Daniël de Kok, Dörte de Kok, and Marie Hinrichs (2014). Build your own
treebank. In: Proceedings of the CLARIN Annual Conference. Soesterberg,
The Netherlands
[7] Wolfgang Lezius (2002). TIGERSearch - Ein Suchwerkzeug für
Baumbanken. In: Proceedings der 6. Konferenz zur Verarbeitung natürlicher
Sprache (KONVENS 2002). Saarbrücken.
[8] Mitchell P. Marcus, Beatrice Santorini, and Mary Ann Marcinkiewicz (1993).</p>
      <p>Building a Large Annotated Corpus of English: The Penn Treebank
Computational Linguistics, 19.2, pp. 313–330.
[9] Scott Martens (2012). TüNDRA: TIGERSearch-style treebank querying as an
XQuery-based web service. In: Proceedings of the joint CLARIN-D/DARIAH
Workshop "Service-oriented Architectures (SOAs) for the Humanities:
Solutions and Impacts" (DH 2012). Hamburg, pp. 41-50.
[10] Scott Martens (2013). TüNDRA: A Web Application for Treebank Search
and Visualization. In: Proceedings of The Twelfth Workshop on
Treebanks and Linguistic Theories (TLT12). Sofia, pp. 133—144. (URL:
http://bultreebank.org/TLT12/TLT12Proceedings.pdf)
[11] Paul Meurer (2012). INESS-Search: A Search System for LFG (and Other)
Treebanks. In: Miriam Butt and Tracy Holloway King, editors, Proceedings
of the LFG2012 Conference. CSLI Publications.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Liesbeth</given-names>
            <surname>Augustinus</surname>
          </string-name>
          , Vincent Vandeghinste, and Frank Van Eynde (
          <year>2012</year>
          ).
          <article-title>Example-based Treebank Querying</article-title>
          .
          <source>In: Proceedings of the 8th International Conference on Language Resources</source>
          and
          <article-title>Evaluation (LREC-</article-title>
          <year>2012</year>
          ). Istanbul, Turkey.
          <source>European Language Resources Association (ELRA).</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Joakim</surname>
            <given-names>Nivre</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marie-Catherine de Marneffe</surname>
            , Filip Ginter, Yoav Goldberg, Jan Hajic,
            <given-names>Christopher D.</given-names>
          </string-name>
          <string-name>
            <surname>Manning</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ryan</surname>
            <given-names>McDonald</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Slav</given-names>
            <surname>Petrov</surname>
          </string-name>
          , Sampo Pyysalo, Natalia Silveira, Reut Tsarfaty and Daniel
          <string-name>
            <surname>Zeman</surname>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>Universal Dependencies v1: A Multilingual Treebank Collection</article-title>
          .
          <source>In: Proceedings of the 10th International Conference on Language Resources and Evaluation (LREC</source>
          <year>2016</year>
          ).
          <article-title>European Language Resources Association (ELRA).</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Petr</given-names>
            <surname>Pajas and Jan Šteˇpánek</surname>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>System for Querying Syntactically Annotated Corpora</article-title>
          .
          <source>In: Proceedings of the ACL-IJCNLP 2009 Software Demonstrations</source>
          , pages
          <fpage>33</fpage>
          -
          <lpage>36</lpage>
          , Suntec, Singapore. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Pontus</surname>
            <given-names>Stenetorp</given-names>
          </string-name>
          , Sampo Pyysalo, Goran Topic´,
          <string-name>
            <surname>Tomoko</surname>
            <given-names>Ohta</given-names>
          </string-name>
          , Sophia Ananiadou and
          <string-name>
            <surname>Jun'ichi Tsujii</surname>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>brat: a Web-based Tool for NLP-Assisted Text Annotation</article-title>
          .
          <source>In: Proceedings of the Demonstrations Session at EACL</source>
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>