<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Biographical data exploration as a test-bed for a multi-view, multi-method approach in the Digital Humanities</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andre´ Blessing</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrea Glaser</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jonas Kuhn</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute for Natural Language Processing (IMS) Universita ̈t Stuttgart Pfaffenwaldring 5b</institution>
          ,
          <addr-line>70569 Stuttgart</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>NDB - Neue Deutsche Biographie</institution>
          ,
          <addr-line>New German Biography</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>1815</year>
      </pub-date>
      <fpage>53</fpage>
      <lpage>60</lpage>
      <abstract>
        <p>The present paper has two purposes: the main point is to report on the transfer and extension of an NLP-based biographical data exploration system that was developed for Wikipedia data and is now applied to a broader collection of traditional textual biographies from different sources and an additional set of structured biographical resources, also adding membership in political parties as a new property for exploration. Along with this, we argue that this expansion step has many characteristic properties of a typical methodological challenge in the Digital Humanities: resources and tools of different origin and with different accuracy are combined for use in a multidisciplinary context. Hence, we view the project context as an interesting test-bed for some methodological considerations.</p>
      </abstract>
      <kwd-group>
        <kwd>information extraction</kwd>
        <kwd>visualization</kwd>
        <kwd>digital humanities</kwd>
        <kwd>exploration system</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        CLARIN1 is a large infrastructure project and has the
mission to advance research in the humanities and social
sciences. Scholars should be able to understand and exploit
the facilities offered by CLARIN
        <xref ref-type="bibr" rid="ref19">(Hinrichs et al., 2010)</xref>
        without technical obstacles. We developed a showcase
        <xref ref-type="bibr" rid="ref11 ref28 ref4">(Blessing and Kuhn, 2014)</xref>
        , which is called TEA2 (Textual
Emigration Analysis), to demonstrate how CLARIN can
be used in a web-based application. The previously
published version of the showcase was based on two data sets:
a data set from the Global Migrant Origin Database, and a
data set which was extracted from the German Wikipedia
edition. The idea for the chosen scenario was to enable
researchers of the humanities access to large textual data.
This approach is not limited to the extraction of
information, it also integrates interaction and visualization of the
results. In particular, transparency is an important aspect to
satisfy the needs of the researcher of the humanities. Each
result must be inspectable. In this work we integrate two
new data sets into our application:
Furthermore we investigate new relations which are of high
interest to researchers of the humanities, for example, if a
person is or was a member of a party, company or a
corporate body.
      </p>
      <p>Next, we view the project context as an interesting test-bed
for some methodological considerations.</p>
      <p>1http://clarin.eu
2http://clarin01.ims.uni-stuttgart.de/geovis/showcase.html</p>
    </sec>
    <sec id="sec-2">
      <title>1.1. The Exemplary Character of Biographical Data</title>
    </sec>
    <sec id="sec-3">
      <title>Exploration</title>
      <p>
        The use of computational methods in the Humanities bears
an enormous potential. Obviously, moving representations
of artifacts and knowledge sources to the digital medium
and interlinking them provides new ways of integrated
exploration. But while this change of medium could be
argued to “merely” speed up the steps a scholar could in
principle take with traditional means, there are
opportunities that clearly expand the traditional methodological
spectrum, (a) through interaction and sharing among scholars,
potentially from quite different fields (e.g., shared
annotations
        <xref ref-type="bibr" rid="ref8">(Bradley, 2012)</xref>
        ), and (b) through scaling to a
substantially larger collection of objects of study, which can
undergo exploration and qualitative analysis, and of course
quantitative analysis
        <xref ref-type="bibr" rid="ref23 ref33">(Moretti, 2013; Wilkens, 2011)</xref>
        .
However, these novel avenues turn out to be very hard to
integrate into established disciplinary frameworks, e.g., in
literary or cultural studies, and from the point of view of
scholarly less erudite computational scientists, it often
appears that the scaling potential of computational analysis
and modeling is heavily under-explored
        <xref ref-type="bibr" rid="ref26 ref27">(Ramsay, 2003;
Ramsay, 2007)</xref>
        . It is important to understand what is
behind this rather reluctant adoption. Our hypothesis is that
humanities scholars perceive a lack of control over the
scalable analytical machinery and should be placed in a
position to apply fully transparent computational models
(including imperfect automatic analysis steps) that invite for
critical reflection and subsequent adaptation.3 An
orthogonal issue lies in the fact that advanced scholarly research
tends to target resources and artifacts that have not
previously been made accessible and studied in detail. So the
digitization process takes up a considerable part of a
typical project and a bootstrapping cycle of computational tools
and models (as it is common in methodologically oriented
projects in the computational sciences) cannot be applied
3The bottom-up approach laid out in
        <xref ref-type="bibr" rid="ref10 ref2 ref3">(Blanke and Hedges,
2013)</xref>
        seems an effective strategy to counteract this situation.
      </p>
      <sec id="sec-3-1">
        <title>Wikipedia ÖBL NDB</title>
        <sec id="sec-3-1-1">
          <title>VIEWS</title>
          <p>NLP
PIPELINE
unstructured sources
structured sources</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>DATA MODEL</title>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>DH scholar geo-centric entity-centric statistic-centric</title>
        <p>on datasets that are sufficiently relevant to the actual
scholarly research question. We believe that biographical data
exploration is an excellent test-bed for pushing forward a
scalability-oriented program in the Digital Humanities: the
compilation of biographical information collections from
heterogeneous sources has a long tradition, and every user
of traditional, printed resources of this kind is aware of
the trade-off between the benefit of large coverage and the
cost of high reliability and depth of individual entries. In
other words, the intricacies that come from scalable
computational models (concerning reliability of data extraction
procedure, granularity and compatibility of data models,
etc.) have pre-digital predecessors, and an exploration
environment may invite to a competent negotiation of these
factors. Here, a very natural multiple view presentation in
a digital exploration platform can bring in a great deal of
transparency: with a brushing-and-linking approach, users
can go back and forth between an entity-centered view on
biographical data (starting out from individuals or a
visualization of tangible aggregates, e.g., by geographical or
temporal affinity) and the sources from which information was
extracted (e.g., natural language text passages or (semi-)
structured information sources). This readily invites to a
critically reflected use of the information. Methodological
artifacts tend to stand out in aggregate presentations along
an independent dimension, and it does not take specialist
knowledge to identify systematic errors (e.g., in an
underlying NLP component), which can then be fixed in an
interactive working environment. Lastly, an important aspect
besides this model character in terms of the interplay of
resources and computational components and the natural
options for multi-view visualization is the relevance of
biographical collections to multiple different disciplines in the
humanities and social sciences. Hence, sizable resources
are already available and are being used, and it is likely that
improved ways of providing access to such collections and
encouraging interactive improvements of reliability,
coverage and connectivity will actually benefit research in
various fields (and will hence generate feedback on the
methodological questions we are raising).</p>
        <p>
          We are not the first who work on the exploration of different
biographical data sets. The BiographyNet project
          <xref ref-type="bibr" rid="ref16 ref24">(Fokkens
et al., 2014; Ockeloen et al., 2013)</xref>
          tackles similar questions
on reliability of resources, significance of derived output,
and how results can be adjusted to improve performance
and acceptance.
        </p>
        <p>System Overview
Figure 1 shows the architecture of our approach. The
system integrates different biographic data sources (top left).
Additional biographic data sources can be integrated if they
are based on textual data. Textual sources are processed
by the NLP pipeline (top middle) which will be explained
in the next section. In addition to textual data, structured
del
mo IMS type system
data</p>
        <p>TCF</p>
        <p>Wrapper
s
e
l
u
d
-om eFUxetIarMatucArteo-r
A
M
I
U</p>
        <p>ClearTK
Converters
Tokenizer</p>
        <p>Tagger</p>
        <p>Parser
Named Entity</p>
        <p>Recognizer</p>
        <p>CLARIN
web services</p>
        <p>TCF exchange
format
data sets (top right) are used to enable real world inference
(e.g. mapping extracted knowledge to a world map). We
discuss the used structured data set in more detail later on.
The data model (middle) central to our system includes the
derived and extracted data and additionally all links to the
sources. This enables transparency by providing access to
the whole processing pipeline. Finally, several views of the
data model (bottom) are provided. These allow the user
to visualize the obtained data in different ways. A specific
view can be used depending on the actual research question.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>2.1. NLP Pipeline</title>
      <p>
        Natural Language Processing (NLP) is typically done by
chaining several tools as a pipeline. The right hand part
of Figure 2 shows some basic tools
        <xref ref-type="bibr" rid="ref21">(Mahlow et al., 2014)</xref>
        which are necessary. This pipeline includes normalization,
sentence segmentation, tokenizing, part-of-speech tagging,
coreference resolution, and named entity recognition. An
important property is that these components are not rigidly
combined. This allows the user to adjust or substitute single
components if the performance of the whole system is not
sufficient. The system is also language independent insofar
as all NLP tools in one language can be replaced by tools in
other languages. Table 1 gives more details about the used
versions. These services are designed to process big data
and do not require local installation of linguistic tools. This
is often time consuming since most tools are using different
input and output formats which have to be adapted.
      </p>
    </sec>
    <sec id="sec-5">
      <title>2.2. Data Model</title>
      <p>The data model of our system has to fit several
requirements: i) store textual data and linguistic annotations; ii)
enable interlinking and exploration of data; iii) aggregate
results for visualization and data export; iv) store process
meta data.</p>
      <p>
        CLARIN-D provides its own data format called TCF
        <xref ref-type="bibr" rid="ref18">(Heid
et al., 2010)</xref>
        which is designed for efficient processing
with minimal overhead. But, such a format is not
adequate as core data model for an application. We decided
to use the Unstructured Information Management
Architecture (UIMA) framework
        <xref ref-type="bibr" rid="ref14">(Ferrucci and Lally, 2004)</xref>
        as data
model. The core of UIMA provides a data-driven
framework for the development and application of NLP
processing systems. It provides a customized annotation scheme
which is called type system. This type system is flexible
and makes it possible to integrate one’s own annotation
on different layers (e.g. part of speech tags, named
entities) in the UIMA framework. It is also possible to keep
track of existing structured information (e.g. hyperlinks in
Wikipedia articles or highlighted phrases in a
biographical lexicon) as the original text’s own annotation in UIMA.
Automatic annotation components are called analysis
engines in the UIMA systems. Each of these engines has to
be defined by a description language which includes the
enumeration of all input and output types. This allows us to
chain different engines including validation checks. UIMA
is a well accepted data model framework, especially since
the most popular UIMA-based application, which is called
Watson
        <xref ref-type="bibr" rid="ref15">(Ferrucci et al., 2010)</xref>
        , won in the US show
“Jeopardy” against human competitors. The flexible type system
also enables the split of content-based annotation and
process meta data annotations
        <xref ref-type="bibr" rid="ref11 ref21 ref28 ref4">(Eckart and Heid, 2014)</xref>
        which
allows keeping track of the processing history including
versioning. Such tracking of process meta data can also
be seen as provenance modeling
        <xref ref-type="bibr" rid="ref24">(Ockeloen et al., 2013)</xref>
        .
The combination of UIMA and TCF is simple since only
a single bridge annotation engine is needed to map both
annotation schemata. ClearTK is used as machine
learning (ML) interface
        <xref ref-type="bibr" rid="ref25">(Ogren et al., 2008)</xref>
        . It integrates
several ML algorithms (e.g. Maximum Entropy
Classification). The extraction of relevant features is a customized
component of the ClearTK framework. The used features
are described in Blessing and Schu¨tze (2010). At the
current stage a standard feature set is used (e.g. part-of-speech
tags, dependency paths, lemma information).
      </p>
    </sec>
    <sec id="sec-6">
      <title>2.3. Textual Emigration Analysis</title>
      <p>After the abstract definition of the requirements and
architecture we give a more detailed view of the the extended
TEA-tool. As mentioned before, we are using the already
PID which refers to the CMDI description of the service
http://hdl.handle.net/11858/00-247C-0000-0007-3736-B
http://hdl.handle.net/11858/00-247C-0000-0022-D906-1
http://hdl.handle.net/11858/00-247C-0000-0007-3735-D
http://hdl.handle.net/11858/00-247C-0000-0022-DDA1-3
http://hdl.handle.net/11858/00-247C-0000-0007-3734-F
deployed web-based application that allows researchers to
make quantitative and qualitative statements on persons
who emigrated to other countries. The visualization of the
results on a map helps to understand spatial aspects of the
emigration paths, for example, if people mostly emigrate to
nearby regions on the same continent or if they are spread
over the whole world. The visualization contains a second
view which aggregates and sums the emigration between
two countries. The aggregated numbers can be inspected in
a third view. Thereby, each number is decomposed by all
persons who are part of the given emigration path. Not only
the person names are shown, but the whole sentence stating
this emigration can be visualized. In the expert mode such
sentences can also be marked as correct or wrong by the
user to increase the performance of the system through
retraining or active learning. For more technical details on
the base system please consider Blessing and Kuhn (2014).
The extended application, which contains the two new data
sets, is shown in Figure 3. In this example the Austrian
Biographical Data is used as data origin. The user selected the
country Germany, and the extended system returned all
persons who emigrated from Germany to other countries. This
information is represented by arcs on the map and as a table
at the bottom of the screen. A key feature of the
application is that each number can be grounded to the underlying
text snippets. This allows users interested in e.g., the two
persons that emigrated from Germany to the US to click on
the details to open an additional view that lists all persons
including the sentence which describes the emigration.
The three view types, geo-driven, text-driven and
quantitative-driven of the TEA-application helps to explore
the data set from different perspectives which allows
researchers to identify inconsistencies. For example, the
geodriven view can be used to compare emigrations in a region
by selecting adjacent countries. Such an analysis helps to
find systematic geo-mapping errors (e.g. former USSR and
the Baltic states). In contrast the text-driven view enables
the identification of errors caused by NLP.</p>
    </sec>
    <sec id="sec-7">
      <title>2.4. Challenges for extension of the TEA-system</title>
      <p>To allow a smooth integration of the new biographic data
sets, a few modifications in the NLP pipeline were needed.
First, the import methods had to be adapted to allow the
extraction of the textual elements from the new XML or
HTML files. Second, the text normalization component had
to be adjusted on biographic texts, because O¨BL or NDB
use a lot more abbreviations which had to be resolved. This
could easily be done using a list of abbreviations provided
by the NDB website.</p>
      <p>
        The integration of a new relation was more challenging: a
new relation extraction component had to be defined and
trained. For the emigration relation the whole process was
done manually which is very time consuming. For the
member-of-party relation we switched to a new system
currently under development called ’extractor creator’. Since
the system is in an early stage of engineering, the
memberof-party relation was used as a development scenario.
Figure 4 shows a screenshot of the extractor creator. Some of
the basic methods of the interactive relation extraction
component were published in Blessing et al. (2012) and
Blessing and Schu¨tze (2010). The novelty in the new system
is that more background knowledge is integrated by using
person identifiers (based on the German Integrated
Authority File - GND) and Wikidata
        <xref ref-type="bibr" rid="ref12">(Erxleben et al., 2014)</xref>
        . This
leads to a more effective filtering in the search which
increases the performance of the whole system. The given
example in Figure 4 shows the lookup of specific persons
and the listing of all mentioned Ko¨rperschaften (corporate
bodies) which are mentioned in the same Wikipedia article.
A click on one of the corporate bodies opens the table on
the right which lists all person who also mention this
corporate body. A mouse-over function allows the user to see the
textual context of the mention. The human instructor can
then add relevant sentences as positive or negative training
examples.
      </p>
      <p>The first results of the novel relation extractor showed that
unlike the emigration relation a more fine-grained syntactic
feature set is needed in the scenario of corporate bodies.
Figure 5 shows a simplified example that includes negations
which occurred only rarely in the emigration scenario.</p>
    </sec>
    <sec id="sec-8">
      <title>2.5. Entity disambiguation</title>
      <p>Along with the extension of the core TEA system, we
perform experiments with special disambiguation techniques
that address named entities with multiple candidate
referents. Often, people playing some role in a biography are
mentioned very briefly, so unless the name is very rare,
machine learning methods for picking the correct person
have a hard time due to the very limited context. Many
approaches rely on extracted features to learn something
specific about people with ambiguous names, which requires
enough training data. In our approach we use topic
models for characteristic properties of the candidate referents.
These properties can be for example nationalities,
professions, or activities a person is involved in. We also
apply topic models to the context of an ambiguous person in
the biography and use the extracted properties to compute
the similarity to the candidate referents. We then create a
target-oriented candidate ranking.</p>
      <p>3.</p>
      <p>Experiments
The largest data set consists of articles about persons which
were extracted from the German Wikipedia edition. It
covers 250,360 persons after filtering by the German Integrated
Authority File (GND). The NDB data set contains 22,149
persons and the O¨ BL data set 18,428 persons. Figure 6
depicts the overlap of the used data sets. Only 1,147 persons
are part of all three data sets. We extracted 12,402 instances
of the emigration relation from the Wikipedia person data
set. For the NBD data set we found 1,932 instances of this
relation and for the O¨ BL data set we extracted 1,188
instances. Most of the persons found in Wikipedia are neither
part of NDB or O¨ BL which lead to the higher number of
Wikipedia emigrations. Moreover, the overlap of all three
data sets is small, meaning that we only have a few cases
in which a person who emigrated is represented in all three
data sets. An automatic comparison of the found instance
for emigration is only possible to a limited extent since the
different textual representations are not parallel for all facts.
The member-of-party extraction is at an early development
stage. Its performance has a high accuracy but the coverage
is low. We started to use Wikidata for evaluation purposes
since it also contains the same relation. However, the first
results showed that Wikidata is not complete enough to be
a sustainable gold standard. This observation was made by
manually evaluating the membership relation in the Social
Democratic Party of Germany. In this evaluation scenario
our extractor found 18 persons which were not represented
in Wikidata. This constitutes 20 percent of the extracted
data. As a consequence, we need a larger manually
annotated data set to enable a valid evaluation on precision.
Both experiments give evidence that we reached our first
goal, which can be seen as a proof-of-concept. The chosen
scenarios are not sufficient to enable an exhaustive
evaluation since we have no well-defined gold standard data sets.
However, components like the relation extraction provide
enough parameters for optimization in the future.</p>
      <p>
        Related Work
Since the Message Understanding Conferences
        <xref ref-type="bibr" rid="ref17">(Grishman
and Sundheim, 1996)</xref>
        in the 1990s, Information Extraction
(IE) is an established field in NLP research. Chiticariu
22N,1D4B9
et al. (2013) presented a study that shows that IE is
addressed in completely different way in research than in
industry. They showed that 75 percent of NLP papers
(20032012) are using machine learning techniques and only 3.5
percent are using rule-based systems. In contrast, 67
percent of the commercial IE systems are using rule-based
approaches
        <xref ref-type="bibr" rid="ref20">(Li et al., 2012)</xref>
        . One reason is the economic
efficiency of rule-based systems which are expensive in
development since the rules are hand crafted but later on the are
very efficient without needing huge computational power
and resources. For researchers such systems are not as
attractive since their goals are different by working on clean
gold standard data sets which allow exhaustive evaluation
by comparing precision and recall numbers. In our system,
we experimented with both, ML-based and rule-based
approaches. Rule-based systems have the big advantage to
provide transparency to the end users. On the other hand,
small changes on the requested relations need a complete
rewriting of the rules. We believe that a hybrid approach
which allows the definition of some rule-based constraints
to correct the output of supervised systems are the systems
which provide the highest acceptance.
      </p>
      <p>
        The drawback of ML-based IE systems
        <xref ref-type="bibr" rid="ref1 ref32">(Agichtein and
Gravano, 2000; Suchanek et al., 2009)</xref>
        is the need of expensive
manually annotated training data. There are unsupervised
approaches
        <xref ref-type="bibr" rid="ref22 ref9">(Mausam et al., 2012; Carlson et al., 2010)</xref>
        to avoid training data but then the semantics of the
extracted information is often not clear. Especially, for DH
researchers, which have a clear definition of the
information to extract, this is not feasible.
      </p>
      <p>
        Another requirement of DH scholars is that they want to use
complete systems which are often called end-to-end
systems. PROPMINER
        <xref ref-type="bibr" rid="ref2">(Akbik et al., 2013)</xref>
        is such a system
which uses deep-syntactic information. For our use case
such a system is not sufficient since they do not provide
several views on the data which also a big factor for the
usability of system in the DH community.
      </p>
      <p>5.</p>
      <p>Conclusion
We presented extensions of an experimental system for
NLP-based exploration of biographical data. Merging data
sources that have non-empty intersections provides an
important access for quality control.</p>
      <p>Offering multiple views for data exploration turns out
useful, not only from a data gathering perspective, but quite
importantly also as a way of inviting users to keep a critical
distance from the presented results. Methodological
artifacts that originate from NLP errors or other problems tend
to stand out in one of the aggregate visualizations.
5.1.</p>
    </sec>
    <sec id="sec-9">
      <title>Outlook</title>
      <p>We are collaborating with scholars of different fields of the
humanities that are interested to use our system.
Common questions are, which persons had certain positions at
what time? Which persons are members of organizations or
smaller groups at the same time? Which persons did their
education at the same institutions? We will incrementally
integrate such relation extractors in our system and observe
the user experience. The mixture of data aggregation and
being transparent is one of the crucial task to gain a high
acceptance from DH scholars. We will also evaluate which
additional factors are relevant for the acceptance of such a
system.</p>
      <p>Acknowledgements
We thank the anonymous reviewers for their valuable
questions and comments. This work is supported by
CLARIND (Common Language Resources and Technology
Infrastructure, http://de.clarin.eu/), funded by the German
Federal Ministry for Education and Research (BMBF) and by
a Nuance Foundation Grant.</p>
      <p>6.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Eugene</given-names>
            <surname>Agichtein</surname>
          </string-name>
          and
          <string-name>
            <given-names>Luis</given-names>
            <surname>Gravano</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Snowball: Extracting relations from large plain-text collections</article-title>
          .
          <source>In Proceedings of the 5th ACM Conference on Digital Libraries</source>
          , pages
          <fpage>85</fpage>
          -
          <lpage>94</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Alan</given-names>
            <surname>Akbik</surname>
          </string-name>
          , Oresti Konomi, and
          <string-name>
            <given-names>Michail</given-names>
            <surname>Melnikov</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Propminer: A workflow for interactive information extraction and exploration using dependency trees</article-title>
          .
          <source>In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics: System Demonstrations</source>
          , pages
          <fpage>157</fpage>
          -
          <lpage>162</lpage>
          , Sofia, Bulgaria, August. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Tobias</given-names>
            <surname>Blanke</surname>
          </string-name>
          and
          <string-name>
            <given-names>Mark</given-names>
            <surname>Hedges</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Scholarly primitives: Building institutional infrastructure for humanities e-science</article-title>
          .
          <source>Future Generation Computer Systems</source>
          ,
          <volume>29</volume>
          (
          <issue>2</issue>
          ):
          <fpage>654</fpage>
          -
          <lpage>661</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Andre</given-names>
            <surname>Blessing</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jonas</given-names>
            <surname>Kuhn</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Textual Emigration Analysis (TEA)</article-title>
          .
          <source>In Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC'14)</source>
          , Reykjavik, Iceland, may.
          <source>European Language Resources Association (ELRA).</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Andre</given-names>
            <surname>Blessing</surname>
          </string-name>
          and Hinrich Schu¨tze.
          <year>2010</year>
          .
          <article-title>Selfannotation for fine-grained geospatial relation extraction</article-title>
          .
          <source>In Proceedings of the 23rd International Conference on Computational Linguistics</source>
          , pages
          <fpage>80</fpage>
          -
          <lpage>88</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Andre</given-names>
            <surname>Blessing</surname>
          </string-name>
          , Jens Stegmann, and
          <string-name>
            <given-names>Jonas</given-names>
            <surname>Kuhn</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>SOA meets relation extraction: Less may be more in interaction</article-title>
          .
          <source>In Proceedings of the Workshop on Serviceoriented Architectures</source>
          (
          <article-title>SOAs) for the Humanities: Solutions and Impacts</article-title>
          , Digital Humanities, pages
          <fpage>6</fpage>
          -
          <lpage>11</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Bernd</given-names>
            <surname>Bohnet</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jonas</given-names>
            <surname>Kuhn</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>The best of bothworlds - a graph-based completion model for transitionbased parsers</article-title>
          .
          <source>In Proceedings of the 13th Conference of the European Chapter of the Association for Computational Linguistics</source>
          , pages
          <fpage>77</fpage>
          -
          <lpage>87</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>John</given-names>
            <surname>Bradley</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Towards a richer sense of digital annotation: Moving beyond a media orientation of the annotation of digital objects</article-title>
          .
          <source>Digital Humanities Quarterly</source>
          ,
          <volume>6</volume>
          (
          <issue>2</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Carlson</surname>
          </string-name>
          , Justin Betteridge, Bryan Kisiel, Burr Settles,
          <string-name>
            <surname>Estevam R. Hruschka</surname>
            Jr., and
            <given-names>Tom M.</given-names>
          </string-name>
          <string-name>
            <surname>Mitchell</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Toward an architecture for never-ending language learning</article-title>
          .
          <source>In Proceedings of the 24th Conference on Artificial Intelligence</source>
          , pages
          <fpage>1306</fpage>
          -
          <lpage>1313</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Laura</given-names>
            <surname>Chiticariu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Yunyao</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Frederick R.</given-names>
            <surname>Reiss</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Rule-based information extraction is dead! long live rule-based information extraction systems</article-title>
          !
          <source>In Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, EMNLP</source>
          <year>2013</year>
          ,
          <volume>18</volume>
          -21
          <source>October</source>
          <year>2013</year>
          , Grand Hyatt Seattle, Seattle, Washington, USA,
          <article-title>A meeting of SIGDAT, a Special Interest Group of the ACL</article-title>
          , pages
          <fpage>827</fpage>
          -
          <lpage>832</lpage>
          . ACL.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Kerstin</given-names>
            <surname>Eckart</surname>
          </string-name>
          and
          <string-name>
            <given-names>Ulrich</given-names>
            <surname>Heid</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Resource interoperability revisited</article-title>
          .
          <source>In Ruppenhofer and Faaß (Ruppenhofer and Faaß</source>
          ,
          <year>2014</year>
          ), pages
          <fpage>116</fpage>
          -
          <lpage>126</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Fredo</given-names>
            <surname>Erxleben</surname>
          </string-name>
          , Michael Gu¨nther, Markus Kro¨tzsch, Julian Mendez, and Denny Vrandecˇic´.
          <year>2014</year>
          .
          <article-title>Introducing wikidata to the linked data web</article-title>
          .
          <source>In Proceedings of the 13th International Semantic Web Conference (ISWC</source>
          <year>2014</year>
          ), volume
          <volume>8796</volume>
          <source>of LNCS</source>
          , pages
          <fpage>50</fpage>
          -
          <lpage>65</lpage>
          . Springer, October.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Manaal</given-names>
            <surname>Faruqui</surname>
          </string-name>
          and Sebastian Pado´.
          <year>2010</year>
          .
          <article-title>Training and evaluating a German named entity recognizer with semantic generalization</article-title>
          .
          <source>In Proceedings of the Conference on Natural Language Processing (KONVENS)</source>
          , pages
          <fpage>129</fpage>
          -
          <lpage>133</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Ferrucci</surname>
          </string-name>
          and
          <string-name>
            <given-names>Adam</given-names>
            <surname>Lally</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>UIMA: an architectural approach to unstructured information processing in the corporate research environment</article-title>
          .
          <source>Natural Language Engineering</source>
          ,
          <volume>10</volume>
          (
          <issue>3-4</issue>
          ):
          <fpage>327</fpage>
          -
          <lpage>348</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>David</given-names>
            <surname>Ferrucci</surname>
          </string-name>
          ,
          <string-name>
            <surname>Eric Brown</surname>
          </string-name>
          , Jennifer Chu-Carroll,
          <string-name>
            <given-names>James</given-names>
            <surname>Fan</surname>
          </string-name>
          , David Gondek,
          <string-name>
            <given-names>Aditya</given-names>
            <surname>Kalyanpur</surname>
          </string-name>
          , Adam Lally,
          <string-name>
            <given-names>William</given-names>
            <surname>Murdock</surname>
          </string-name>
          , Eric Nyberg, John Prager, Nico Schlaefer, and
          <string-name>
            <given-names>Christopher</given-names>
            <surname>Welty</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Building Watson: An Overview of the DeepQA Project</article-title>
          .
          <source>AI Magazine</source>
          ,
          <volume>31</volume>
          (
          <issue>3</issue>
          ):
          <fpage>59</fpage>
          -
          <lpage>79</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Antske</given-names>
            <surname>Fokkens</surname>
          </string-name>
          , Serge ter Braake, Niels Ockeloen, Piek Vossen, Susan Legeˆne, and
          <string-name>
            <given-names>Guus</given-names>
            <surname>Schreiber</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Biographynet: Methodological issues when nlp supports historical research</article-title>
          .
          <source>In Proceedings of the 9th International Conference on Language Resources and Evaluation (LREC</source>
          <year>2014</year>
          ), Reykjavik, Iceland, May
          <volume>26</volume>
          - 31.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Ralph</given-names>
            <surname>Grishman</surname>
          </string-name>
          and
          <string-name>
            <given-names>Beth</given-names>
            <surname>Sundheim</surname>
          </string-name>
          .
          <year>1996</year>
          .
          <article-title>Message understanding conference-6: a brief history</article-title>
          .
          <source>In Proceedings of the 16th conference on Computational linguistics</source>
          , pages
          <fpage>466</fpage>
          -
          <lpage>471</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Ulrich</given-names>
            <surname>Heid</surname>
          </string-name>
          , Helmut Schmid, Kerstin Eckart, and
          <string-name>
            <given-names>Erhard</given-names>
            <surname>Hinrichs</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>A corpus representation format for linguistic web services: the D-SPIN Text Corpus Format and its relationship with ISO standards</article-title>
          .
          <source>In Proceedings of LREC-2010</source>
          ,
          <string-name>
            <given-names>Linguistic</given-names>
            <surname>Resources</surname>
          </string-name>
          and Evaluation Conference, Malta. [CD-ROM].
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Marie</given-names>
            <surname>Hinrichs</surname>
          </string-name>
          , Thomas Zastrow, and
          <string-name>
            <given-names>Erhard</given-names>
            <surname>Hinrichs</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Weblicht: Web-based lrt services in a distributed escience infrastructure</article-title>
          .
          <source>In Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC'10)</source>
          .
          <source>electronic proceedings.</source>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Yunyao</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Laura</given-names>
            <surname>Chiticariu</surname>
          </string-name>
          , Huahai Yang,
          <string-name>
            <surname>Frederick R. Reiss</surname>
          </string-name>
          , and
          <string-name>
            <surname>Arnaldo</surname>
          </string-name>
          Carreno-fuentes.
          <year>2012</year>
          .
          <article-title>Wizie: A best practices guided development environment for information extraction</article-title>
          .
          <source>In Proceedings of the ACL 2012 System Demonstrations, ACL '12</source>
          , pages
          <fpage>109</fpage>
          -
          <lpage>114</lpage>
          , Stroudsburg, PA, USA. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Cerstin</given-names>
            <surname>Mahlow</surname>
          </string-name>
          , Kerstin Eckart, Jens Stegmann, Andre´ Blessing, Gregor Thiele, Markus Ga¨rtner, and
          <string-name>
            <given-names>Jonas</given-names>
            <surname>Kuhn</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Resources, tools, and applications at the CLARIN center stuttgart</article-title>
          .
          <source>In Ruppenhofer and Faaß (Ruppenhofer and Faaß</source>
          ,
          <year>2014</year>
          ), pages
          <fpage>127</fpage>
          -
          <lpage>137</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Mausam</surname>
            ,
            <given-names>Michael</given-names>
          </string-name>
          <string-name>
            <surname>Schmitz</surname>
            , Robert Bart, Stephen Soderland, and
            <given-names>Oren</given-names>
          </string-name>
          <string-name>
            <surname>Etzioni</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Open language learning for information extraction</article-title>
          .
          <source>In Proceedings of Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CONLL).</source>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Franco</given-names>
            <surname>Moretti</surname>
          </string-name>
          .
          <year>2013</year>
          . Distant Reading. Verso, London.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Niels</given-names>
            <surname>Ockeloen</surname>
          </string-name>
          , Antske Fokkens, Serge Ter Braake,
          <string-name>
            <surname>Piek T. J. M. Vossen</surname>
          </string-name>
          , Victor de Boer, Guus Schreiber, and Susan Legeˆne.
          <year>2013</year>
          .
          <article-title>Biographynet: Managing provenance at multiple levels and from different perspectives</article-title>
          . In Paul T. Groth, Marieke van Erp,
          <string-name>
            <surname>Tomi Kauppinen</surname>
            ,
            <given-names>Jun</given-names>
          </string-name>
          <string-name>
            <surname>Zhao</surname>
          </string-name>
          , Carsten Keßler, Line C.
          <article-title>Pouchard, Carole A</article-title>
          .
          <string-name>
            <surname>Goble</surname>
          </string-name>
          , Yolanda Gil, and Jacco van Ossenbruggen, editors,
          <source>Proceedings of the 3rd International Workshop on Linked Science</source>
          <year>2013</year>
          - Supporting Reproducibility,
          <article-title>Scientific Investigations and Experiments (LISC2013) In conjunction with the 12th</article-title>
          <source>International Semantic Web Conference 2013 (ISWC</source>
          <year>2013</year>
          ), Sydney, Australia, October
          <volume>21</volume>
          ,
          <year>2013</year>
          ., volume
          <volume>1116</volume>
          <source>of CEUR Workshop Proceedings</source>
          , pages
          <fpage>59</fpage>
          -
          <lpage>71</lpage>
          . CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <surname>Philip</surname>
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Ogren</surname>
            , Philipp G. Wetzler, and
            <given-names>Steven</given-names>
          </string-name>
          <string-name>
            <surname>Bethard</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>ClearTK: A UIMA toolkit for statistical natural language processing</article-title>
          .
          <source>In UIMA for NLP workshop at Language Resources and Evaluation Conference</source>
          , pages
          <fpage>32</fpage>
          -
          <lpage>38</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <given-names>Stephen</given-names>
            <surname>Ramsay</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Toward an algorithmic criticism</article-title>
          .
          <source>Literary and Linguistic Computing</source>
          ,
          <volume>18</volume>
          :
          <fpage>167</fpage>
          -
          <lpage>174</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>Stephen</given-names>
            <surname>Ramsay</surname>
          </string-name>
          ,
          <year>2007</year>
          . Algorithmic Criticism, pages
          <fpage>477</fpage>
          -
          <lpage>491</lpage>
          . Blackwell Publishing, Oxford.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <given-names>Josef</given-names>
            <surname>Ruppenhofer</surname>
          </string-name>
          and Gertrud Faaß, editors.
          <source>2014. Proceedings of the 12th Edition of the Konvens Conference</source>
          , Hildesheim, Germany, October 8-
          <issue>10</issue>
          ,
          <year>2014</year>
          . Universita¨tsbibliothek Hildesheim.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <given-names>Helmut</given-names>
            <surname>Schmid</surname>
          </string-name>
          and
          <string-name>
            <given-names>Florian</given-names>
            <surname>Laws</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Estimation of conditional probabilities with decision trees and an application to fine-grained POS tagging</article-title>
          .
          <source>In Proceedings of the 22nd International Conference on Computational Linguistics (Coling</source>
          <year>2008</year>
          ), pages
          <fpage>777</fpage>
          -
          <lpage>784</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <given-names>Helmut</given-names>
            <surname>Schmid</surname>
          </string-name>
          .
          <year>1995</year>
          .
          <article-title>Improvements in part-of-speech tagging with an application to German</article-title>
          . In
          <source>In Proceedings of the ACL SIGDAT-Workshop</source>
          , pages
          <fpage>47</fpage>
          -
          <lpage>50</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <string-name>
            <given-names>Helmut</given-names>
            <surname>Schmid</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Unsupervised learning of period disambiguation for tokenisation</article-title>
          .
          <source>Technical report</source>
          , IMS, University of Stuttgart.
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <string-name>
            <surname>Fabian M. Suchanek</surname>
            , Mauro Sozio, and
            <given-names>Gerhard</given-names>
          </string-name>
          <string-name>
            <surname>Weikum</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>SOFIE: A Self-Organizing Framework for Information Extraction</article-title>
          .
          <source>In Proceedings of the 18th International Conference on World Wide Web</source>
          , pages
          <fpage>631</fpage>
          -
          <lpage>640</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <string-name>
            <given-names>Matthew</given-names>
            <surname>Wilkens</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Canons, close reading, and the evolution of method</article-title>
          . In Matthew K. Gold, editor,
          <source>Debates in the Digital Humanities</source>
          . University of Minnesota Press, Minneapolis.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>