<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Digital Editions beyond XML - Graph-based Digital Editions</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andreas Kuczera</string-name>
          <email>andreas.kuczera@adwmainz.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Academy of Science and Literature</institution>
          ,
          <addr-line>Mainz</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>XML has been the de facto standard for digital editions for years, but its serious limitations include an inability to represent overlapping markup and the encoding of multiple annotation hierarchies. With emerging graph database technologies we have the opportunity to develop new approaches. In this paper the advantages and modelling principles of graph-based digital editions will be discussed. XML in combination with TEI has become the standard format for digital editions. Most digital research environments, such as TextGrid or Ediarum, use TEI-XML for encoding sources and research data. But XML has some limitations which restrict researchers working on digital editions in certain ways. In this paper these limitations will be discussed and proposals for a graph-based digital edition will be presented.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>The use of XML for digital editions is widespread and has become a standard over the last
years.</p>
      <p>
        XML has some inherent limitations, however, such as its inability to represent overlapping
markup as well as to encode multiple annotation hierarchies [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. An easy-to-understand example
of overlapping markup is if someone wants to encode the formal structure of a source. In that
case he might need overlapping structures for pages and chapters. TEI solves this problem by
using empty elements like &lt;lb&gt; and &lt;pb&gt; instead of &lt;div&gt; and &lt;p&gt;-tags. Another example is
different sources of the same text. If you want to encode them in a single document, you quickly
encounter overlapping markup-structures on one level. The situation becomes even worse when
you try to encode diverging interpretations of the same source by different researchers in a
single file.
      </p>
      <p>
        To solve this problem several approaches have been developed. Daniel Jettka presents
some of these problems in his paper [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. A first approach is to use additional structures for
the representation of multiple hierarchies in standard XML with the help of milestones [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] or
fragmentation, or a mix of both.
      </p>
      <p>What all these approaches have in common is that the complexity of the XML increases
to a point where it becomes very hard to handle from an editor’s or application’s perspective.
In a second approach, Stand-off markup can be used to separate the source from the related
information. This means that the source’s XML content is indexed on a character basis and
the related information is stored in another file with pointers to the index-numbers. With this
approach, the simultaneous display of multiple annotation hierarchies is possible. A
disadvantage of this solution is the rigidity of the index. If characters in the source file are changed after
indexing and enriching the material, all indexes are worthless as long as the whole document is
not re-indexed. There is currently no editing tool available for solving this problem that also
takes basic aspects of usability into account.
3
3.1</p>
    </sec>
    <sec id="sec-2">
      <title>New perspectives for digital editions</title>
      <sec id="sec-2-1">
        <title>The text as a chain of nodes</title>
        <p>
          My paper at http://mittelalter.hypotheses.org [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] proposes graph-based digital editions as a
posibility to handle multiple annotation hierarchies. Figure 1 shows an example from this paper
where a record from the Regesta Imperii database (www.regesta-imperii.de) was modeled in the
Neo4j graph database. As a first step, the text is encoded as word nodes connected by edges.
3.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Adding annotations to the graph</title>
        <p>The second step is to append annotation information. Figure 2 shows overlapping markup of a
reference to another database record, represented by the yellow node, and another annotation,
represented by the green node.
As mentioned initially, TEI-XML represents a standard for encoding sources in digital
editions. If graph-based digital editions become a realistic alternative for digital editions, the
graph database must be able to handle TEI-XML data. For this purpose a GitHub project for
importing an TEI-XML document into Neo4j was started.</p>
        <p>For a first feasibility study a TEI-XML document containing a Latin letter from the 17th
century with a German abstract of the text has been used. The file includes additional
information about the persons, places and other entities in the letter. In an elaborated virtual research
environment this information would of course be separated from the text of the letter but for
demonstrational purposes of the study this information was also added to the example XML
file.</p>
        <p>Listing 1 presents some parts of the identifying information about persons, places, and
entities are listed.
...</p>
        <p>Listing 1: Persons, places and items in the xml file
&lt; p a r t i c D e s c &gt;
&lt; l i s t P e r s o n &gt;
&lt; p e r s o n xml : id =" P006 " &gt;
&lt; idno type =" gnd " &gt; http :// d - nb . info / gnd / 1 1 9 3 5 7 1 0 0 &lt; / idno &gt;
&lt; p e r s N a m e type =" reg " &gt;
&lt; surname &gt; L u b i e n i e c k i &lt;/ surname &gt;
&lt; forename &gt; S t a n i s Å Ć a w &lt;/ forename &gt;
&lt;/ persName &gt;
&lt;/ person &gt;
&lt; p l a c e xml : id =" DE - HAM " &gt;</p>
        <p>&lt; p l a c e N a m e type =" reg " &gt; Hamburg &lt;/ p l a c e N am e &gt;</p>
        <p>&lt;idno type =" gnd "&gt; http ://d-nb. info / gnd /4023118 -5 &lt;/ idno &gt;
&lt;/ place &gt;
&lt;textClass &gt;
&lt;keywords &gt;
&lt;list &gt;
&lt;item xml :id =" C_1664_W1 "&gt;
&lt;idno type =" iau "&gt;C /1664 W1 &lt;/ idno &gt;
&lt;label &gt;C /1664 W1 &lt;/ label &gt;
&lt;rs type =" objectType " key =" astro_comet "/ &gt;
&lt;/item &gt;
&lt;item xml :id =" Venus "&gt;
&lt;idno type =" gnd "&gt; http ://d-nb. info / gnd /4062527 -8 &lt;/ idno &gt;
&lt;label &gt; Venus &lt;/ label &gt;
&lt;rs type =" objectType " key =" astro_planet "/ &gt;
&lt;/item &gt;</p>
        <p>&lt;/item &gt;
&lt;item xml :id =" hypothesis "&gt;
&lt;idno type =" hyp "&gt; hypothesis &lt;/ idno &gt;
&lt;label &gt; hypothese about comets &lt;/ label &gt;
&lt;rs type =" objectType " key =" hypothesis "/ &gt;
&lt;/item &gt;
&lt;/list &gt;
&lt;/ keywords &gt;</p>
        <p>Listing 2 shows the German abstract of the Latin letter. Within the abstract, persons,
places and entities are identified and linked to the other entities shown in the first listing.</p>
        <sec id="sec-2-2-1">
          <title>Listing 2: The abstract of the letter</title>
          <p>&lt;abstract &gt;
&lt;p&gt;&lt; name ref ="# P007 "&gt; Langius &lt;/ name &gt; entschuldigt sich , dass er
die Anfrage &lt;name ref ="# P006 "&gt; Lubienietzki &lt;/ name &gt; zu dem
Kometen des Jahres 1664 versp ä tet beantwortet : Den
&lt;rs ref ="# C_1664_W1 "&gt; Kometen &lt;/rs &gt; habe er im ausgehenden
Dezember 1664 und Anfang Januar 1665 im &lt;name key ="# Hya "&gt;
Sternbild Hydra &lt;/ name &gt; beobachtet , und ein ä hnliches Phä nomen
sei im ausgehenden März bzw . Anfang April in &lt;name ref ="# Peg "&gt;
Pegasus &lt;/ name &gt; und &lt;name ref ="# And "&gt; Andromeda &lt;/ name &gt; erschienen .
In Anlehnung an &lt;name ref ="# P001 "&gt; Demokrit &lt;/ name &gt; äu ß ert &lt;name
ref ="# P007 "&gt; Langius &lt;/ name &gt; die &lt;rs ref =" hyp "&gt; Hypothese , dass
Kometen aus Atompartikeln gebildet werden : Sie entst ü nden im
Kegelschatten der &lt;name ref ="# Terra "&gt;Erde &lt;/ name &gt;
und wü rden erst dann sichtbar , wenn sie aus diesem hervor - und
in das Licht der &lt;name ref ="# Sol "&gt; Sonne &lt;/ name &gt; hineintr äten &lt;/rs &gt;.
Seine Hypothese veranschaulicht &lt;name ref ="# P007 "&gt; Langius &lt;/ name &gt;
...</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>The root of the document</title>
        <p>If we select an XML-like approach to the graph the "xml-root" of the document is represented
by the pink node (figure 3). All other informations can be reached from that node but they
are not ordered in an hierarchical way as it would be in an XML-Document. Four nodes are
directly connected to the root node. Those will be discussed below.
3.5</p>
      </sec>
      <sec id="sec-2-4">
        <title>Hyperedges for the processes of sending and receiving</title>
        <p>The two empty nodes linked to the root node are hyperedges.1 One hyperedge represents the
sending information of the letter with a combined link to the author of the letter, and the
1For the concept of a hyper-edge: http://neo4j.com/docs/stable/cypher-cookbook-hyperedges.html.
place from which it was sent. The other hyperedge represents the recipient information with a
combined link to the person who received the letter and the place where it was received. This
construct facilitates the traversal of the graph if there are many letters in the database.
3.6</p>
      </sec>
      <sec id="sec-2-5">
        <title>From the root to the abstract and the body text paragraphs</title>
        <p>Beside the root node there are two other linked nodes next to the hyperedge nodes. One leads
to the beginning of the abstract, the other leads to the Latin text of the letter. Following
the HAS_ABSTRACT-edge the abstract node can be reached. From this node, following the
NEXT-edge we can get the abstract word by word. You can explore the abstract by
doubleclicking on the word nodes and see the connected nodes for the annotation of persons and
places. All word-nodes of the abstract are directly connected to a p-node too, which represents
their belonging to the paragraph represented by the p-node, while the order of the words is
represented by the NEXT edges.
3.7</p>
      </sec>
      <sec id="sec-2-6">
        <title>Traverse the graph to get the abstract text</title>
        <p>Besides the possibility of exploring the abstract in the graph, a Cypher query can be used to
get the text of the abstract as a result. Listing 3 shows the query and figure 4 shows the results
in the Neo4j frontend.</p>
        <sec id="sec-2-6-1">
          <title>Listing 3: Get the text of the abstract</title>
          <p>// Get Abstract Text
MATCH p = ( t : Tag { name : ’ abstract ’}) -[: NEXT *] - &gt;( e )
WHERE NOT ( e ) -[: NEXT ] - &gt;()
RETURN REDUCE ( s ="" , x in tail ( nodes ( p )) | s +" " + x . text )</p>
          <p>Let us now examine the cypher query. With the MATCH-p-statement we start a traversal
query starting from a tag-node with the name "abstract" connected to the next node with
NEXT-edges and to the following node with NEXT-edges, until we come to a node that has no
departing NEXT-edge. The ’reduce’ function reduces the result to one single line of text only.
The rest of the query converts the resulting chain of nodes to the represented words.</p>
        </sec>
      </sec>
      <sec id="sec-2-7">
        <title>The hierarchy of the TEI file in the graph database</title>
        <p>Querying for the tag-node the Neo4j frontend shows the structure of the imported TEI-XML
as shown in figure 6. Starting from the pink root-node you can follow the abstract-node to
the abstract or the body-node to the body of the TEI-XML file. All paragraphs in the body
are connected to the body-node and their own hierarchy is represented by NEXT_TAG edges.
One IS_CHILD_OF-edge leads to the head-node with all the entity information connected to
it.</p>
      </sec>
      <sec id="sec-2-8">
        <title>Example of further annotation and exploration perspectives</title>
        <p>
          Beside the planets and comets a scientific hypothesis is mentioned in the letter and encoded in
the XML document as an item.2 In the following example, the statement of the hypothesis will
be connected with the parts of the text in the abstract in which it is explained, as well as with
the planets and persons mentioned. Modelling Digital Editions in Graph-Databases changes the
way pieces of information are explored. My paper [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] explains that most search cases are with
concrete interest need (cin-cases) were as explorational approaches usually are not supported.
So starting point for queries are search-interfaces on the web. The usual starting point for a
2The last item of the list in listing 1 represents the hypothesis which is annotated in the abstract.
query in a graph is an entity. Coming from that entity all connected nodes with the different
edges are explored. So not only the way how things are encoded will change [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] but also the
way things are queried.
        </p>
        <sec id="sec-2-8-1">
          <title>Listing 4: Create annotations in the graph database</title>
          <p>// C r e a t e E d g e s from text n o d e s to the h y p t h e s i s node
M A T C H ( from { id : ’ h y p o t h e s i s ’}) , ( to { text : ’ H y p ot h e s e , ’})</p>
          <p>C R E A T E from -[: S T A R T _ H Y P O T H E S I S ] - &gt; to ;
M A T C H ( from { id : ’ h y p o t h e s i s ’}) , ( to { text : ’ h i n e i n t r ä ten ’})</p>
          <p>C R E A T E from -[: E N D _ H Y P O T H E S I S ] - &gt; to ;
// C r e a t e e d g e s to the p r e s e n t e r of the h y p o t h e s i s
M A T C H ( from { id : ’ h y p o t h e s i s ’}) , ( to { s u r n a m e : ’ Lange ’})</p>
          <p>C R E A T E from -[: P R E S E N T E R _ O F _ H Y P O T H E S I S ] - &gt; to ;
// T h i n g s c o v e r e d by the h y p o t h e s i s
M A T C H ( from { id : ’ h y p o t h e s i s ’}) , ( to { l a b e l : ’ Sol ’})</p>
          <p>C R E A T E from -[: I S _ A B O U T ] - &gt; to ;
M A T C H ( from { id : ’ h y p o t h e s i s ’}) , ( to { l a b e l : ’ Terra ’})</p>
          <p>C R E A T E from -[: I S _ A B O U T ] - &gt; to ;
M A T C H ( from { id : ’ h y p o t h e s i s ’}) , ( to { l a b e l : ’ C / 1 6 6 4 W1 ’})</p>
          <p>C R E A T E from -[: I S _ A B O U T ] - &gt; to ;
In this paper a first glimpse of the great potential of graph-based digital editions was
demonstrated even if we must concede that we lack of powerful tools und user-interfaces. The
conceptual development is still in its early stages, but by using graph databases we could:
seemlessly include multiple annotation hierachies within one dataset
handle a scenario where several researchers are working together on the same source with
a transparent separation of their different markup needs
query the various annotational layers within one query language and
changes to the source text will neither destroy the word-order nor the meaning of existing
annotations.</p>
          <p>Many of the discussed issues could certainly also be solved with specific markup strategies or
with Stand-off markup, but this solution usually involves great effort and often produces XML
that is very hard to handle and not human-readable at all. "Ease of use" has always been a
major reason cited for the choice of XML as a (standard) format. Due to the complexity of future
digital editions, there is no reason we should not take the step to graph-based digital editions.
The contents of a graph database – its graphs – are human-readable and can represent very
complex annotational hierarchies, while also removing the worry about boundary demarcation.
Moreover everything can be exported to XML or graphml files for long-term archival purposes.</p>
          <p>This becomes possible in a model which brings the way things are structured in the human
mind and the database structure closer together. Graph-based digital editions are the next
important step for the Digital Humanities, as the editing disciplines acquire a very flexible,
easy-to-use but powerful tool. To achieve this goal it is important that we find standardized
ways for modelling events, actions, political aims, etc., in order to explore and compare different
sources from different contexts. The second important task is the programming of user interfaces
which can be easily used and adjusted to the researchers needs and can explore more than one
datasource on the internet. Developing this will be a big challenge.
5</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Acknowledgments</title>
      <p>The Neo4j-TEI-Importer was developed by Stefan Armbruster from Neo Technology (
stefan.armbruster@neotechnology.com) and I want to thank him for his help. He published
the importer on GitHub (https://github.com/sarmbruster/neo4j-tei-importer). Thanks
also to Sascha Kaufmann from the Digital Humanities Department of the University of Bern
for fruitful discussions on how to model XML data in a graph database environment.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>TEI</given-names>
            <surname>Consortium</surname>
          </string-name>
          .
          <article-title>P5: Richtlinien für die Auszeichnung und den Austausch elektronischer Texte</article-title>
          ,
          <year>2016</year>
          . http://www.tei-c.org/release/doc/tei-p5-doc/de/html/ref-milestone.html,
          <source>last viewed April</source>
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Jettka</surname>
          </string-name>
          . Repräsentation,
          <article-title>Verarbeitung und Visualisierung multipler Hierarchien mit XStandoff und</article-title>
          XSLT,
          <year>2011</year>
          . http://www.daniel-jettka.de/pdf/Multiple_Hierarchien_mit_ XStandoff_und_XSLT.pdf,
          <source>last viewed April</source>
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Andreas</given-names>
            <surname>Kuczera</surname>
          </string-name>
          .
          <article-title>Digitale Perspektiven mediävistischer Quellenrecherche</article-title>
          , in: Mittelalter.
          <source>Interdisziplinäre Forschung und Rezeptionsgeschichte</source>
          ,
          <year>2014</year>
          . http://mittelalter.hypotheses.org/ 3492.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Andreas</given-names>
            <surname>Kuczera</surname>
          </string-name>
          .
          <article-title>Graphbasierte digitale Editionen</article-title>
          , in: Mittelalter.
          <source>Interdisziplinäre Forschung und Rezeptionsgeschichte</source>
          ,
          <year>2016</year>
          . http://mittelalter.hypotheses.org/7994.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Sperberg-McQueen</surname>
            ,
            <given-names>C M</given-names>
          </string-name>
          ;
          <article-title>Huitfeldt, Claus. GODDAG: A Data Structure for Overlapping Hierarchies</article-title>
          . Lecture Notes in Computer Science (
          <year>2023</year>
          ),
          <year>2000</year>
          .
          <fpage>139</fpage>
          -
          <lpage>160</lpage>
          . http://cmsmcq.com/
          <year>2000</year>
          / poddp2000.html,
          <source>last viewed April</source>
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Amir</given-names>
            <surname>Zeldes Thomas Krause. ANNIS3</surname>
          </string-name>
          :
          <article-title>A new architecture for generic corpus query and visualization</article-title>
          .
          <source>Digital Scholarship in the Humanities</source>
          , Vol.
          <volume>31</volume>
          , No.
          <volume>1</volume>
          ,
          <year>2016</year>
          .
          <fpage>118</fpage>
          -
          <lpage>139</lpage>
          . http://dsh. oxfordjournals.org/content/31/1/118.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>