<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>PaleOrdia: Semantically Describing (Cuneiform) Paleography using Paleographic Linked Open Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Timo Homburg</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Mainz University Of Applied Sciences</institution>
          ,
          <addr-line>Lucy-Hillebrand Straße 2, Mainz</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This publication describes PaleOrdia, a web application developed to visualize (cuneiform) paleographic sign variants in Wikidata and the data model developed in Wikidata to represent paleography. Modeling paleographic sign variants of (ancient) scripts in linked open data is a relatively new development. It will enable better descriptions of digital scholarly editions with paleographic annotations supported by established web annotation data model vocabularies. As a use case for showcasing the capabilities of PaleOrdia, the cuneiform annotation tool Cuneur is presented as one way to harness the paleographiclinked open data for digital scholarly editions.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;PaleOrdia</kwd>
        <kwd>Paleography</kwd>
        <kwd>Cuneiform</kwd>
        <kwd>Annotation</kwd>
        <kwd>Paleographic Linked Open Data</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Describing the paleography of inscriptions on cultural heritage objects is a common task in
many digital scholarly editions [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] projects. In a digital scholarly edition, a set of texts is
commonly transcribed, annotated, and finally translated or interpreted so that the respective
scholar can address the targeted research question. While the contents of the respective textual
materials are likely the main focus of the scholar’s work, a closer inspection of the stylistic
choices made in writing the text is necessary in many disciplines. Such analysis might hint at
the detection of authors of texts by writing style, the identification of particularities of writing
in a specific time and space, and finally, may give hints about the circumstances in which a
text has been written. To make an accurate assessment of authorship, a detailed knowledge of
not only the preferred choice of words but also the shape of the characters the author uses for
writing, i.e., paleographic features, is important. An automatic analysis of paleographic data
needs, at best, accurately described training data, which may assist in automated analysis of
the given work of text at hand. This advocates for knowledge graphs of paleographic features,
which may be reused in diferent research contexts. This publication wants to highlight the
need for paleographic-linked open data by discussing the cases of cuneiform digital scholarly
editions, which will serve as the primary, but not exclusive, application case made possible
by this work. Section 2 will give some background on cuneiform signs and the terminologies
used in paleography, Section 3 explains how paleographic linked open data can be represented
in Wikidata. Finally, the tool PaleOrdia, a tool to manage paleographic linked open data in
Wikidata, is introduced in Section 4, and the usefulness of paleographic linked open data is
shown in an application example in Section 5.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Foundations and related work on the cuneiform LOD cloud</title>
      <p>The cuneiform script is one of the earliest writing systems used in Ancient Mesopotamia
as the script for many languages such as Sumerian, Akkadian, and Hittite. Throughout its
existence, for 3000 years, cuneiform signs have considerably evolved and been simplified.
They usually depict a transformation from a pictograph to increasingly simplified cuneiform
wedge configurations—the basis of all known cuneiform signs. Figure 1 shows the evolution of
cuneiform signs throughout space and time. For example, the cuneiform sign SAG for head starts
as a pictograph of a head in the Late Uruk period and is subsequently simplified. Starting from
the Early Dynastic period, cuneiform signs are comprised of cuneiform wedges to represent
cuneiform signs. Cuneiform wedges as atomic components of every cuneiform sign allow
for a more simple and formalized drawing of the grapheme and change in composition and
positioning of the cuneiform wedges in the subsequent centuries.</p>
      <sec id="sec-2-1">
        <title>2.1. Sign variant terminology</title>
        <p>This section briefly introduces terminology that may be used to describe paleographic sign
variants.</p>
        <p>Definition 1.</p>
        <p>Character A unit of information that often corresponds to a grapheme or a symbol.</p>
        <p>The definition of a character can be seen as equivalent to something that can be represented
by a Unicode codepoint. In contrast, a Unicode code point stands for a variety of shapes of
the same grapheme, e.g., may be depicted by diferent fonts, which may visualize a Unicode
codepoint to the user.</p>
        <p>Definition 2. Sign A cuneiform sign is a set of shapes/graphemes of cuneiform wedge
configurations across time and space that have been classified under a common identifier.</p>
        <p>A sign identifier for a cuneiform sign may be represented in diferent ways. The cuneiform
research community defines so-called sign names for cuneiform signs, often derived from the
most common reading of the sign in the Sumerian language. Many, but not all, sign names are
equivalent to a Unicode codepoint. Figure 1 shows the three sign names SAG, NINDA and GU7,
which each have a corresponding Unicode codepoint in the Unicode standard. However, there
are several sign names that have no Unicode codepoint attested at the time of writing. Most
cuneiform signs without a Unicode codepoint are combinations of already existing cuneiform
signs.</p>
        <p>Definition 3. Sign Variant A sign variant is a class of representations of a cuneiform sign that
has been attested in time and space.</p>
        <p>A sign variant is a set of distinct configurations of cuneiform wedges that have been classified
by a sign name and have been assigned a time period and/or location. Figure 2 shows a set
of cuneiform signs that difer in the number of wedges, their positioning, and the shapes of
the cuneiform wedges used to build the signs. The shapes of the individual cuneiform wedges
shown in Figure 2 difer slightly due to font variations and due to semantic expressions font
creators might want to convey. For example, a filled wedge head of a cuneiform sign might hint
at the sign variant being created on a stone surface vs. on a clay surface.</p>
        <p>Definition 4. Stylistic Sign Variant A stylistic sign variant is a stylistic change of a sign variant
that does not constitute a change of its characteristic elements.</p>
        <p>However, in reality, subtle changes in writing on the cuneiform clay tablet, expressed, for
example, in the length of individual strokes, the angle of individual strokes in certain cuneiform
sign variants, or the pointiness of wedge heads, might reveal the style of a particular writer or
even writing school. The aforementioned changes would typically be part of a stylistic sign
variant that hints at a specific author of cuneiform texts.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Capturing sign variants using character encodings</title>
        <p>
          Various approaches have been researched in the past to capture the essence of cuneiform sign
variants. Diferent encodings such as [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] and [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] try to capture the number of cuneiform wedges
per wedge type, and in the case of PaleoCodage, the positioning of wedges towards each other in
a String encoding to create a unique identifier for each of the cuneiform sign variants in existence.
These codes purposefully do not capture the elements that describe a stylistic sign variant but
can be used as elements in knowledge graphs [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] capturing paleographic information. Character
encodings such as this form the building blocks for identifying and classifying cuneiform signs
in knowledge graphs.
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Related Work</title>
        <p>
          The first ideas of modeling paleography with linked open data technologies emerged in the
digital humanities community in 2020. [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] proposed modeling paleographic features with
linked open data vocabularies and creating a formalized vocabulary. [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] described approaches
to generalize a paleographic vocabulary which can be used in conjunction with the
OntolexLemon [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] model to express the paleographic variants in which diferent words can be written.
In particular, this model also introduces the connection between lexemes and paleographic
descriptions and the concept of a paleographic sign variant occurrence for annotation purposes.
Further approaches to capture features of inscriptions include the CIDOC CRMtex extension
[
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], which allows the description of characters on inscriptions of surfaces of cultural heritage
objects. In general, though, digital humanities and computational linguistics approaches rarely
use paleographic information for classifications for the time being.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Modeling paleography in Wikidata</title>
      <p>In the previous sections, foundations for modeling paleographic-linked open data were defined.
This section explains how Wikidata can be used to represent paleographic features such as
cuneiform signs. Many cuneiform signs have been described in the Unicode standard, which
Wikidata adopts as QIDs. Hence, a starting point for a paleographic description are the Unicode
signs themselves. However, coverage of the Unicode codepoints for cuneiform signs would not
be suficient. Cuneiform signs may appear as ligatures, represented as more than one Unicode
codepoint, or cuneiform signs may not occur suficiently often to be considered for addition to
the Unicode standard. These signs will be added as new items to Wikidata and referenced in
respective literature. Figure 3 shows the data model for paleographic data adopted in Wikidata.
The data model relates cuneiform signs to their paleographic sign variants, which are classified
by time periods and - if applicable - encodings for their description. Lexicographical data
within Wikidata, such as the Sumerian word "a" (water) (L228723) may link to paleographic
sign variants used within their respective forms. In this way, scholars can express not only that
a Lexeme has been occurring in a specific time period, at a specific place, and in a specific text
source but also in which paleographic shape the Lexeme form has been attested. As shown
in fig:cuneiformsignhalcompoenents, many cuneiform signs consist of other cuneiform signs.
These can be modeled using Wikidata has parts (wdt:P527) relations. Finally, the semantics
of the original pictographs can be captured as their meanings in Wikidata using the depicts
(wdt:P180) relation. This leads to the capture of cuneiform sign meanings not only on the level
of the grapheme but also on a semantic level.</p>
    </sec>
    <sec id="sec-4">
      <title>4. PaleOrdia: A tool to visualize paleographic linked open data</title>
      <p>
        PaleOrdia1 is a fork of the tool Ordia [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], which has been developed as a view on Lexeme data
in Wikidata. PaleOrdia difers from the original Ordia tool in two fundamental ways:
• PaleOrdia is a static web application which runs on a Github Page, as opposed to the
original Ordia, which needed to be run on a web server
1https://situx.github.io/paleordia/script?q=Q401&amp;qLabel=cuneiform
      </p>
      <p>• PaleOrdia combines the view of Wikidata Lexemes with Wikidata Paleography data
PaleOrdia ofers the following functionalities to highlight cuneiform paleographic data:
• Listing of cuneiform sign variants by time period and reference work
• Identification of cuneiform signs which have not been included into Unicode 2
• Listing of cuneiform signs by type (compound3 and allograph signs4)
• Representation of cuneiform sign etymology and compounds</p>
      <sec id="sec-4-1">
        <title>4.1. Character Data / Cuneiform Signs</title>
        <p>PaleoOrdia allows users to view information about a cuneiform sign, including the reference
works and reference databases in which it is attested, its readings (phonetic values), its
classification as a compound sign, allograph, or cuneiform sign, and whether a single Unicode
codepoint represents it. Besides metadata such as the attestations of a cuneiform sign in sign
lists and on actual cuneiform tablet texts, each PaleOrdia page for a cuneiform sign also lists
the diferent sign variants 5 the sign has been attested with, as shown in Figure 4 Compound
signs may be shown per cuneiform sign and cuneiform sign variant. Figure 56 shows compound
signs containing the cuneiform sign HAL in a sign variant common in the Old Babylonian and
2https://situx.github.io/paleordia/no_unicode/?q=Q401&amp;qLabel=cuneiform&amp;qb=Q401
3https://situx.github.io/paleordia/compoundsigns/?q=Q401&amp;qLabel=cuneiform&amp;qb=Q401
4https://situx.github.io/paleordia/allographs/?q=Q401&amp;qLabel=cuneiform&amp;qb=Q401
5https://situx.github.io/paleordia/c/?q=Q87554995&amp;qLabel=%F0%92%80%80
6https://situx.github.io/paleordia/cf/?q=Q120671708&amp;qLabel=Cuneiform%20Sign%20Variant%20HAL%
20(Akkadian)
Akkadian periods. Similar visualizations exist for other time periods, and a generic visualization
based on Unicode codepoints is generated for the Unicode sign itself. Finally, users can list the
Lexemes in which the Unicode codepoint appears and the attested readings for a cuneiform
sign, which may be used to search for the sign and the sign variants in a research context. In
essence, PaleOrdia hereby acts as the view on a linked data-based paleographic sign registry
that can be extended collaboratively and provides a basis for discussion by scholars.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Application Case: Digital Editions of Cuneiform Tablets</title>
      <p>A digital edition of cuneiform tablets encompasses a variety of steps by a scholar but usually
requires the following components:
1. A transliteration of the written contents of the cuneiform tablet’s sides into the Latin
alphabet
2. The annotation of interesting text passages
3. The annotation of interesting features on image media depicting the cuneiform tablet
(e.g., broken parts, cuneiform signs, or seal impressions)
7https://fcgl.gitlab.io/annotator-showcase/
and with URIs, which allow for the reusage of paleographic sign variants across the boundaries
of a single cuneiform digital edition.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusions</title>
      <p>This publication introduced the application of a paleographic-linked open data model in
Wikidata. The model was tested using the cuneiform script as an example use case and has been
used to describe actual cuneiform sign variants found on images and renderings of cuneiform
tablet surfaces. The tool PaleOrdia gives an overview and can manage the entered cuneiform
sign variants using only a static homepage on Github. This allows cuneiform scholars to get
an overview of available cuneiform signs, allows them to compare these sign variants to their
ifndings on the cuneiform clay tablets, and create a linked open data graph of image annotations
that are linked to this sign variant registry, which has therefore emerged within Wikidata. The
results of such annotations can not only prove valuable for the cuneiform scholar community
but may also provide the basis and training data for a variety of machine-assisted classification
methods. Finally, this case study based on the cuneiform script might inspire the modeling of
sign variants of other scripts in Wikidata, contributing to an interconnected linked open data
cloud for paleography.</p>
      <sec id="sec-6-1">
        <title>6.1. Future Work</title>
        <p>
          Future work should enhance the paleographic linked open data cloud with metrics that allow
the calculation of the similarity of diferent graphemes in the linked open data cloud either by
semantic similarity, by image similarity metrics, or by metrics built from encodings which allow
the expression of a characters characteristics, such as a PaleoCode [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] or a Gottstein code [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]
for cuneiform. Expressing these metrics will allow for the implementation of a better semantic
search for paleographic sign variants and may prove valuable for approaches for automatic
annotation of cuneiform signs.
        </p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>P.</given-names>
            <surname>Sahle</surname>
          </string-name>
          ,
          <article-title>What is a scholarly digital edition?, Digital scholarly editing: Theories and practices 1 (</article-title>
          <year>2016</year>
          )
          <fpage>19</fpage>
          -
          <lpage>39</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S. V.</given-names>
            <surname>Panayotov</surname>
          </string-name>
          ,
          <article-title>The gottstein system implemented on a digital middle and neo-assyrian palaeography</article-title>
          , CDLN, London (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>T.</given-names>
            <surname>Homburg</surname>
          </string-name>
          ,
          <article-title>Paleocodage-enhancing machine-readable cuneiform descriptions using a machine-readable paleographic encoding</article-title>
          ,
          <source>Digital Scholarship in the Humanities</source>
          <volume>36</volume>
          (
          <year>2021</year>
          )
          <fpage>ii127</fpage>
          -
          <lpage>ii154</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>T.</given-names>
            <surname>Homburg</surname>
          </string-name>
          , T. Declerck,
          <article-title>Towards the integration of cuneiform in the ontolex-lemon framework</article-title>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>T.</given-names>
            <surname>Homburg</surname>
          </string-name>
          ,
          <article-title>Towards paleographic linked open data (plod): A general vocabulary to describe paleographic features</article-title>
          .,
          <source>in: DH</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bosque-Gil</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gracia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Buitelaar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Cimiano</surname>
          </string-name>
          ,
          <article-title>The ontolex-lemon model: development and applications</article-title>
          ,
          <source>in: Proceedings of eLex 2017 conference</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>19</fpage>
          -
          <lpage>21</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Doerr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Murano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Felicetti</surname>
          </string-name>
          , Definition of the crmtex,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>F.</given-names>
            <surname>Å</surname>
          </string-name>
          . Nielsen,
          <article-title>Ordia: A web application for wikidata lexemes</article-title>
          , in: The Semantic Web:
          <article-title>ESWC 2019 Satellite Events: ESWC 2019 Satellite Events</article-title>
          , Portorož, Slovenia, June 2-6,
          <year>2019</year>
          ,
          <source>Revised Selected Papers 16</source>
          , Springer,
          <year>2019</year>
          , pp.
          <fpage>141</fpage>
          -
          <lpage>146</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>R.</given-names>
            <surname>Sanderson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Ciccarese</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Young</surname>
          </string-name>
          ,
          <article-title>Web annotation data model</article-title>
          ,
          <source>W3C recommendation 23</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>