<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>STOLE: A Reference Ontology for Historical Research Documents</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Laura Pandolfo</string-name>
          <email>laura.pandolfo@edu.unige.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DIBRIS, Universita di Genova</institution>
          ,
          <addr-line>Via Opera Pia, 13</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Historical documents are a relevant part of cultural heritage. It is well-established that this domain is very complex: data are often heterogeneous, semantically rich, and highly interlinked. For this reason, searching and linking them with related contents represents a challenging task. The use of Semantic Web technologies can provide innovative methods for more e ective search and retrieval operations. In this paper we present stole, a reference ontology that provides a vocabulary of terms and relations that model the history of Italian public administration' domain. stole can be considered a rst step towards the development of a knowledge exchange standard on this speci c domain.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Historical texts and documents are considered an important component of
cultural heritage. Among other sources, historical newspapers and journals archived
in libraries represent a valuable source of information for historians.</p>
      <p>
        In the last years, several portals and digital libraries in the cultural heritage
eld have been enhanced with Semantic Web (SW) technologies in order to
implement methods for more e ective search and retrieval operations. In fact, they
can o er e ective solutions about design and implementation of user-friendly
ways to access and query content and meta-data { see [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] for a survey in the
historical research domain. The use of SW technologies can improve the
productivity of research in historical domain, since they help to identify implicit
and explicit knowledge included in the documents, e.g., reference to a historical
person contained in a historical source can be discovered and related to other
entities, for example events in which that person has been involved in.
      </p>
      <p>One of the main elements of the SW infrastructure are ontologies. They
provide a shared and clear representation for a speci c domain and they may play
a major role in supporting knowledge extraction, knowledge discovery, and data
integration processes. An ontology is usually de ned as a formal speci cation of
domain knowledge conceptualization. A Knowledge Base (KB) is composed of
two elements, i.e., a Terminological Box (TBox) and an Assertional Box (ABox).
The TBox represents the intensional knowledge of the KB and it is made up of
classes of data and relations among them. In the following, we use the terms
TBox and ontology interchangeably to denote the conceptual part of a KB. The
ABox contains extensional knowledge, which is speci c to the individuals of the
domain.</p>
      <p>In this paper we present stole1, a reference ontology that provides a
vocabulary of terms and relations with which it is possible to model the history of
Italian public administration' domain. stole represents a rst step towards the
development of a knowledge exchange standard on this speci c domain. To the
best of our knowledge, this is the rst time that ontological modeling has been
undertaken for this speci c domain.</p>
      <p>
        The paper is organized as follows: in Section 2 we describe the speci c case
study that motivated this research. In Section 3 we present the stole ontology
and its design process. This ontology is built on an extended version of stole [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
We conclude the paper in Section 4 by identifying stimulating directions and
challenging issues for continued and future research in this application domain.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Case Study</title>
      <p>In the last years the way of doing historical research has deeply changed. From
traditional research in archives and libraries, which necessarily required the
presence on site and had precise limits and constraints, the digitization of historical
documents allows new perspectives and ways to conduct research. The use of
SW technologies has upset times, costs, methods of historical research, and is
also changing the so-called culture of the document.</p>
      <p>The heritage of the history of public administration represents a fundamental
element to understand the history of the Italian institutions as well as the history
of the country in general. One of the main sources used in this eld of research is
represented by historical text and documents, including journals and newspapers
of the age. Recently, several web sites concerning speci c domain databases on
Italian institutions, e.g., the Bibliography of Italian Parliament 2 and Internet
Archive 3, o er to the scholars the opportunity to easily access the sources, which
are sometimes totally unknown.</p>
      <p>The development of the stole ontology responds to the needs of some
researchers of Department of History of the University of Sassari which, since the
1980s, have been involved in a project designed to collect and digitalize historical
journals regarding the origin and the evolution of institutions, customs, usages,
and rules in the Italian public administration. As a result, the arap4 digital
archive of the University of Sassari was created and it actually collects a large
amount of information about some of the most relevant journal articles published
between 1848 and 1946 concerning the legislative history of public
administration in Italy. The main goal of arap was to o er to the scienti c community
1 stole is the acronym for the Italian \STOria LEgislativa della pubblica
amministrazione italiana", that means \Legislative History of Italian Public Administration".
2 http://bpr.camera.it
3 https://archive.org
4 arap is the acronym for the Italian \Archivio di Riviste sull'Amministrazione
Pubblica", that means \Archive of Journals on Italian Public Administration".
interested in those documents, such as historians, lawyers, political scientists,
and sociologists, a repository of important sources, otherwise hardly accessible.</p>
      <p>The journal articles included in arap have a remarkably value for the wealth
of information they contain. In fact, starting the research from the journals
could make a positive and signi cant contribution to the current knowledge. For
example, through the study of these documents it is possible to establish
connections between authors, institutions, persons and historical events mentioned
in an article. Speci cally, the link between an author and the people cited can
reveal a lot, such as the political and cultural reference of the author. Clearly,
the relationship between an institution and an historical event can o er support
to understand the evolution of public administration. The use of certain
concepts or the recurring of names in relation to some historical events can be an
indicator of a trend.</p>
      <p>In this context, stole represent the core element of the arap digital archive
and its main goal is to to model historical concepts and gain insights into this
speci c eld in order to support historians in their research tasks.
3</p>
    </sec>
    <sec id="sec-3">
      <title>The STOLE Ontology</title>
      <p>The main steps of the stole ontology design process concerned the identi cation
both of the key concepts for this speci c domain, and the proper language for
the TBox implementation. Moreover, we populated the ontology, i.e., lling the
ABox with semantic annotations. A team of domain experts was involved during
the whole ontology development process. In particular, they contributed at the
early stage in order to de ne the key issues related to the application domain.</p>
      <p>In the rst phase, we detected the main categories of data expressed in the
considered historical documents. The results of this process enabled us to detect
three categories of elements: 1) Data concerning the authors of the articles,
e.g., name, surname and biography; 2) Data concerning the journal and the
article, e.g., article title, journal name, date and topics raised in the article; 3)
Data concerning some relevant facts and people cited in the article, e.g, people,
historical events, institutions. In this speci c domain, historical analysis is based
on these categories of information and focused on the interrelations among these
data. For instance, the relation between an author and the people mentioned in
an article could provide valuable information to historians, e.g., if an author has
often referred to Giuseppe Mazzini then it could easily be interpreted that this
author was favourable to the republic.</p>
      <p>During the second phase, the TBox of the stole ontology has been designed
building on some existing standards and meta-data vocabularies, such as
Bibliographic Ontology (BIBO) 5, Bio Vocabulary (BIO) 6, Dublin Core (DC) 7, Friend
5 http://bibliontology.com
6 http://vocab.org/bio/0.1/.html
7 http://dublincore.org
of A Friend (FOAF) 8, and Ontology of the Chamber of Deputies (OCD) 9. In
details, BIBO has been used to describe information about documents, for instance
bibo:volume, bibo:issue, bibo:pageStart and bibo:pageEnd, denoting
values of volume, issue, page start and page end of an article, respectively. Both
BIO and FOAF have been used in order to describe information about people.
In the case of FOAF, we reused terms such as foaf:person, foaf:firstName,
foaf:surname, foaf:gender. Other information about people, their
relationships and the events in their lives have been described using BIO concepts, such
as bio:birth, bio:death, bio:event, bio:place, bio:biography. From DC we
reused concepts related to the structure and characteristics of a document, e.g,
dc:title, dc:isPartOf, dc:publisher, dc:description. Considering OCD,
we noticed that it is a relevant source for stole ontology since, for example,
most part of the authors of articles related to the history of the public
administration topics during their lives were also involved in government activities. In
particular, OCD provided us valuable concepts in order to give detailed
information about people involved in political o ces. Finally, the usage of these core
ontologies allows both extensibility and interoperability of stole with other
resources and applications.</p>
      <p>
        In dealing with the modeling language, we decide to use OWL2 DL [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] since it
allows to properly model the knowledge for our application domain by means of
constructs like cardinality restrictions and other role constraints, e.g., functional
properties. The TBox is composed of 268 axioms, 19 classes, 34 object properties,
and 33 data properties.
      </p>
      <p>In the following, we describe main classes of the stole ontology.
Article represents our library, namely the collection of historical journal
articles. Every instance of this class has data properties such as articleTitle,
articleDate, pageStart, and pageEnd.</p>
      <p>Institution is used to represent the public institutions cited in the articles.</p>
      <p>This class contains four subclasses, namely Central, Local,
PoliticalInstitution and EconomicInstitution. Notice that Central
and Local are disjoint classes. The same holds for PoliticalInstitution
and EconomicInstitution. In this way, for example, an institution can be
both a local institution and a political institution at the same time.
Jurisprudence is a subclass of Article and contains a series of verdicts which
are entirely written in the articles. Every individual of this subclass has the
following data properties: verdictDate, verdictTitle, and byCourt.
Law is also a subclass of Article, and it contains a set of principles, rules, and
regulations set up by a government or other authority which are entirely
written in the articles. This subclass has data properties such as lawDate
and lawTitle.</p>
      <p>Event denotes relevant events. It contains ve subclasses modeling di erent
kinds of events: Birth and Death are subclasses related to a person's life;
8 http://www.foaf-project.org
9 http://data.camera.it/data</p>
      <p>BeginPublication and EndPublication represent the publication period
of a journal; HistoricalEvent contains the most relevant events that have
marked the Italian history.</p>
      <p>Journal denotes the collection of historical journals. This class has data
properties such as journalArticle, publisher, and issn.</p>
      <p>Person is the class representing people involved in the Italian legislative and
public administration history. This class contains one subclass, Author, that
includes the contributors of the articles. Every instance of this class has some
data properties as firstName, surname, and biography.</p>
      <p>Place represents cities and countries related to people and events.
Subject is a class representing topics tackled in the historical journals.</p>
      <p>
        In conclusion, there are other examples of ontologies designed for the cultural
heritage domain, e.g., CIDOC CRM [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], however these are rather generic to
be applied in our context where a detailed modeling of the singular domain
is needed. Despite its speci c nature, stole could be used in several elds of
research, e.g., admnistrative law, political science, history of institutions.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Current Work and Open Problems</title>
      <p>
        Currently, we are extending and improving in many ways the implementation
of this work. First, we are dealing with a key issue for historians, namely how
to disambiguate individuals and how to manage changing names, e.g., di erent
people with the same name or institutions that changed name retaining the
same functions. This point still represents an open challenge in this application
domain{ see [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>
        Another open problem relates the ontology population process. For the
current version, stole has been populated leveraging a set of annotated historical
documents comprised into the arap archive. Semantic annotations were
provided by a team of domain experts and individuals were added to the ontology
by means of a JAVA program built on top of the OWL APIs [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. This activity
requires specialized expertise, it is time consuming and resource-intensive. Given
these reasons, we are studying to nd solutions for its automatization on the
basis of some recent contributions { see, e.g.,[
        <xref ref-type="bibr" rid="ref10 ref13 ref6 ref7">7, 10, 13, 6</xref>
        ]. Most approaches for
automatic or semi-automatic ontology population process from texts are based
on the following techniques: Natural Language Processing [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], Machine Learning
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], and Information Extraction [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Once the ontology will be fully populated,
we are planning to perform an experimental analysis on the stole ontology
involving state of the art DL reasoners on both classi cation and query answering
tasks.
      </p>
      <p>Acknowledgments I would to thank the anonymous reviewers for their valuable
suggestions, which were helpful in improving the nal version of the paper.
Moreover, I would also like to thank my supervisors, Giovanni Adorni and Luca
Pulina, for their support, and Salvatore Mura and Prof. Francesco Soddu for the
valuable discussions about the application domain.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Adorni</surname>
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maratea</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pandolfo L. Pulina L</surname>
          </string-name>
          .
          <article-title>An Ontology for Historical Research Documents</article-title>
          .
          <source>Web reasoning and Rule Systems</source>
          . Springer (
          <year>2015</year>
          )
          <volume>11</volume>
          {
          <fpage>18</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bishop</surname>
            ,
            <given-names>C.M.</given-names>
          </string-name>
          :
          <source>Pattern Recognition and Machine Learning</source>
          . Springer (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kobilarov</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Becker</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cyganiak</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hellmann</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Dbpedia - A Crystallization Point for the Web of Data</article-title>
          .
          <source>Web Semantics: Science, Services and Agents on the World Wide Web</source>
          <volume>7</volume>
          (
          <issue>3</issue>
          ) (
          <year>2009</year>
          )
          <volume>154</volume>
          {
          <fpage>165</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Dale</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moisl</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Somers</surname>
            ,
            <given-names>H.L.</given-names>
          </string-name>
          :
          <article-title>Handbook of Natural Language Processing</article-title>
          . CRC (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Doerr</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>The CIDOC Conceptual Reference Module: An Ontological Approach to Semantic Interoperability of Metadata</article-title>
          .
          <source>AI Mag</source>
          .
          <volume>24</volume>
          (
          <issue>3</issue>
          ) (
          <year>2003</year>
          )
          <volume>75</volume>
          {
          <fpage>92</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Faria</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serra</surname>
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Girardi</surname>
          </string-name>
          , R.:
          <article-title>A Domain-Independent Process for Automatic Ontology Population from Text</article-title>
          .
          <source>Journal of Science of Computer Programming</source>
          <volume>95</volume>
          (
          <year>2014</year>
          )
          <volume>26</volume>
          {
          <fpage>43</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Fernandez</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cantador</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lopez</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vallet</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Castells</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motta</surname>
          </string-name>
          , E.:
          <article-title>Semantically Enhanced Information Retrieval: An Ontology-based Approach</article-title>
          .
          <source>Web Semantics: Science, Services and Agents on the World Wide Web</source>
          <volume>9</volume>
          (
          <issue>4</issue>
          ) (
          <year>2011</year>
          )
          <volume>434</volume>
          {
          <fpage>452</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Grau</surname>
            ,
            <given-names>B.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Horrocks</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motik</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parsia</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patel-Schneider</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sattler</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          :
          <article-title>Owl 2: The next step for owl</article-title>
          .
          <source>Web Semantics: Science, Services and Agents on the World Wide Web</source>
          <volume>6</volume>
          (
          <issue>4</issue>
          ) (
          <year>2008</year>
          )
          <volume>309</volume>
          {
          <fpage>322</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Horridge</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bechhofer</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>The OWL API: A Java Api for OWL Ontologies</article-title>
          .
          <source>Semantic Web</source>
          <volume>2</volume>
          (
          <issue>1</issue>
          ) (
          <year>2011</year>
          )
          <volume>11</volume>
          {
          <fpage>21</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Kara</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alan</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sabuncu</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Akp nar</surname>
          </string-name>
          , S.,
          <string-name>
            <surname>Cicekli</surname>
            ,
            <given-names>N.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alpaslan</surname>
            ,
            <given-names>F.N.:</given-names>
          </string-name>
          <article-title>An Ontology-based Retrieval System Using Semantic Indexing</article-title>
          .
          <source>Information Systems</source>
          <volume>37</volume>
          (
          <issue>4</issue>
          ) (
          <year>2012</year>
          )
          <volume>294</volume>
          {
          <fpage>305</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Meron</surname>
          </string-name>
          <article-title>~o-Pen~uela,</article-title>
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Ashkpour</surname>
          </string-name>
          , A., van Erp,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Mandemakers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Breure</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Scharnhorst</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Schlobach</surname>
          </string-name>
          , S., van
          <string-name>
            <surname>Harmelen</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Semantic Technologies for Historical Research: A survey</article-title>
          .
          <source>Semantic Web Journal</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Moens</surname>
            ,
            <given-names>M.F.</given-names>
          </string-name>
          :
          <article-title>Information Extraction: Algorithms and Prospects in a Retrieval Context</article-title>
          . Springer Netherlands (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Sanchez</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Batet</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Isern</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Valls</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Ontology-based Semantic Similarity: A New Feature-based Approach</article-title>
          .
          <source>Expert Systems with Applications</source>
          <volume>39</volume>
          (
          <issue>9</issue>
          ) (
          <year>2012</year>
          )
          <volume>7718</volume>
          {
          <fpage>7728</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>