<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>An OWL framework for rule-based recognition of places in Italian non-structured text</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Domenico Cantone</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrea Fornaia</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marianna Nicolosi-Asmundo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniele Francesco Santamaria</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Emiliano Tramontana</string-name>
          <email>tramontanag@dmi.unict.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Catania, Dept. of Mathematics and Computer Science</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we present a technique for recognising location names in non-structured texts. Our approach is based on grammar rules devised for the Italian language and semantic web tools such as geographic linked datasets, OWL ontologies, and SWRL rules, to handle data and reason about them, even in presence of name ambiguities. To the best of our knowledge, this is the rst attempt of addressing the problem of location recognition in the context of Italian texts with such an ontological support.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Recognising location names of geographical places and of public or private
buildings inside non-structured text documents is an important issue with several
practical applications. For instance, in the investigative eld it is important to
reveal a place named in the transcription of an interception (i.e., by means of
wiretapping) and in the social media context to reveal the places visited by users
to provide targeted advertisements. This problem falls both in the category of
information extraction from unlabeled texts and of named entity recognition,
where most of the existing approaches are focused on the English language and
are in general not applicable to the Italian language, or they lead to
unsatisfactory results due to the peculiarities of the language.</p>
      <p>
        The problem of location name recognition has been addressed in various
ways [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], e.g., using maximum entropy models [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], or with Conditional
Random Fields [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], or with automatic learning techniques to infer the rules for the
named entities identi cation inside free texts [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], or, in the last decade, also
with linked data and ontologies [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] (a survey of the main geographical
ontologies and datasets can be found in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]). However, many of such approaches have
been tested, or developed, vertically on top of the English language, making
them hard to generalize to the Italian language. In addition, in some our
preliminary experiment, widespread applications such as Stanbol [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] turned out to
be unsatisfactory from the point of view of success ratio when applied to Italian
places drawn from non-structured text written in Italian.
      </p>
      <p>
        In this paper we focus on the problem of recognising names of geographical
places in the Italian country, which appear in non-structured Italian text
documents. Our approach consists in extracting location names from Italian texts
according to an extended version of the algorithm presented in [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], and then
storing data and making inferences, even in presence of name ambiguities, with
semantic web tools such as geographic linked datasets, OWL ontologies [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], and
SWRL rules [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>
        We recall that the algorithm presented in [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] relies on a set of three
nite state machines, each designed to recognise several sentence patterns for the
Italian language, in which location names are typically found. Such an analysis
takes as input a non-structured text written in Italian and yields as output a
HTML text, where candidate location names have been automatically marked
by a label.
      </p>
      <p>
        In this contribution, we uni ed the nite state machines presented in [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] in
a single automaton, reducing the overall number of states involved during the
extraction of location names from the text; each location detected by the
algorithm is then searched in the OpenStreetMap dataset [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Such a search yields a
list of possible matches of real places, each with its degree of reliability. Then,
the data retrieved by OpenStreetMap are inserted in a novel ontology, namely
OntoLocEstimation, to handle ambiguous geographical names.
OntoLocEstimation uses the ontology OntoLuoghi, introduced in [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], that contains a detailed
description of the administrative model of Italian places.
      </p>
      <p>
        The use of open datasets, such as OpenStreetMap, allows a widespread and
detailed coverage of Italian geographical places and provides a high precision
in the detection of real places. In addition, the introduction of Semantic Web
Rule Language (SWRL) rules [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] allows inferences on knowledge, implicitly
contained in the novel ontologies which are more re ned than other currently
available ontologies on geographical places.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Background</title>
      <p>
        The semantic web is a vision of the World Wide Web in which information
carries an explicit meaning, so it can be automatically processed and integrated by
machines, and data can be accessed and modi ed at a global level, resulting in
increased coherence and dissemination of knowledge. Moreover, thanks to suitable
procedures of automated reasoning, it is possible to extract implicit
information present in data, leveraging a deeper knowledge of the domain. The domain
is speci ed by means of expressions describing statements about web shared
resources. Such expressions are given as triples of the form
subject-predicateobject. The subject denotes the resource to describe, namely the actor of the
statement, the object denotes the recipient or the result of the action, and the
predicate denotes traits or aspects of the resource, i.e., a relationship between
the subject and the object. For instance, if we wanted to express the concept
\Michelangelo made the Sistine Chapel ceiling", we would write
&lt;#Michelangelo&gt;&lt;#makes&gt;&lt;#Sistine Chapel Ceiling&gt;, where \#" indicates that the
local context is used.1 Every time we want to add some information about a
1 For space reasons, we express triples using a simpli ed version of the Turtle notation
[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
speci c resource, we reuse the resource as subject or object of a new triple,
depending on the type of relationship. For example, if we additionally wanted to
express the creation date of the Sistine Chapel ceiling, we would write
&lt;#Sistine Chapel Ceiling&gt;&lt;#hasCreationDate&gt;\1512".
      </p>
      <p>Sharing data is an important feature of the semantic web. Because of the
uniqueness of the URI of the resources and of the document, we can refer to them
without ambiguity. For instance, assuming that the URL http://wwww.unict.
it/art.rdf refers to the document containing information about
\Michelangelo", every time we need to state something about him, we use the
document http://wwww.unict.it/art.rdf, and the resource http://wwww.unict.
it/art.rdf#Michelangelo as a resource in our statements.</p>
      <p>
        Semantic web data are usually published using the Resource Description
Framework (RDF) [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] (or its extension RDF Schema (RDFS) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]) and the
Ontology Web Language (OWL) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. An important feature of the semantic web is
the capability to extract implicit information from the described data. This is
why the language RDFS has been introduced. Unlike the RDF language, RDFS
is powerful enough to enable such a feature since it allows to express subclass
and subproperty relationships. Elements of the domain sharing common
characteristics can be grouped in particular sets called classes. Classes, in their turn,
can be organised in hierarchies. A hierarchy of classes is called a taxonomy.
Properties can be organised in hierarchies too. For instance, the relation \has father"
can be modelled as a subproperty of the relation\has ancestor". If an element of
the domain is related to another element by a subproperty, then it is related to
the superproperty as well. RDFS provides other interesting inference capabilities
that, for space reasons, we do not report here. However, RDFS is far away from
allowing complex reasoning. For instance, let us consider the relation \uncle" in
a domain of persons. If we express that \Marta" is daughter of \Frank", and
\Frank" is brother of \John", we are making no assumption on the relationship
between \Marta" and \John". However, in the knowledge domain we have to
model we might be aware of the fact that \the brother of my father is my uncle"
as a result of the combination of sibling-daughter properties. Relationships of
this type can be expressed in OWL using a multitude of properties.
      </p>
      <p>
        The task of writing OWL data requires the de nition of an ontology.
Informally, in computer science an ontology de nes a set of representational primitives
(classes and properties) apt to model a domain of knowledge or of discourse [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ].
When an OWL reasoner (namely, a given software system able to extract
information from OWL les) is executed on OWL data, we can perform the task
of mining data from resources. Although an RDFS schema can be potentially
converted into an OWL ontology (such a task is known as RDF alignment [
        <xref ref-type="bibr" rid="ref6 ref7">6,7</xref>
        ]),
this is not su cient to turn an RDF document into an OWL ontology. RDF and
OWL partially share the syntax but not the semantics. In fact, OWL allows
one to specify far more about the properties and classes of an RDFS schema
by means of a formal description of the data. Expressiveness of such
ontological model depends on the OWL pro le adopted. Pro les are fragments of the
language that trade o some expressive power for the e ciency of reasoning,
introduced in order to deal with several types of application domains. Some of
these pro les include SWRL rules. Such rules have the form of an implication
between an antecedent (body) and consequent (head). The intended meaning
can be read as: whenever the conditions speci ed in the antecedent hold, then
the conditions speci ed in the consequent must hold as well. The desirable
features of the OWL language shortly outlined above strongly motivate our interest
in using OWL ontologies and related reasoning tools for the location recognition
task. More details about OWL reasoning capabilities, semantics, and pro les can
be found in [
        <xref ref-type="bibr" rid="ref11 ref4 ref7">4, 7, 11</xref>
        ]. In particular, in [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], a reasoning system based on a
fragment of set-theory is proposed particularly suitable for the ontologies presented
in this paper.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Rule-Based Location Extraction</title>
      <p>
        In this section we brie y illustrate the approach we used to extract
geographical and administrative places from Italian non-structured text. Unlike machine
learning approaches, both unsupervised and supervised, we proposed a
rulebased approach built from simple grammar rules of the Italian language
complemented by a dictionary, where each rule identi es a di erent pattern that
characterises sentences at the end of which we usually nd a location name. The
devised rules are supported by a speci cally compiled Italian lexicon, containing
the classes of words Articles, Verbs, and Descriptors, and a list of Non-places
words that are known false positives; in order to improve the overall accuracy
of the extraction tool, the lexicon can also be extended by other, user-detected
false positives [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. As shown in [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], these rules provide a large coverage of the
Italian grammar for what concerns statements about places.
      </p>
      <p>Figure 1 depicts the entire work ow used for the location extraction: in the
rst sentence splitting step, an input text T is separated into a list of sentences
using occurrences of punctuation marks, i.e., full-stop, ellipsis,
exclamationmark, question-mark. Any other non-letter symbol is ignored, e.g., dollar sign,
percent sign, etc. Each sentence is further segmented into words using the space
character as a separator. The tokens (words) are then fed to a nite state
machine implementing three di erent rules, i.e., grammar cases possibly implying
the use of a place name at the end of a sentence: if the accepting state (6) of
the automaton is reached, the current token is marked as a location candidate.
The result is a list of candidate words that must pass through a con gurable
sequence of lters before being actually labelled as a place name.
start$
0"</p>
      <p>D$
A/P$
V$</p>
      <p>D$
V$
1"
2"
3"
4"
5"</p>
      <p>A/P$</p>
      <p>A/P$
A/P$
else$
else$
else$
else$
else$
A/P$</p>
      <p>A/P$</p>
      <p>6"
A:$ar2cle $
P:$preposi2on
D:$descriptor</p>
      <p>Tokens$
$V:$place5related$verb$
$L:$candidate$place$
$X:$anything$else$</p>
      <p>
        Figure 2 shows the uni ed automaton used in this contribution to implement
the three grammar rules de ned in [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]; those rules were devised to accommodate
the mentioning of a place name in sentences such as the rule names suggest,
namely, Da Roma (from Rome), Vicino a Roma (Near Rome) and Andando a
Roma (Going to Rome). To give a simple detection example, Figure 3 shows an
excerpt of the uni ed automaton giving a nite state machine expressing only the
rst rule, (Da Roma), whose name recalls the grammar pattern responsible for
the matching. The automaton scans the tokens of a given sentence and remains
in state 0 until a preposition (P) or an article (A) is found, at which point the
automaton changes its internal state to 1. Subsequently, the automaton remains
in state 1 until a di erent kind of word is encountered, in which case the nal
state is reached and a new candidate word is found. However, several candidates
will be actually dropped afterwards by the lters, e.g., known false positives or
conjugated verbs.
      </p>
      <p>A/P$
3"</p>
      <p>A/P$
else$
6"</p>
      <p>All the candidates found by the automaton are then given as input to a
sequence of lters in order to remove trivial false positives that may have been
selected:
{ Filter0: the candidate for a location name must begin with a capital letter;
even if it may considerably improve the detection accuracy in several cases,
this is not a mandatory lter. In fact it is applied only when we can assume
that location names are written with a leading capital letter (e.g., if we are
analysing newspaper articles);
{ Filter1: remove all known false positives using the devised lexicon of
nonplaces;
{ Filter2: remove conjugated verbs.</p>
      <p>Any word surviving the above lters is labelled as a place name, and is
given to the ontologic support to store the results and automatically retrieve the
information concerning the algorithm and dataset used, and the spotted place.
4</p>
    </sec>
    <sec id="sec-4">
      <title>An ontology for reasoning with places</title>
      <p>In this section we rst describe the ontology OntoLocEstimation, developed with
the purpose of reasoning with geographical and administrative places. Then we
show how the ontology is populated by means of a Java framework.
OntoLocEstimation is associated to the algorithm introduced in Section 3 and it allows us
to manage the identi cation of a speci c location even in presence of uncertainty.
4.1</p>
      <sec id="sec-4-1">
        <title>The ontological model</title>
        <p>
          We illustrate how the ontology OntoLocEstimation is structured.
OntoLocEstimation extends the ontology OntoLuoghi [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] with OWL constructs allowing
us to deal with the administration of Italian places and with the algorithm
described in Section 3. In its turn, OntoLuoghi reuses some concepts and properties
of LinkedGeoData [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
        </p>
        <p>Locations are modelled by means of a taxonomy of OWL classes. Association
among locations is performed through a taxonomy of object-properties. The path
allowed, namely the hierarchy of such classes, is shown in Figure 5.
Doublehoop entities in Figure 5 are considered as optional. Names of the OWL classes
and object-properties involved are shown in Figure 4.2 Reasoning capabilities
concerning locations are strengthened using the SWRL rules shown in Figure 6.</p>
        <p>The entity \Localisation" is equivalent to \LinkedGeoData:Place"
(equivalences among subclasses of \Localisation" and of \Place" are not reported here
for space reasons). The object-property \hasLocalisation" is de ned as an OWL
transitive property. Thus, if the pairs of objects (x; y) and (y; z) are in the
property \hasLocalisation", then the pair (x; z) is included in the property
\hasLocalisation" too. The property \hasLocalisation" can be used together with its
subproperties to infer the administrative hierarchy of a location providing only the
top level of a place (i.e., the superproperty). For instance, if we write the
statements (Sicily hasState Italy), (Catania hasRegion Italy), (Acireale hasProvince
Catania), then (Sicily hasLocalisation Italy), (Catania hasLocalisation Italy),
(Acireale hasLocalisation Catania), (Acireale hasLocalisation Sicily), (Acireale
hasLocalisation Italy) are inferred thanks to the subproperty relationships and
to the transitivity of \hasLocalisation". In addition, the set of SWRL rules
depicted in Figure 6 allows one to infer the statements (Catania hasState Italy),
(Acireale hasRegion Sicilia), (Acireale hasState Italy). Moreover, subproperties
are de ned as functional properties so as to guarantee that places can not be
associated to di erent locations. For instance, Acireale can not be associated to
the province of Catania and Palermo at the same time.</p>
        <p>
          Usually, an algorithm recognising a geographical location from a keyword
in a non-structured text yields several candidates, each in combination with a
degree of belief. The degree of belief indicates how much the association between
the keyword and the candidate geographical place is considered reliable by the
approach in use. The ontology should model this situation and also manage
additional information such as the algorithm applied for extracting location names
from the text and the dataset adopted. Such information is relevant if one wishes
to compare the accuracy that di erent algorithms and datasets provide for the
task of recognising location names. Our approach works as follows. Every time
the algorithm nds a keyword, an instance of the class \TextKey" is added
to the dataset. The class \LocEstimation" models the fact that the algorithm
of Section 3 is executed on the keyword and some results are provided. The
instances of \Textkey" and \LocEstimation" are related to each other by the
2 Images in Figures 4, 7, 8, 9, and 10 are drawn from the interface of the editor
Protege [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
object-property \hasLocEstimation". For every estimation there may be zero,
one, or more matches. Each match of the algorithm is modelled by the
\GuessedLocation" class. Instances of the class \GuessedLocation" are associated to the
ones of the class \LocEstimation" by the object-property \hasGuessedLocation".
Each instance of the class \GuessedLocation" provides information concerning
the candidate geographical place and the relative degree of belief. The degree
of belief is introduced by means of the data-property \hasGuessedValue"
having the data-type double as range. The geographical place is speci ed with the
object-property \hasReferredLoction" having as range every instance of the class
\Localisation". Figure 7 illustrates classes and properties introduced in order to
model the recognition of places, an example is shown in Figure 8.
        </p>
        <p>The ontology provided takes also into account information concerning the
algorithm for the extraction of the location and the dataset used. Classes and
object-properties de ned for this purpose are shown in Figure 9.</p>
        <p>A class \Algorithm" provides information about the algorithm applied.
Currently, we identify two types of algorithms, but others can be added. The class
\DetectionAlgorithm" includes all the algorithms used to establish in a
nonstructured text whether a word represents a geographical place while the class
\EstimationAlgorithm" includes information about the algorithm used to choose
a set of possible geographical places and to assign them a degree of belief. In
addition, it is possible to specify the dataset used by the algorithm by means of the
class \Dataset". The subclasses of the class \Dataset" describe the type of the
dataset. For instance, the class \LocationGeneralDataset" is used to represent
the datasets containing general geographical information.</p>
        <p>In order to keep track of the algorithms used, we provide the class
\AlgorithmSet" that is associated to an instance of the class \LocEstimation" by
means of the object-property \hasAlgorithmSet". To each instance of the class
\AlgorithmSet", one or more instances of the class \AlgorithmCore" are
associated by means of the object-property \hasAlgorithmCore". The latter class
relates an algorithm to the datasets used. In order to indicate the datasets used,
the \hasDataSet" object-property is provided, having as range the \DataSet"
class. Analogously, to associate an algorithm to an instance of the class
\DataCore", the object-property \hasAlgorithm" is provided, having as range the
class \Algorithm". In particular, two subproperties of the class \hasAlgorithm"
are provided: the property \hasDetectionAlgorithm", that associates an instance
of the class \AlgorithmCore" to an instance of the class \DetectionAlgorithm",
and the property \hasEstimationAlgorithmCore" for the instances of the class
\EstimationAlgorithm". How these classes and relations are used is shown in
the example of Figure 10. Note that, by means of the class \AlgorithmCore",
additional information about why, how, and when the algorithm set is used can
be tracked.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Populating the ontology</title>
        <p>
          We describe shortly how the ontology OntoLocEstimation is populated by means
of a built ad-hoc Java framework. The rst step of the process is to retrieve an
open dataset of locations, that is as rich as possible, and to map such knowledge
inside the ontology. We take into account the OpenStreetMap dataset and
implement a Java parser in order to populate the ontology with the OpenStreetMap
entries. The parser exploits the OWL API library [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] together with Jena
Ontology API [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] that we adopted to perform SPARQL [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] queries. We used the
reasoner Pellet [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ] to carry out inferences on our ontology. As far as we know,
Pellet provides the best deal between e ciency and data-type reasoning
capabilities.
        </p>
        <p>In a preliminary phase of this work we also considered the LinkedGeoData
dataset that provides a semi-automated conversion of a subset of the
OpenStreetMap dataset in RDF format. However, as outlined at the beginning of this
paper, conversion of RDF in OWL is not straightforward. Thus, being mainly
interested in OWL reasoning, we opted to provide the mapping of OpenStreetMap
data in OWL statements by means of a simple parser.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>We have presented an application that automatically recognises locations inside
non-structured texts written in Italian and that is supported by an OWL
ontology for locations management. The ontology also keeps track of the algorithm
applied for detecting places and stores the degree of belief of each candidate
location. Di erent algorithmic approaches can be stored in the ontology and,
consequently, compared. Data of geographical locations are retrieved from an
open dataset, i.e. OpenStreetMap, and adapted to the logical model of the
ontology.</p>
      <p>
        The results presented here can nd applications in several contexts, such as
the digital humanities eld. The techniques illustrated can be adapted to
recognise other kinds of elements inside non-structured text, such as descriptors of
archaeological ndings and nding places. This process simpli es the task of the
automatic processing of archaeological archives, digitalisation and subsequent
conversion in linked data format. In [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] we proposed an ontology for pottery
classi cation, cataloguing, and reasoning that lends itself particularly well to
such a task. This approach, combined with OWL reasoning capabilities, allows
one to gain a deeper knowledge and better dissemination of the considered
application domain.
      </p>
      <p>We plan to introduce additional lters and rules in the algorithm so as to
gain a larger coverage of the Italian grammar. In particular, we aim to make
the algorithm sensitive to the contexts from which words are drown, namely
non-structured texts. Finally, we plan to abandon external datasets, i.e.
OpenStreetMap, in favour of internal built-in datasets that can be integrated by
information provided by local governments or nal users.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>Work partially supported by the project PRIME: Piattaforma di Reasoning
Integrata, Multimedia, Esperta within PO FESR Sicilia 2007/2013, and by the FIR
projects COMPACT: Computazione a dabile su testi rmati, code D84C46, and
Organizzazione e trattamento di trascrizioni e testi in scenari di security, code
375E90.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>1. \Apache stanbol," https://stanbol.apache.org.</mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. \Jena
          <string-name>
            <surname>Ontology</surname>
            <given-names>API</given-names>
          </string-name>
          ,
          <article-title>"</article-title>
          https://jena.apache.org.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>3. \LinkedGeoData," http://linkedgeodata.org.</mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>4. \Ontology Web Language," http://www.w3.org/2001/sw/wiki/OWL.</mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>5. \OpenStreetMap," http://www.openstreetmap.org.</mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. \OWL 2
          <string-name>
            <surname>Mapping to</surname>
            <given-names>RDF</given-names>
          </string-name>
          ," http://www.w3.org/TR/owl2-mapping
          <string-name>
            <surname>-</surname>
          </string-name>
          to-rdf/.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>\OWL Semantics Mapping</surname>
          </string-name>
          ," http://www.w3.org/TR/owl-semantics/mapping. html.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>8. \Protege," http://webprotege.stanford.edu.</mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>9. \The Resource Description Framework," http://www.w3.org/RDF/.</mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10. \
          <article-title>The Resource Description Framework Schema,"</article-title>
          http://www.w3.org/TR/ rdf-schema/.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>\Semantic Web Rule Language</surname>
          </string-name>
          ," http://www.w3.org/Submission/SWRL/.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12. \
          <article-title>SPARQL Query Language for RDF,"</article-title>
          https://www.w3.org/TR/ rdf-sparql-query/.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>13. \Turtle Notation," http://www.w3.org/TR/turtle/.</mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <given-names>E.</given-names>
            <surname>Agichtein</surname>
          </string-name>
          and
          <string-name>
            <given-names>L.</given-names>
            <surname>Gravano</surname>
          </string-name>
          , \Snowball:
          <article-title>Extracting relations from large plaintext collections,"</article-title>
          <source>in Proceedings of the fth ACM conference on Digital Libraries. ACM</source>
          ,
          <year>2000</year>
          , pp.
          <volume>85</volume>
          {
          <fpage>94</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <given-names>A.</given-names>
            <surname>Ballatore</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. C.</given-names>
            <surname>Wilson</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Bertolotto</surname>
          </string-name>
          , \
          <article-title>A survey of volunteered open geo-knowledge bases in the semantic web," in Quality issues in the management of web information</article-title>
          . Springer,
          <year>2013</year>
          , pp.
          <volume>93</volume>
          {
          <fpage>120</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <given-names>D.</given-names>
            <surname>Cantone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Longo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nicolosi-Asmundo</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D. F.</given-names>
            <surname>Santamaria</surname>
          </string-name>
          , \
          <article-title>Web ontology representation and reasoning via fragments of set theory,"</article-title>
          <source>in Web Reasoning and Rule Systems - 9th International Conference, RR 2015</source>
          , Berlin, Germany,
          <source>August 4-5</source>
          ,
          <year>2015</year>
          , Proceedings,
          <year>2015</year>
          , pp.
          <volume>61</volume>
          {
          <fpage>76</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>D.</given-names>
            <surname>Cantone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nicolosi-Asmundo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. F.</given-names>
            <surname>Santamaria</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Trapani</surname>
          </string-name>
          , \
          <article-title>Ontoceramic: an OWL ontology for ceramics classi cation,"</article-title>
          <source>in Proceedings of the 30th Italian Conference on Computational Logic</source>
          ,
          <string-name>
            <surname>CILC</surname>
          </string-name>
          <year>2015</year>
          , vol.
          <volume>1459</volume>
          .
          <source>Genova: CEUR Electronic Workshop Proceedings, July 1-3</source>
          <year>2015</year>
          , pp.
          <volume>122</volume>
          {
          <fpage>127</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <given-names>D.</given-names>
            <surname>Caruso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Giunta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Messina</surname>
          </string-name>
          , G. Pappalardo, and E. Tramontana, \
          <article-title>Rulebased location extraction from Italian unstructured text,"</article-title>
          <source>in Proceedings of XVI Workshop "</source>
          From Object to Agents
          <source>"(WOA)</source>
          , vol.
          <volume>1382</volume>
          ,
          <string-name>
            <surname>July</surname>
          </string-name>
          17-19
          <year>2015</year>
          , pp.
          <volume>46</volume>
          {
          <fpage>52</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <given-names>M.</given-names>
            <surname>Horridge</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Bechhofer</surname>
          </string-name>
          , \
          <article-title>The OWL API: A Java API for OWL ontologies,"</article-title>
          <source>Semantic Web</source>
          , vol.
          <volume>2</volume>
          , no.
          <issue>1</issue>
          , pp.
          <volume>11</volume>
          {
          <issue>21</issue>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <given-names>J.</given-names>
            <surname>Kleb</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Volz</surname>
          </string-name>
          , \
          <article-title>Ontology based entity disambiguation with natural language patterns,"</article-title>
          <source>in Fourth International Conference on Digital Information Management</source>
          ,
          <year>2009</year>
          .
          <article-title>ICDIM 2009</article-title>
          . IEEE,
          <year>2009</year>
          , pp.
          <volume>1</volume>
          {
          <fpage>8</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21. J.
          <article-title>La erty, A. McCallum, and</article-title>
          <string-name>
            <given-names>F. C.</given-names>
            <surname>Pereira</surname>
          </string-name>
          , \
          <article-title>Conditional random elds: Probabilistic models for segmenting and labeling sequence data,"</article-title>
          <source>in Proceedings of the Eighteenth International Conference on Machine Learning (ICML)</source>
          ,
          <year>2001</year>
          , pp.
          <volume>282</volume>
          {
          <fpage>289</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22. L. Liu and
          <string-name>
            <given-names>M. T.</given-names>
            <surname>Zsu</surname>
          </string-name>
          ,
          <source>Encyclopedia of database systems</source>
          . Springer Publishing Company, Incorporated,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <given-names>K.</given-names>
            <surname>Nigam</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          <article-title>La erty, and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>McCallum</surname>
          </string-name>
          , \
          <article-title>Using maximum entropy for text classication,"</article-title>
          <source>in IJCAI-99 Workshop on Machine Learning for Information Filtering</source>
          , vol.
          <volume>1</volume>
          ,
          <issue>1999</issue>
          , pp.
          <volume>61</volume>
          {
          <fpage>67</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24. S. Sarawagi, \
          <article-title>Information extraction," Foundations and trends in databases</article-title>
          , vol.
          <volume>1</volume>
          , no.
          <issue>3</issue>
          , pp.
          <volume>261</volume>
          {
          <issue>377</issue>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25. E.
          <string-name>
            <surname>Sirin</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Parsia</surname>
            ,
            <given-names>B. C.</given-names>
          </string-name>
          <string-name>
            <surname>Grau</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Kalyanpur</surname>
            , and
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Katz</surname>
          </string-name>
          , \
          <article-title>Pellet: a practical OWL-DL reasoner," Web Semantics: science</article-title>
          ,
          <source>services and agents on the World Wide Web</source>
          , vol.
          <volume>5</volume>
          , no.
          <issue>2</issue>
          , pp.
          <volume>51</volume>
          {
          <issue>53</issue>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>