<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Machine Reader for the Semantic Web</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Aldo Gangemi</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesco Draicchio</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Valentina Presutti</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrea Giovanni Nuzzolese</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Diego Reforgiato</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>FRED is a machine reading tool for converting text into internally well-connected and quality linked-data-ready ontologies in webservice-acceptable time. It implements a novel approach for ontology design from natural language sentences, combining Discourse Representation Theory (DRT), linguistic frame semantics, and Ontology Design Patterns (ODP). The current version of the tool includes Earmark-based markup, and enrichment with word sense disambiguation (WSD) and named entity resolution (NER) o -the-shelf components.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The problem of knowledge extraction (KE) from text is still insu ciently
addressed from a semantic web (SW) perspective. Being able to automatically
produce quality linked data and ontologies from natural language text would be
a breakthrough as it would enable the development of applications that
automatically produce machine-readable information from Web content as soon as
it is edited and published by generic Web users. A rather detailed landscape
analysis of the currently available tools for KE, and their exploitation for SW
basic tasks is presented in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]: it shows the substantial lacking of tools for
creating RDF graphs that are connected enough to perform application tasks such
as event extraction, fact detection, story mining, etc.
      </p>
      <p>
        FRED4 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is an exception, since it is intended to produce semantic data
and ontologies with a quality closer to what is expected at least from average
linked datasets and vocabularies: FRED candidates as a deep version of a
machine reader [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] for the Semantic Web. The following requirements have inspired
the design of FRED: (i) ability to capture accurate semantic structures (i.e.
compliant to formal semantics); (ii) representing complex relations (i.e. n-ary,
multigrade relations); (iii) exploitation of sophisticated lexical resources (e.g.
VerbNet, FrameNet); (iv) no need of large-size domain-speci c text corpora and
training sessions (i.e. we address open information extraction); (v) minimal time
of computation; (vi) ability to map natural language to RDF/OWL
representations; (vii) ability to link the extracted knowledge to both lexical linked data
and linked datasets (for maximal interoperability).
4 http://wit.istc.cnr.it/stlab-tools/fred
      </p>
      <p>
        Related works and comparison to other tools for knowledge extraction are
detailed in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], and [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], where FRED seems to outperform (though on a limited
test) the other tools in sophisticated tasks such as relation and factoid extraction,
frame detection, and taxonomy induction.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>FRED at work</title>
      <p>
        In this section we present an overview of the system and a scenario that shows
the output resulting from FRED. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] shows that detecting the most appropriate
frames from the input text leads to improve the design quality of the
resulting ontology because frames can be directly mapped to an important variety
of ontology design patterns based on n-ary relations. On the above
consideration, FRED makes use of Boxer [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], a deep semantic parser based on categorial
grammar and Discourse Representation Theory (DRT), which generates formal
semantic representation of text through an event (neo-Davidsonian) semantics.
      </p>
      <p>Boxer frame-based approach supports FRED in automatically design an
ontology by following good modeling practices based on ontology patterns. However,
Boxer tranforms natural language to a logical form compliant with DRT
(substantially a variety of rst-order logic) which di ers a lot from RDF or OWL,
and the heuristics that it implements for interpreting a natural language and
transforming it to a DRT-based structure can be sometimes awkward when
directly translated to ontologies for the SW, because it obeys pure FOL-oriented
design style. For this reason, Boxer DRT-based output is transformed by FRED
to OWL/RDF ontologies by means of a set of heuristics, some of them are showed
in Figure 1. Figure 2 depicts the main components of FRED: Communication:
exposes APIs for querying the system; Refactoring : transforms Boxer output5
into a convenient data structure to be passed to the Reenginering component;
Reengineering : applies a collection of ad-hoc mapping rules and heuristics for
producing logically consistent OWL/RDF like triples.</p>
      <p>
        In addition, FRED architecture is open to be easily integrated with other
components that exploit the text span markup speci cation supported its current
version, i.e. Earmark [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. This solution, which is similar to architectures such as
      </p>
      <sec id="sec-2-1">
        <title>5 Boxer is an external component.</title>
        <p>NIF and NERD, makes it trivial to augment FRED graphs with o -the-shelf
components for e.g. NER, WSD, etc. Our demo, available online6, allows a user
to enter any text (or select one from the list of examples provided), and by simply
clicking the Read It! button, to receive a RDF/OWL representation of it7. Some
features can be customized, e.g., type of output, NER or WSD activation,
tenserelation between events activation, etc. By default, the output is in the form
of Graphviz-like graphs (showing only a core subset of triples), to allow human
users to quickly check the OWL/RDF representation.</p>
        <p>FRED output consists in RDF triples including either xed properties: rdf:type,
rdfs:subClassOf, owl:sameAs, dul:associateWith, owl:equivalentTo,
Earmark properties, thematic roles for extracted events, etc., or customized
properties produced by the automatic (machine) reading of the sentence.</p>
      </sec>
      <sec id="sec-2-2">
        <title>6 http://wit.istc.cnr.it/stlab-tools/fred 7 A graphical output is provided for human users</title>
        <p>As an example, consider the sentence \The Black Hand assassinated Franz
Ferdinand during his visit to Sarajevo." Figure 3 shows FRED output for this
sentence: an instance of dul:Event, fred:assassinate 1, is used to represent
the assassination of Franz Ferdinand. It is typed as fred:Assassinate, which
is disambiguated by the VerbNet frame vn.data:Assassinate 42010000. Such
an event involves the individual fred:Black hand as agent, and the
individual fred:Franz ferdinand (who is recognized and resolved as the same
individual as dbpedia:Archduke Franz Ferdinand of Austria) as patient.
Furthermore, the RDF graph expresses that such an event happened fred:during
fred:visit 1. FRED heuristically assigns a type fred:Visit to fred:visit 1;
such a type is disambiguated by means of alignments to WordNet, which in turn
is aligned to other ontologies, so that FRED can infer e.g., that fred:Visit is a
d0:Activity. The location of the visit is also identi ed and correctly resolved to
dbpedia:Sarajevo. Notice that FRED assigns types to all identi ed individuals
either by using classes from existing ontologies e.g., DOLCE, Schema.org, etc.,
when the entity can be resolved e.g., as a DBpedia entity, or by creating new
classes based on the terms used in the input text and disambiguating them on
WordNet.</p>
        <p>FRED also supports more sophisticated constructs, e.g. propositional
referents, called situations, full- edged negation on events or situations, and basic
modalities over events.
3 Conclusion
We have presented FRED, a machine reader for the Semantic Web, which
automatically extracts rich and connected knowledge from text, represents it as
OWL/RDF, and links it to other resources: VerbNet, FrameNet, WordNet,
DBpedia, foundational ontologies, etc. The current research is on mainly evaluating
it on vertical tasks, and extending its internal components for multilinguality.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Johan</given-names>
            <surname>Bos</surname>
          </string-name>
          .
          <article-title>Wide-Coverage Semantic Analysis with Boxer</article-title>
          .
          <source>In Johan Bos and Rodolfo Delmonte</source>
          , editors,
          <source>Semantics in Text Processing</source>
          , pages
          <volume>277</volume>
          {
          <fpage>286</fpage>
          .
          <string-name>
            <surname>College</surname>
            <given-names>Publications</given-names>
          </string-name>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Bonaventura</given-names>
            <surname>Coppola</surname>
          </string-name>
          , Aldo Gangemi, Al o Massimiliano Gliozzo, Davide Picca, and
          <string-name>
            <given-names>Valentina</given-names>
            <surname>Presutti</surname>
          </string-name>
          .
          <article-title>Frame detection over the semantic web</article-title>
          . In Lora Aroyo et al., editor,
          <source>ESWC</source>
          , volume
          <volume>5554</volume>
          <source>of LNCS</source>
          , pages
          <volume>126</volume>
          {
          <fpage>142</fpage>
          . Springer,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Oren</given-names>
            <surname>Etzioni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Michele</given-names>
            <surname>Banko</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Michael</given-names>
            <surname>Cafarella</surname>
          </string-name>
          .
          <article-title>Machine reading</article-title>
          .
          <source>In Proceedings of the 21st National Conference on Arti cial Intelligence (AAAI)</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Aldo</given-names>
            <surname>Gangemi</surname>
          </string-name>
          .
          <article-title>A comparison of knowledge extraction tools for the semantic web</article-title>
          .
          <source>In Proceedings of ESWC2013</source>
          . Springer,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Valentina</given-names>
            <surname>Presutti</surname>
          </string-name>
          , Francesco Draicchio, and
          <string-name>
            <given-names>Aldo</given-names>
            <surname>Gangemi</surname>
          </string-name>
          .
          <article-title>Knowledge extraction based on discourse representation theory and linguistic frames. In EKAW: Knowledge Engineering and Knowledge Management that matters</article-title>
          . Springer,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Peroni</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gangemi</surname>
            <given-names>A.</given-names>
          </string-name>
          , and Vitali F.
          <article-title>Dealing with markup semantics</article-title>
          .
          <source>In Proceedings of the 7th International Conference on Semantic Systems</source>
          , Graz,
          <source>Austria (i-Semantics2011)</source>
          . ACM,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>