<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>TRENTINOMEDIA: Exploiting NLP and Background Knowledge to Browse a Large Multimedia News Store?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Roldano Cattoni</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesco Corcoglioniti</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christian Girardi</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bernardo Magnini</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luciano Serafini</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roberto Zanoli</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DISI, University of Trento</institution>
          ,
          <addr-line>Via Sommarive 14, 38123 Trento</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Fondazione Bruno Kessler</institution>
          ,
          <addr-line>Via Sommarive 18, 38123 Trento</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>TRENTINOMEDIA provides access to a large and daily updated repository of multimedia news in the Trentino Province. Natural Language Processing (NLP) techniques are used to automatically extract knowledge from news, which is then integrated with background knowledge from (Semantic) Web resources and exploited to enable two interaction mechanisms that ease information access: entity-based search and contextualized semantic enrichment. TRENTINOMEDIA is a real multimodal archive of public knowledge in Trentino and shows the potential of linking knowledge and multimedia and applying NLP on a large scale.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>Finding information about an entity in a large news collection using standard
keywordbased search may be time-consuming. Searching for a specific person, for instance,
may return a large list of news about homonymous persons, that need to be checked and
filtered manually. Also, understanding the contents of a news can be expensive, if the
user is not familiar with the entities mentioned and needs information about them.</p>
      <p>The presented TRENTINOMEDIA system shows how the use of NLP and Semantic
Web techniques may help in addressing these problems. TRENTINOMEDIA supports the
“smart” access to a large and dynamic (daily updated) repository of multimedia news in
the Italian Trentino Province. “Smart” means that NLP techniques are used to
automatically extract knowledge about the entities mentioned in the news. Extracted knowledge
is then integrated with background knowledge about the same entities gathered from
(Semantic) Web resources, so to build a comprehensive knowledge base of entity
descriptions linked to the news of the collection. Exploiting the interlinking of knowledge
and multimedia, two interaction mechanisms are provided to ease information access:
– entity-based search, enabling a user to find exactly the news about a specific entity;
– contextualized semantic enrichment, consisting in the visualization of additional
knowledge about a mentioned entity that may ease a user’s understanding of a news.</p>
      <p>Two main usages are foreseen for TRENTINOMEDIA: (i) a professional usage,
restricted to the news providers and aimed at addressing internal needs, including
automatic news documentation, support tools for journalists and integration of advanced
? This work was supported by the LiveMemories project (http://www.livememories.org) funded
by the Autonomous Province of Trento (Italy) under the call “Major Project”.
Resource
type (e.g. video)
DC metadata
multimedia file</p>
      <p>▲contained in
Mention
position
extent
extracted attributes
coreference cluster</p>
      <p>▼ denotes
Entity
name</p>
      <p>▼ described by
Triple
subject (implicit)
predicate
object
inter-resource
relations (part
of, caption of,
related to,
derived from)
matching
▼ context
Context
time
location
topic
▲holds in</p>
      <p>Linking &amp;
entity creation
Coreference
resolution
Mention
extraction</p>
      <p>Resource
preprocessing</p>
      <p>Content
acquisition</p>
      <p>GeoNames</p>
      <p>Coop.</p>
      <p>Trentina
NEWS</p>
      <p>BACKGROUND KNOWLEDGE
functionalities in existing editorial platforms; and (ii) an open use by citizens through
on-line subscriptions to news services, possibly delivered to mobile devices.</p>
      <p>The presented work has been carried out within the LiveMemories project, aimed
at automatically interpreting heterogeneous multimedia resources, transforming them
into “active memories”. The remainder of the paper presents the system architecture in
section 2 and the demonstrated user interface in section 3, while section 4 concludes.
2</p>
    </sec>
    <sec id="sec-2">
      <title>System architecture</title>
      <p>The architecture of TRENTINOMEDIA is shown in figure 1 and includes three
components: the KNOWLEDGESTORE, a content processing pipeline and a Web frontend.</p>
      <p>
        The KNOWLEDGESTORE [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] builds on Hadoop3 and Hbase4 to provide a scalable
storage for multimedia news and background knowledge, which are represented
according to the (simplified) schema of figure 2. News are stored as multimedia resources,
which include texts, images and videos. Knowledge is stored as a contextualized
ontology. It consists of a set of entities (e.g., “Michael Schumacher”) which are described
by hsubject, predicate, objecti RDF [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] triples (e.g., hMichael Schumacher, pilot of,
Mercedes GPi). In turn, each triple is associated to the htime, space, topici context the
represented fact holds in (e.g., h2012, World, Formula 1i). The representation of
contexts follows the Contextualized Knowledge Repository approach [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and permits to
accommodate “conflicting” knowledge holding under different circumstances (e.g. the
fact that Schumacher raced for different teams). Resources and entities are linked by
mentions, i.e. proper names in a news that refer to an entity. They permit to navigate
from a news to its mentioned entities and back, realizing the tight interlinking of
knowledge and multimedia at the basis of the interaction mechanisms of TRENTINOMEDIA.
      </p>
      <p>Concerning the content processing pipeline, it integrates a number of NLP tools
to load, process and interlink news and background knowledge, resulting in the full</p>
      <sec id="sec-2-1">
        <title>3 http://hadoop.apache.org 4 http://hbase.apache.org</title>
        <p>
          population of the schema in figure 2. Apart from the loading of background
knowledge, which is bootstrapped by manually selecting and wrapping the relevant
knowledge sources, the pipeline works automatically and incrementally, processing news as
they are collected daily. The rest of this section describes the processing steps of the
pipeline, while the user interface of the frontend is described in the next section.
Content acquisition. News are supplied daily by a number of news providers local to
the Trentino region. They are in Italian, cover a time period from 1999 to 2011 and
consist of text articles, images and videos. Loading of news is performed automatically and
table 1 shows some statistics about the news collected so far. Background knowledge
is collected manually through a set of ad-hoc wrappers from selected (Semantic) Web
sources, including selected pages of the Italian Wikipedia, sport-related community
sites and the sites of local and national public administrations and government bodies.
Overall, it consists of 352,244 facts about 28,687 persons and 1,806 organizations.
Resource preprocessing. Several operations are performed on stored news with the
goal of easing their further processing in the pipeline. Preprocessing includes the
extraction of speech transcription from audio and video news, the annotation of news
with a number of linguistic taggers (e.g., part of speech tagging and temporal
expression recognition, performed using the TextPro tool suite5 [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]) and the segmentation of
complex news in their components (e.g., the separation of individual stories in a news
broadcast or the extraction of texts, figures and captions from a complex XML article).
Mention extraction. Textual news are processed with TextPro to recognize mentions
of three types of named entities: persons, organizations and geo-political / location
entities. For each mention, a number of attributes is extracted from the mention and its
surrounding text. Given the text “the German pilot Michael Schumacher”, for instance,
the system recognizes “Michael Schumacher” as a person mention and annotates it with
FIRSTNAME “Michael”, LASTNAME “Schumacher”, ROLE “pilot” and NATIONALITY
“German”. Attributes are extracted based on a set of rules (e.g., to split first and last
names) and language-specific lists of words (e.g., for nationalities, roles, . . . ). Statistics
about the mentions recognized so far are reported in the second column of table 2.
Coreference resolution. This step consists in grouping together in a mention cluster
all the mentions that (are assumed to) refer to the same entity, e.g., to decide that
mentions “Michael Schumacher” and “Schumi” in different news denote the same person.
Two coreference resolution systems are used. Person and organization mentions are
processed with JQT2 [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], a system based on the Quality Threshold algorithm [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] that
5 http://textpro.fbk.eu
compares every pair of mentions and decides for coreference if their similarity score is
above a certain dynamic threshold; similarity is computed based on a rich set of features
(e.g., mention attributes and nearby words), while the threshold is higher for ambiguous
names, requiring more “evidence” to assume coreference. Location mentions are
processed with GeoCoder [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], a system based on geometric methods and on the idea that
locations in the same news are likely to be close one to another; it exploits the Google
Maps geo-referencing service6 and the GeoNames geographical database7. Statistics
about the mention clusters identified so far are reported in the third column of table 2.
Linking and entity creation. The last step consists in linking mention clusters to
entities in the background knowledge and to external knowledge sources. Clusters of
location mentions are already linked to well-known GeoNames toponyms by GeoCoder.
Clusters of person and organization mentions are linked to entities in the background
knowledge by exploiting the representation of contexts. The algorithm [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] firstly
identifies the htime, space, topici contexts most appropriate for a mention cluster among the
ones in the KNOWLEDGESTORE, based on the mentions attributes and the metadata of
the containing news (e.g., the publication date). Then, it searches for a matching entity
only in those contexts, improving disambiguation. The fourth column of table 2 reports
the fraction of mention clusters linked by the systems, i.e. the linking coverage:
coverage is low for clusters (10.74%), but increases in terms of mentions (31.03%), meaning
that the background knowledge mainly consists of popular (and thus frequently
mentioned) entities. New entities are then created and stored for unlinked mention clusters,
as they denote real-world entities unknown in the background knowledge; the last
column of table 2 reports the total number of entities obtained so far. All the entities are
finally associated to the corresponding Wikipedia pages using the WikiMachine tool8.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>User Interface</title>
      <p>The entry point of the TRENTINOMEDIA Web interface is a search page supporting
entity-based search. The user supplies a proper name which is looked up among the
entities in the KNOWLEDGESTORE and a list of matching entities is returned for
disambiguation; entities are listed by type and distinguished with short labels generated
from stored information, as in figure 3, left side. By selecting an entity, the user is
presented with the list of news mentioning that entity, retrieved based on the
associations between entities, mentions and resources stored in the KNOWLEDGESTORE. A
descriptive card is also displayed, as shown in the right side of figure 3. It contains all the
information known about the entity, including: (i) background knowledge, (ii)
information carried by the attributes of the entity mentions and (iii) frequently co-occurring and
likely related entities. The example in figure 3 shows the potential but also the
weaknesses of processing noisy, real world data with automatic NLP tools. In particular,
typos and the use of different names for the same entity (e.g., acronyms, abbreviations)
may cause coreference resolution to fail and identify multiple entities in place of one,
as happens with “F1” and “Formula 1”, “Raikkonen” and “Kimi Raikkonen”, “Micheal
Schumacher” and “Michael Schumacher”. Still, the use of additional information
ex</p>
      <sec id="sec-3-1">
        <title>6 http://code.google.com/apis/maps/ 7 http://www.geonames.org/ 8 http://thewikimachine.fbk.eu/</title>
        <p>tracted from texts (e.g., keywords) can often overcome the problem, as happens with
the correct coreference of “Schumacher”, “Michael” and “Michael Shumacher”).</p>
        <p>The other interaction mechanism supported by TRENTINOMEDIA—contextual
semantic enrichment—is accessed through the SmartReader interface shown in figure 4,
which is displayed by selecting a news. The SmartReader allows a user to read news or
watch videos while gaining access to related information linked in the
KNOWLEDGESTORE to the news and its mentioned entities. The interface is organized in two panels.
The left panel displays the text of the news or the video with its speech transcription
and permits to selectively highlight the recognized mentions of named entities. The
right panel provides contextual information that enriches the news or a selected
mention. It can display a cloud of automatically extracted keywords, each providing access
to related news. It can also show additional information about the selected mention, by
presenting: (i) the Wikipedia page associated to the mentioned entity, (ii) a map
displaying a location entity and (iii) a descriptive card with information about the entity in the
background knowledge. In the latter case, only facts which are valid in the htime, space,
topici context of the news are shown, e.g., that “Schumacher is a pilot of Mercedes GP
in 2010”, so to avoid to overload and confound the user with irrelevant information.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4 Conclusions</title>
      <p>TRENTINOMEDIA shows how the application of NLP techniques and the interlinking of
knowledge and multimedia resources can be beneficial to users accessing information
contents. In particular, two mechanisms to exploit this interlinking are demonstrated:
entity-based search exploits links from knowledge (entities) to resources, while
semantic enrichment exploits links in the opposite direction. TRENTINOMEDIA also shows
that NLP and Semantic Web technologies are mature enough to support the large scale
extraction, storage and processing of knowledge from multimedia resources.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Buscaldi</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Magnini</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Grounding toponyms in an Italian local news corpus</article-title>
          .
          <source>In: Proc. of 6th Workshop on Geographic Information Retrieval</source>
          . pp.
          <volume>15</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>15</lpage>
          :
          <fpage>5</fpage>
          . GIR '
          <volume>10</volume>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Carroll</surname>
            ,
            <given-names>J.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klyne</surname>
          </string-name>
          , G.:
          <article-title>Resource description framework (RDF): Concepts and abstract syntax</article-title>
          .
          <source>W3C recommendation</source>
          (
          <year>2004</year>
          ), http://www.w3.org/TR/2004/REC-rdf-concepts-
          <volume>20040210</volume>
          /
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Cattoni</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corcoglioniti</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Girardi</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Magnini</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serafini</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zanoli</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>The KNOWLEDGESTORE: an entity-based storage system</article-title>
          .
          <source>In: Proc. of 8th Int. Conf. on Language Resources and Evaluation</source>
          . LREC '
          <volume>12</volume>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Heyer</surname>
            ,
            <given-names>L.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kruglyak</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yooseph</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Exploring expression data: Identification and analysis of coexpressed genes</article-title>
          .
          <source>Genome Research</source>
          <volume>9</volume>
          (
          <issue>11</issue>
          ),
          <fpage>1106</fpage>
          -
          <lpage>1115</lpage>
          (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Homola</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tamilin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serafini</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Modeling contextualized knowledge</article-title>
          .
          <source>In: Proc. of 2nd Int. Workshop on Context, Information And Ontologies. CIAO '10</source>
          , vol.
          <volume>626</volume>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Pianta</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Girardi</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zanoli</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>The TextPro tool suite</article-title>
          .
          <source>In: Proc. of 6th Int. Conf. on Language Resources and Evaluation</source>
          . LREC '
          <volume>08</volume>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Tamilin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Magnini</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serafini</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Leveraging entity linking by contextualized background knowledge: A case study for news domain in Italian</article-title>
          .
          <source>In: Proc. of 6th Workshop on Semantic Web Applications and Perspectives</source>
          . SWAP '
          <volume>10</volume>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Zanoli</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corcoglioniti</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Girardi</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Exploiting background knowledge for clustering person names</article-title>
          .
          <source>In: Proc. of Evalita</source>
          <year>2011</year>
          <article-title>- Evaluation of NLP and Speech Tools for Italian (</article-title>
          <year>2012</year>
          ), to appear
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>