<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Hacking History: Automatic Historical Event Extraction for Enriching Cultural Heritage Multimedia Collections!</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Roxane Segers</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marieke van Erp</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lourens van der Meij</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lora Aroyo</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Guus Schreiber</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bob Wielinga</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jacco van Ossenbruggen</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Johan Oomen</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Geertje Jacobs</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centre for Mathematics and Computer Sciences</institution>
          ,
          <addr-line>CWI</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Netherlands Institute for Sound and Vision</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Rijksmuseum Amsterdam</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>VU University Amsterdam</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Within cultural heritage collections, objects are often grounded in a particular historical setting. This setting can currently not be made explicit, as structured descriptions of events are either missing or not marked up explicitly. This poster reports a study on automatic extraction of an historical event thesaurus from unstructured texts. We also present a demo in which relations between events and museum objects are visualised to accommodate event- and object-driven search and browsing of two cultural heritage collections.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Events have recently gained attention in the knowledge representation
community as valuable constructs [
        <xref ref-type="bibr" rid="ref4 ref7 ref8">4, 7, 8</xref>
        ] that can help tie together relevant but yet
unrelated elements of information. In the cultural heritage domain, knowledge
about historical events is often concealed in textual descriptions that can only
be accessed via keyword search. As such, the available knowledge can not be
reused across collections as it is not part of the shared metadata and controlled
vocabularies.
      </p>
      <p>
        In this study, we investigate how historical events in unstructured text
collections can be captured and modeled to create an event thesaurus for enriching
metadata in cultural heritage collections. We adopt the SEM event model [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] to
distinguish event types, actors, locations, and dates. We experiment with
natural language processing (NLP) techniques to extract event names and their
associated actors, dates and locations. Additionally, we show how this resulting
preliminary event thesaurus is employed in a new platform for event- and object
driven search and browsing of the collections of the Rijksmuseum Amsterdam
(RMA) and the Netherlands Institute for Sound and Vision (S&amp;V).
      </p>
      <p>!"#$%&amp;'%&amp;()*)%+",)
!"#$%&amp;'(()*+,-).##/&amp;,0''*,0),1)0)*"#20+),&amp;*')3)2,'3,4)-,2##*,5'-6#/#*&amp;# !789:)$&amp;;
-."#(&amp;'//)01%,)*""$&amp;%2''0%2)%3)2)0."421)%&amp;0')+)4%'+%5),
4""0%6',7"$"0&amp;"8%9:42)01#(:.2)0:;%&lt;"4%2)%:44"*)%&lt;"4%6',7"$"0&amp;"
&amp;:;2)41%2)%&amp;5))2)%+'.:&amp;:'4).)%"#&amp;:)=%2)#)*&gt;)0%?@AB8
3CD?@@BDED?F
"##$%&amp;'()*!+,)-(#
.)/&amp;%(#+,)-(0 G5))2)%+'.:&amp;:'4).)%"#&amp;:)
&gt;:',0"+(:#".%"1+)#&amp;1
12)'($20G'("%H2:*:2;';'=%I'("**)2 !"#%%.'()0?@ABD?JD?@ !$#%?@A@DFKDLF !$#%JF)%))M5 !%&amp;#%&amp;5))2)%$5"0&amp;%JF)%))M5 !%'#
*338)4"'%/&amp;())5)02-0:&amp;"&amp;'6.9%0"71)+0)#&amp;1 "N(M""00"2)N&gt;.M)'"0")00)42. !!!$$"###&amp;)$)4:4, !$# 1:)7*;:&lt;')&amp;%:#(0%"1+)#*&amp;1:.:&amp;":0)%,)6?1'#@,(A7:B)"D2$?)"J40D&amp;:?"1@ !%!!""(###G5))2)%+'.:&amp;:'?4@)A.@)D%F"K#&amp;D:L)? !!%'##
"##$%&amp;'()*!=;&lt;)%(#,7&lt;=; O%+0)&lt; % ? J L &gt; P % 4)Q&amp;%R
S0)H1:121)'4#&amp;:%"-&amp;)'2)%$S"00)41'1%,888 -:4$:4H%+4"'44;"7*4,'%M*1)&amp;%1888 X"U&amp;4&amp;"2='%4I)'1(:V"%*&lt;0*:;W"2 -#(:.2H%&lt;4"'44%7)*)4'M%H1&amp;;)()0 H"4$'*H14&amp;'%&lt;4"74*%'YM"14%-+:888 X)!&amp;)%$)"#1$&amp;*))".4%&lt;=%"H44%!20":)&amp;"1888
?$72!@',&amp;A'(&amp;$-!B'(5! O%+0)&lt; % C % 4)Q&amp;%R
Fig. 1. Screenshot of object page in the Agora Event Browsing Demonstrator
3"&lt;:,"&amp;:'4%S"&amp;(%T)&amp;":.1
2</p>
    </sec>
    <sec id="sec-2">
      <title>Event Extraction from Text</title>
      <p>
        As no annotated historical document collections exist in Dutch, our approach
is focused on extracting named events with minimal manual effort. For this
study we selected 3,724 historical Wikipedia articles as a test set. The event
extraction process consists of three steps: in the first step, we recognize actor
names and locations using the Stanford Named Entity Recognition system [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
adapted for Dutch historical texts. Dates were recognized via regular expressions.
This step resulted in 18,623 candidates for actors (F-measure of 0.77), 7,023
locations (F-measure of 0.66) and 7,981 dates. In the second step, we use a
pattern-based method for recognizing event names such as French Revolution.
We harvest patterns from the Web (e.g., destroyed during the, before the) using
the Yahoo! search API 5 and a seed set of one hundred historical events. Patterns
are ranked by frequency of co-occurrence with two or more seed events [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. To
retrieve event candidates, we applied the patterns to the Wikipedia corpus. The
event candidates are then filtered, based on a threshold on the pattern score,
resulting in a set of 2,444 unique events. The precision score of this set is 56.3%.
      </p>
      <p>
        In the third step, we associate events with actors, locations and dates.
We experiment with both redundancy and co-occurrence of data on the Web,
inspired by the work of Geleijnse et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and Cilibrasi &amp;Vitanyi[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Each
combination of an event name and actor/location/date is sent to Yahoo! and for each
pair a score is computed. We discovered 392 event names that were paired with
an actor, a location and a date. Through manual evaluation we conclude the
following: 71.9% (323) are correct event names, 45.6% (179) are correct actors,
41.1% (161) are correct locations and 51.5% (202) are correct dates.
5 http://developer.yahoo.com/search
      </p>
    </sec>
    <sec id="sec-3">
      <title>Enrichment by Events</title>
      <p>The extracted events are linked to the RMA and S&amp;V collections. In total 35
unique events provide direct relations from 435 S&amp;V objects to 675 RMA objects.
An additional 34 unique events provide links from 391 S&amp;V objects to 362 RMA
objects, but this link exists indirectly through the event instance (e.g., S&amp;V
object - Actor - RMA object). We hypothesize that these links are potentially
useful for navigating cultural heritage collections.
4</p>
    </sec>
    <sec id="sec-4">
      <title>The Agora demonstrator</title>
      <p>The automatically generated event thesaurus is applied in a new historical event
browser called Agora6 which provides an integrated access route to museum
objects and audio-visual material from RMA and S&amp;V respectively. It is a first
step towards a platform to investigate the added value of historical events and
narratives for the exploration of integrated collections. For each event and object
there is an automatically generated page that shows (1) all associated objects,
e.g., museum and audio-visual objects; (2) all associated events and the type of
their relationship, e.g., previous-in-time event, sub-event; (3a) the event
descriptive metadata, e.g., actors, place, period; or (3b) object descriptive metadata
organized in three groups, e.g., biographical, material and semiotic dimensions
– see figure 1 for a screenshot –and finally (4) the navigation path. The
current version of the event thesaurus will be extended further to accommodate
searching for relations between events such as temporal inclusion, causality and
meronymy.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Discussion</title>
      <p>
        In this paper, we presented a modular pipeline for capturing knowledge about
historical events from Dutch texts. Compared with previous approaches (i.e., [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]),
it relies on a minimum of manual annotation and can be repurposed for other
languages. To the best of our knowledge, this is the first work to extract events
from unstructured Dutch text. Although our results are promising, more
sophisticated techniques are necessary to obtain more fine-grained extractions and
define measures for the historic relevance of the extracted events. Additionally,
we also aim to find and represent relations between events such as causality,
meronymy and correlation.
6
      </p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>This research was funded by the CAMeRA Institute of the VU University
Amsterdam and by the CATCH programme, NWO grant 640.004.801.
6 http://agora.cs.vu.nl/demo</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>R.</given-names>
            <surname>Cilibrasi</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Vitanyi</surname>
          </string-name>
          .
          <article-title>The google similarity distance</article-title>
          .
          <source>IEEE Trans. Knowledge and Data Engineering</source>
          ,
          <volume>19</volume>
          (
          <issue>3</issue>
          ):
          <fpage>370</fpage>
          -
          <lpage>383</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Finkel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Grenager</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <article-title>Incorporating non-local information into information extraction systems by gibbs sampling</article-title>
          .
          <source>In Proceedings of the 43nd Annual Meeting of the Association for Computational Linguistics (ACL</source>
          <year>2005</year>
          ),
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>G.</given-names>
            <surname>Geleijnse</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Korst</surname>
          </string-name>
          , and V. de Boer.
          <article-title>Instance classification using co-occurrences on the web</article-title>
          .
          <source>In Proceedings of the ISWC 2006 workshop on Web Content Mining (WebConMine)</source>
          , Athens, GA, USA,
          <year>November 2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>N.</given-names>
            <surname>Gkalelis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Mezaris</surname>
          </string-name>
          ,
          <string-name>
            <surname>and I. Kompatsiaris.</surname>
          </string-name>
          <article-title>Automatic event-based indexing of multimedia content using a joint content-event model</article-title>
          .
          <source>In ACM Events in MultiMedia Workshop (EiMM10)</source>
          ,
          <year>Oct 2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>N.</given-names>
            <surname>Ide</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Woolner</surname>
          </string-name>
          .
          <article-title>Exploiting semantic web technologies for intelligent access to historical documents</article-title>
          .
          <source>In Proceedings of the Fourth Language Resources and Evaluation Conference (LREC)</source>
          , pages
          <fpage>2177</fpage>
          -
          <lpage>2180</lpage>
          , Lisbon, Portugal,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>E.</given-names>
            <surname>Riloff</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Jones</surname>
          </string-name>
          .
          <article-title>Learning dictionaries for information extraction by multilevel bootstrapping</article-title>
          .
          <source>In Proceedings of AAAI '99</source>
          , pages
          <fpage>474</fpage>
          -
          <lpage>479</lpage>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>R.</given-names>
            <surname>Shaw</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Troncy</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Hardman</surname>
          </string-name>
          . Lode:
          <article-title>Linking open descriptions of events</article-title>
          .
          <source>In 4th Annual Asian Semantic Web Conference (ASWC'09)</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>W. R. van Hage</surname>
          </string-name>
          , V. Malais´e, G. de Vries, G. Schreiber, and M. van Someren.
          <article-title>Abstracting and reasoning over ship trajectories and web data with the Simple Event Model (SEM)</article-title>
          .
          <source>Multimedia Tools and Applications</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>