<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>June</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Exploring Entity-Centric Methods in the UK Government Web Archive</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Philip Webstery</string-name>
          <email>philip.webster@shef</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paul Cloughy</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gianluca Demartiniy</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tom Storrarz</string-name>
          <email>Tom.Storrar@nationalarchives.gsi.gov.uk</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sonia Ranadez</string-name>
          <email>Sonia.Ranade@nationalarchives.gsi.gov.uk</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Graham Seamanz</string-name>
          <email>Graham.Seaman@nationalarchives.gsi.gov.uk</email>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Cultural Heritage, Web Archives, Entity-Based Access</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2016</year>
      </pub-date>
      <volume>22</volume>
      <issue>2016</issue>
      <abstract>
        <p>Being able to explore large digital collections e ectively is of interest to both academics and practitioners alike. The need to go beyond the provision of keyword-driven functionality to features that support exploration and discovery is widely recognised. In addition, providers are seeking to support more diverse groups of users with varying information needs and tasks. Increasing amounts of cultural heritage are being stored in web archives that present unique challenges as a form of digital cultural heritage. This paper describes a collaboration between the University of She eld and the UK National Archives to investigate entity-based methods for exploring the UK Government Web Archive.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        There is a clear need for cultural heritage institutions
(museums, libraries and archives) to provide systems that
go beyond keyword-based search and support more diverse
information seeking behaviours, such as browsing and
exploration [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
        ]. Whitelaw [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] points out that many digital
cultural collections are only accessible via keyword search
and suggests that users can often feel constrained by the
search box as it limits their ability to browse and explore
the collection. He calls for the design of more \generous
interfaces" that provide richer browsable user experiences and
allow the scale and complexity of digital cultural heritage
collections to be revealed, especially to non-specialist users.
Similar calls are being made to support exploration by
enabling navigation, interpretation and use of items in
digital collections and aiding users' analytical and sense-making
processes more generally [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ].
      </p>
      <p>
        Web archives are an area of digital cultural heritage
gaining increasing attention from researchers. One of the key
challenges of such collections is the sheer volume of content
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. In this paper we describe a recently instigated
collaborative research project between the UK National Archives1
1http://www.nationalarchives.gov.uk/
and the Information School (University of She eld) to
investigate the use of entity-based methods for supporting user's
exploration the UK Government Web Archives. In
particular we are focusing on issues of scaling up approaches for
entity extraction and disambiguation. There is a need to
assist users with navigating the content of large digital archives
and help them to better understand how resources are
interconnected over di erent dimensions, such as time, entities
and events. Ultimately the users of the Government Web
archive will bene t from improved features for exploration
and discovery. Section 2 describes related work; Section 3
provides background to the UK Government Web Archive;
and Section 4 describes planned research, including research
aims and challenges.
2.
2.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>Access to digital cultural heritage</title>
      <p>
        Increasingly cultural heritage portals are encouraging user
participation by o ering people opportunities to interact
with content, for example encouraging them to tag resources,
making recommendations to other users and personalisation
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Users served by providers of cultural heritage commonly
can range from expert user groups (e.g., scholars and
curators) accessing the content for professional purposes through
to novices engaging with materials for leisure purposes and
enjoyment [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ]. There may also be a wide range of users
in between, such as students or hobbyists, accessing cultural
heritage to learn and discover [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>
        Johnson [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] argues that users attempting to access
online cultural heritage resources face at least three challenges:
(1) knowing where to look; (2) knowing what to say; and
(3) making sense of archival material (e.g., interpreting and
forming connections between items). This is particularly
pertinent for non-expert users who often lack the
knowledge and skills to engage with cultural heritage resources
[
        <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
        ]. Common search tasks by users of digital cultural
heritage include fact- nding and those of a more exploratory
or information gathering nature. Fact- nding and
knownitem tasks tend to revolve around search, whilst
information gathering tasks lend themselves more to browsing and
exploration. There is a clear need for cultural heritage
institutions to provide systems that go beyond keyword-based
search functions [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]; something that is also recognised by the
wider search community [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>Searching web archives</title>
      <p>
        Increasing amounts of cultural heritage are being stored
in web archives. However, similar to cultural heritage more
generally, \unlocking the potential of web archives requires
tools that support exploration and discovery of captured
content" [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] (p. 851). The development of such tools
typically requires the use of natural language processing. For
example, the ARCOMEM project investigates the
extraction and enrichment of entities, topics, opinions and events
over time [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
      </p>
      <p>
        Providing access to web archives presents a variety of
challenges. One issue is coping with scale as typically collections
run into billions of documents. In light of this, Lin et al.
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] describe a Big Data architecture suitable for large-scale
infrastructure. Another issue is performance. Inverted
indexes are considered to be an essential technique for the
provision of timely search results [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Indexes are widely
used to reference archive data held in ARC (ARChive) or
WARC2 (Web ARChive) le formats [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Indexes are held
in memory for performance reasons [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], but the increase
in the amount of archived data makes this increasingly
difcult. Thus, research has also investigated techniques for
reducing the size of the index, such as de-duplication [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
      </p>
      <p>
        Another key challenge is how to provide e ective access
to web archives, especially as users typically expect a user
experience similar to live-web search engines [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. Search
log analysis performed by the Portuguese National Archive
shows that web archive users typically have a brief
relationship with archival search [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. A typical session consists of
either a fulltext search request or a URL search request.
Fulltext search accounts for around two thirds of queries
with around 60% of fulltext sessions lasting only 1 minute.
Costa et al. [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] also show that 85% of fulltext sessions
contain up to 3 queries, with 44% of queries modifying existing
queries. The di culties that users have with Web archives
are made clearer still by the results of user studies. Users
of Web archives do not often search historical versions of
archived Web pages [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. When archived documents are
accessed via a search interface temporal restrictions are
infrequently used as users often seek the oldest documents in the
archive only rather than discovering how documents have
evolved over time [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ].
      </p>
      <p>
        Dealing with particular structural characteristics is
another challenge, such as time. For example, Berberich et al.
[
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] describe approaches for providing text search over
temporally versioned document collections, such as web archives.
This is achieved through adapting an inverted index to
support temporal search. Their `time-travel text search'
approach supports the exploration of digital collections over
time by enabling the evolutionary history of document
collections, such as Wikipedia or the Web, to be indexed and
exposed.
      </p>
      <p>.
2.3</p>
    </sec>
    <sec id="sec-4">
      <title>Entity-based information access</title>
      <p>
        Identifying named entities (e.g., people, organisations and
locations) has been the focus of research for many years.
More recent e orts have attempted to link entities to linked
data resources and knowledge bases, such as DBPedia3,
Freebase4 and YAGO2 [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. A major challenge is
disambigua2http://www.digitalpreservation.gov/formats/fdd/
fdd000236.shtml
3http://wiki.dbpedia.org/
4https://www.freebase.com/
tion: identifying the correct entity among a number of
entities with the same name. Many techniques exist for entity
extraction, disambiguation and linking them to linked data
resources [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]5. Researchers have investigated entity-based
methods for speci c domains. For example, relevant to this
paper Van Hooland et al. describe the use of entity
recognition and disambiguation methods for cultural heritage
collections [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. Named entities are commonly used to analyse
and provide access to web archives [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. However,
adapting general methods to speci c domains and applications
remains an enduring challenge.
      </p>
    </sec>
    <sec id="sec-5">
      <title>UK GOVERNMENT WEB ARCHIVE</title>
      <p>The UK National Archives seeks to collect and secure the
future of the public record in all its forms and to make it
as accessible as possible. One of the largest digital
collections is the UK Government Web Archive6 (UKGWA). The
UKGWA is the preservation and access solution for digital
content that is published online by UK Government. It is
free to use and is one of the largest and most heavily used
web archives in the world receiving approximately 15
million page views per month. The collection is comprised of
the contents of over 3,000 websites and social media
channels, including 2.5 billion web pages dating from 1996 to the
present.</p>
      <p>From work undertaken by the National Archives to
identify and analyse users of the UKGWA the various user groups
include academics (e.g., researchers), National Archives sta ,
Central Government sta , professional services (e.g., law
rms) and the general public. The purposes for using the
archive are varied and include family history research by the
general public, locating government reports (e.g., by school
teachers), nding information about government procedures
(e.g., computing tax), viewing the evolution of websites over
time, and locating previous versions of current documents.</p>
      <p>The current interface to the UKGWA provides traditional
keyword-based search functionalities. Figure 1 shows an
ex5See also: http://jho .de/wp-content/papercite-data/pdf/
ho art-2015wk.pdf
6http://www.nationalarchives.gov.uk/webarchive/
ample of results presented for the query \Margaret Thatcher".
The aim of our collaborative research project is to develop
interfaces that better support exploration and browsing through
the use of named entities extracted from the content.
Entities, such as people, places, organisations and events can
be extracted from the archive and linked to form a network
that users can explore in addition to navigating the content
directly.
4.1</p>
    </sec>
    <sec id="sec-6">
      <title>PLANNED RESEARCH</title>
    </sec>
    <sec id="sec-7">
      <title>Research goals</title>
      <p>The aim of our research project is to investigate
entitycentric methods for supporting users as they navigate the
UK Government Web Archive. This would allow users to
explore the archive based on entities (e.g., people, locations,
events, etc.), as well as allowing connection with existing
linked data resources, such as DBPedia and Freebase, and
knowledge bases such as YAGO2. The project also provides
opportunities to investigate the success of applying
alternative methods to domain-speci c collections.</p>
      <p>
        The UK National Archives have already explored
largescale entity extraction and linking and this work will utilise
existing annotations and ontological resources [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], in
addition to exploring new methods and especially focusing on
approaches that can be scaled to large collections. This
research project consists of two main strands: (1) investigating
entity-centric techniques for entity extraction,
disambiguation and linking; and (2) investigating how entity-based
networks can best support user browsing and exploration. Use
cases developed in prior studies that go beyond known-item
search tasks will help inform the development of prototype
systems (e.g., a user wanting to view the evolution of web
pages over time). Originality of the work will include: (1)
investigating the e ectiveness of various entity-centric
techniques for a large digital archive; (2) investigating the
preferred way of surfacing the entity network to users; and
(3) developing suitable evaluation methodologies to
establish the success of entity-centric approaches for navigating
digital archives.
4.2
      </p>
    </sec>
    <sec id="sec-8">
      <title>Challenges</title>
      <p>The challenges faced in the research are similar to those
faced by web archives in general7 and include:</p>
      <p>
        Scale: as with any web archive we face the problem of
applying Natural Language Processing techniques, developing
interactive systems and providing storage solutions at scale.
The UKGWA contains around 80TB of data, but some of
this is duplicated. Even so, the non-duplicated portions of
the archive present a potential overhead of many months
of computational e ort to process the corpus to extract and
disambiguate sets of relevant entities. As a result, recent
advances in GPU acceleration of NLP algorithms and database
management systems will be explored to maximise
processing e ciency and to enable deeper analyses that would
otherwise be di cult or impossible due to time constraints.
GPU acceleration has been the focus of recent studies in
NLP [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ].
      </p>
      <p>Parallelisation: the size of the corpus does make one
7http://www.netpreserve.org/sites/default/ les/resources/
2011 06 IIPC WebArchives-TheFutures.pdf
aspect of the processing much simpler { the parallelisation of
such a collection is trivial as it consists of a large number of
documents that can be processed completely independently
of each other. This allows greater exibility when designing
a system to process the data - options for such a dataset
include clusters of machines, multicore CPUs, and the use
of massively parallel processing using GPU cores. However,
questions around how to make parallel abound.</p>
      <p>
        User interaction: designing e ective features and
interfaces for users to support entity-based exploration of the
archive. In particular, surfacing large networks of entities
will present challenges in making them accessible and
usable, particularly for non-experts. Developing techniques
that engage users and allow the surfacing of interesting
content from the archive in accessible ways will require careful
thought [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ].
      </p>
      <p>Data: not only is the scale of data a challenge, its
heterogeneous nature is also problematic as it requires handling
web pages, PDF les and many other formats commonly
found in web archives. Many PDF les contain only scanned
images of textual content that will require pre-processing
stages, including OCR and error correction. Web pages
require boilerplate removal prior to entity extraction.
Historical proprietary word processor le formats are another
potential source of technical issue. Open Source libraries
will be used wherever possible, but a processing pipeline
will be developed to manage handling diverse data formats
and associated processing steps. Such a pipeline would need
to demonstrate both vertical and horizontal scalability in
order to be relevant to archive-scale use cases.</p>
      <p>Structure: capturing and surfacing structural properties
and characteristics of the archive data, such as time, will be
important. This will help users to contextualise and
navigate web content and support user tasks such as viewing the
evolution of archive content and entities over time.</p>
      <p>Domain: an enduring challenge in entity-based
methods adapting techniques and resources (e.g., knowledge bases
and vocabularies) to speci c domains. For example, in the
UKGWA there will likely exist many entities that may occur
frequently (e.g., political leaders of less signi cant parties)
that cannot be found and therefore linked to in general
resources, such as YAGO2 or DBPedia.</p>
      <p>Longevity: of practical importance to TNA is long term
support of entity knowledge bases. For example, Freebase
has been bought by Google and on 21 August 2016 the
Freebase API will be shutdown. On the other hand WikiData,
run by Wikimedia, may be more stable over time. However,
longevity of knowledge bases used for entity extraction and
disambiguation is still an issue with deploying entity-based
solutions in archival contexts that needs attention.
5.</p>
    </sec>
    <sec id="sec-9">
      <title>CONCLUSIONS</title>
      <p>This paper describes a collaborative project to investigate
the use of entity-centric methods to support exploration and
discovery within the UK Government Web Archive. Aspects
of the work will involve investigating methods that can
operate at scale for the identi cation, disambiguation and linking
of named entities, as well as developing e ective interfaces
and functionalities to support the wide range of users
accessing the archive with varying information needs, goals and
tasks. There is a clear need to transfer exploratory research
into practice and this project provides a unique opportunity
to assist with this transfer.</p>
    </sec>
    <sec id="sec-10">
      <title>ACKNOWLEDGEMENTS</title>
      <p>Work partially supported by the UK Arts and Humanities
Council (AHRC) and the UK National Archives.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Benjamins</surname>
            ,
            <given-names>V.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Contreras</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blazquez</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dodero</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garcia</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Navas</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hernandez</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wert</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Cultural Heritage and the Semantic Web</article-title>
          .
          <source>In: The Semantic Web: Research and Applications: First European Semantic Web Symposium</source>
          , ESWS 2004 Heraklion, Crete, Greece, May
          <volume>10</volume>
          -12,
          <year>2004</year>
          . Proceedings. Springer Berlin Heidelberg, Berlin, Heidelberg (
          <year>2004</year>
          )
          <volume>433</volume>
          {
          <fpage>444</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Whitelaw</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Generous interfaces for digital cultural collections</article-title>
          .
          <source>Digital Humanities Quarterly</source>
          <volume>9</volume>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Johnson</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Users, use and context: Supporting interaction between users and digital archives</article-title>
          . In Craven, L., ed.:
          <article-title>What are Archives? Cultural and Theoretical Perspectives: A Reader. (Ashgate Publishing Ltd</article-title>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4] van den Akker, C.,
          <string-name>
            <surname>van Nuland</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          .,
          <string-name>
            <surname>van der Meij</surname>
          </string-name>
          , L.,
          <string-name>
            <surname>van Erp</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , Leg^ene,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Aroyo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Schreiber</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          :
          <article-title>From information delivery to interpretation support: Evaluating cultural heritage access on the web</article-title>
          .
          <source>In: Proceedings of the 5th Annual ACM Web Science Conference. WebSci '13</source>
          , New York, NY, USA, ACM (
          <year>2013</year>
          )
          <volume>431</volume>
          {
          <fpage>440</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Koolen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kamps</surname>
          </string-name>
          , J.:
          <article-title>Searching cultural heritage data: Does structure help expert searchers? In: Adaptivity, Personalization and Fusion of Heterogeneous Information</article-title>
          . RIAO '
          <volume>10</volume>
          (
          <year>2010</year>
          )
          <volume>152</volume>
          {
          <fpage>155</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gholami</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rao</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Infrastructure for supporting exploration and discovery in web archives</article-title>
          .
          <source>In: Proceedings of the 23rd International Conference on World Wide Web. WWW '14 Companion</source>
          , New York, NY, USA, ACM (
          <year>2014</year>
          )
          <volume>851</volume>
          {
          <fpage>856</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Ardissono</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , Ku ik, T.,
          <string-name>
            <surname>Petrelli</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Personalization in cultural heritage: The road travelled and the one ahead. User Modeling and User-Adapted Interaction 22 (</article-title>
          <year>2012</year>
          )
          <volume>73</volume>
          {
          <fpage>99</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Amin</surname>
            ,
            <given-names>A.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hardman</surname>
          </string-name>
          , L.,
          <string-name>
            <surname>van Ossenbruggen</surname>
            ,
            <given-names>J.R.:</given-names>
          </string-name>
          <article-title>Searching in the cultural heritage domain: Capturing cultural heritage expert information seeking needs</article-title>
          .
          <source>Information Systems [INS]</source>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Vilar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sauperl</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Archival literacy: Di erent users, di erent information needs, behaviour and skills</article-title>
          .
          <source>In: Information Literacy. Lifelong Learning and Digital Citizenship in the 21st Century</source>
          . Springer (
          <year>2014</year>
          )
          <volume>149</volume>
          {
          <fpage>159</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Skov</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Hobby-related information-seeking behaviour of highly dedicated online museum visitors</article-title>
          .
          <source>Information Research</source>
          <volume>18</volume>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Skov</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ingwersen</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Exploring information seeking behaviour in a digital museum context</article-title>
          .
          <source>In: Proceedings of the Second International Symposium on Information Interaction in Context. IIiX '08</source>
          , New York, NY, USA, ACM (
          <year>2008</year>
          )
          <volume>110</volume>
          {
          <fpage>115</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Hardman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aroyo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ossenbruggen</surname>
            ,
            <given-names>J.v.</given-names>
          </string-name>
          , Hyvonen, E.:
          <article-title>Using ai to access and experience cultural heritage</article-title>
          .
          <source>IEEE Intelligent Systems</source>
          <volume>24</volume>
          (
          <year>2009</year>
          )
          <volume>23</volume>
          {
          <fpage>25</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Marchionini</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Exploratory search: From nding to understanding</article-title>
          .
          <source>Commun. ACM</source>
          <volume>49</volume>
          (
          <year>2006</year>
          )
          <volume>41</volume>
          {
          <fpage>46</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Risse</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Demidova</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dietze</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peters</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Papailiou</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Doka</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stavrakas</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Plachouras</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Senellart</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carpentier</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          , et al.:
          <article-title>The arcomem architecture for social- and semantic-driven web archiving</article-title>
          .
          <source>Future Internet</source>
          <volume>6</volume>
          (
          <year>2014</year>
          ) 688a^
          <fpage>AS716</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Gomes</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nogueira</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miranda</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Costa</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Introducing the portuguese web archive initiative</article-title>
          . (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Gomes</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miranda</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Costa</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>A survey on web archiving initiatives. Research and Advanced Technology for Digital Libraries (</article-title>
          <year>2011</year>
          )
          <volume>408</volume>
          {
          <fpage>420</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Gomes</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Costa</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cruz</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miranda</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fontes</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Creating a billion-scale searchable web archive</article-title>
          .
          <source>In: Proceedings of the 22nd international conference on World Wide Web companion, International World Wide Web Conferences Steering Committee</source>
          (
          <year>2013</year>
          )
          <volume>1059</volume>
          {
          <fpage>1066</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Gomes</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cruz</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miranda</surname>
            ,
            <given-names>J.a.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Costa</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fontes</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <article-title>a.: Search the past with the portuguese web archive</article-title>
          .
          <source>In: Proceedings of the 22Nd International Conference on World Wide Web. WWW '13 Companion</source>
          , New York, NY, USA, ACM (
          <year>2013</year>
          )
          <volume>321</volume>
          {
          <fpage>324</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Costa</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Silva</surname>
            ,
            <given-names>M.J.:</given-names>
          </string-name>
          <article-title>Characterizing search behavior in web archives</article-title>
          .
          <source>In: TWAW</source>
          . (
          <year>2011</year>
          )
          <volume>33</volume>
          {
          <fpage>40</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Berberich</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bedathur</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neumann</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weikum</surname>
          </string-name>
          , G.:
          <article-title>A time machine for text search</article-title>
          .
          <source>In: Proceedings of the 30th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval. SIGIR '07</source>
          , New York, NY, USA, ACM (
          <year>2007</year>
          )
          <volume>519</volume>
          {
          <fpage>526</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Ho</surname>
            <given-names>art</given-names>
          </string-name>
          , J.,
          <string-name>
            <surname>Suchanek</surname>
            ,
            <given-names>F.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berberich</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weikum</surname>
          </string-name>
          , G.:
          <article-title>Yago2: A spatially and temporally enhanced knowledge base from wikipedia</article-title>
          .
          <source>Artif. Intell</source>
          .
          <volume>194</volume>
          (
          <year>2013</year>
          )
          <volume>28</volume>
          {
          <fpage>61</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Ho</surname>
            <given-names>art</given-names>
          </string-name>
          , J.:
          <article-title>Discovering and disambiguating named entities in text</article-title>
          .
          <source>PhD thesis</source>
          , University of Saarland, Postfach
          <volume>151141</volume>
          , 66041 Saarbrucken (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Van Hooland</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>De Wilde</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verborgh</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Steiner</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , Van de Walle, R.:
          <article-title>Exploring entity recognition and disambiguation for cultural heritage collections</article-title>
          .
          <source>Digital Scholarship in the Humanities</source>
          <volume>30</volume>
          (
          <year>2015</year>
          )
          <volume>262</volume>
          {
          <fpage>279</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Maynard</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Greenwood</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          :
          <article-title>Large scale semantic annotation, indexing and search at the national archives</article-title>
          . In Chair),
          <string-name>
            <given-names>N.C.C.</given-names>
            ,
            <surname>Choukri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Declerck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            , DoA xan, M.U.,
            <surname>Maegaard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Mariani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Moreno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Odijk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Piperidis</surname>
          </string-name>
          , S., eds.
          <source>: Proceedings of the Eight International Conference on Language Resources and Evaluation (LREC'12)</source>
          , Istanbul, Turkey,
          <source>European Language Resources Association (ELRA)</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Hall</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berg-Kirkpatrick</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Canny</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klein</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Sparser, better, faster gpu parsing</article-title>
          .
          <source>In: ACL</source>
          . (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Mercun</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zumer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aalberg</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Presenting and Exploring the Complexity of Bibliographic Relationships</article-title>
          . In:
          <article-title>The Outreach of Digital Libraries: A Globalized Resource Network: 14th International Conference on Asia-Paci c Digital Libraries</article-title>
          ,
          <string-name>
            <surname>ICADL</surname>
          </string-name>
          <year>2012</year>
          , Taipei, Taiwan,
          <source>November 12-15</source>
          ,
          <year>2012</year>
          , Proceedings. Springer Berlin Heidelberg, Berlin, Heidelberg (
          <year>2012</year>
          )
          <volume>63</volume>
          {
          <fpage>66</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>