<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Understanding Web Archives</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Helen Hockx-Yu</string-name>
          <email>Helen.Hockx-Yu@bl.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Head of Web Archiving British Library</institution>
          ,
          <addr-line>London</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <abstract>
        <p>This talk provides an insight into web archives by examining the "unknown" aspects beyond the archived web pages, or the "text". It argues that web archives have a rich set of semantics which when explored o ers a new way of understanding their characteristics. It showcases examples of British Library's work beyond the \document-centric" approach of providing access.</p>
      </abstract>
      <kwd-group>
        <kwd>Web archives</kwd>
        <kwd>exploration</kwd>
        <kwd>semantics</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>The e ort to archive the web started in the mid-1990s, a few years after the web
was born. This was initiated by the Internet Archive in the US. Many national
libraries and archives, which traditionally have the duty to preserve a nation's
cultural and scienti c heritage, followed the suite and started actively collecting
web content. Internet Archive's Wayback Machine1 is the earliest and most
comprehensive web archive to date, containing over 435 billion web pages archived
from 1996. Many national heritage organisations have established collections
covering their respective national web domain or subsets of it.</p>
      <p>There are however issues related to the access and use of web archives: it is
often restricted by legal requirements on one hand, in exchange for reproducing
copyrighted material for the purpose of cultural heritage, and by the (single)
envisaged use case on the other [HY14]. The latter is based on the assumption
of web archives consisting of historical documents (web pages) used for reference.
Researchers access previous states of individual web pages and websites in a web
archive, which are selected, described and grouped together by curators, in the
same way as printed books and journals. The over-focus on \documents" or
\text" means contexts of archived material tend to be ignored or regarded as
irrelevant.
A common assumption is that web archives contain copies of older versions
of websites which are no longer current and have been replaced by the \live"
version. Brugger and Finnemann argue that archived web resources are \reborn",
di erent from digitised and born digital collections and from the live web in many
ways [BF13]. Using the British Library's web archive as an example, this talk
examines in detail the many boundaries and imitations related to web archive,
determined by purpose, strategy, legal requirements and technological choices.
It also points out a fundamental oversight which impacts users' interpretation
or understanding of web archives: very little is explained or made clear to the
users beyond the actual HTML pages (or the \text"). A typical example of this
is the common error message \Resource Not in Archive", which is presented to
the end-users when a requested URL cannot be found in the archive. This could
be caused by many reasons: some are intended, introduced by things like data
limitation at crawl time or content beyond the scope of the crawl; others relate
to technical limitations, e.g. dynamic content which the web crawlers are not
capable of collecting.
3</p>
    </sec>
    <sec id="sec-2">
      <title>More to \text"</title>
      <p>E ort started to emerge in recently years which moves away from the level of
single webpages or websites to the entire web archive collection. Using
visualisation and data analytic techniques, new ways have been developed to view web
archives, o ering opportunity to unlock embedded patterns and trends,
relationships and contexts, which are not possible by consulting websites individually.
This is in alignment with the changes in scholarly practices as researchers
increasingly take advantage of new possibilities o ered by technology. New methods
of scholarship are emerging, which challenges the primacy of \text" as object
of study. This talk references the concepts of \paratexts"[Nie10] and \distant
reading"[Mor00], as theoretical basis for using web archives as scholarly sources.
The role of web archives is to provide services supporting scholars who read texts
di erently.</p>
      <p>This talk focuses on a range of non-text attributes of web archives (including
an example visualisation or demo for each), explored by the British Library or
others, as additional ways of understanding web archives. Scholars are
encouraged to explore other contextual or \para-textual" content in the web archives,
such as viral content and crawl logs.</p>
      <p>{ Statistical overview, scale and distribution of a web national domain
{ Size: bytes
{ Space: geo location, postcodes
{ Type of content, e.g. le format, language
{ Structure, linked entities and networks
{ Evolution, pattern of change over time, e.g. domain names
{ Correlation, e.g. between certain term and historical event</p>
      <p>This talk also discusses the general issues related to analytical access, such
as researchers' scepticism or suspicion about hidden algorithms behind analysis,
and how biases in data and how data collection decisions lead to variances in
outputs.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [BF13]
          <article-title>Niels Brugger and Niels Ole Finnemann. The web and digital humanities: Theoretical and methodological concerns</article-title>
          .
          <source>Journal of Broadcasting &amp; Electronic Media</source>
          ,
          <volume>57</volume>
          (
          <issue>1</issue>
          ):
          <volume>66</volume>
          {
          <fpage>80</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [HY14]
          <string-name>
            <given-names>Helen</given-names>
            <surname>Hockx-Yu</surname>
          </string-name>
          .
          <article-title>Access and scholarly use of web archives</article-title>
          .
          <source>Alexandria</source>
          ,
          <volume>25</volume>
          (
          <issue>1</issue>
          ):
          <volume>113</volume>
          {
          <fpage>127</fpage>
          ,
          <year>August 2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Mor00]
          <string-name>
            <given-names>Franco</given-names>
            <surname>Moretti</surname>
          </string-name>
          . Conjectures on world literature. http://newleftreview. org/II/1/franco-moretti
          <source>-conjectures-on-world-literature (accessed on 17 November</source>
          <year>2014</year>
          ),
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [Nie10]
          <article-title>Niels Brugger. Website analysis: Elements of a conceptual architecture</article-title>
          . http://cfi.au.dk/fileadmin/www.cfi.au.dk/publikationer/cfis_ skriftserie/012_brugger.
          <source>pdf (accessed on 17 November</source>
          <year>2014</year>
          ),
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>