<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Humorous View into the Past: The Old Jokes Archive</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Department of Computer Science</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martin-Luther-University Halle-Wittenberg</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Germany mark.hall@informatik.uni-halle.de</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of English and History, Edge Hill University</institution>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <fpage>161</fpage>
      <lpage>165</lpage>
      <abstract>
        <p>Jokes represent one of the most understudied sources about nineteenth century society. Due to their ephemeral nature they slipped from attention as soon as they were no longer funny or topical. Digitisation of newspapers and books has made them available again, but due to their short nature they are not easily accessible through current generic keyword-based newspaper search systems. In this paper we present the Old Jokes Archive, which aims to provide a digital archive focused solely on jokes. The archive will support the full process from initial text acquisition to search and nally re-use by both academic and general public users.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Jokes represent one of the most ephemeral spoken (or written) interactions [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
They might provoke brief laughter before the conversation moves on. They might,
if they are particularly funny, even be re-told to friends, family, or colleagues.
But these exchanges typically go unrecorded. While more substantial works of
art and literature are carefully preserved for posterity by libraries and museums,
even the most rib-tickling gags are usually disposed of and forgotten when they
lose their capacity to provoke a laugh.
      </p>
      <p>
        While jokes are often treated in a disposable fashion, they have nevertheless
played important roles in many historical cultures. In nineteenth-century Britain,
for instance, the possession of a good new joke represented signi cant cultural
capital, for to be a true wit was a position of social distinction [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. This appetite
for humour was fed by the popular press and, by the 1880s, most of the
bestselling newspapers and magazines in Britain featured a regular column of jokes,
puns, and comic stories.
      </p>
      <p>
        The availability of jokes in newspapers also means that, unlike longer,
humorous novels and stories, jokes reached a much wider audience spanning all
social classes. Thus a single joke could in one week reach as many readers as
a best-selling novel might throughout its author's life-time (see for example [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]
for circulation numbers on Mark Twain's work). This means that joke-based
humour potentially represents a more democratic and representative view of
nineteenth-century society's tastes.
      </p>
      <p>
        At the same time, while the second half of the nineteenth century saw the
introduction of copyright laws [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], this had little e ect on the re-use of jokes across
newspapers. Editors continued to crib jokes from other newspapers, sometimes
verbatim, sometimes adapting the joke's setting or context to suit local tastes.
Due to this jokes can serve as an illustration of what and how ideas spread via
the nineteenth century equivalent of viral memes [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>The jokes themselves are a potentially invaluable source of information about
historical cultures, as they are typically built upon an assumption of shared
knowledge; a belief that an audience will immediately understand how a stock
character (such as a mother-in-law) behaves, how a familiar social situation
should ordinarily play out, or what to respond with when somebody says
`knockknock.' Historians can reverse-engineer these jokes in order to uncover the ideas
and attitudes that joke-writers and their editors assumed were widely held at
the time. The subjects of jokes, and the dynamics of laughter (who is laughing
at/with who), can also reveal valuable new insights into the power relations at
work in historical communities.</p>
      <p>Even though they represent such a rich data-source, jokes are not widely
used by most historians. Many Victorianists, for instance, rarely venture beyond
the cartoons of Punch magazine when attempting to make sense of the period's
comic culture. Millions of historical jokes have never been examined by historians
and are therefore ripe for further exploration.</p>
      <p>The chief obstacle to this research centres on the di culty of nding and
accessing historical jokes. Many were never recorded and have now been lost to
us, while even those that were written down and preserved in books and
newspapers are tricky to uncover. The large-scale digitisation of newspapers, books, and
other print archives presents new opportunities for solving this problem.
However, even with the help of keyword searches it remains di cult for researchers
to locate speci c jokes pertaining to their interests. For example, a keyword
search for the word 'lawyer' in a typical digital newspaper archive will return a
jumble of millions of news stories, adverts, editorials, letters, serialised stories,
and poetry, within which we might also nd jokes about the legal profession. At
present there is no straightforward way to search speci cally for historical jokes.
2</p>
    </sec>
    <sec id="sec-2">
      <title>The Old Jokes Archive</title>
      <p>The Old Jokes Archive (OJA) aims to address this gap, by providing the rst,
large-scale digital repository of historical humour, targeted both at academic
researchers and the wider public. To support this the OJA provides a range of
functionalities centred around two areas: the acquisition of the joke data and
search and (re-)use functionality.
2.1</p>
      <sec id="sec-2-1">
        <title>Acquisition</title>
        <p>The joke acquisition starts based on the existing digitisation of newspapers and
joke books, with the nal goal being annotated text versions that can then be
used by researchers and the general public. Throughout the process a
semiautomatic approach is used, combining automatic image and text processing
with manual validation and post-processing using a crowdsourcing approach. In
the OJA's precursor { the Victorian Jokes Archive { we have tested a number
of crowdsourcing methods, in particular to improve the accuracy of
crowdsourcing results, the experience of which will be integrated into the OJA. The joke
acquisition process follows these ve steps:
1. Identi cation As the digitisation has focused on whole newspaper pages,
the rst step is in the identi cation of those areas of the scanned pages that
contain jokes and then the splitting of those areas into the individual jokes.
We are investigating automated methods for this, but have initially adapted
techniques from the Digital Playbills project.
2. Transcription OCR will be used to create an initial transcription.
Newspaper paper and typeface can be relatively poor (see Figure 1), resulting in
comparatively high error rates. To deal with this we have developed error
classi cation heuristics, that allow us to determine the quality of the OCR
output. Based on this crowdsourcing users can be o ered a choice to work
on a transcription that has a low error rate { mostly typo-correction { or a
high rate { essentially transcription from scratch. Using this both users who
have signi cant time to invest or those who just want to do something quick
and simple can be o ered appropriate crowdsourcing jobs.
3. Classi cation The resulting corrected transcription is then classi ed using
automated heuristics developed previously. The classi cation works at a high
level, distinguishing categories such as question-and-answer, dialog, or puns.
The classi cation is then used by the following steps to apply
categoryspeci c heuristics.
4. Segmentation The joke's text is segmented into chunks. The exact chunks
depend on the joke category determined in the previous step, but for example
for question-and-answer jokes this would be segmenting the question and the
answer element, while for dialog jokes it includes identifying speakers, spoken
text, and asides. Also more generally we have developed heuristics to identify
joke titles and attribution.
5. Annotation The chunks are then, where appropriate, annotated with
speci c meta-data extracted from the chunk. Among others we have developed
heuristics to identify dialogue speaker gender and are currently working to
automatically identify social class, age, and jobs of speakers.</p>
        <p>After running through all ve steps, the joke texts will be publicly available
through the online archive. Jokes that have been OCRed, but lack the error
correction and further processing will also be available, if the user also wishes to
see those.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Search and Use</title>
        <p>
          In order to make the resulting archive available, the OJA will be available online
and provide a range of access points to explore, search, and use the jokes:
{ Search The OJA will provide a state-of-the-art faceted search system
enabling users to narrow their search for jokes via keywords and via the
categories and annotations created in the acquisition stage.
{ Exploration While search works well for users who know what they want,
for the general public a more open-ended, browsing based interface will be
developed. Based on previous work [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] we will be developing a virtual
museum of jokes through which the user can explore the available jokes.
{ Related Jokes As described above, jokes were frequently copied from one
publication to another, often with minor modi cations. The OJA will provide
the necessary tools to trace such copies across multiple publications and to
visualise joke distribution networks.
{ Export The OJA will provide export functionality at all points in the
system, whether they be search results, exploration pages, or individual jokes.
For interoperability reasons a joke-speci c TEI schema will be used.
Additionally the OJA will provide a workspace allowing users to save jokes to
their own work area and then export that.
{ Re-interpretation The vast majority of historic jokes tend not to have
aged well in the way they are presented. The OJA will provide space for
users to re-interpret the joke's text. As part of this we have also developed
algorithms to automatically convert jokes into a single-panel comic image.
In both cases the aim is to have these shared via social media, in part to
increase the project's visibility, but also to investigate overlap and di erences
between popular jokes in the nineteenth century and now.
        </p>
        <p>While initially the focus will be on English-language jokes, the project will be
built with multi-lingual content in mind. Additionally content will be available
through a permissive open-culture license to encourage use and re-use both
academically and for private reasons.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Conclusion</title>
      <p>Jokes represent a so far largely untapped resource for investigating a range of
historical questions, from language use to the spread of ideas. The Old Jokes
Archive will act as a central data-source point, starting with nineteenth
century English-language jokes, but with an aim to expanding both temporally and
linguistically. The tools, methods, and practices developed in the course of this
project will be applicable to future archival projects that aim to 'remix' existing
digitised content by extracting, organising, and re-presenting it in new ways.</p>
      <p>While the initial focus will be on the joke's texts, the long-term aim is to also
consider visual aspects in the source data, such as wehn jokes were accompanied
by drawings.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Hall</surname>
            ,
            <given-names>M.M.:</given-names>
          </string-name>
          <article-title>Digital museum map</article-title>
          . In: Mendez,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Crestani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Ribeiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>David</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Lopes</surname>
          </string-name>
          , J.C. (eds.)
          <article-title>Digital Libraries for Open Knowledge</article-title>
          . pp.
          <volume>304</volume>
          {
          <fpage>307</fpage>
          . Springer International Publishing,
          <string-name>
            <surname>Cham</surname>
          </string-name>
          (
          <year>2018</year>
          ). https://doi.org/10.1007/978-3-
          <fpage>030</fpage>
          -00066- 0 28
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Nicholson</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>`you kick the bucket; we do the rest!': Jokes and the culture of reprinting in the transatlantic press1</article-title>
          .
          <source>Journal of Victorian Culture</source>
          <volume>17</volume>
          (
          <issue>3</issue>
          ),
          <volume>273</volume>
          {
          <fpage>286</fpage>
          (
          <year>2012</year>
          ). https://doi.org/10.1080/13555502.
          <year>2012</year>
          .
          <volume>702664</volume>
          , http://dx.doi.org/10.1080/13555502.
          <year>2012</year>
          .702664
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Nicholson</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          : Capital Company - Writing and Telling Jokes in Victorian Britain. forthcoming (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Seville</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Literary Copyright Reform in Early Victorian England: The Framing of the 1842 Copyright Act</article-title>
          . Caombridge University Press (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Stone</surname>
            ,
            <given-names>A.E.</given-names>
          </string-name>
          : Review:
          <article-title>Mark twain in england, by dennis welland</article-title>
          .
          <source>Nineteenth-Century Fiction</source>
          <volume>34</volume>
          (
          <issue>9</issue>
          ),
          <volume>357</volume>
          {
          <fpage>359</fpage>
          (
          <year>1979</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>