<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>April</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>The Digital Archiving of Historical Political Cartoons: An Introduction</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Junte Zhang</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kees Ribbens</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rob Zeeman</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Erasmus University Rotterdam, &amp; NIOD Institute for War</institution>
          ,
          <addr-line>Holocaust and Genocide, Studies</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Meertens Institute</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Meertens Institute &amp; NIOD, Institute for War</institution>
          ,
          <addr-line>Holocaust, and Genocide Studies</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2013</year>
      </pub-date>
      <volume>26</volume>
      <issue>2013</issue>
      <abstract>
        <p>Political (editorial) cartoons often capture the Zeitgeist of society and convey a message. Increasingly, historians study them to understand commentaries of past events or personalities. Visual culture as an academic subject could be greatly enhanced if this information can be digitally archived. We employ crowdsourcing to obtain valuable metadata by guiding volunteers' feedback using an online survey with 31 targeted questions. We provide intellectual access to a set of about 300 cartoons of a single creator spanning over multiple years in a highly interactive search engine.</p>
      </abstract>
      <kwd-group>
        <kwd>metadata</kwd>
        <kwd>crowdsourcing</kwd>
        <kwd>e-Humanities</kwd>
        <kwd>cartoons</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        Newspapers often have political (editorial) cartoons that
contain a commentary about events or personalities [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] which
is being disseminated. For historians, these capture the
Zeitgeist of the period of time of their study, and become an
invaluable source of information. These print newspapers are
stored in libraries and get digitally archived { for example
by the National Library of the Netherlands { for long-term
preservation to continued access. Digital archiving is the
management of the life cycle of digital assets (records) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ],
from preservation to continued use.
      </p>
      <p>
        In the Radical Political Representation project, we aim
to digitally archive historical political cartoons created by a
single cartoonist and published before and during the
Second World War, so we gain insight into di erent points of
view and support the study of visual culture using a
computational approach. This is made possible because newspaper
pages have been digitized as images, which contain cartoons.
These cartoons are not yet machine-readable, therefore
providing intellectual access is the best option. It has been
proposed in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] to detect the text lines in cartoons using
OCR, but this is di cult because it involves handwritten
texts. In [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] it is pointed out that \more descriptive areas
by which images might be accessed are largely neglected,"
and argued that subject indexing as a eld of academic work
is aboutness { and VRA Core 4.0 is referred to as a metadata
schema to record bibliographic information.
      </p>
      <p>Our aim is to transcribe a cartoon, and move beyond
standard bibliographic information by comprehensively
capturing its meaning(s) for historical research by eliciting user
feedback using crowdsourcing. So we address the following
question: How can we provide intellectual access to, and
allow for, advanced use of these cartoons?
2.</p>
    </sec>
    <sec id="sec-2">
      <title>CROWDSOURCING OF CARTOONS</title>
      <p>The objects of our study are so far 286 cartoons
published by Maarten Meuldijk in the weekly Volk en Vaderland
(VoVa) of the National Socialist Movement in the
Netherlands from 1937 to 1942. Pages on which they occur have
been digitized by the National Library. To obtain
descriptions about the cartoons, we experiment with crowdsourcing
to see whether crowdsourcing is applicable in our context.</p>
      <p>The search tasks that we have in mind are more complex,
therefore we created a comprehensive survey that captures
the questions historians typically would ask about a cartoon.
This also requires more contextual knowledge. Fig.1(a) shows
the VoVa Annotation Editor developed in Adobe Flex, where
we guide users through a set of 31 targeted questions in
8 stages, and aid them by o ering answers of these
questions with pre-de ned multiple choice answers in
combination with open answers. Users can zoom in/out on a cartoon,
but also read contextual information related to the cartoon
in the articles on the page { a strategy used by a number of
users. There are no time limits and a cartoon is randomly
assigned and stays assigned to a user until completion. To
control for the completion of a cartoon description, we
validate all questions for at least 1 given answer.</p>
      <p>We invited interested volunteers online and in printed
national media. In total 189 users registered, where eventually
83 volunteers participated with at least 1 completed
description of a cartoon and with 5 users completing more than 10.
(a) The VoVa Annotation Editor, where volunteers can
provide valuable metadata about the cartoon, ranging from plain
descriptions to their opinion of a cartoon.
(b) The VoVA Search Engine, which is used to gain
intellectual (advanced) access to the cartoons.</p>
    </sec>
    <sec id="sec-3">
      <title>3. SERENDIPITY IN CONTEXT</title>
      <p>
        Having obtained the metadata, we want to use it. Since
the search engine should serve historians, we design it to
support serendipitous search and be highly interactive in order
to focus on a high recall (rather than precision). The system
has been designed to maximize the user's ability to explore.
We have proposed search features to support serendipitous
and focused access in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], and these features have been
reimplemented here. The search features primarily deal with
query expansion, recommendation, and interactive
visualizations of aggregated results. The former is based on using
ternary search trees for spellchecking, returning the top term
vectors related to the original query, and returning the top
terms that have the original query as substring. The latter
is based on charts, maps and word clouds.
      </p>
      <p>
        A user can improve the searching in a session by e ectively
reducing the information space step by step, i.e.
incrementally combining questions. This con rms with the
Berrypicking model of [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] { queries are not static, but rather evolve,
and users \gather information in bits and pieces instead of in
one grand best retrieved set." These steps are stored as part
of the search trail, so the overview is kept. The user
interface of the system is depicted in Fig 1(b). In this example,
someone looked for a cartoon about a \Jood" (Jew) used as
a main keyword to describe a cartoon, with captions under
it, and a \ster" (star) depicted as a symbol. The search
engine treats the questions asked in the survey as facets, and
is therefore a straight-forward question-answering system.
Facets that always appear are the date of publication of a
cartoon, and the education and knowledge levels of the
volunteers who provided the descriptions. We show in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] that
making the credibility of the source transparent gives users
greater con dence in their selection. We think historians
will be aided with this part of the search process.
      </p>
      <p>There are di erent search strategies possible. Users can
search by full-text or focused (within the answers of
questions). The query gets highlighted in context given the
fulltext and the survey question. A dynamic word cloud widget
that supports query expansion is not activated, unless the
autocompletion is used. Using the Advanced Search option,
users can look up a question and then enter a keyword also
with the autocompletion feature. Wildcard (empty) queries
can be used to obtain the distribution of words of the
answers given a question in a word cloud for a quick summary.
4.</p>
    </sec>
    <sec id="sec-4">
      <title>CONCLUSIONS</title>
      <p>We have presented { in a compressed version { the mission
statement and some results of the Radical Political
Representation project. We completed the rst phase of
crowdsourcing, and pending further releases of data by the
National Library, we can further digitally archive the complete
series of Meuldijk cartoons. The technical infrastructure to
digitally archive political cartoons has been set-up.</p>
      <p>This means we can expand our scope to other cartoonists
in di erent times { there is no shortage of cartoons. We
can re ne our survey to allow for more di erent information
needs of historians, or embed our survey as part or
extension of a formal metadata schema like VRA Core. We will
improve the UI and further implement useful information
visualization of results, and evaluate the search engine. It
can be used at www.meertens.knaw.nl/vova/search.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Bates</surname>
          </string-name>
          .
          <article-title>The design of browsing and berrypicking techniques for the online search interface</article-title>
          .
          <source>Online Review</source>
          ,
          <volume>13</volume>
          (
          <issue>5</issue>
          ):
          <volume>407</volume>
          {
          <fpage>424</fpage>
          ,
          <year>1989</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>G. M.</given-names>
            <surname>Hodge</surname>
          </string-name>
          .
          <article-title>An information life-cycle approach : Best practices for digital archiving</article-title>
          .
          <source>Journal of Electronic Publishing</source>
          ,
          <volume>5</volume>
          (
          <issue>4</issue>
          ):1{
          <fpage>14</fpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C.</given-names>
            <surname>Landbeck</surname>
          </string-name>
          .
          <article-title>Issues in subject analysis and description of political cartoons</article-title>
          .
          <source>Advances in Classi cation Research Online</source>
          ,
          <volume>19</volume>
          (
          <issue>1</issue>
          ),
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>C.</given-names>
            <surname>Sterling</surname>
          </string-name>
          .
          <source>Encyclopedia of Journalism. Sage</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          .
          <article-title>Searching digital political cartoons</article-title>
          .
          <source>In Proceedings of the 2010 IEEE International Conference on Granular Computing, GRC '10</source>
          , pages
          <fpage>541</fpage>
          {
          <fpage>545</fpage>
          , Washington, DC, USA,
          <year>2010</year>
          . IEEE Computer Society.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>J. Zhang.</surname>
          </string-name>
          <article-title>Supporting serendipitous and focused search</article-title>
          .
          <source>In EuroHCIR</source>
          , volume
          <volume>909</volume>
          <source>of CEUR Workshop Proceedings</source>
          , pages
          <volume>79</volume>
          {
          <fpage>82</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Amin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. S. M.</given-names>
            <surname>Cramer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Evers</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Hardman</surname>
          </string-name>
          .
          <article-title>Improving user con dence in cultural heritage aggregated results</article-title>
          .
          <source>In SIGIR</source>
          , pages
          <volume>702</volume>
          {
          <fpage>703</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>