<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Introducing a User Interface with an Entity-Strategy- based Approach for Exploring Document Collections</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Daniel Hienert</string-name>
          <email>daniel.hienert@gesis.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wilko van Hoek</string-name>
          <email>wilko.vanhoek@gesis.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>GESIS - Leibniz Institute for the Social Sciences</institution>
          ,
          <addr-line>Cologne</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we present a first sketch of an alternative approach for searching and exploring document collections. The traditional approach applied in Digital Libraries and Web Search Engines is based on search forms and result lists. The user enters a keyword and is presented with a list of document metadata with authors, titles and descriptions. We propose an alternative approach that is based on entities in a document collection like authors, documents and topics. The user can search for these entities and can then choose from a set of highly abstracted search strategies, e.g. to get highly cited papers from an author. The approach is applied in a zoomable and infinite user interface that enables the user to explore freely and where the search history is always present.</p>
      </abstract>
      <kwd-group>
        <kwd>Visual Interface</kwd>
        <kwd>Exploratory Search</kwd>
        <kwd>Visual Exploration</kwd>
        <kwd>Search Strategies</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Today’s Digital Libraries (DLs) still make use of the standard paradigm of
queryresponse. Users can enter a query and are presented with a list of relevant documents
which they have to inspect and filter according to their information need. Already
Bates presented a list of alternative search strategies from the real-world such as
‘Citation Searching’ or ‘Journal Run’ [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] which partly have been adopted in modern
scholarly database systems such as Scopus or Web of Science. However, this is far
from being usual practice in DLs, where full-text indexing of documents opens new
possibilities. Exploratory Search [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] proposes a model beyond query-response, with a
focus on the learn and investigation step in the search process. Highly-interactive
search systems can support these steps whereby interactive visual search systems are
a part of it. Therefore already a number of visual search tools have been proposed for
the exploration of DL content. Early attempts experimented with different visual
metaphors or tried to gain insight with the visualisation of the distribution of information
facets. More recent tools are for example the INVISQUE system [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] which supports
the search and manipulation of results on an infinite panel or PivotPaths [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] which
show relations between concepts, resources and people. Another important aspect for
learning and investigation while searching is the visualisation of the search process
itself. Scientists spend much time with literature search; their search history can
enlarge quickly over months to even years. Today’s DLs only support to save search
results or documents, other artefacts such as document inspection are lost. Research
showed that search histories support revisitation [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] in web search and support the
user’s orientation within a search session [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In the following we want to present a
User Interface (UI) concept which combines visual information search with different
search strategies and a visible search history.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Concept Overview</title>
      <p>Based on document
• Cited by (Citation Searching)
• Referenced by
• Similar Topics (Subject search)
• Main/co-authors (Author searching)
• Documents from footnotes (Footnote</p>
      <p>chasing)
• Same Journal (Journal run)
• Same category/classification (Area</p>
      <p>scanning)
• Related/Similar papers (e.g. by
content, topics, references, subject etc.)
Figure 1 shows the core idea of our approach applied to an example from the field of
the social sciences. The mockup shows one possible exploration path based on a
realworld document collection from the social science portal Sowiport1.</p>
      <p>First, the user searches for the author ‘Ulrich Beck’ and can then choose
‘Highlycited papers’ from the strategies menu to get an overview of his most influencing
work. As a result the most-cited papers are presented in a small list. After inspecting
the abstracts in the document view, the user classifies the third paper as interesting
and chooses ‘Cited by’ from the methods menu to show the latest papers which
influenced it. Based on this paper the user is interested in the topic ‘Cosmopolitanism’ and
wants to see highly-cited papers for this topic. She/he arrives at the author ‘Esref
Aksu’ and grabs the keywords ‘Cosmopolitan Democracy’ from the abstract and initiates
a new search. Choosing ‘Main Authors’ from the strategies menu shows two authors
and their papers with which the search process can continue.</p>
      <p>Because the search history is always visible on the user interface, the user can
return to a previous search step such as a search, a person or a document and can
continue the search there. In the above stated example, the user returned to the topic
‘Globalization’ from Beck’s ‘Cosmopolitical Realism’ and initiated a new search.
1 http://sowiport.gesis.org</p>
    </sec>
    <sec id="sec-3">
      <title>Discussion &amp; Future Work</title>
      <p>
        The proposed approach has several benefits over the standard search form/result list
paradigm. Based on the four core ideas of our approach these are:
1. UI as infinite Panel: The use of an UI panel with infinite space is a prerequisite
that has been used in other visual exploration tools as well [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. In our context it
removes the limitation of showing only one search step, but can represent whole
search sessions and multiple sessions.
2. Entities: In a standard DL, search results are limited to a list of documents
ordered by relevance. The use of entities such as persons, documents or topics
allows the intuitive application of search strategies.
3. Search Strategy: Complex search strategies are encapsulated in one-click UI
elements and can be applied easily. This allows alternative exploration and views
on the document collection. New search strategies can be implemented easily.
4. Search Graph: Every step in the search process is visible on the UI and forms a
search graph over time. That allows an overview of the whole search session, but
also over a set of search sessions. Therefore a prior search path can be continued,
but also a search session can be shared with another person, e.g. among
colleagues in a research group.
      </p>
      <p>However, most search strategies require complex computation and a rich data set. For
example, “Highly-cited papers” for an author needs a separate citation index, which
may not always be present in nowadays DLs or the real-time computation of metrics
like author centrality can be a challenge. In a next step we want to implement a
system prototype that can be used for exploring different document collections such as
the arXiv2 corpus for the natural sciences or Sowiport for the social sciences which
contain this rich information. Based on that, we will perform various user tests to
verify the basic plausibility of our approach.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bates</surname>
            ,
            <given-names>M.J.:</given-names>
          </string-name>
          <article-title>The Design of Browsing and Berrypicking Techniques for the Online Search Interface</article-title>
          .
          <source>Online Rev</source>
          .
          <volume>13</volume>
          ,
          <issue>5</issue>
          ,
          <fpage>407</fpage>
          -
          <lpage>424</lpage>
          (
          <year>1989</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Dörk</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          et al.:
          <article-title>PivotPaths: Strolling through Faceted Information Spaces</article-title>
          .
          <source>IEEE Trans Vis Comput Graph</source>
          .
          <volume>18</volume>
          ,
          <issue>12</issue>
          ,
          <fpage>2709</fpage>
          -
          <lpage>2718</lpage>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Imko</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          et al.:
          <article-title>Semantic History Map: Graphs Aiding Web Revisitation Support</article-title>
          .
          <source>Presented at the August</source>
          (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Kodagoda</surname>
          </string-name>
          , N. et al.:
          <article-title>Using Interactive Visual Reasoning to Support Sense-Making: Implications for Design</article-title>
          .
          <source>IEEE Trans. Vis. Comput. Graph</source>
          .
          <volume>19</volume>
          ,
          <issue>12</issue>
          ,
          <fpage>2217</fpage>
          -
          <lpage>2226</lpage>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Marchionini</surname>
          </string-name>
          , G.:
          <article-title>Exploratory search: from finding to understanding</article-title>
          .
          <source>Commun. ACM</source>
          .
          <volume>49</volume>
          ,
          <issue>4</issue>
          ,
          <fpage>41</fpage>
          -
          <lpage>46</lpage>
          (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Mayer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Web History Tools and Revisitation Support: A Survey of Existing Approaches and Directions</article-title>
          .
          <source>Found. Trends® Hum.-Comput. Interact. 2</source>
          ,
          <issue>3</issue>
          ,
          <fpage>173</fpage>
          -
          <lpage>278</lpage>
          (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>