<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Journal of
Conference on Web and Social Media</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.13140/2.1.1341.1520</article-id>
      <title-group>
        <article-title>Towards Supporting Complex Retrieval Tasks Through Graph-Based Information Retrieval and Visual Analytics</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Aleksandar Bobic</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jean-Marie Le Gof</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christian Gütl</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CERN</institution>
          ,
          <addr-line>Espl. des Particules 1, Meyrin, 1211</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Graz University of Technology</institution>
          ,
          <addr-line>Rechbauerstraße 12, Graz, 8010</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>3</volume>
      <fpage>15</fpage>
      <lpage>18</lpage>
      <abstract>
        <p>The retrieval result analysis approaches of existing retrieval solutions tend to be either too simple, provide too few features for exploring retrieval results or are very narrowly focused. We present an enhanced approach that attempts to address these issues and help the wider community to get more insight from their retrieved data. To this end, this paper presents an enhanced graph-based retrieval prototype built on the Collaboration Spotting platform. It combines information retrieval and visual analytics concepts to provide an advanced solution for data retrieval and exploration. It enables users to retrieve information, explore it from diferent perspectives using a graph representation and perform further searches based on their navigation and selection interactively. Compared to traditional retrieval solutions, a search action in CS can reveal more detailed aspects/techniques when visually analysing the search output. To gain initial feedback, we interviewed five domain experts in related fields. Findings reveal that the developed retrieval approach provides users with helpful ways of exploring search results and provides mechanisms of connecting features that are not explicitly linked otherwise. Furthermore, several research directions and improvements have been identified for future work, which should be addressed.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;information retrieval</kwd>
        <kwd>visual analytics</kwd>
        <kwd>knowledge discovery</kwd>
        <kwd>visualization system</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        port analysing correlations between papers. This ordered
list format does not help users to extract complex
relationWith the recent digitalisation eforts and steadily growing ships and gain deeper insights from large retrieval results
data piles, the amount of generated information rapidly [
        <xref ref-type="bibr" rid="ref1 ref6 ref7">6, 7, 1</xref>
        ]. In the context of bibliometric data, examples of
increased over a short period. This increase in data quan- data retrieval insights might include identifying author
tity made the need for eficient retrieval and visual analyt- collaboration networks, identifying trending research
arics tools apparent. This need is also reflected in multiple eas in recent years, and discovering common concepts
works which identified the necessity for IR applications shared among fields. User-centred interactive analysis
that would enable users to carry out complex retrieval of bibliometric data can lead to better insights, novel
tasks, visualise hidden connections by leveraging interac- research projects, and more informed decision-making
tion and visualisation and extract implicit insights from [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ].
retrieved data automatically [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. Examples of such com- A variety of visual analytics (VA) tools and
visuplex retrieval tasks could include retrieval of institution alisation approaches were created as a result of the
collaborating in a specific field, identification of author above-outlined needs for supporting bibliometric data
collaboration networks, retrieval of upcoming research exploration, and analysis workflows by diferent interest
topics connected to existing topics and more. groups [
        <xref ref-type="bibr" rid="ref10 ref6 ref8">10, 8, 6</xref>
        ]. A straightforward and broad division
      </p>
      <p>
        As one example, the need for the above-mentioned can be made between solutions created for
bibliometfeatures to analyse data and grasp connections is also ric mapping and general-purpose VA tools [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Both
present in bibliometric data. For this application scenario, groups leverage multiple visualisation techniques to
prodata are traditionally gathered, indexed and made acces- vide users with an insightful exploration process and
sible by services such as Google Scholar [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], Microsoft reveal hidden connections which can not be easily
inAcademic [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and ArXiv [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] which present search results ferred from an ordered list of retrieval results. A common
as an ordered list based on assumed relevance and do not approach to representing large connected datasets is
disofer advanced analytics approaches which would sup- playing and analysing them as a connected graph. The
potential of the graph representation has been apparent
to researchers and tool creators for quite some time [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>
        Another example of a graph-based representation is
Collaboration Spotting (CS). It is a graph-based visual
analytics (VA) platform created to address the limitations
of existing graph-based exploration tools such as limited
leveraging of interactivity and network visualisations,
and visualisation of explicit and implicit connections researchers and the broader community. Even though
between features [12]. It enables users to explore sizeable various natural language processing (NLP) and IR
apconnected datasets by navigating through or changing proaches can be applied to bibliometric data, they might
perspectives1 and contexts2. not produce insightful results [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Therefore,
multi
      </p>
      <p>
        To enable users to execute complex retrieval tasks and ple visualisation techniques, VA tools and
bibliometricgain further insight into their retrieval results, and based oriented solutions were created to provide better insights
on existing work, we develop an enhanced CS-based re- into the increasing amount of bibliometric data.
trieval system as a prototype. However, as an example A commonly used graph-based visual analysis tool
and due to large amounts of available data, we focus on with a broad application range, including analysis of
bibliometric data. As our main contribution, we integrate bibliometric data, is Gephi [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. An example of a more
an enhanced retrieval mechanism in the CS platform’s narrowly focused bibliometric data tool is Galex which
main version. We combine graph-based VA and informa- represents disciplines, areas and institutions as an
intertion retrieval (IR) by introducing an enhanced IR system active galaxy [22]. BiblioViz focuses specifically on table
that retrieves data from a search provider and leverages and graph visualisation to enable users to investigate
biban interactive graph representation to display the search liometric data from various perspectives [23]. Another
results. It also provides a mechanism for further search tool for analysing publication data is VISPubComPAS
refinement through simple graph interactions. Further- which focuses on the analysis of institutions and authors
more, to identify the needs of experts, understand how [24]. Additionally, a solution for exploring university
bibto develop a system supporting users at multiple steps of liometric data for driving strategic decisions is presented
their retrieval tasks and potentially expanding the sys- by [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
tem for broader use, we interview five experts with a As a result of this research area’s growing popularity,
semi-structured approach. multiple surveys were created covering diferent aspects
      </p>
      <p>
        This paper is structured in the following manner: Sec- and solutions. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] provides an overview of interactive
tion 2 introduces briefly related concepts and related VA approaches for patent and publication data. Next, [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]
work. Section 3 describes the requirements, architec- report on approaches for extracting and visualising
bibture, technical details and the user interface (UI) of the liometric data. Finally, [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] identify multiple solutions and
retrieval system. Section 4 describes three sample case two common workflows for processing and visualising
studies with real-world data and presents how the re- publication data. These surveys indicate the potential
trieval solution in CS could be used to gain further insight of reviewed approaches but also identify multiple open
into bibliometric data. Additionally, it also describes the challenges, such as lack of applications leveraging user
feedback gathered from experts and discusses potential interaction for analysis, lack of empirical research
regardfuture research directions. The paper concludes with ing the efectiveness of visualisation techniques and tools,
Section 5 where we discuss the current implementation visualisation of relationships between diferent data
feaand future work. tures and more.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <sec id="sec-2-1">
        <title>2.2. Bibliometric-Oriented Search</title>
      </sec>
      <sec id="sec-2-2">
        <title>Systems</title>
      </sec>
      <sec id="sec-2-3">
        <title>2.1. Visual Analysis Bibliometric</title>
      </sec>
      <sec id="sec-2-4">
        <title>Approaches</title>
        <sec id="sec-2-4-1">
          <title>Although some of the aforementioned search engines</title>
          <p>
            and repositories provide further insight into author
inA variety of modern solutions such as search engines [3, fluence, relations between papers, and more, their main
4, 13], repositories [
            <xref ref-type="bibr" rid="ref5">14, 15, 5</xref>
            ] and services [16, 17] collect, focus is still related to representing content as an ordered
create and retrieve large amounts of bibliometric data list. This almost never-ending list of results ranked by
which can potentially provide new insights. This data assumed relevance does not provide a way of gaining
can be analysed using VA, which is a science that aims in-depth insights into data [
            <xref ref-type="bibr" rid="ref1 ref6 ref7">6, 7, 1</xref>
            ]. As identified by [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ] IR
to provide explainable insight into large abstract data systems should enable the execution of elaborate retrieval
through interactive data visualisation [18]. It can be tasks, which might lead to more significant insights and
combined with IR approaches to provide a deeper insight drive decision making processes by leveraging
visualiinto retrieval results by visualising them and enabling sation methods to display connections in the retrieved
the use of advanced tools for their analysis [19, 20, 21]. data. Multiple approaches have been created to mitigate
          </p>
          <p>Bibliometric data analysis is usually demanding, te- the issues of traditional bibliometric search engines by
dious and time-consuming and can overwhelm novice combining VA with IR. An example that leverages the
above-mentioned connections is Rexplore, an analytics
1Data features represented as the graph nodes. tool that enables retrieval of research publication data
2Data features represented as graph edges. via facets and sorting of results [19]. Another example</p>
        </sec>
        <sec id="sec-2-4-2">
          <title>The approaches above are, to the authors’ knowledge,</title>
          <p>either not actively developed anymore, are not
accessible, cannot be used on large scale data, or too simple to
provide users with advanced analytics insights.</p>
          <p>As a possible alternative, CS is a graph-based VA
platform that enables users to analyse large quantities of
connected data through the use of filters, facets, and
contexts [12]. Unlike other approaches, it can be used to
analyse a wide variety of datasets and enables users to
change the graph structure dynamically. A separate CSC
version was developed to explore how to provide users
with complete retrieval and analytics experience [25].
However, this version did not enable users to manipulate
their subsequent searches with a finer granularity (for
example by combining their selected nodes that
represent the search result features with Boolean operators)
since it relied on document embeddings and was never
implemented in the primary CS version. Furthermore, it
did not explicitly combine graph interactions with the
retrieval process.</p>
          <p>Based on insights and data analysis requirements, we
aim to incorporate a prototype IR system into the primary
CS platform to support users in performing complex
retrieval tasks. As part of this process, we introduce a novel
way of performing searches by exploring intrinsic graph
patterns and selecting graph nodes from combinations
of diferent features using the prototype. Furthermore,
we add new connections to external services in CS and
an analytics integration to perform empirical evaluation
studies. Finally, we discuss use cases in bibliometric data
analysis, describe possible approaches to analysing such
data with CS, report feedback from expert interviews and
discuss potential future research directions.</p>
        </sec>
        <sec id="sec-2-4-3">
          <title>3https://www.connectedpapers.com/</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Design and Implementation</title>
      <p>is PivotSlice, which focuses on searching and analysis of
retrieval results using a combination of filters and facets
[20]. 3.1. Prototype Requirements</p>
      <p>
        An example from industry is Connected Papers3 which
visualises retrieval results as connected graphs where pa- Based on the identified gaps and needs outlined in the
pers are connected based on their similarity. Another previous sections, our goal is to build an enhanced
graphsimilar solution is Open Knowledge Maps which visu- based retrieval and exploration prototype based on the
alises retrieval results as a multi-level bubble chart where existing CS system. As an example application scenario,
papers are grouped based on text similarity [21]. Al- we chose to use bibliometric data due to its vast
accessibilthough there are many existing approaches and services, ity. To this end, the retrieval system should provide
sufi[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] identify the limits of these tools, focusing on provid- cient flexibility to enable CS users to search via multiple
ing search results in the form of individual papers or queries and through a wide variety of data. Additionally,
focusing on bibliometric analysis and provide a concep- the retrieval system should leverage users’ interactions
tual solution. to provide an eficient retrieval and exploration workflow.
Furthermore, the system should enable the investigation
of implicit connections between the entities of a dataset
2.3. Collaboration Spotting (e.g. Institution collaborations based on co-authorship).
Finally, the expanded CS system should be ready for
empirical analysis studies and gather interaction data.
      </p>
      <p>High-level requirements can be summarized as:
1. Support integration of multiple datasets and</p>
      <p>search providers.
2. Support exploration of implicit and explicit entity</p>
      <p>connections.
3. Collect user interaction data for empirical studies.
4. Visualise complex search results using various</p>
      <p>visual cues.
5. Enable exploration of search results using graph</p>
      <p>interactions.
6. Enable search query refinement through graph</p>
      <p>interactions.
7. Support complex search query creation.
8. Provide explainable report generation.
9. Enable visual creation of retrieval queries and</p>
      <p>ifltering steps.
10. Enable graph analysis approaches to gain further</p>
      <p>insight.
11. Enable usage of graphs for knowledge retrieval.</p>
      <sec id="sec-3-1">
        <title>As part of the initial prototype we focus on requirements 1 to 6.</title>
        <sec id="sec-3-1-1">
          <title>3.2. Prototype Architecture</title>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>To address the novel combination of graph interaction</title>
        <p>and retrieval concepts described in this work and based
on the above-listed requirements we enhanced the
existing architecture seen in Fig. 1 with new components. The
architecture is split into multiple conceptual components
for clarity. However, in reality, the Graph Calculation,
API Request Handlers and the Search are one module.
This architecture is set to change once the move from
a prototype to a production system is made. The Graph
Vis. &amp; Interaction (Fig. 1 a) component and the Menus &amp;</p>
      </sec>
      <sec id="sec-3-3">
        <title>Side-Panels (Fig. 1 b) component handle interactions such</title>
        <p>as selecting a search source, entering search queries,
navigating graphs and selecting graph elements for search
refinement. Once users start a new search, the
Search</p>
      </sec>
      <sec id="sec-3-4">
        <title>Handler (Fig. 1 c) sends a request with their query and se</title>
        <p>lected dataset to the API Request Handler (Fig. 1 d), which
forwards the information to the Search Source &amp; Provider</p>
      </sec>
      <sec id="sec-3-5">
        <title>Selector (Fig. 1 e) component. Here the request is parsed,</title>
        <p>and the appropriate data search provider is selected based
on a project environment variable.</p>
      </sec>
      <sec id="sec-3-6">
        <title>The currently supported data search providers include</title>
        <p>Elasticsearch4, Whoosh5 and the ArXiv API. Users who
aim to perform an initial shallow exploration with a small
amount of data and no advanced pre-processing can use
the ArXiv API or an API from another existing hosted
search provider. However, the introduction of new search
providers would require implementing a new Python
search component that would communicate with the
search providers. On the other hand, users who aim to
get a deeper insight into their data and perform a more
thorough exploration can use an existing search provider
like Elasticsearch or Whoosh. Furthermore, the latter
search providers enable the use of time-demanding
preprocessing and pre-analytic steps outside of the expanded</p>
      </sec>
      <sec id="sec-3-7">
        <title>CS system. For example, a user might wish to extract named entities or add additional data features before importing them into the system.</title>
      </sec>
      <sec id="sec-3-8">
        <title>The selected dataset, query and search operator are</title>
        <p>sent to the Search Handler (Fig. 1 f) component, where the
search is executed using the previously selected search
provider (Fig. 1 g), and the results are transformed into a
CS-specific format. Results represent a network of data
out of which a graph corresponding to users selection
is built using the Graph Calculation (Fig. 1 h) module,
which retrieves the graph id from the newly generated
graph. The id is then sent back to the Front-end to
retrieve the newly generated graph. The users’ retrieval
and exploration process is enhanced by features such as
navigation through the result graph and multiple
iterations on their search. The interactions users perform on
the Front-end are tracked using Matomo6 as part of the</p>
      </sec>
      <sec id="sec-3-9">
        <title>Analytics System (Fig. 1 i) component for user behaviour</title>
        <p>and engagement analysis.</p>
        <sec id="sec-3-9-1">
          <title>3.3. User Interface</title>
          <p>form searches by opening the search modal seen in Fig. 2.</p>
        </sec>
      </sec>
      <sec id="sec-3-10">
        <title>Once they enter their queries, they can select one of the</title>
        <p>available data sources visible in the left drop-down (Fig. 2
k) and select a binding Boolean operator for the queries
5https://whoosh.readthedocs.io/en/latest/index.html
To retrieve information on an initial dataset, users per- they selected the relevant nodes, they can open the search
pahGr Calcution</p>
        <sec id="sec-3-10-1">
          <title>3.4. Data Preparation</title>
          <p>The dataset should be appropriately pre-processed to
leverage the prototype’s features efectively. The ex- What is your occupation, and what are your daily tasks?
ample dataset is retrieved from the Journal of Univer- Where do you see the strengths of CS?
sal Computer Science (J.UCS) [26] since the authors Where do you see the weaknesses of CS?
had full access to it’s detailed metadata. The data in- What could be improved in CS?
cludes the doi, title, abstract, authors, afiliations, author Did you identify any other uses-cases for the system?
keyphrases and publication categories. Since the
authordefined keyphrases might be biased and reflect only
on a subset of the paper content, we extract additional location. They then explore institutions that are
conkeyphrases using the keyphrase extraction tool YAKE! nected if their representatives wrote a joint paper. Using
[27] to provide an alternative view on the paper con- this view, the executive can identify institutions where
tent. Additionally, we split the publication categories they might know someone and establish a collaboration.
into level 1, 2 and 3 to provide users with the possibility
of exploring categorical data through graph navigation. 4.1.3. Introduction to a New Topic
We extend this data with data from Scopus7 by extracting
the afiliation name, afiliation city and afiliation country.</p>
          <p>Finally, we convert the data into a format that the CS
platform can process. Once a graph is generated from the
data, users can explore which authors and institutions
collaborate, identify authors’ focus categories, and more.
A novice researcher in software engineering explores
the J.UCS categories using the prototype system. They
further explore the author keyphrase of papers in the
software engineering category to identify points of
interest relevant to their research. They notice that software
engineering is connected to formal methods and decide
to investigate both topics’ authors. Only a few authors
published in J.USC about these topics, so they return to
the previous author keyphrase graph and search for the
same topics using the ArXiv API. The search results
represent a more diverse set of documents that can be used
to identify prominent authors in the field of interest by
observing the node sizes.</p>
        </sec>
        <sec id="sec-3-10-2">
          <title>4.2. Expert Evaluation</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Case Studies and Evaluation</title>
      <sec id="sec-4-1">
        <title>The focus of the case studies is on bibliometric data from the J.UCS journal as described above.</title>
        <sec id="sec-4-1-1">
          <title>4.1. Case Studies</title>
          <p>4.1.1. Potential Reviewers
A journal editor would like to identify potential reviewers 4.2.1. Study Environment
for an IR and NLP paper. Using the prototype system,
they search for "IR" and "NLP" on the J.UCS dataset. They To identify further potential users’ needs, we organised
ifrst filter out the resulting journal categories that do not individual interviews with five experts from diferent
dofall into one of the two above mentioned topics. Next, mains who could benefit from using CS. The interviews
they navigate to a new key phrase graph where they were semi-structured to gain quick feedback that will
select phrases closely related to NLP or IR and navigate guide further research and development eforts and
poto the author view. The authors are connected if they tentially enable the discovery of additional edge cases
have joint publications. The editor can now identify that the authors might not have identified yet.
Furtherpotential candidates who are likely knowledgeable in more, we aimed to identify how to implement future
the fields mentioned above and avoid authors who have versions of CS in particular in a way which will enable
previously published papers with the submission author. users to perform complex retrieval and analysis tasks,
support users at multiple steps of the retrieval process
4.1.2. Identification of Potential Collaborators and gain potential users’ view for shaping future system
features. As part of the interview, which was held as
A company executive searches for online education us- an online meeting, we presented the enhanced CS
sysing the prototype system to identify potential collabora- tem, discussed the three use cases mentioned earlier and
tors in online education. They explore the results from demonstrated how users could use CS for the first use
keyphrases’ perspective to identify relevant phrases and case using a dataset from J.UCS as an example through
use them to perform a search. Next, they explore and fil- screen sharing. Finally, the experts were asked the five
ter out countries that are not easily accessible from their questions depicted in Table 1. During the interview, they
could ask to view specific sections of CS again and asked
further questions about how the system works.</p>
          <p>
            7The data was downloaded from Scopus in winter of 2020-2021
using the Python library Pybliometrics [28]
4.2.2. Study Participants
The first participant was a librarian with more than 30
years of experience who also had experience in database
usage and is leading the library services for the last 11
years. The next participant was a computer scientist
and doctoral student focusing on learning environments
and learning analytics. The third participant was a
postdoctoral researcher focusing on computer science and
psychology who participated in research projects
focusing on VA, UI design, mitigation of cognitive biases and
more. The fourth participant was a senior data scientist
who analyses literature based on clients’ requirements
and implements machine learning algorithms for various
datasets based on this analysis. The final participant was
a Knowledge Transfer Oficer, who, among other things,
focuses on patent and research paper exploration and
retrieval. All participants were previously vaguely familiar
with the project but did not know how it works or the
details of how it can be used and what are its features.
4.2.3. Study Results
more, similarly to what was concluded in [
            <xref ref-type="bibr" rid="ref10">10</xref>
            ] experts
suggested the use of other data types such as source code
and multimedia attached to scientific work. Menus could
be improved by including wording which calls for action9
and is understandable for the general public.
Additionally, it was proposed that they should take up less space.
          </p>
          <p>
            To simplify the graph search and exploration, the
system should support natural language queries that can
be automatically translated into search and exploration
actions. The UI could be additionally improved by
providing an onboarding tutorial with short introductory
examples, introducing an advanced UI mode with the
complete set of features and a simple UI mode that can
be used to navigate through predefined templates and
presenting a traditional list view of results alongside the
graph view. The accommodation of novice users was
recognised as a critical feature also by [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ] who suggested
that the amount of data shown should be adjustable in
order not to overwhelm novice users. Furthermore, it was
mentioned that creating reports based on the performed
actions and enabling easy graph export with the search
and navigation history and the option to customise the
background colour to better fit in professional reports
would be beneficial.
          </p>
          <p>We also identified additional use cases such as
creating yearly reports about larger institutions’ publications,
code analysis evaluation where concepts used and bugs
encountered by each user could be visualised, and
analysis of personal email corpora. A use case that two experts
mentioned is the visualisation and exploration of
employee skills and project participation inside companies.</p>
          <p>In conclusion, the combination of IR and VA helps
facilitate user exploration through graph navigation and
helps avoid fine-tuning keyphrases for relevant results.</p>
          <p>A commonly identified strength of the prototype
compared to traditional web search systems is that users can
explore results eficiently and avoid fine-tuning precision
and recall through keyphrases by navigating through
the graph. Additional strengths include the ability to
identify relations between fields and authors, the visual
feedback provided through the node sizes, the ability
to make sense of information that would be dificult to
analyse with simpler representations and the ability to
explore implicit connections. An expert also mentioned
that "Navigation is the door to serendipity". In the context
of the prototype system, navigation is well supported by
enabling diferent perspectives and contexts.</p>
          <p>
            We also identified much room for improvements. Sug- 4.3. Future Research Directions
gestions include visualising other data relationships such
as the impact of papers on diferent fields, using a wider Based on the expert feedback, literature survey, initial
variety of visual cues to display new dimensions and requirements and our own experience, we identified
sevavoid node label overlap. The need to handle visualisa- eral future research directions. Some of the identified
tions of multidimensional datasets was also identified directions are listed below.
by [
            <xref ref-type="bibr" rid="ref10 ref8">10, 8</xref>
            ]. Moreover, data should also be presented with
traditional charts to give the user a familiar overview of IR aspects include the use of retrieved graphs not
the data. Furthermore, more quantitative details about only for gaining analytical insights but also for advanced
the retrieved data and more insightful details such as the knowledge retrieval for example by exploiting graph
patlargest clusters and what they include were among the terns for further retrieval processes. Furthermore, we
suggestions. Experts also proposed exploring ways of need to identify how to support user groups to perform
integrating financial data and general impact data 8 to multi-user retrieval and analysis tasks together. Another
increase the added value of data exploration. A similar broad question identified by [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ] is how to support users
conclusion was reached by [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ] who suggest using social in complex retrieval tasks.
media for the expansion of scientific datasets.
FurtherGraph analysis aspects
may include content
summa8For example, if a solution is mentioned in news articles with- rization of larger graph clusters, entity generation from
out an explicit citation it should still count as a mention which
contributes to the general impact of a work. 9For example "Select by:"
graph patterns and identification of improved clustering
and layout techniques which might be more appropriate
for the dynamic nature of the graphs in this work.
          </p>
          <p>
            Machine learning aspects contain an exploration of
conversational IR approaches to enhance users analytical
abilities of result graphs as well as generate user models
based on user interactions which could aid users in the
retrieval process [
            <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
            ].
          </p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>Engineering aspects of future work include im</title>
        <p>proved connection generation and system refactoring.
The current system is not scalable and should be
rewritten in modern technologies with modularity in mind.
Furthermore, the connection calculation process should
be refactored to avoid implying connections between
points that might not directly connect in the retrieved
dataset.</p>
      </sec>
      <sec id="sec-4-3">
        <title>Evaluation aspects which represent the final key as</title>
        <p>
          pect and are a prevalent issue in VA systems are
concerned with eficient quantitative evaluation, which will
provide a clearer picture about the usefulness of the
system [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion and Future Work</title>
      <sec id="sec-5-1">
        <title>This paper describes a graph-based visual analytics and</title>
        <p>IR prototype that enables the search and exploration of
data through a combination of IR and VA approaches.
The solution is built as an enhancement to the CS
system. As part of the IR process, users perform a traditional
search whose results are then presented as an interactive
graph that can be explored or used to perform multiple
additional searches. To investigate how the introduced
solution could help users in their retrieval process,
identify users needs and ideas for future system development,
we held interviews with five experts. Their answers
indicate that the prototype does provide a helpful workflow
for analysing data but that there is also room for
improvement. Among the areas of improvement, we identified
enrichment of the dataset using data from other domains,
UI simplifications, the introduction of new interaction
approaches and displaying the search result data in
traditional and graph form. Furthermore, visualisations could
be enhanced by additional visual cues. We also discuss
future research directions that would be beneficial for the
proposed system. We plan to improve and refactor the
system and conduct an empirical study to gain further
insight into how this approach can help support users in
their retrieval process.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <sec id="sec-6-1">
        <title>We want to thank the five interviewed experts for their time and for contributing valuable feedback. We would also like to thank André Rattinger for scraping and supplying the primary J.UCS dataset.</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          , X. Cheng, S. Dong,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Dou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Li</surname>
          </string-name>
          , T.-Y. Liu, et al.,
          <article-title>Information retrieval: a view from the chinese ir community</article-title>
          ,
          <source>Frontiers of Computer Science</source>
          <volume>15</volume>
          (
          <year>2021</year>
          )
          <fpage>1</fpage>
          -
          <lpage>15</lpage>
          . doi:
          <volume>10</volume>
          .1007/s11704-020-9159-0.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Culpepper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Diaz</surname>
          </string-name>
          , M. D. Smucker,
          <article-title>Research frontiers in information retrieval: Report from the third strategic workshop on information retrieval in lorne (swirl 2018), SIGIR Forum 52 (</article-title>
          <year>2018</year>
          )
          <fpage>34</fpage>
          -
          <lpage>90</lpage>
          . URL: https://doi.org/10.1145/3274784. 3274788. doi:
          <volume>10</volume>
          .1145/3274784.3274788.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>P.</given-names>
            <surname>Jacsó</surname>
          </string-name>
          ,
          <article-title>Google scholar: the pros and the cons</article-title>
          ,
          <source>Online information review 29</source>
          (
          <year>2005</year>
          )
          <fpage>208</fpage>
          -
          <lpage>214</lpage>
          . doi:
          <volume>10</volume>
          . 1108/14684520510598066.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Sinha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Ma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Eide</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.-J.</given-names>
            <surname>Hsu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>An overview of microsoft academic service (mas) and applications</article-title>
          ,
          <source>in: Proceedings of the 24th international conference on world wide web</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>243</fpage>
          -
          <lpage>246</lpage>
          . doi:
          <volume>10</volume>
          .1145/2740908. 2742839.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>P.</given-names>
            <surname>Ginsparg</surname>
          </string-name>
          , Arxiv at 20,
          <source>Nature</source>
          <volume>476</volume>
          (
          <year>2011</year>
          )
          <fpage>145</fpage>
          -
          <lpage>147</lpage>
          . doi:
          <volume>10</volume>
          .1038/476145a.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M. E.</given-names>
            <surname>Bales</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. N.</given-names>
            <surname>Wright</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. R.</given-names>
            <surname>Oxley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. R.</given-names>
            <surname>Wheeler</surname>
          </string-name>
          ,
          <article-title>Bibliometric visualization and analysis software: State of the art, workflows, and best practices (</article-title>
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Bascur</surname>
          </string-name>
          ,
          <string-name>
            <surname>N. J. van Eck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Waltman</surname>
          </string-name>
          ,
          <article-title>An interactive visual tool for scientific literature search: Proposal and algorithmic specification</article-title>
          .,
          <source>in: BIR@ ECIR</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>76</fpage>
          -
          <lpage>87</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Kong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Xia</surname>
          </string-name>
          ,
          <article-title>A survey of scholarly data visualization</article-title>
          ,
          <source>Ieee Access</source>
          <volume>6</volume>
          (
          <year>2018</year>
          )
          <fpage>19205</fpage>
          -
          <lpage>19221</lpage>
          . doi:
          <volume>10</volume>
          .1109/ACCESS.
          <year>2018</year>
          .
          <volume>2815030</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosenthal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Bolte</surname>
          </string-name>
          ,
          <article-title>Visual analytics of bibliographical data for strategic decision support of university leaders: A design study</article-title>
          .,
          <source>in: VISIGRAPP (3: IVAPP)</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>297</fpage>
          -
          <lpage>305</lpage>
          . doi:
          <volume>10</volume>
          .5220/0007396302970305.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>P.</given-names>
            <surname>Federico</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Heimerl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Koch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Miksch</surname>
          </string-name>
          ,
          <article-title>A survey on visual approaches for analyzing scientific literature and patents</article-title>
          ,
          <source>IEEE transactions on visualization and computer graphics 23</source>
          (
          <year>2017</year>
          )
          <fpage>2179</fpage>
          -
          <lpage>2198</lpage>
          . doi:
          <volume>10</volume>
          .1109/TVCG.
          <year>2016</year>
          .
          <volume>2610422</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Bastian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Heymann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jacomy</surname>
          </string-name>
          , Gephi: an open
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>