<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Emerging Knowledge Extraction and Visualization in Medical Document Corpora</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Christian Nawroth</string-name>
          <email>christian.nawroth@fernuni-hagen.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marc Herrmann</string-name>
          <email>marc.herrmann@studium.fernuni-hagen.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Felix Engel</string-name>
          <email>fengel@ftk.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paul Mc Kevitt</string-name>
          <email>pmckevitt@ftk.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matthias Hemmje</string-name>
          <email>mhemmje@ftk.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Chair of Multimedia and Internet Applications, Faculty of Mathematics and Computer Science, University of Hagen</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Nawroth</institution>
          ,
          <addr-line>Herrmann, Engel</addr-line>
          ,
          <institution>Mc Kevitt and Hemmje</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Research Institute for Communication and Cooperation</institution>
          ,
          <addr-line>Dortmund</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <fpage>236</fpage>
      <lpage>254</lpage>
      <abstract>
        <p>In this paper, we demonstrate our concept of emerging Named Entities (eNEs) for the tasks of Emerging Knowledge Extraction and Visualization in medical document corpora. We derive four use cases that utilize eNEs in medical document corpora to support medical expert users accessing emerging knowledge. We design the visual Emerging Named Entity Recognition and Information Retrieval System (visual eNER-IRS), supporting three of these use cases. We demonstrate proofof-concept emerging knowledge visualizations for the different use cases. Finally, we present a detailed user evaluation of our visualization approach. The evaluation concludes that our approach helps users utilize eNEs on a corpus and a single document level. Overall, this paper demonstrates the benefits of our approach for the related project RecomRatio by providing recent and emerging knowledge for evidence-based medical use cases. Hence, the main contribution is a visualization of new medical concepts and emerging knowledge in literature for medical experts for supporting medical information, decision-making, and reporting in medical research and treatment.</p>
      </abstract>
      <kwd-group>
        <kwd>Emerging Named Entity Recognition (eNER)</kwd>
        <kwd>Emerging Knowledge Visualiazation</kwd>
        <kwd>emerging Named Entities (eNEs)</kwd>
        <kwd>Emerging Named Entity Recognition and Information Retrieval System (visual eNER-IRS)</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        2
Argument Entities available for Information Retrieval (IR) in medical document
corpora supporting medical argumentation engineering. Following our previous
work on this topic [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ][
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], medical eNEs are names for medical entities (e.g., for
diseases, drugs) that are in use in a medical document corpus (e.g., PubMed /
MEDLINE). Yet, they are not formally acknowledged through the expert
community, i.e., by adding them to a medical vocabulary. In addition, to support
the underlying medical argumentation use cases within RecomRatio we defined
an emerging Argument Entity (eAE) as an argument that contains an eNE in
its premise or the conclusion element [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. We argue that eNEs usually
represent the most recent knowledge in a domain. Hence, emerging Named Entity
Recognition (eNER) aims to recognize them in the document corpus and make
them available for medical IR use cases. To recognize eNEs, we propose a hybrid
approach combining textual Natural Language Processing (NLP) with Machine
Learning (ML) techniques on temporal features [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Whilst our previous
publications focused on the recognition of eNEs, here we explain why and how we
provide eNEs to the user through visual interfaces that support four use cases for
medical expert users. For our tasks presented here, we use two corpora: PubMed
MEDLINE Baseline 20203 (MEDLINE) and PubMed Open Access (PMC OA)
Subset4. Whilst the former generally only consists of the title and abstract, the
latter also contains the full texts, so we decided to use both for our Document
Engineering project. Between 1970 and 2019, the number of citations added to
MEDLINE grew from 219.337 entries per year to 1.406.789, based on our corpus
index statistic derived from our experimental corpora. So the yearly growth rate
increased by a factor 6.4 within 50 years. Furthermore, we outline how our work
links knowledge from these two corpora to the ClinicalTrials5 document corpus.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>State-of-the-Art and Related Work</title>
      <p>
        Our work is related to the task of realtime Emerging Topic Detection in
Microblogs as presented by Chen et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], which also utilizes ML techniques on
non-textual features to detect emerging topics within microblogs. Our approach
differs as it does not focus on realtime detection, but long-term eNEs in a
scientific text corpus and therefore, it uses different non-textual temporal features
compared to Chen et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Furthermore, our approach for eNER aims to
recognize eNE in scientific corpora and hence combine Document Engineering
techniques from traditional NER and ML. In recent work, Wang et al. [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] apply hot
topic detection to the field of academic big data, which they call Academic Hot
Topic Detection. Like our approach, they combine a textual NER approach in
the first stage with a feature learning approach. Their main features are a
cooccurrence graph and word embeddings amongst additional document related
features. In contrast, we focus on eNEs in a solely temporal way, not yet
analyzing whether these topics are “hot”, i.e. setting a trend of popular information
3 https://www.nlm.nih.gov/databases/download/pubmed medline.html
4 https://www.ncbi.nlm.nih.gov/pmc/tools/openftlist/
5 https://clinicaltrials.gov/
need/interest. The design of the visualization subsystem generally follows the
IVIS4BigData Framework introduced by Bornschlegl et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The framework
describes a method to transform raw data from big data sources into“advanced
visual user interfaces for Big Data Analysis” to allow“efficient and effective”
Human-Computer Interaction (HCI). A major component of the IVIS4BigData
is a pipeline that consists of different steps to provide data insight and
effectuation based on raw data. These steps are Raw Data Collection, Data Structures,
Visual Structures, and Views. Between Data Collection and Data Structures, the
framework includes an analytics layer that comprises an analytical component
of the underlying big data use cases.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Visual eNER-IRS System Design</title>
      <p>
        To design the visual Emerging Named Entity Recognition and Information
Retrieval System (visual eNER-IRS), we apply a user-centered design approach
[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Therefore, we introduce the four use cases of the visual eNER-IRS, give
a brief overview of the general system architecture, and derive the architecture
of the visualization subsystem that will become the basis of the later Argument
Visualization System. Here, we focus on the visualization subsystem and explain
the underlying architecture briefly to ensure general understanding. A detailed
description of the eNER pipeline is given in [
        <xref ref-type="bibr" rid="ref11 ref12">12, 11</xref>
        ].
3.1
      </p>
      <sec id="sec-3-1">
        <title>Use Cases of the Visual eNER-IRS</title>
        <p>
          The visual eNER-IRS is intended to support four different information retrieval
use cases [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. These are eNE retrieval support, document linking through NEs,
emerging Knowledge Discovery, and (later) emerging Argument Entity
discovery, as shown in Fig. 1. These four use cases are supported by one or two generic
visual use cases provided through the visualization subsystem. In the
following, we briefly introduce the four general use cases summarizing [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. The first,
eNE Retrieval Support (see Fig. 1) aims at providing functionality that utilizes
eNEs to enhance and support several standard retrieval methods, like query
completion, filtering, faceted search, and boosting of ranking results depending
on eNEs. The associated visual use case is visual eNE Retrieval Support. This
visualization use case is intended to highlight eNEs during several steps of user
interaction with the retrieval system. The second general use case supported by
one visualization use case is document linking through eNEs. In this use case,
eNEs are utilized to provide a link between documents from different corpora.
For example, a user finds a clinical trial in the ClinicalTrials (CT) Corpus that
contains eNEs that represent new medical knowledge in the respective clinical
trial. Then these eNEs can be used to search for documents in another text
corpus, e.g., MEDLINE, to retrieve new and emerging knowledge from that text
corpus too. The associated visual use case Visual Linking through eNEs is
intended to provide an interactive graphical representation of that use case, i.e.,
a network graph showing links between documents from different corpora based
4
on eNEs. The third general use case is visual emerging Knowledge Discovery.
This use case has an exploratory characteristic and enables users to explore new
knowledge on document level and the single emerging entity level. The
associated visual use cases (visual eNE highlighting in documents, visual eNE detailed
views) provide views for both exploratory aspects, which means visual
highlighting of eNEs in selected papers and providing detailed information on selected
eNEs based on the textual and temporal analysis of the eNER-IRS. The fourth
use case is emerging Argument Entity Discovery. Based on emerging Argument
Entities (e.g., from a survey article) in arguments’ premises or conclusions, the
expert medical users can retrieve, link, and visualize arguments that cover the
most recent medical knowledge. This use case is not covered by this paper but
published in [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ].
3.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>General Conceptual Architecture</title>
        <p>
          Following the motivation and the three initial use cases, our architectural
modelling approach (see Fig. 2) for recognizing eNEs in a medical document and
query corpus combines methods from NLP, NER, IR, and ML [
          <xref ref-type="bibr" rid="ref11 ref12">12, 11</xref>
          ]. Our
approach follows the Model View Controller (MVC) paradigm [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
        </p>
        <p>
          Here, we focus on the conceptual design of the view layer that contains the
visualization components of the eNER-IRS. As the View layer interacts with the
controller layer, we introduce the controller layer for a general understanding.
A more detailed description and evaluation of the underlying eNER pipeline in
the controller layer is published in [
          <xref ref-type="bibr" rid="ref11 ref12">12, 11</xref>
          ]. The core components of the
controller layer that are referenced in the View layer are the medical document
corpus, the baseline NLP and NER, the temporal features search engine, and
the temporal eNER Classifier. The medical document corpus is the indexed
document corpus containing all medical documents within the system. The baseline
NLP and NER provides the extraction of eNER candidates based on textual
features. The temporal features search engine extracts temporal features of the
eNE-candidates from the medical document corpus (e.g., year of first use). The
temporal eNER Classifier is the ML component that finally classifies eNE
candidates based on these temporal features. Based on the extracted eNEs and the
baseline preprocessed documents, the emerging Argumentation Extractor aims
at identifying eAEs in arguments. It will rely on a state-of-the-art argumentation
mining framework, e.g., ArgumenText [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. Following this overview of the core
eNER components in the controller layer in the next subsection, we introduce
the emerging knowledge visualization subsystem’s conceptual design.
3.3
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>Conceptual Design of the Visual Subsystem</title>
        <p>The conceptual design of the visualization subsystem has to consider challenges
posed by the big data characteristics of the underlying document corpora and
vocabularies. Hence, to address these challenges for the visualization
subsystem’s conceptual design, we applied the IVIS4BigData Framework. In general,
IVIS4BigData aims to make big data resources available and beneficial for users
through visualization. Fig. 3 shows how we use the IVIS4BigData Framework to
transform raw textual data from medical corpora into views that allow medical
expert users’ visual data insight and emerging knowledge effectuation.
Compared to the full IVIS4BigData Framework in our work, the pipeline part and a
feedback channel (user empowerment) are implemented. Furthermore, we focus
on the end-user’s view, but we do not implement the views on the first three
components of the full IVIS4BigData pipeline. Textual Big Data Sources for our
6
system are the two corpora PubMed and MEDLINE, as introduced above. The
raw textual data is collected in the medical document corpus from the system
architecture design (see Fig. 2). In the IVIS4BigData, the medical document corpus
refers to the“Data Collection” component. The analytics layer of IVIS4BigData
in our system design is represented by the two components that perform analytic
tasks (eNER) on the raw data and turn the raw data into data structures: The
baseline NLP and NER and the temporal eNER classifier in the controller layer
as described above. The following IVIS4BigData component“Data Structures” in
our work is represented by a JSON structure that encodes the mapping between
a corpus document and the automatically recognized eNEs. These mappings are
then transformed into a visual structure that is stored in a search engine index
whose Entity-Relationship Model (ERM) is shown in Fig. 4. The Article entity
has several attributes derived directly from the underlying corpus metadata,
such as Author, Abstract, or its corpus ID (e.g., for PUBMED / MEDLINE:
PMID). In contrast, Entity itself and its attributes are not taken from metadata
but extracted through the eNER-IRS. For example, those attributes contain the
Date Created (time of first use in the corpus), the name, and a possible category
of it. Hence, more generally speaking, the mapping introduced above maps
already existing corpus knowledge to new knowledge extracted by the eNER-IRS.
The final component of the IVIS4BigData Framework is the eNER-IRS-GUI
that we discuss in the remainder of this paper. We present a prototypical
proofof-concept implementation following the three initial use cases (eNE retrieval
support, document linking through NEs, emerging Knowledge Discovery) in the
following section.
Emerging Knowledge Extraction and Visualization in ...
7</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Proof-of-Concept Implementation</title>
      <p>This section describes the prototypical proof-of-concept implementation of the
visual eNER-IRS use cases based on the IVIS4BigData framework using several
methods and packages. First, we give a brief overview of the technical
implementation followed by descriptions of each of the visual eNER-IRS use cases followed
by a description of the graphical user interface (GUI) aspects.
4.1</p>
      <sec id="sec-4-1">
        <title>Technical Implementation</title>
        <p>In the following, we describe the prototypical technical implementation of the
previously introduced visual eNER-IRS use cases based on eNEs. The
underlying concept of the Visualization Subsystem (see Fig. 1) can be considered as two
independent software systems based on the IVIS4BigData Framework as follows:
To transform data structures into visual structures (see Fig 4) as a first step,
we developed a batch application based on the Spring Batch Java-Framework. It
converts all the eNER-IRS output JSON files, including data about eNEs such as
its name and occurrences in medical documents, into processed and consolidated
XML files. These files are then indexed by Apache Solr to create the eNEs visual
data structures (IVIS4BigData: Visual Mappings). These files are also used in
another batch processing step that reads the raw data of the different
medical corpus (PMC, ClinicalTrials), extracts relevant data attributes, and enriches
the data by adding information about entities such as MeSH concpets and eNEs.
Also, these XML output files are indexed by Apache Solr, and as a result, two
visual data structures for medical documents and eNEs are created. The second
system is a client-server architecture software based on the Spring Web
ModelView-Controller (MVC) Java-Framework. It generally reads the visual structure
data from Apache Solr and displays it on different web pages to provide the
described visual eNER-IRS uses cases (IVIS4BigData: View Transformations).</p>
        <p>The web page design and functionality is based on frameworks such as
Boot8
strap6 and jQuery7. Additionally, for the specific use case of visual linking and
highlighting of eNEs, we rely on the JS-Library D3.js8 for diagram and PDF.js9
for document visualization. To transport data from server to client, multiple
REST API endpoints are available and tailored for the specific visual eNER-IRS
use case. In the following, the prototypical implementations of the use cases are
explained.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Visual eNE Retrieval Support</title>
        <p>The document search enables users to browse through the different document
collections (PubMed, ClinicalTrials) easily with two filter categories (eNEs and
Medical Subject Headings): The main search field and the result of the document
search are displayed in the right view area. Above, next to the search field, there
is an additional button to control on which document collection the search should
be performed. Below, the matching documents, including title, publication date,
unique identifiers, entity categories, and source document collection, are listed
separately. The detailed view of a paper shows the authors, the assigned
entities. If available, the abstract in which the assigned entities are highlighted, is
shown. The left sidebar contains all necessary controls to conveniently browse the
dataset of Emerging Named Entities and Medical Subject Headings. As a result,
the documents are filtered based on the selected entities. Furthermore, the
sorting of the search result can be influenced by the green Learning To Rank (LTR)
button next to the main search field. Generally speaking, the documents are
sorted by an individual score in descending order. This score is a measurement
for how relevant each document is for the given user query. The gray shaded
numerical value reflects the default sorting, whereas the green shaded value also
considers the information about eNEs. In detail, this score increases based on
the number of assigned eNEs and on whether the user query also matches these
entities. To improve user search experience, the main search field is extended by
an auto-completion functionality (see Fig. 5) that lists query suggestions from
multiple datasets based on the current user input. Depending on the selected
document collections, the first suggestions are made by matching document
titles and eNEs. Also, further suggestions derive from the entity categories eNEs
and Medical Subject Headings (see Fig. 6). Selecting an entity suggestion results
in automatically adding the entity as a filter criterion in the left sidebar.
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Visual Linking through eNEs</title>
        <p>The relationships between entities are created once multiple entities appear in
a single document. The greater the number of documents in which two related
entities appear, the closer their relationship is. Such relationships can be
researched interactively with the help of the network graph accessible under the
6 https://getbootstrap.com/
7 https://jquery.com/
8 https://d3js.org/
9 https://mozilla.github.io/pdf.js/getting started/
Emerging Knowledge Extraction and Visualization in ...
9
navigation item Emerging Named Entity Graph (see Fig. 7). In the left sidebar,
all entities with at least one relationship are displayed. The list can be filtered
regarding type, reference, and name. Additionally, the list automatically updates
once an entity is selected by only showing entities in a direct relationship. The
network graph itself is shown in the right view area and refreshes automatically
once the entities’ selection is updated. A node represents each chosen eNE, and
its size depends on the total number of documents from the different collections
it is assigned to. The links between nodes display the connecting documents,
and their amount is represented by edge width. By clicking on the link details
of the relationship are revealed.
4.4</p>
      </sec>
      <sec id="sec-4-4">
        <title>Visual eNE Highlighting in Documents</title>
        <p>The detailed view of a document contains all mandatory attributes such as
title, unique identifiers, document collection and also optional attributes such
as authors, assigned entities (eNEs and non-eNEs (MeSH)) and the abstract or
original PDF-document. The availability of the optional attributes depends on
the collection source of the document. For example, only for documents from
the PubMed Central Open Access (PMCOA) collection, the original document
in PDF-format can be displayed. In the left sidebar, the assigned eNEs and
noneNEs are listed as interactive buttons. Also, these entities are highlighted in the
continuous text of the abstract or the PDF-document, if available (see Fig. 8).
In case of viewing a PDF-document, additional buttons to page backwards and
forwards and download the document are displayed. For eNEs, the interactive
button can be expanded to a drop-down list showing all the document pages on
which the term appears.
10
This visual eNER-IRS use case provides the highest interaction between expert
medical users and the eNER-IRS processes. Hence, the user can acknowledge or
reject an eNE candidate suggested by the eNER-IRS. The requested feedback
(and the respective interface) is intentionally binary (ACKNOWLEDGE /
REJECT), keeping in mind that expert medical users may lack data science
knowledge to give a more differentiated assessment. However, for those expert medical
users with data science / ML skills, the visualization provides two metadata
parameters from the eNER-IRS for their decision process. The data types of the
result set of the ML process for this visual eNER-IRS use case are terms/tokens
representing eNEs and temporal (statistical) metadata derived from the big data
analysis of the temporal feature search engine. Besides the temporal metadata,
the classification threshold from the underlying eNER-IRS component is
displayed (see Fig. 9).</p>
        <p>The detailed view of an entity contains all mandatory attributes. These are
unique identifiers, category, type, name, and also optional attributes such as
references to documents from different collections, overall frequency, and multiple
dates related to the creation, revision, and establishment (See Fig. 10). The
buttons in the left sidebar grouped by document collection reflect the related
entities. An additional bar chart showing the frequency of occurrence on an
annual basis is displayed depending on data availability. Two drop-down lists
can change the plotted range of years on the right side above the diagram.
Additionally, related entities’ occurrence data can be added interactively by the
buttons shown on the right-hand side.
Emerging Knowledge Extraction and Visualization in ...
11
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Evaluation</title>
      <sec id="sec-5-1">
        <title>Evaluation Methodology</title>
        <p>
          The evaluation primarily follows a task-oriented evaluation approach which
additionally includes the UMUX methodology to assess the perceived usability [
          <xref ref-type="bibr" rid="ref5 ref8">5,
8</xref>
          ]. It is based on a 20-page questionnaire which is included as auxiliary
material. Nine participants contributed to the evaluation. They belong to the user
stereotypes Medical User, Information Retrieval Expert, Science and Engineering
Expert and other.
5.2
        </p>
      </sec>
      <sec id="sec-5-2">
        <title>Test Questions</title>
        <p>Based on the first preparatory study results, we designed three medical test
scenarios described in the questionnaire. The test scenarios aim to figure out
to which extent users can use the eNER-IRS visualization to fulfill particular
use case scenarios. For each scenario, the questionnaire provides a detailed task
description:
1. Medical Document Search In this scenario, an exemplary search for medical
documents is conducted. The search includes, on the one hand, the filtering
of search results and on the other hand the visual highlight of Emerging
Named Entities. In particular, the highlighting of eNEs enables the user to
perform a context-sensitive and professional evaluation of the terms.
2. Details of eNEs This scenario covers the detailed consideration of all the
available information about Emerging Named Entities.
3. Relationships between eNEs This scenario shows how the use and
configure the network graph to identify and investigate the relationships between
eNEs.
12</p>
        <p>
          Fig. 8. Visual eNE Highlighting In Documents [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] (Article: [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]).
        </p>
        <p>For each of the scenarios, we defined 2-5 multiple choice test questions (TQs).
For each TQ, one or more answer options are correct. Users’ answers with all
correct answer options are classified as correct, with some correct answer options
as partly correct and with no correct answer options as wrong. Fig. 11 shows the
evaluation for the test questions (overall results and results per question). turns
out that there is a share of ≥ 32 of correct or partially correct answers for all
three scenarios. This finding concludes that, in general, users were able to fulfill
the three test scenarios defined above. However, within the questions of the
particular scenarios, a variance regarding the outcome can be observed. In scenario
(1), the TQ04, and scenario (2), the TQ06 has an outcome of correct answers of
&lt; 0.5. TQ04 deals with the highlighting of single eNEs in a document, TQ06 is
about ranking statistics. This finding concludes that only the visualization
details must be improved while the overall system is usable and performs well.users
were able
5.3</p>
      </sec>
      <sec id="sec-5-3">
        <title>Usability Metrics</title>
        <p>
          The goal of this step is to assess the perceived usability of the user based on ISO
9241-11. Therefore, the Usability Metric for User Experience (UMUX) with its
four-item Likert scale is used [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. with the following questions:
– Q01 The system’s capabilities meet my requirements.
– Q02 Using [this system] is a frustrating experience.
– Q03 [This system] is easy to use.
– Q04 I have to spend too much time correcting things with [this system].
        </p>
        <p>First, Fig. 12 shows the results of Q01 and Q03.</p>
        <p>It turns out that a more than 0.5 of the participants gave a positive or neutral
assessment for both questions. However, for Q03, turns out that a significant
Emerging Knowledge Extraction and Visualization in ...
13
portion of participants do not think that the system is easy to use. We argue
that this is not surprising but emphasize that the system is an expert system
that may require more extensive training to be used beneficially. Secondly, Fig.
13 shows the results of Q02 and Q04, now with a reversed colour scale compared
to Fig. 12 as explained before.</p>
        <p>Again, turns out that more than half of the participants have a positive or
neutral assessment regarding Q02 and Q04.</p>
        <p>Thirdly, we plot the mean results and the standard deviations for all
questions (See Fig. 14). The plots of the mean and standard deviation show reflect
the results introduced earlier. For Questions Q01 and Q03, the mean is greater
or equal to the neutral assessment, while for Q02 and Q04, it is below. We argue
that the relatively strong standard deviations result from the heterogeneous
participant group evaluating our system, with different experiences in the medical
domain, and using expert retrieval systems.</p>
        <p>Overall the usability evaluation showed a positive outcome, leading to the
conclusion that the system, in general, has reliable usability whilst there are
again improvements in visualization details. Furthermore, it shows the need for
sufficient training on the system for users inexperienced in using expert systems
or who are new to the medical domain.
5.4</p>
      </sec>
      <sec id="sec-5-4">
        <title>Added Value in Professional Terms</title>
        <p>
          The following questions are intended to assess the added value in professional
terms related to certain areas in the prototypical application. In contrast to
UMUX, here the 5-point Likert scale is used to express how much the tester
agrees (5) or disagrees (1) with a particular statement [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], (Neutral: (3), N/A:
(0)):
– Q05 The possibility to filter medical documents with the help of eNEs can
generate a true added value during the research process.
– Q06 The visually highlighted eNEs in continuous text and original medical
documents support the assessment in terms of their quality and professional
relevance.
– Q07 The possibility to download the original medical document with visually
highlighted eNEs allows distributing the information with expert colleagues
simply.
– Q08 The interactive bar chart is a valuable visualization to display the
occurrences of eNEs.
– Q09 The interactive network graph is a valuable visualization to display the
relationships between eNEs.
– Q10 The visualized relationships between eNEs support the assessment in
terms of their quality and professional relevance.
        </p>
        <p>Fig. 15 envsisions the results of the questions above.</p>
        <p>Here, turns out that again all questions have a neutral or better outcome
in more than half of the answers given. Except for Q08, more than 50% have
a better than neutral outcome. The high ratings for Q05, Q06, and Q09 are
promising. These questions reflect the core of our work and our use case. They
show that our concept of eNE and its utilization within information retrieval
use cases and their visualization are seen as beneficial by most participants. In
contrast, the questions Q07 and Q08 focus on detail visual implementations.
They emphasize that there is room for improvement regarding the aspects of
the visualization. Again, we plotted the mean results per question, including the
standard deviation (see Fig. 16).</p>
        <p>For all questions, it shows mean values significantly above the neutral value
of (3). Question seven has the strongest standard deviation, while the other
standard deviations are more moderate. That again reflects a strong variance among
Emerging Knowledge Extraction and Visualization in ...
15
Fig. 11. Test Questions’ Results</p>
        <p>Fig. 12. Usability Assessment I
the participants when it comes to using a detailed implementation feature. The
high means for questions Q05, Q06, and Q09 demonstrate that our concept of
eNEs and their visualization is useful and beneficial amongst the participants.
6</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusion and Discussion</title>
      <p>In this paper, we outlined a complete workflow utilizing and visualizing our
concept of emerging knowledge represented by eNEs. We introduced four visual
use cases to utilize eNEs by medical experts in document corpora. We derived
a system design for recognizing and visualizing them to support medical
retrieval and argumentation use cases. We designed a visualization subsystem on
the IVIS4BigData Framework, visualizing eNEs in documents and document
corpora and demonstrated a proof-of-concept implementation. Our evaluation
indicated that our three use cases and their four visualization could help expert
medical users searching for document-based medical evidence. Furthermore, we
identified items for improving our concept in visual details. We showed the
necessity for sufficient user training with the complex task of eNER in a medical
document corpus. Concerning future work, we outlined the emerging
Argumentation Extraction system.
18</p>
      <p>Nawroth, Herrmann, Engel, Mc Kevitt and Hemmje
Emerging Knowledge Extraction and Visualization in ...
19</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <fpage>5</fpage>
          -Point Likert Scale. In: Preedy,
          <string-name>
            <given-names>V.R.</given-names>
            ,
            <surname>Watson</surname>
          </string-name>
          ,
          <string-name>
            <surname>R.R</surname>
          </string-name>
          . (eds.)
          <source>Handbook of Disease Burdens and Quality of Life Measures</source>
          , pp.
          <fpage>4288</fpage>
          -
          <lpage>4288</lpage>
          . Springer, New York, NY (
          <year>2010</year>
          ). https://doi.org/10.1007/978-0-
          <fpage>387</fpage>
          -78665-0 6363, https://doi.org/10.1007/978-0-
          <fpage>387</fpage>
          -78665-0 6363
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. of Bielefeld, U.: Rationalizing
          <string-name>
            <surname>Recommendations (RecomRatio): ProjectHomepage. Bielefeld</surname>
          </string-name>
          (
          <year>2017</year>
          ), http://www.spp-ratio.de/de/projekte/ratiorec/
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bornschlegl</surname>
            ,
            <given-names>M.X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berwind</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaufmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Engel</surname>
            ,
            <given-names>F.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Walsh</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hemmje</surname>
            ,
            <given-names>M.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Riestra</surname>
            ,
            <given-names>R.:</given-names>
          </string-name>
          <article-title>IVIS4BigData: A Reference Model for Advanced Visual Interfaces Supporting Big Data Analysis in Virtual Research Environments</article-title>
          . In: Bornschlegl,
          <string-name>
            <given-names>M.X.</given-names>
            ,
            <surname>Engel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.C.</given-names>
            ,
            <surname>Bond</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Hemmje</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.L</surname>
          </string-name>
          . (eds.)
          <article-title>Advanced Visual Interfaces</article-title>
          .
          <source>Supporting Big Data Applications</source>
          . pp.
          <fpage>1</fpage>
          -
          <lpage>18</lpage>
          . Lecture Notes in Computer Science, Springer International Publishing,
          <string-name>
            <surname>Cham</surname>
          </string-name>
          (
          <year>2016</year>
          ). https://doi.org/10.1007/978-3-
          <fpage>319</fpage>
          -50070-6 1
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Amiri</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chua</surname>
            ,
            <given-names>T.S.:</given-names>
          </string-name>
          <article-title>Emerging topic detection for organizations from microblogs</article-title>
          .
          <source>In: SIGIR 2013 - Proceedings of the 36th International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          . pp.
          <fpage>43</fpage>
          -
          <lpage>52</lpage>
          (
          <year>2013</year>
          ). https://doi.org/10.1145/2484028.2484057
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Finstad</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>The usability metric for user experience</article-title>
          .
          <source>Interacting with Computers</source>
          <volume>22</volume>
          (
          <issue>5</issue>
          ),
          <fpage>323</fpage>
          -
          <lpage>327</lpage>
          (
          <year>2010</year>
          ), publisher: Oxford University Press Oxford, UK
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Herrmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Visualisierung von Emerging Named Entities im Rahmen von Information Retrieval in medizinischen virtuellen Forschungsumgebungen</article-title>
          . Masterthesis, FernUniversita¨t in Hagen, Hagen (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Krasner</surname>
            ,
            <given-names>G.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pope</surname>
          </string-name>
          , S.T.,
          <article-title>others: A description of the model-view-controller user interface paradigm in the smalltalk-80 system</article-title>
          .
          <source>Journal of object oriented programming 1(3)</source>
          ,
          <fpage>26</fpage>
          -
          <lpage>49</lpage>
          (
          <year>1988</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Laubheimer</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Beyond the NPS: Measuring Perceived Usability with the SUS, NASA-TLX, and the Single Ease Question After Tasks</article-title>
          and Usability
          <string-name>
            <surname>Tests</surname>
          </string-name>
          (
          <year>2018</year>
          ), https://www.nngroup.com/articles/measuring-perceived-usability/
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Nawroth</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Engel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eljasik-Swoboda</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hemmje</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Towards enabling emerging named entity recognition as a clinical information and argumentation support</article-title>
          .
          <source>In: DATA 2018 - Proceedings of the 7th International Conference on Data Science, Technology and Applications</source>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Nawroth</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , Duttenh¨ofer,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Hemmje</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          :
          <article-title>Argumentationsunterstu¨tzung durch emergentes Wissen in der Medizin</article-title>
          . In: to appear in: Wilhelm Bauer, Joachim Warschat,
          <source>Innovation durch Natural Language</source>
          Processing - Mit Ku¨
          <article-title>nstlicher Intelligenz die Wettbewerbsfa¨higkeit verbessern (</article-title>
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Nawroth</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Engel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hemmje</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Emerging Named Entity Recognition in a Medical Knowledge Management Ecosystem</article-title>
          .
          <source>In: Accepted for: Proceedings of 12th International Conference on Knowledge Engineering and Ontology Development</source>
          . Budapest (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Nawroth</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Engel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mc</surname>
            <given-names>Kevitt</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Hemmje</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.L.</surname>
          </string-name>
          :
          <article-title>Emerging Named Entity Recognition on Retrieval Features in an Affective Computing Corpus</article-title>
          .
          <source>In: 2019 IEEE International Conference on Bioinformatics and Biomedicine (BIBM)</source>
          . pp.
          <fpage>2860</fpage>
          -
          <lpage>2868</lpage>
          (
          <year>Nov 2019</year>
          ). https://doi.org/10.1109/BIBM47256.
          <year>2019</year>
          .8983247
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Norman</surname>
            ,
            <given-names>D.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Draper</surname>
            ,
            <given-names>S.W.:</given-names>
          </string-name>
          <article-title>User centered system design: New perspectives on human-computer interaction</article-title>
          . CRC Press (
          <year>1986</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Palau</surname>
            ,
            <given-names>R.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moens</surname>
            ,
            <given-names>M.F.</given-names>
          </string-name>
          :
          <article-title>Argumentation mining: the detection, classification and structure of arguments in text</article-title>
          .
          <source>In: Proceedings of the 12th international conference on artificial intelligence and law</source>
          . pp.
          <fpage>98</fpage>
          -
          <lpage>107</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Siroy</surname>
            ,
            <given-names>A.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boland</surname>
            ,
            <given-names>G.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Milton</surname>
            ,
            <given-names>D.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roszik</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frankian</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Malke</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haydu</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prieto</surname>
            ,
            <given-names>V.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tetzlaff</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ivan</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>W.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Torres-Cabala</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Curry</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roy-Chowdhuri</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Broaddus</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rashid</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stewart</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gershenwald</surname>
            ,
            <given-names>J.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Amaria</surname>
            ,
            <given-names>R.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patel</surname>
            ,
            <given-names>S.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Papadopoulos</surname>
            ,
            <given-names>N.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bedikian</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hwu</surname>
            ,
            <given-names>W.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hwu</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Diab</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Woodman</surname>
            ,
            <given-names>S.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aldape</surname>
            ,
            <given-names>K.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Luthra</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patel</surname>
            ,
            <given-names>K.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shaw</surname>
            ,
            <given-names>K.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mills</surname>
            ,
            <given-names>G.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mendelsohn</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meric-Bernstam</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>K.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Routbort</surname>
            ,
            <given-names>M.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lazar</surname>
            ,
            <given-names>A.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Davies</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          :
          <article-title>Beyond BRAFV600: Clinical Mutation Panel Testing by Next-Generation Sequencing in Advanced Melanoma</article-title>
          .
          <source>Journal of Investigative Dermatology</source>
          <volume>135</volume>
          (
          <issue>2</issue>
          ),
          <fpage>508</fpage>
          -
          <lpage>515</lpage>
          (
          <year>Feb 2015</year>
          ). https://doi.org/10.1038/jid.
          <year>2014</year>
          .
          <volume>366</volume>
          , http://www.sciencedirect.com/science/article/pii/S0022202X15371189
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Stab</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daxenberger</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stahlhut</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schiller</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tauchmann</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eger</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gurevych</surname>
          </string-name>
          , I.:
          <article-title>ArgumenText: Searching for Arguments in Heterogeneous Sources</article-title>
          .
          <source>In: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations</source>
          . pp.
          <fpage>21</fpage>
          -
          <lpage>25</lpage>
          . Association for Computational Linguistics, New Orleans,
          <source>Louisiana (Jun</source>
          <year>2018</year>
          ). https://doi.org/10.18653/v1/
          <fpage>N18</fpage>
          -5005, https://www.aclweb.org/anthology/N18- 5005
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
          </string-name>
          , H.:
          <article-title>Detecting Hot Topics From Academic Big Data</article-title>
          .
          <source>IEEE Access 7</source>
          ,
          <fpage>185916</fpage>
          -
          <lpage>185927</lpage>
          (
          <year>2019</year>
          ), publisher: IEEE
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>