<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>M. Baldauf);</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Document Tagging - Exploring Adaptation Efects among Domain Experts</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sebastian Müller</string-name>
          <email>sebastian.mueller@ost.ch</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matthias Baldauf</string-name>
          <email>matthias.baldauf@ost.ch</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Peter Fröhlich</string-name>
          <email>peter.froehlich@ait.ac.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>St.Gallen, Switzerland</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>AIT Austrian Institute of Technology, Center for Technology Experience</institution>
          ,
          <addr-line>Giefinggasse 4, 1210 Wien</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>OST - Eastern Switzerland University of Applied Sciences, IPM Institute for Information and Process Management</institution>
          ,
          <addr-line>Rosenbergstrasse 59, 9001</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>ing Automation Experiences</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>1876</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>Keeping a professional knowledge database with scientific publications up to date requires continuous scanning and annotating of newly published articles by domain experts. To shorten this time-consuming process, we study experts' assistance through a domain-specific tag recommender. We introduce the real-life case of a knowledge management system for nursing practitioners and present its architecture and user interface for creating and assigning tag recommendations. While the original tagging interface was thought to assist experts in the tagging process and possibly challenge them to reconsider their tag selections or assign more tags to a document, a preliminary evaluation shows an uncritical adoption of the provided recommendations by experts. We conclude that future design iterations of the recommendation user interface should try to prevent blind trust in the system and encourage reflection on tag suggestions.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <sec id="sec-1-1">
        <title>Healthcare professionals are required to ensure safe and</title>
        <p>
          cost-eficient care. One way to achieve this is through
evidence-based practice, consisting of evidence from
research, context, patient preferences, and clinical
expertise. To provide evidence-based knowledge to nursing
practitioners at the point of care, the Eastern Switzerland
University of Applied Sciences ofers a knowledge
management system that informs about the latest scientific
work for practical use [
          <xref ref-type="bibr" rid="ref1 ref2">1</xref>
          ].
        </p>
        <p>While the system has been attracting great interest
among nursing practitioners, scanning and preparing
relevant scientific work is challenging. One of the core tasks
of the editorial team is the annotation of scientific
publications which is both time-consuming and cost-intensive.
To this day, the team labeled 1’515 nursing care or medical
publications with one or several of 24 tags to categorize
the works for later recommendation to practitioners.</p>
        <p>
          In order to support the work of the editorial team, we
have been implementing a tag recommender system for
health- and care-related scientific publications. Using a
BERT-based feature engineering approach [
          <xref ref-type="bibr" rid="ref3">2</xref>
          ] combined
with a standard radial basis function support vector
ma
        </p>
        <p>Besides such performance measures, we particularly
(P. Fröhlich)</p>
      </sec>
      <sec id="sec-1-2">
        <title>In this position paper, we present the current version</title>
        <p>focused on the user interaction with the tag recom- ering potential tag selections. We followed user-centered
© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License of the tag recommender system and report on first
expeAttribution 4.0 International (CC BY 4.0).
riences of the experts involved with the corresponding asked to re-tag 60 publications. Besides the functional
tagging user interface. recommender described above, we implemented an
additional version of the recommender which picked one
of the 24 tags at random. Both experts were provided
2. System Architecture and recommendations from both variants during the tests,
Tagging Interface without knowing which system recommended the tags
to the publication they were currently reviewing. For
The knowledge management system FIT-Nursing Care1 is each publication they tagged, they were asked to identify
a TYPO3-based website that makes relevant publications which of the two systems generated the notification.
available to nursing practitioners. The website as well as Their success rate in identifying the functional
recomour Python-based recommender system are hosted on a mender was 77% and 85%, respectively. This showed that
Linux server. Data such as relevant scientific publications the non-randomness of the recommendations is
noticeare managed in a MariaDB database and made accessible able and the experts can assess whether they are shown
via a REST-API. Novel suitable publications are identified a solid recommendation or not.
by the editorial team and added to the system’s literature After the system had been in use for three months, we
database. The recommender system is implemented as an interviewed the key expert mainly tasked with tagging
external component that checks the database for newly documents about his experiences. He reported that he
added publications once a day and determines tags to be had noticed a definite change in his tagging approach.
recommended. While originally he scanned a publication first and then</p>
        <p>
          Nursing care experts administrate these publications chose one or several adequate tags out of the list, he now
in the backend of the website. This includes uploading started to first check on the three recommendations made
new publications and assigning one or several of the and then check their plausibility against the publication,
24 available tags, selected from an alphabetically sorted in many cases only by examining the title of the
publicalist. For the last task, the recommender system comes tion. This process of quickly validating the tags
recominto play. Based on results from a previous co-design mended led to ignoring potential other suitable tags. In
workshop with the experts and as a consequence of the many cases, the option of assigning a non-recommended
distribution of the number of tags per document, three tag was overlooked.
tags are recommended and presented in red with the Furthermore, we found that the current version of
associated probabilities (see Fig. 1) [
          <xref ref-type="bibr" rid="ref4">3</xref>
          ]). the probability display did not fulfill its purpose. The
        </p>
        <p>The design shown in fig. 1 was deliberately chosen and expert reported to rather rely on the red highlighting, yet
implemented to not mislead experts into just confirming neglecting the actual probability presented. Whether the
the assignment of the tags recommended. With them still percentage provided was above 90% or between 60% and
having to select tags actively, we aimed at supporting 90% would result in the same outcome - the assignment of
the experts’ decision-making process without the system the tag. As described by the expert interviewed, this was
being too authoritarian. In the human factors literature, mainly done out of eficiency and convenience. This is
communicating the uncertainty (or reliability) of a sys- an interesting finding, which tends to be obtained much
tem’s perceptions or predictions has proven to support more often in real-world trials like the one presented
the formation of long-term trust [4, 5]. Ease of use and here, as compared to lab studies. Findings pointing in
seamless integration into the existing tagging process this direction have been presented [6], but only little
was achieved, in order to foster the continuous use of the empirical evidence has so far been gathered.
system without any additional eforts for the experts.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>4. Design Challenges</title>
    </sec>
    <sec id="sec-3">
      <title>3. First Evaluations</title>
      <sec id="sec-3-1">
        <title>While our first evaluation of the implemented system</title>
        <p>Throughout the implementation of the recommender was informal and only involved the key expert mainly
resystem, its performance was continuously evaluated on a sponsible for organizing and tagging the documents, we
statistical level using common machine learning metrics still found a crucial adaptation efect in his work
behav(such as precision, recall, map@3). ior. Despite his knowledge on the statistical evaluation</p>
        <p>Besides that, we were interested in how the novel tag results and the overall functionality of the recommender,
recommendation is used by the experts and how their he quickly started to have blind faith in the system. Only
original tagging behavior is impacted by the novel func- focusing on the top three recommended tags and
ignortion. We designed and conducted a small user study ing the probabilities shown, the expert accepted passing
with two experts from the editorial team, which were on the responsibility for assigning the correct tags to the
system.</p>
        <p>For further iterations of the tag recommender user prototype and evaluate alternatives for the current user
interface, we identify two design challenges: interface, while particularly trying to encourage
reflec</p>
        <p>Prevent blind trust: Users seem to trust the recom- tion on the tag suggestions.
mender regarding the top three tags suggested while
ignoring the probabilities displayed. How can we better
point out uncertainties and create awareness for more References
accurate manual checking? A potential solution could
include diferent color codes for visualizing the diferent
levels of certainty. Another approach could be to also
show the uncertainties of the other, non-recommended
tags, in order to increase users’ sensemaking of the data.</p>
        <p>Also, diferent levels of trust indications could be
experimented with: additionally to the tag-specific reliability,
also explanations for the quality of tag results could be
provided, as well as overall system reliability [7].</p>
        <p>Encourage reflection: The recommender was
supposed to assist in the tagging process, yet not to overrule
or replace the experts’ opinion. However, recommended
tags seem to influence the experts’ choice very strongly.</p>
        <p>How can we encourage reflection on recommendations
and combine automated recommendations and expert
knowledge in the best possible way?</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R.</given-names>
            <surname>Ranegger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Haug</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Vetsch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Baumberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bürgin</surname>
          </string-name>
          ,
          <article-title>Providing evidence-based knowledge on nursing interventions at the point of care: findings from a mapping project</article-title>
          ,
          <source>BMC Medical Informatics and Decision Making</source>
          <volume>22</volume>
          (
          <year>2022</year>
          )
          <article-title>308</article-title>
          . URL: https://doi.org/10.1186/s12911-022
          <article-title>-02053-8</article-title>
          . doi:
          <volume>1</volume>
          <fpage>0</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <source>1 1 8 6 / s 1 2</source>
          <volume>9 1 1 - 0 2 2 - 0 2 0 5 3 - 8</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , Bert:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          ,
          <year>2019</year>
          .
          <article-title>a r X i v : 1 8 1 0 . 0 4 8 0 5</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Tödtli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Vetsch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rickenmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Haug</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Baldauf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Fröhlich</surname>
          </string-name>
          ,
          <article-title>Designing experts' interactions with a semi-automated document tagging system</article-title>
          ,
          <source>Proceedings of the Workshop on Engaging with Automation co-located with the ACM Conference on Human Factors in Computing Systems (CHI</source>
          <year>2022</year>
          )
          <article-title>(</article-title>
          <year>2022</year>
          ). URL: https://ceur-ws.
          <source>org/ 5. Conclusion and Outlook</source>
          Vol-
          <volume>3154</volume>
          /short4.pdf.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>