<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Explainable Robo-Advisors: Empirical Investigations to Specify and Evaluate a User-Centric Taxonomy of Explanations in the Financial Domain</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sidra Naveed</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gunnar Stevens</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dean-Robin Kern</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>Even though Recommender Systems (RS) have been widely applied in various financial domains such as Robo-advisors (RA), these systems still operate as a black box with no or limited explanations. Even in cases, where explanations are provided, such systems are mostly designed from the developers' perspective where the user needs and perspective of explanations are not taken into account. In this work, we aim to address the challenges of designing eXplainable Robo-Advisors (XRA) - by adopting a user-centric methodology. For this purpose, we applied a mixed-method approach, in which we conducted three qualitative focus group discussions (FGD) and supplemented the results with a quantitative survey insight. More specifically, we made two major contributions: 1) We extended the existing explanation categories to contextualize it for the financial domain - by identifying the user's specific needs for explainability in the context of the financial domain, 2) We quantified the user preferences of specific explanations with regard to the financial domain and explainable RA - by evaluating the user's personal relevance (PRE) and perceived quality (PQE) of explanations.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Recommender Systems in Financial Domain</kwd>
        <kwd>Robo-Advisor</kwd>
        <kwd>Explainable Robo-Advisor (XRA)</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        With recent advances in artificial intelligence (AI) and machine learning (ML), a growing number
of complex decision-making tasks are delegated to software systems and applications. In this
context, recommender systems (RS) have been widely used in various financial services domains
(such as online banking, loan, stocks, asset allocation, and portfolio management) – to support
people in complex financial decisions e.g., financial investment or retirement planning [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        Compared to other domains, such as movie or song recommendations, the financial domain
has several peculiarities: the finance world is complex where the risk involved in wrong decisions
is high and the average financial literacy of common users might be quite low [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. For instance,
most people might be familiar with the term Action Movie, but technical terms such as Bonds
or ETFs are typically known to domain experts only. Another issue refers to the long-term
validation of the recommendation quality. In the case of movies, a user can directly evaluate
the quality of the recommended movie by watching it. This is diferent in the case of financial
recommendations, where the actual recommendation quality might only be evaluated in the
long run e.g., years after investing money in buying a house or investing in funds. Furthermore,
the consequences of accepting the recommendations are diferent, where an unsuitable movie
recommendation might be bothersome, but an unsuitable financial recommendation can have a
dramatic impact on the life of a private investor. Thus, applying a RS in the financial domain is
a challenging task.
      </p>
      <p>
        In recent years, Robo-Advisors (RA) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] have become a popular financial RS that provides a
digital alternative for human financial advisors. Such technologies are now being considered the
"new wealth management interface of the 21st century". In general, RAs ask a series of questions
about the financial situation, risk tolerance, and investment preferences to create a personalized
portfolio recommendation [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. However, unlike in-person advice, users of RA cannot ask for
explanations and clarification as the technology still operates as a black box [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        To deal with the black box situation, various authorities such as the European Commission
(EC) and finance regulators have demanded better transparency and explainability of such
system [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Moreover, recent studies [
        <xref ref-type="bibr" rid="ref3 ref5">5, 3</xref>
        ] have also shown that such technologies are less
accepted by users due to the lack of transparency, explanation, and balancing of information
asymmetries.
      </p>
      <p>
        Currently, explanations of such financial RS have been commonly designed based on the
developer’s intuition for a "good explanation" [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] – without considering the user’s perspective
on the issue. In this work, we argue for a user-centric approach by considering the user’s need
and understanding of explanations in the specific context, to make explanations meaningful
and usable for the common user.
      </p>
      <p>
        To address the challenges of designing eXplainable Robo-Advisors (XRA) [
        <xref ref-type="bibr" rid="ref3 ref7">7, 3</xref>
        ] from a holistic
user perspective, we adopted a user-centric approach to mainly address the following research
questions:
      </p>
      <p>
        RQ1: What are the domain-specific user needs for explanations w.r.t. RA systems?
RQ2a: How important are the explanations for the users when interacting with RA
systems?
RQ2b: How the quality of explanations are perceived by the users when interacting with
RA systems?
These research questions aim to address the following: (1) the specification of explanation
taxonomies [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], which takes the domain-specific needs and peculiarities into account, and
(2) evaluate the taxonomy from a user’s perspective regarding the two aspects mentioned in
the literature: the Personal Relevance of Explanations (PRE) [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and the Perceived Quality of
Explanations (PQE) [
        <xref ref-type="bibr" rid="ref10 ref7">10, 7</xref>
        ].
      </p>
      <p>In this work, we adopted a mixed-method approach: To answer the "What" aspect of
explanations, we first conducted three qualitative focus groups discussion to explore the domain-specific
need for explanations by users (RQ1).</p>
      <p>To answer the "How" aspect of explanations, in the second step, we applied a quantitative
online survey to evaluate the personal importance (RQ2a) and the perceived quality (RQ2b) of
explanations identified in the first step.</p>
      <p>Overall, this work contributes in the following ways:
• Contextualizing and extending the theoretically-driven general taxonomies of
explanations with regard to the financial domain and XRA. This theoretical contribution of a
user-centric taxonomy provides in-depth insights into the user’s understanding of what
constitutes a good explanation and why it is needed.
• Quantifying user preferences of explanations with regard to the financial domain and
XRA. This practical contribution aims to help designers to identify and prioritize the
user need and relevance of specific explanations in the context of designing explainable
ifnancial RS.</p>
      <p>In the following, we will discuss relevant related work. Next, we will describe the
methodologies for both the qualitative study using focus group discussions and a quantitative online
survey. Afterward, we will present the results and insights from both qualitative and
quantitative studies. Finally, we will critically discuss the findings of both studies and conclude the
work by providing an outlook for future research.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Literature Review</title>
      <sec id="sec-2-1">
        <title>2.1. Explanations in Recommender Systems</title>
        <p>
          In the recommender systems (RS) literature, a number of explanation approaches for RS have
been proposed [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. The existing research in explainable RS has shown that explanations
can be beneficial for the success of RS in diferent ways e.g., providing the reasoning behind
a recommendation, enhancing the system acceptance by providing users with negative and
positive consequences of recommendations, helping users in making well-informed decisions,
or enabling communication between the service provider and the user [
          <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
          ]. Despite the
considerable amount of research in RS for generating and presenting explanations, providing
adequate explanations from the user’s perspective, and addressing the user’s needs for
explanations in a specific context or domain, remain under-explored. Providing diferent types and
levels of explanations could impact the user’s perception of the system in various forms e.g.,
lack of explanations could result in dificulty in understanding recommendations, which could
negatively afect the overall system acceptance [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ].
        </p>
        <p>
          Diferent types and classifications of explanations have been proposed in the RS literature.
A work presented in [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] provided a detailed overview of diferent explanation types along
with the question that can be addressed by each explanation type. These explanation types
are case-based, contextual, contrastive, counterfactual, everyday, scientific, simulation-based,
statistical, and trace-based explanations. Others have classified explanations in terms of the
underlying algorithm (e.g., collaborative filtering, content-based, or hybrid) or diferent sources
of information that influence the style of explanations generated [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ].
        </p>
        <p>
          In line with the type of information sources that generate explanations, a more recent work,
presented in [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], has provided a systematic literature review on explanations of RS. The authors
categorized explanations into four main categories based on the type of information presented
in the explanations. These explanation types are:
1. User Preferences or Input-Output Explanation – explanation exploits the content
related to the user’s provided inputs, which is further measured in terms of decisive input
values, preference match, feature importance analysis, and suitability estimate.
2. Decision Output or Outcome Explanation – explanation is provided by analyzing the
features of the alternative decisions, which might include a list of features and pros and
cons of each alternative, or the decisive features used in the inference process.
3. Decision Inference Process or Procedural Explanation – explanation based on the
inference process of a specific decision problem, which could be provided in terms of the
inference trace, inference and domain knowledge, decision method side-outcomes, and
self-reflective statistics.
4. Background and Complementary Information or Knowledge-Based Explanation
– explanation based on the additional background information related to the specific
decision problem, information about the knowledge sources used in the inference process,
past suggestions, user choices in similar situations, etc.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Financial Recommender Systems</title>
        <p>
          RS for stock investments have existed since the 1990s. While early work focused on individual
stocks [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ], there is a trend in the literature towards recommending entire portfolios [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ],
where the portfolio management shall consider diferent assets such as stocks, bonds, ETFs,
crypto-assets to provide the broadest possible risk diversification, match the risk appetite and
ifnancial situation of the investor, and consider social-ethical constraints. Algorithmic portfolio
management traditionally uses statistical procedures, but in recent years machine learning,
especially deep learning methods have also become popular.
        </p>
        <p>
          Robo Advisors (RA) have become the most prominent example of portfolio management
in the mass market [17]. As they currently work as black boxes, the call for explainable RAs
becomes a serious issue in society and as well as in academia [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
        </p>
        <p>
          To address the issue of the black box situation, in the last years technical studies have been
carried out [
          <xref ref-type="bibr" rid="ref3">3, 18</xref>
          ]. For instance, Babaei et al. [18] presented a portfolio optimizer, where the
input was a set of cryptocurrencies, and the output is the recommended portfolio. To provide
input-output explanations, they used the framework SHAP, which computes the Shapley values
of each cryptocurrency in the recommended portfolio. Krishnan et al. [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] used the data from
the RA system Paytm Money. This study aimed to show the link between the answers provided
for the initial questions asked by the RA and the resulting risk classes determined by the RA.
For this, they explored diferent XAI frameworks such as LIME, SHAP, and DeepLift, to provide
input-output explanations.
        </p>
        <p>
          Both works [
          <xref ref-type="bibr" rid="ref3">3, 18</xref>
          ] did not run any user study, but only focused from an engineering
perspective on the technical feasibility. Addressing the human desirability for explanations from an
HCI perspective, studies such as [
          <xref ref-type="bibr" rid="ref10 ref7">10, 7, 19</xref>
          ] have evaluated the efect of providing explanations
in the financial domain.
        </p>
        <p>
          Ben et al. [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] conducted an Amazon Mechanical Turk study, where the authors compared
the efects of providing diferent types of explanations on user’s perception (“no explanation”,
“human expert”, “global”, “feature-based”, and “performance-based”). In this study, participants
made financial decisions in a fictive game scenario, selling lemonade with the help of an
advisory system. Based on the results of their experiment, the authors concluded that there
was a significant diference in perception of explanations, and that the willingness to accept
non-human advice was higher on average than the so-called human advice.
        </p>
        <p>Schemmer et al. [19] implemented an AI advice system, which estimates the price of a
property to prevent buying overvalued houses. Their system estimates the price as the output
by various input features, such as the year of construction or living area. The system gives
inputoutput explanations by showing the feature importance calculated by the LIME framework [20].
The researchers conducted two focus groups, where the quality feedback given indicates that
explanations, in general, will be useful but for instance, the concept of feature importance was
dificult to understand by users and some also found it confusing to get multiple explanations.</p>
        <p>
          Deo et al. [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] have conducted a user study regarding XRA. For this reason, the authors used a
replica RA system. This system assigns a risk category based on ten questions answered by the
user. In the second step, a recommender system algorithm matches the user’s risk profile to a
fund risk profile to give funds recommendations. Users got additional explanations about local
and global feature importance using a SHAP-like visualization. In addition, explanations based
on features that afect the user and fund risk were provided. Using an online survey, Deo et al.
[
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] measured the user preferences and how the user comprehends the provided explanations.
The authors concluded that users benefit from explanations that relate the user input to the
system outcome. Yet, explanations should be simple, clearly stated, and meaningful so as not to
disrupt the user experience.
        </p>
        <p>
          Overall, even though the application of RS can be found extensively in various financial
domains e.g., online banking, loan, insurance, real estate, stocks, and asset allocation, to the
best of our knowledge, providing explanations in financial RS is underexplored in current
explainable RS literature. This holds especially true for RAs, which have become popular in
recent years. There is not just a lack of quantitative studies evaluating XRA, but also qualitative
studies understanding users’ perceptions and the need for explanations in the financial domain.
Regarding this, our work complements the research presented in [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], in which we relied on
the above-mentioned theoretical classification of explanations presented in [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] – used as a
sensitizing concept to investigate the user’s perception, need, or demand for explainability in
the financial domain.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>In this work, we used a mixed-method approach, supplementing the qualitative focus group
discussions with a quantitative online survey. In the following, we will describe both methodologies
in detail.</p>
      <sec id="sec-3-1">
        <title>3.1. Qualitative Focus Group Discussion</title>
        <p>By its very nature, users’ needs and personal meanings are subjective and open-ended. For this
reason, we used a qualitative research approach using Focus Group Discussions (FGD) [21]. As
a qualitative research technique, FGD has the ability to explore topics in-depth and determine
the constituting elements of the phenomenon (the "What") and their meaning (the "Why") from
a user perspective [22]. Such an explorative approach is particularly suitable when there is a
lack of predetermined hypotheses, theoretical concepts need to be contextualized regarding a
specific domain, or diferent perspectives on the issue need to be explored.</p>
        <p>We conducted three focus group discussions with the following stakeholders to consider
various perspectives about designing and using an eXplainable Robo-Advisor:
1. Domain Experts – participants with background and/or expertise in the financial domain.
2. HCI Experts – participants with background and/or expertise in Human-Centered</p>
        <p>Interaction and eXplainable AI (XAI).
3. Common Users – participants with no or limited background and/or expertise in the
ifnancial domain and XAI.</p>
        <p>We used the convenience sampling method [23] to recruit participants for the diferent focus
groups. In total, we recruited 13 participants (Domain Experts: 5, HCI Experts: 4, Common
Users: 4). All of them were well-educated (at least a university degree), but only the domain
experts had long experience with the financial domain (see Table 1).</p>
        <p>The domain experts and common users focus groups were conducted online using video
conferencing software with cameras switched on. The whole session was video recorded with
the consent of all participants. To conduct the sessions online for the two groups, for each group
we organized a virtual whiteboard using Miro1. The HCI experts focus group was conducted in
presence and audio of the entire session was recorded with the consent of all participants.</p>
        <sec id="sec-3-1-1">
          <title>3.1.1. Study Procedure</title>
          <p>The procedure for all focus groups was similar, but we had to make some adaptations for the
special needs of the online FGD. To conduct the focus group discussions, we followed the steps
and procedures described below which are also shown in Figure 1.</p>
          <p>1. Welcome: Participants were welcomed and were asked to introduce themselves in terms
of their background knowledge as well as their expertise in the financial domain.
2. Introduction I: A short introduction on the purpose of the focus group discussion was
provided as well as the expected outcomes of the study were explained to the participants.
3. Introduction II: In case of the Onsite-FGD, participants were given a brief explanation of
how to use Post-Its to record their thoughts. In the case of the Online-FGD, participants
were given a brief introduction to the Miro Board to collect thoughts and structure the
discussion.
4. Walk through: Participants were given a walk-through to a replica of a real (German)
Robo-Advisor bevestor2. Our RA replica was an interactive mock-up that was built using
Axure3, with no real functionality but allowing participants to interact with it.
5. Evaluate and Discussion: Participants were then asked to explore and critically analyze
the prototype. The following session was designed for brainstorming and reflects a
discussion covering three steps:
2https://bevestor.de/
3https://ancnjj.axshare.com
• Brainstorming – in which participants were required to write down any questions
they want the system to answer or explain to them, related to any aspect of the
system that is unclear to them.
• Presenting – in which participants were asked to present and explain the motivation
behind each written question.
• Clustering – in which participants were then asked to create clusters of similar
questions and label the clusters at the end.</p>
          <p>Each focus group was carried out by two researchers. One researcher took on the role
of a moderator. Her task was to explain the individual steps and stimulate the discussion.
Beyond this, she refrained from interfering. For instance, she took care not to participate in the
discussion herself or to evaluate the opinions and views of participants. The second researcher
took on the role of an observer by taking supplementary field notes in addition to the audio
and video recordings.</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>3.1.2. Data Analytics</title>
          <p>All FGDs were audio recorded and transcribed using amberscript4. The analysis was done
with the help of MAXQDA5 following the principles of the Thematic Analysis [24]. During
analysis, each FGD data set was coded by two independent researchers. Subsequently, the two
researchers discussed and analyzed each FGD jointly. The codes were consolidated and used to
define themes, which capture and summarize the core issue of coherent and meaningful pattern
[24]. Themes were discussed in joint interpretation workshops by the group of authors to gain
a mutual understanding of the material.</p>
          <p>Regarding the coding and defining themes, three principles have been discussed in the
literature: First, the principle of emergent, inductive, data-driven, and button-up coding [25],
where the researcher starts with a blank sheet and all codes resulted only from reading and
interpreting the empirical data. Second, the principle of deductive, theory-driven, and top-down
coding [25] starts with a pre-defined category system, where the aim of the analysis is mainly to
validate and illustrate the category system based on the empirical data. Third, the principle of
abductive, reflexive, and counter-current coding that integrates the elements of both approaches
[24]. The approach seems to sensitize researchers in their empirical work but needs to be flexible
enough to allow for new experiences.</p>
          <p>In our research, we adopted this third, abductive coding principle, where the theoretical
concepts outlined in Section 2.1 served as a sensitizing lens to analyze our data but allowed us
to open to new experiences that could not be well-explained by existing explanation categories.
As a result, our coding represents a dialog between the theory and the empirical data, by going
through the data and themes several times to refine them.</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Quantitative Online Survey</title>
        <p>The FGD is an established methodology in qualitative research for stimulating group discussion
and exploring a topic from diferent perspectives [ 26]. However, the method is limited in terms
4https://www.amberscript.com/en/
5https://www.maxqda.com/
of assigning specific statements to a person and quantitatively measuring personal perceptions
and preferences at an individual level. Therefore, we conducted a follow-up online survey, to
quantitatively evaluate the user perception of the RA replica in terms of its explainability and to
quantify the user relevance for diferent types of explanations they want in the RA. To address
both RQ2a and RQ2b, we asked our participants to evaluate both, the Personal Relevance of
Explanations (PRE) as well as Perceived Quality of Explanations (PQE). Both dimensions have
been evaluated in terms of the explanation categories identified in the FGD (see Section 4), i.e.
Recommender Explanation, Domain-Specific Information , and Shared Understanding.</p>
        <p>We adopted items from the User-Centric Evaluation framework [27] to get a comprehensive
evaluation of the Recommender Explanation category. As this framework does not cover the
aspects of Domain-Specific Information and Shared Understanding, for these categories we created
items on our own, where we tried to use the wordings from the FGDs as close as possible. For
evaluating the PRE, the questionnaire items were rated on a five-point scale from "Not at all
important" to "Very important" (see Table 3 and Table 4). For evaluating PQE, all questionnaire
items were rated on a five-point Likert response scale from "Strongly Disagree" to "Strongly
Agree" (see Table 5).</p>
        <sec id="sec-3-2-1">
          <title>3.2.1. Study Procedure</title>
          <p>To conduct the online survey, the following steps and procedures were followed:
1. Introduction: A short introduction on the purpose of the online survey and the procedure
of the survey was provided to the participants.
2. Prototype Exploration: Participants were presented with the same RA replica that they
saw before in the FGD.
3. Evaluation: Participants were then asked to explore and critically analyze the prototype
and were asked to return to the questionnaire after exploring the prototype, to answer a
series of questions.</p>
          <p>Except for P09, all the other 12 participants from the FGDs completely filled out the survey.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Qualitative FGD Results and Insights</title>
      <p>We identified the following major findings from the focus group studies: 1) The users have
multiple perceptions and understanding of explanations, 2) The empirical view of explanations
is diferent from the theoretical view of explanations, and 3) Overall, there is a general structure
of the empirical view about explanations for all three focus groups. These findings are reported
in the following sections.</p>
      <sec id="sec-4-1">
        <title>4.1. A User-Centered Taxonomy of Domain-Specific Explanations</title>
        <p>
          To address the RQ1, we followed the labels of clusters given by each focus group during the
Clustering phase. The labeled clusters were: "Explanation" (given by domain experts and
common users) and "Definition and Explanation of Terms" (given by HCI experts). In addition,
common users clustered several cards that explicitly used the term "Explaining", but the cluster
Recommender Explanations
16 7 6 3 IImmppaacctt ooff itnhpeuintpountroisnkpcolarstfsoifliicoagtieonneration
UOsuetrpuPtreEfxeprelannceatoiornIn[p8u]t12 6 4 2 SAyssstuemmp’tsiopnosrtmfoalidoegteongereanteioranteprtohceepssortfolio Procedural Explanation [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]
10 5 5 - PAodrdtiftoiolinoailnifnofromrmataiotinonw.ur.ste.dotfhoerrpuosretfroslio generation EKxnpolwanleadtigoen-B[8a]sed
8 - 7 1 PRoeartsfoonliiongchbaerhaicntdertihsteicpwor.rt.fto.lriiosko,pcthimanizgaetsio,fnuture options, etc. Outcome Explanation [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]
Domain-Specific Information
51 17 23 11 Definition of domain-specific terms and concepts Information [28]
Shared Understanding
10 2 7 1 uMsuertuaanldutnhdeerssytsatnedmin’sginbteetrwpereetnaatinosnwoefrsangsivweenrsby the Shared Understanding [29]
was labeled as "Information". Due to this, we also considered the aspect of "Information" in our
thematic analysis.
        </p>
        <p>By analyzing the codes in detail, we found that the user’s definition of explanation is quite
broad, but not arbitrary. From a user’s point of view, an explanation should help the user to
understand and make sense of the system and the provided recommendations. From such a
user perspective, we coded 107 total responses in the context of requesting an explanation from
the system (see Table 2).</p>
        <p>The further thematic analysis reveals that these responses can be grouped into three main
categories: "Recommendation Explanation" (43 of 107 codes), "Domain-Specific Information" (54
of 107 codes), and "Shared Understanding" (10 of 107 codes). In the following, we discuss these
categories in more detail.</p>
        <p>
          Recommendation Explanation. In our analysis, we used the explanation taxonomy
presented in [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] and as described in Section 2.1, but for our coding, we used the term "Recommender
Explanation" to make this category conceptually distinguishable from the other forms of
explanations we discovered from the FGD. We used the term Recommender to stress that the
requested explanations are directly related to the recommendation made by the system. We
further used the sub-categories from the explanation taxonomy presented in [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], to get a more
ifne-grained coding for this category:
        </p>
        <p>User Preference Explanation is used by us, to classify all responses where participants wanted
an explanation that shows a link between the recommendation and their preferences or the
answers given to the questions. For example, "What if I give a wrong answer?" (P07), "What was
the impact of each answer I gave on the result?" (P02), "What would be the impact of answering it
not honestly?" (P02).</p>
        <p>Output Explanation is used by us to classify all responses where participants wanted an
explanation about the recommended portfolio, but was not directly linked to the input in terms
of their answered questions, such as knowing more about risk, expected revenues, or reasons
for the specific portfolio composition. For example, "How high is the chance of having less than
200 % of my investment after 10 years?" (P03), "Why the system is not highlighting the aspects
that were considered or the reason why it suggested certain shares or funds?" (P02), "Why am I not
100 % invested?" (P03).</p>
        <p>Procedural Explanation is used by us to classify all responses where participants wanted
an explanation about the internal logic and reasoning of the system to recommend a specific
portfolio. For example, "What are the assumptions based on?" (P10), "Why are the specific assets
selected for me?" (P04), "Where do the values (proportion) come from?" (P10), "Why should I use
this system and not other systems like Trading View?" (P10).</p>
        <p>Knowledge-Based Explanation is used by us, to classify all responses where participants
wanted an explanation about the data used by the system to generate the portfolio or additional
information about other users or situations to answer their questions. For example, "What is
the standard duration for investment"? (P09), "What is the rate of success for these investment
portfolios?" (P06), "What is the database used to derive the results?" (P08), "What is the average
amount of trades of users within my peers?" (P03).</p>
        <p>
          Only 46 of the 107 coded requests for explanations are classified in the category of
"Recommendation Explanation". This indicates that in most cases, explanations requested by participants
do not fall into the categories provided by the theoretical taxonomy of explanations [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
        </p>
        <p>Domain-Specific Information. According to [28], the information relates to a specific
context that the user is not familiar with. For example, some domain-specific terminologies or
concepts, provide knowledge about something comprising either facts or details about a subject,
event, or situation, or provide knowledge that adds a value to a situation in a particular context
to make it understandable.</p>
        <p>We adopt this understanding to define the category of Domain-Specific Information , to classify
all responses accordingly. For example, "What is the risk-return profile?" (P04), "What exactly is
meant by the number of transactions?" (P12), "What does security assets mean?" (P09).</p>
        <p>According to the FGD responses, this category was the most prominent one, as 51 of the 107
coded requests for explanations are classified into this category.</p>
        <p>Shared Understanding. According to [29], Shared Understanding can be defined as
elaborating the mutual knowledge, beliefs, or assumptions or it can be an elaboration of how the system
interpreted the user’s goals and objectives. We adopt this understanding to define the category
of Shared Understanding, to classify all responses where participants wanted an explanation
from the system in order to know, if there is mutual knowledge or how the system interpreted
their answers to the questions. For example, "What if I understand small loss diferently from
you?" (P07), "How can I know, how you interpreted my risk assessment answers?" (P08), "What
about ethical aspects?" (P02).</p>
        <p>This category was the least prominent, as only 10 out of the 107 coded requests for
explanations were classified in this category.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Inter-Group Diferences</title>
        <p>By comparing the focus groups (see Table 2), we further observed that HCI Experts group was
the most responsive one (52 codes), followed by the Domain Experts group (37 codes), where
the Common Users group was least responsive (18 codes).</p>
        <p>In most cases, the responses from the three groups were related to all three aspects, namely
Recommender Explanation, Domain-Specific Information , and Shared Understanding. The only
diferences between the groups were that there were no responses related to Knowledge-based
Explanation from the Common Users group and Outcome Explanation from the Domain Experts.
Overall, the results depict the need for integrating all three aspects for users in the system
design to have an explainable financial RS.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Quantitative Results and Insights</title>
      <p>We identified the following major findings from the online survey: 1) The domain-specific
aspects of explainability, i.e. Recommendation Explanation, Domain-Specific Information , and
Shared Understanding, are equally relevant for participants without any significant diferences,
2) The domain-specific taxonomy is significantly more important for participants as compared
to the domain-general taxonomy adopted from the literature, 3) Evaluation of the RA replica is
not well-assessed by the participants w.r.t. its PQE. In the following, we describe these findings
in detail.</p>
      <sec id="sec-5-1">
        <title>5.1. Personal Relevance of Explanations (PRE)</title>
        <p>To address our RQ2a, we identified the users’ importance for explanations in the context of
the RA system in two steps: 1) Personal relevance of explanations w.r.t. the empirically-driven,
domain-specific categories, and 2) Personal relevance of explanations w.r.t. theory-driven,
domain-general categories.</p>
        <sec id="sec-5-1-1">
          <title>5.1.1. PRE w.r.t. Domain-Specific Categories</title>
          <p>We first computed the overall mean score of the combined three domain-specific categories and
we saw that the personal importance of the combined categories for all participants is quite high
on average (Avg. Mean = 4.08). This is also true for individual categories, especially regarding
Recommendation Explanation (Mean = 4.33) and Domain-Specific Information (Mean = 4.0) (as
shown in Table 3). The results are also in line with the focus group discussion from Section
4.2, where most of the responses were related to these two categories (see Table 2). Moreover,
the category Shared Understanding (Mean = 3.91) was also rated relatively high, indicating the
user importance to have explanations about how the system interpreted the user’s input.</p>
          <p>To statistically check the inter-category diferences, we utilized an ANOVA (  = 0.05) on the
mean scores of all participants. The result shows that there is no significant diference among
the three domain-specific categories regarding personal importance (  (2, 33) = 1.21,  = .310,
 2 = .07). Despite the insignificant diference, the high mean scores for all three categories still
indicate that these categories are important to participants.</p>
          <p>We further analyzed the inter-group diferences. We observed that Recommendation
Explanation was rated higher (Mean = 4.40) by the Domain Experts group. In the case of Domain-Specific
Information (Mean = 4.55) and Shared Understanding (Mean = 4.33), the HCI Experts had higher
ratings as compared to other groups.</p>
          <p>We checked the observed inter-group diferences statistically by applying one-way MANOVA
( = 0.05) on aggregated categories. The result showed that the categories are not significantly
rated diferently among all three groups (  (6, 14) = 0.69,  = 0.66,  2 = 0.22). To further
determine individual efects regarding each category, we ran univariate tests. The results shown
in Table 3 indicate that there is no significant diference among the groups in terms of individual
category i.e. Recommendation Explanation ( = 0.95), Domain-Specific Information ( = 0.23),
and Shared Understanding ( = 0.56). As the results were not significant, therefore the observed
inter-group diferences should not be overrated.</p>
        </sec>
        <sec id="sec-5-1-2">
          <title>5.1.2. PRE w.r.t. Generally Defined Categories</title>
          <p>
            In addition to the empirically grounded categories, to obtain a comprehensive evaluation, we
also adopted the well-established, generally defined categories from the explanation taxonomy
presented in [
            <xref ref-type="bibr" rid="ref8">8</xref>
            ]. This taxonomy covers additional areas, which have been proven to be relevant
in other domains. We used the same five-point scale as in Section 5.1.1, to evaluate the PRE
w.r.t. the categories theoretically defined by [
            <xref ref-type="bibr" rid="ref8">8</xref>
            ], namely: Input-Output Explanation, Outcome
Explanation, Procedural Explanation, and Knowledge-Based Explanation.
          </p>
          <p>We found that the personal importance of the combined categories for all participants is
between "moderately important" to "important" (Avg. Mean = 3.56). This is also true for
individual categories, where the Input-Output Explanation seems to be more important for
the participants (Mean = 3.89), followed by Outcome Explanation (Mean = 3.58), Procedural
Explanation (Mean = 3.54), and Knowledge-based Explanation (Mean = 3.21). To statistically
check the inter-category diferences, we utilized an ANOVA (  = 0.05) on the mean scores of
all participants. The result shows that there is a significant diference among the four categories
regarding personal importance ( (2, 45) = 5.44,  = .008,  2 = .19). We further ran individual
T-tests to check which of these categories are significantly diferent from others and found that
Input-Output Explanation received a significantly higher rating (  = 0.01) than Knowledge-Based
Explanation. We further found that Procedural Explanation had a significantly higher rating
than Knowledge-Based Explanation ( = 0.009).</p>
          <p>Moreover, we analyzed the inter-group diferences. We observed that HCI Experts have
higher ratings for all four categories as compared to other groups i.e. Input-Output Explanation
(Mean = 4.25), Outcome Explanation (Mean = 4.0), Procedural Explanation (Mean = 4.16),
Note: 5-point Likert Scale coding (1 = "Not at all important", 3 = "Moderately Important", 5: "Very Important"). Significant
diferences measured with (  = 0.05), are marked with *. Higher values (highlighted in bold) indicate better results.
Explanation Category and Items
For me, it is important to know ...
and Knowledge-Based Explanation (Mean = 3.93). Another rating pattern that can be observed
is that the Common Users have worse ratings for all categories except the Knowledge-Based
Explanation, as compared to other groups.</p>
          <p>To statistically check the inter-group diferences, we applied one-way MANOVA (  = 0.05)
on aggregated categories. The result showed that all four categories were not rated significantly
diferently among all three focus groups (  (8, 12) = 1.26,  = 0.34,  2 = 0.45). To determine
individual efects regarding each category, we ran univariate tests. The results, shown in Table
4, indicate that there was no significant diference among the groups in terms of individual
category, i.e. Input-Output Explanation ( = 0.07), Outcome Explanation ( = 0.40), Procedural
Explanation ( = 0.28), and Knowledge-Based Explanation ( = 0.27). As the results were not
statistically significant, the observed inter-group diferences should not be overrated.</p>
        </sec>
        <sec id="sec-5-1-3">
          <title>5.1.3. Empirically-driven Vs. Theory-driven Taxonomies</title>
          <p>In our study, we combined two types of taxonomy: the domain-specific (as shown in Table 3)
and the domain-general (as shown in Table 4).</p>
          <p>
            With our analysis, we further wanted to identify if there is any inter-taxonomy diference
regarding personal importance. For this, we first computed the overall mean score of combined
categories for each taxonomy and found that on average the personal importance of the general
categories taken from theory [
            <xref ref-type="bibr" rid="ref8">8</xref>
            ] was rated lower (Mean = 3.56) as compared to the
domainspecific categories grounded in the empirical data (Mean = 4.08). This is also true for individual
categories in both taxonomies, where each domain-general category (See Table 4) was rated
less important than each domain-specific category (See in Table 3). Analyzing both taxonomies
together, a general pattern can also be observed that for most categories, the HCI Experts have
higher personal importance as compared to other groups.
          </p>
          <p>To statistically check the inter-taxonomy, we conducted T-test and observed that
intertaxonomy diference is significant ( (82) = 3.20,  = .002,  = .70). This significant
intertaxonomy diference indicates that the empirically-grounded categories seem to be more relevant
for the participants than the theoretically adopted ones, in the context of the financial domain.</p>
        </sec>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Perceived Quality of Explanation (PQE)</title>
        <p>The previous section showed that XRA is an important issue for the participants, especially
regarding the categories identified in the FDGs. To further address our RQ2b, we asked the
participants to evaluate the perceived quality of explanations provided by the RA replica w.r.t.
the domain-specific categories i.e. Recommendation Explanation, Domain- Specific Information ,
and Shared Understanding using a five-point scale from "Strongly Disagree" to "Strongly Agree".</p>
        <p>We first computed the overall mean score of the combined three domain-specific categories
and saw that the perceived quality of the combined categories for all participants, is quite low
(Avg. Mean = 2.16) which is equal to "Disagree". This is also true for individual categories (as
shown in Table 5), where the Recommendation Explanation (Mean = 2.08) performed the worst,
Domain Specific Information (Mean = 2.19), and Shared Understanding (Mean = 2.22) are rated
higher, but still below 3 ("Disagree"). Overall, the results indicate that the explanations provided
by the RA replica are not well-perceived by participants w.r.t. all categories.</p>
        <p>To statistically check the inter-category diferences, we utilized ANOVA (  = 0.05) on the
mean scores of all participants. The result shows that there is no significant diference among
the categories regarding their perceived quality ( (2, 33) = .06,  = .943,  2 &lt; .01). This
indicates that participants perceived the explanation quality for all categories as similarly worse.</p>
        <p>We further analyzed the inter-group diferences. The results shown in Table 5 indicate that
Common Users perceived the quality of explanations in all cases as higher than the other groups.
To check the inter-group diferences statistically, we applied one-way MANOVA (  = 0.05)
on aggregated categories. The result showed that all categories are not significantly rated
diferently among the three focus groups (  (6, 14) = 2.09,  = 0.11,  2 = 0.47). To determine
individual efects regarding each category, we ran univariate tests. The results shown in Table 5
indicate that Common Users perceived the quality of explanations in all cases as higher than
the other groups. However, only in the case of "Recommendation Explanation", the rating is
significantly higher for Common Users as compared to other groups ( = .013).</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Discussion and Outlook</title>
      <p>To address the challenges of designing an eXplainable Robo-Advisor (XRA), we applied a
mixedmethod approach to qualitatively explore the user’s need and understanding of explanations in
ifnancial RS through FGD, which we then quantitatively verified and supplemented the findings
with an online survey.</p>
      <p>
        First, we addressed our RQ1, by conducting three qualitative FGD and identified a
usercentered taxonomy of domain-specific explanations. With this taxonomy, we demonstrated
that general explanation frameworks as presented in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], need to be adapted to take the
domainspecific needs into account. This is highlighted in our FGD insights which revealed that in
addition to Recommendation Explanation, the aspects of Domain-Specific Information and Shared
Understanding seem to be highly relevant for the users w.r.t. system explainability (See Table
2). The results are also in line with the study presented in [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], where the insights showed
the diferences in the users’ perception of explainable RS between Digital Cameras and Music
domains – indicating the efect of the domain on the user’s need and perception of explanations.
      </p>
      <p>
        We further quantitatively verified the FGD results and addressed our RQ2a by evaluating
the personal relevance of explanations (PRE) in two phases: 1) PRE w.r.t. our domain-specific
categories, and 2) PRE w.r.t. the general categories presented in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. For the former case, we
did not find any significant diferences in ratings for the categories. However, the personal
relevance was high for all three categories. The results also showed no significant diference in
the ratings of the groups (See Section 5.1.1). For the latter case, we also found no significant
diference between the rating of categories as well as the ratings made by the diferent groups
(see Section 5.1.2). Even though the results are not statistically significant, the insights from
both cases, showed an interesting pattern – where in all cases except for Recommendation
Explanation category, HCI Experts have a higher rating as compared to other groups (see Table
3 and Table 4). We believe that the reason for this could be twofold: 1) It has been shown that
in general, HCI Experts have higher expectations from the system to be self-descriptive [30].
This might also be the case w.r.t. explainability, but the limited explanations provided in the RA
replica might have resulted in a higher need for explanations, 2) the HCI Experts in our sample
have no experience or knowledge of the finance domain. This might have also afected the
results triggering them to have higher needs and relevance for all explanation categories.
      </p>
      <p>Our study further reveals that all general categories have received lower mean scores
compared to the domain-specific categories. The further quantitative comparison reveals that this
inter-taxonomy diference is significant. This means on average, the domain-specific categories
have higher personal relevance compared to the domain-general categories. These enumerative
depictions of the results thus verify the importance of domain-specific explanation needs in the
context of financial RS to perceive the system as explainable.</p>
      <p>
        We further addressed our RQ2b by evaluating the replica of an existing RA in terms of
the perceived quality of the explanations (PQE) (See Section 5.2). The participants rated the
explanation quality as rather low, which highlights the existing need from academia [
        <xref ref-type="bibr" rid="ref1 ref9">9, 1</xref>
        ] to
improve the explainability of financial RS. An interesting pattern we observed, however, is
that Common Users rated the PQE higher for all three categories as compared to other groups.
This could be explained under the assumption that all three groups are diferent in terms of
domain knowledge and their ability to perceive and understand system-provided information.
In this context, previous works on explainable RS in complex domains have also shown that
the complexity of the domain and decision task, also afect the user’s need to see explanations
at certain levels of detail [31, 32, 33]. It has been previously shown that novice users seem to
benefit more from the RS in a complex domain that provides simple or no explanations, to
avoid cognitive overload [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. This might be the reason that, despite limited or no technical
background, Common Users have a higher perceived quality of explanations when interacting
with the RA replica, which provides limited explanations. In addition, in the case of Recommender
Explanation the initial questions might serve as a placebo explanation [34] for the Common Users.
They might have implicitly assumed that the recommended portfolio and its corresponding
explanations were the results of the answers given by users, – which was not the case in our
experiment. Compared to this, it seems that Domain Experts are more skeptical about such
kinds of placebo explanations, thus reflecting in their lower perceived quality.
      </p>
      <p>Overall, the mixed-method approach of our study provides novel insights. However, the
approach has its limitations in terms of the small sample size used for both studies. Despite this
limitation, the results still shed a positive light on taking the domain-specific user’s needs into
account, to design the complex financial RS explainable from the user’s perspective. Future work
will validate the quantitative findings on a large sample size and will further focus on providing
the design implications from the user’s perspective to make the financial RS explainable for
users.
rule-base evidential reasoning using level 2 quotes, Expert Systems with Applications 39
(2012) 7150–7157.
[17] F. Abraham, S. L. Schmukler, J. Tessada, Robo-advisors: Investing through machines,</p>
      <p>World Bank Research and Policy Briefs (2019).
[18] G. Babaei, P. Giudici, E. Rafinetti, Explainable artificial intelligence for crypto asset
allocation, Finance Research Letters (2022) 102941.
[19] M. Schemmer, P. Hemmer, N. Kühl, S. Schäfer, Designing resilient ai-based robo-advisors:
A prototype for real estate appraisal, in: 17th International Conference on Design Science
Research in Information Systems and Technology, 1st-3rd June 2022, St. Petersburg, FL,
USA, 2022.
[20] M. T. Ribeiro, S. Singh, C. Guestrin, " why should i trust you?" explaining the predictions
of any classifier, in: Proceedings of the 22nd ACM SIGKDD international conference on
knowledge discovery and data mining, 2016, pp. 1135–1144.
[21] D. L. Morgan, Focus groups as qualitative research, volume 16, Sage publications, 1996.
[22] M. Bloor, Focus groups in social research, Sage, 2001.
[23] A. S. Acharya, A. Prakash, P. Saxena, A. Nigam, Sampling: Why and how of it, Indian</p>
      <p>Journal of Medical Specialties 4 (2013) 330–333.
[24] V. Braun, V. Clarke, Thematic analysis., American Psychological Association, 2012.
[25] J. Fereday, E. Muir-Cochrane, Demonstrating rigor using thematic analysis: A hybrid
approach of inductive and deductive coding and theme development, International journal
of qualitative methods 5 (2006) 80–92.
[26] T. O. Nyumba, K. Wilson, C. J. Derrick, N. Mukherjee, The use of focus group discussion
methodology: Insights from two decades of application in conservation, Methods in
Ecology and evolution 9 (2018) 20–32.
[27] P. Pu, L. Chen, R. Hu, A user-centric evaluation framework for recommender systems, in:</p>
      <p>Proceedings of the fifth ACM conference on Recommender systems, 2011, pp. 157–164.
[28] A. D. Madden, A definition of information, in: Aslib Proceedings, MCB UP Ltd, 2000.
[29] E. A. C. Bittner, J. M. Leimeister, Why shared understanding matters–engineering a
collaboration process for shared understanding to improve collaboration efectiveness in
heterogeneous teams, in: 2013 46th Hawaii International Conference on System Sciences,
IEEE, 2013, pp. 106–114.
[30] J. Prümper, Software-evaluation based upon iso 9241 part 10, in: Vienna Conference on</p>
      <p>Human Computer Interaction, Springer, 1993, pp. 255–265.
[31] S. Naveed, An Interactive Hybrid Approach to Generate Explainable and Controllable</p>
      <p>Recommendations, Ph.D. thesis, University of Duisburg-Essen, 2021.
[32] S. Naveed, B. Loepp, J. Ziegler, On the use of feature-based collaborative explanations:
An empirical comparison of explanation styles, in: Adjunct Publication of the 28th ACM
Conference on User Modeling, Adaptation and Personalization, 2020, pp. 226–232.
[33] S. Naveed, J. Ziegler, Featuristic: An interactive hybrid system for generating explainable
recommendations–beyond system accuracy, system 18 (2020) 33.
[34] M. Eiband, D. Buschek, A. Kremer, H. Hussmann, The impact of placebic explanations on
trust in intelligent systems, in: Extended abstracts of the 2019 CHI conference on human
factors in computing systems, 2019, pp. 1–6.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D.</given-names>
            <surname>Zibriczky</surname>
          </string-name>
          ,
          <article-title>Recommender systems meet finance: a literature review</article-title>
          ,
          <source>in: Proc. 2nd Int. Workshop Personalization Recommender Syst</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L.</given-names>
            <surname>Klapper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lusardi</surname>
          </string-name>
          ,
          <string-name>
            <surname>P. Van Oudheusden</surname>
          </string-name>
          ,
          <article-title>Financial literacy around the world</article-title>
          , World Bank. Washington DC: World Bank (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Krishnan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Deo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Sontakke</surname>
          </string-name>
          ,
          <article-title>Operationalizing algorithmic explainability in the context of risk profiling done by robo financial advisory apps (</article-title>
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>E. B.</given-names>
            <surname>Authority</surname>
          </string-name>
          ,
          <source>Eba report on big data and advanced analytics</source>
          ,
          <year>2020</year>
          . URL: https://www.eba.europa.eu/sites/default/documents/files/document_library/Final% 20Report
          <source>%20on%20Big%20Data%20and%20Advanced%20Analytics.pdf, accessed = 2022-07-28.</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>D.</given-names>
            <surname>Shin</surname>
          </string-name>
          ,
          <article-title>User perceptions of algorithmic decisions in the personalized ai system: perceptual evaluation of fairness, accountability, transparency, and explainability</article-title>
          ,
          <source>Journal of Broadcasting &amp; Electronic Media</source>
          <volume>64</volume>
          (
          <year>2020</year>
          )
          <fpage>541</fpage>
          -
          <lpage>565</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>R.</given-names>
            <surname>Guidotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Monreale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ruggieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Turini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Giannotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Pedreschi</surname>
          </string-name>
          ,
          <article-title>A survey of methods for explaining black box models, ACM computing surveys (CSUR) 51 (</article-title>
          <year>2018</year>
          )
          <fpage>1</fpage>
          -
          <lpage>42</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Deo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. S.</given-names>
            <surname>Sontakke</surname>
          </string-name>
          ,
          <article-title>Usability, user comprehension, and perceptions of explanations for complex decision support systems in finance: A robo-advisory use case</article-title>
          ,
          <source>Computer</source>
          <volume>54</volume>
          (
          <year>2021</year>
          )
          <fpage>38</fpage>
          -
          <lpage>48</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>I.</given-names>
            <surname>Nunes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Jannach</surname>
          </string-name>
          ,
          <article-title>A systematic review and taxonomy of explanations in decision support and recommender systems, User Modeling and User-Adapted Interaction 27 (</article-title>
          <year>2017</year>
          )
          <fpage>393</fpage>
          -
          <lpage>444</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>T.</given-names>
            <surname>Butler</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.</surname>
          </string-name>
          <article-title>O'Brien, Artificial intelligence for regulatory compliance: Are we there yet?</article-title>
          ,
          <source>Journal of Financial Compliance</source>
          <volume>3</volume>
          (
          <year>2019</year>
          )
          <fpage>44</fpage>
          -
          <lpage>59</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>D. Ben</given-names>
            <surname>David</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. S.</given-names>
            <surname>Reshef</surname>
          </string-name>
          , T. Tron,
          <article-title>Explainable ai and adoption of financial algorithmic advisors: An experimental study</article-title>
          ,
          <source>in: Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>390</fpage>
          -
          <lpage>400</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>F.</given-names>
            <surname>Gedikli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Jannach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ge</surname>
          </string-name>
          ,
          <article-title>How should i explain? a comparison of diferent explanation types for recommender systems</article-title>
          ,
          <source>International Journal of Human-Computer Studies</source>
          <volume>72</volume>
          (
          <year>2014</year>
          )
          <fpage>367</fpage>
          -
          <lpage>382</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>N.</given-names>
            <surname>Tintarev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Masthof</surname>
          </string-name>
          ,
          <article-title>Explaining recommendations: Design and evaluation</article-title>
          , in: Recommender systems handbook, Springer,
          <year>2015</year>
          , pp.
          <fpage>353</fpage>
          -
          <lpage>382</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>N.</given-names>
            <surname>Tintarev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Masthof</surname>
          </string-name>
          ,
          <article-title>Designing and evaluating explanations for recommender systems</article-title>
          ,
          <source>in: Recommender systems handbook</source>
          , Springer,
          <year>2011</year>
          , pp.
          <fpage>479</fpage>
          -
          <lpage>510</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Millecamp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Naveed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Verbert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ziegler</surname>
          </string-name>
          ,
          <article-title>To explain or not to explain: the efects of personal characteristics when explaining feature-based recommendations in diferent domains</article-title>
          ,
          <source>in: Proceedings of the 6th Joint Workshop on Interfaces</source>
          and
          <article-title>Human Decision Making for Recommender Systems</article-title>
          , volume
          <volume>2450</volume>
          , CEUR; http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2450</volume>
          /paper2. pdf,
          <year>2019</year>
          , pp.
          <fpage>10</fpage>
          -
          <lpage>18</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>S.</given-names>
            <surname>Chari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Seneviratne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. M.</given-names>
            <surname>Gruen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Foreman</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. K. Das</surname>
          </string-name>
          ,
          <string-name>
            <surname>D. L. McGuinness</surname>
          </string-name>
          ,
          <article-title>Explanation ontology: a model of explanations for user-centered ai</article-title>
          , in: International Semantic Web Conference, Springer,
          <year>2020</year>
          , pp.
          <fpage>228</fpage>
          -
          <lpage>243</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>L.</given-names>
            <surname>Dymova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Sevastianov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kaczmarek</surname>
          </string-name>
          ,
          <article-title>A stock trading expert system based on the</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>