<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Joint Workshop on Interfaces and Human Decision Making for Recommender Systems, October</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Comparing User Interfaces for Customizing Multi-Objective Recommender Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Patrik Dokoupil</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ludovico Boratto</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ladislav Peska</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Faculty of Mathematics and Physics, Charles University</institution>
          ,
          <addr-line>Prague, Czechia</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Cagliari</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>18</volume>
      <issue>2024</issue>
      <fpage>0000</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>The goal of Multi-Objective Recommender Systems (MORSs) is to adapt to the needs and preferences of the users from diferent beyond-accuracy perspectives. When a MORS operates at the local level, it tailors its results to the needs of each individual user. Recent studies have highlighted that, however, the self-declared propensity of the users towards the diferent objectives does not always match with the characteristics of the accepted recommendations. Therefore, in this study, we delve into diferent ways for users to express their preference toward multi-objective goals and observe whether they have some impact on declared propensities and overall user satisfaction. In particular, we explore four diferent user interface (UI) designs and perform a user study focused on the interactions with both the UI and the recommendations. Results show that multiple UIs lead to similar results w.r.t. usage statistics, but users' perceptions of these UIs often difer. These results highlight the importance of examining MORSs from multiple perspectives to accommodate the users' actual needs when producing recommendations. Study data and detailed results are available from https://osf.io/pbd54/.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Multi-objective recommender systems</kwd>
        <kwd>User study</kwd>
        <kwd>Recommender systems UI</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Multiple-Objective Recommender Systems (MORSs) produce results that account for the efectiveness
perspective but also go beyond it so as to tackle perspectives such as novelty, diversity, and fairness (to
name a few) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The optimization for these objectives can happen at the aggregate level so that the system
can guarantee certain properties (e.g., all providers receive a certain exposure in the recommendation
lists). Another alternative is to build MORSs that operate at the local (individual) level so as to shape
results towards the prominence of diferent goals for individual users (e.g., each user would receive
recommendations with a diferent level of diversity) [ 2]. Having local MORS, one may aim to provide
users with additional control over the recommendations and allow them to set their propensities towards
individual objectives [3]. This is in line with the general trend of a growing need for understanding and
control over recommendations as illustrated, e.g., by the recent EU’s Digital Services Act1. The regulation
requires that the main driving forces of the recommendation process are disclosed and also that users
should be allowed to select their preferred options (Article 27). However, recent literature revealed a
mismatch between the self-declared propensity of users for the diferent objectives and the characteristics
of the recommendations they accept (i.e., the items they choose among the recommendations are less
novel or diverse than what they believe they would like) [3].
      </p>
      <p>In this work, we explore the issue of self-declared propensities from the perspective of UI design.
Specifically, we focused on a widely used combination of relevance, diversity, and novelty criteria and
designed four diferent UIs that allow users to express their propensity toward the objectives. In a user
study (Section 3), we allowed users to interact with both the UI and the recommendations themselves
and evaluated the impact of diferent customization UIs. In particular, we observed whether the UI
designs afect how users perceive individual objectives, how they interact with recommendations, and
whether there is some impact on perceived recommendation quality and overall satisfaction.</p>
      <p>Our results (Section 4) show that there is a trade-of between the perceived usability of the diferent
UIs and their efectiveness at indicating user propensity. Moreover, no UI has clearly shown to be the
most efective, as users exploited three of our designs with similar efectiveness.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Background and Related Work</title>
      <sec id="sec-2-1">
        <title>2.1. Customization UIs in Recommenders</title>
        <p>We are not aware of any previous studies focusing on the comparison of UI designs for local MORS
setting. However, in the context of MORS, the work that most closely aligns with ours is by Harper et
al. [4], where the authors propose an algorithm allowing users to control item popularity and recency.
In addition to the algorithm and its ofline evaluation, the authors conducted a user study in the movie
domain, finding that the tuned recommendations were rated more positively by users. They also
highlighted the importance of individual-level optimization, as no single global setting worked equally
well for all users. The tuning was done using buttons labeled neutrally as “left” and “right.” While this
choice was intentional and justified, users responded negatively when asked about the ease of use of
the tuning interface. Therefore, we focused on diferent UI designs for RS tuning in our work.</p>
        <p>Several additional UI variants were considered for value setting in RS as well as other HCI tasks
[5, 3, 6, 7]. In web design praxis, sliders are considered to be a primary design choice for values
specification as long as these do not have to be very precise [ 7]. This is well-reflected in UIs used for RS
tuning, as illustrated, e.g., by the work of Liang and Willemsen [5] on tuneable exploration-oriented
music RS. Similarly, sliders were also used in [3] for the customization of local MORS. Nonetheless,
some researchers pointed out the slider’s inferior performance (e.g., w.r.t. response times) in situations
with limited options and advocated standard HTML radio buttons instead [6].</p>
        <p>However, unlike in other scenarios, the particular value of a MORS objective does not carry an inherent
meaning for the user (compared, e.g., to a price setting in an e-shop’s faceted search). Therefore, users
can only perceive the values relative to their previous settings (i.e., incremental increase/decrease; also
denoted as “relative” in the literature) or through the comparison with the weights of other criteria
( also denoted as “absolute” [8]). Naturally, UIs can be tailored to better reflect one of these views.
Another open question is the optimal level of response granularity [9] so that the task complexity is
minimized while the UI expressive power is still suficient.</p>
        <p>From these points of view, we can understand sliders as fine-grained UI collecting absolute feedback.
To cover other design options, we propose and evaluate two alternatives to the sliders UI. In options
UI, we provide users with several prompts to relatively increase/decrease the importance of individual
criteria (as such, coarse-grained feedback with relative answers is received). In a way, this layout is
most similar to the left/right buttons described in [4], but without the obfuscated labeling. The
plusminus buttons UI is inspired by common RPG gaming designs for character stats and, as such, provides
coarse-grained absolute feedback. Finally, in [3], authors reported on an extensive over-weighting of
beyond-relevance criteria by the users, and so the particular interpretation of user-provided weights
can be questioned as well. This was the main driving force for sliders_shifted UI variant, which reduces
the weights of beyond-relevance criteria.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Objectives in MORS</title>
        <p>MORS typically aim to balance recommendations’ relevance with various beyond-accuracy objectives,
including diversity, novelty, serendipity, or fairness [10]. Out of the available options, we adopted the
approach from [3], focusing on the following variants of relevance, novelty, and diversity.
• Estimated relevance  of recommendation list  was set to the mean of estimated relevance
scores (ˆ,) predicted by the relevance-only baseline: () = |1| ∑︀∈ ˆ,.</p>
        <sec id="sec-2-2-1">
          <title>Preference Elicitation t Recommendations</title>
          <p>a</p>
        </sec>
        <sec id="sec-2-2-2">
          <title>Consent and</title>
        </sec>
        <sec id="sec-2-2-3">
          <title>Demographics</title>
          <p>• Gender
• Age
• …</p>
        </sec>
        <sec id="sec-2-2-4">
          <title>Search Load More</title>
          <p>e
p
e
r
x
6
S
O
R
S</p>
        </sec>
        <sec id="sec-2-2-5">
          <title>Customization UI</title>
          <p>M
R
S
O
n
e
r
i
a
n
o
i
t
s
e
u
Q
y
d
u
t
S
t
s
o
P</p>
          <p>where (, ) is cosine similarity on items’ ratings.
• Novelty was defined as mean popularity complement: () = |1| ∑︀
where , is the feedback of user  on item  and  is the set of all users.
∈ 1</p>
          <p>−
1
• Diversity was set to collaborative intra-list diversity: CF-ILD() = ||* (||− 1)
︁(
|∈ :, exists| )︁ ,</p>
          <p>| |
∑︀
∀,∈,̸= (, ),</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. User Study</title>
      <p>The study was conducted on a movies domain using the EasyStudy framework [11] and the experimental
setup was largely based on [3]. In particular, we utilized the same filtered MovieLens Latest [ 12] dataset,
preference elicitation process, objective criteria definitions, items presentation, and task definition. In
the rest of this section, we provide details about the data pre-processing, describe the study flow, and
specify the customization UI variants we evaluated.</p>
      <sec id="sec-3-1">
        <title>3.1. Dataset</title>
        <p>For the purpose of the study, we utilized an augmented version of the MovieLens-Latest dataset [12].
The dataset was selected for its relative novelty and the general familiarity and popularity of the movie
domain among the general public. Both factors should contribute to the realisticness of the study. The
dataset was utilized both to train the collaborative filtering algorithms and as a starting point to gather
necessary item metadata. As the feedback collected during the study was binary, we binarized the
dataset as well (4* and above counts as positive). Furthermore, we only considered the more recent and
less obscure portion of the data. In particular, we removed movies released before 1990, ratings older
than 2010, movies that have less than 50 ratings per year, and users with less than 100 ratings. This
resulted in 9K users, 2K movies, and 1.5M ratings. In order to properly visualize the items, additional
metadata were collected from respective IMDb profiles: movie descriptions, posters, and links to movie
trailers.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Study flow</title>
        <p>The user study was organized in four phases: informed consent, preference elicitation, recommendation
comparison, and post-study questionnaire. The schematic of the study flow is depicted in Figure 1, while
the detailed description of individual steps follows.
3.2.1. Pre-study
Prior to the study commencement, users were shown a study mission statement and detailed instructions
and were asked for informed consent on the publication of anonymized data. Since all participants were
recruited using the Prolific service, we relied on the demographic information participants submitted
there and did not ask participants for additional demographics.
3.2.2. Preference elicitation
After the initial step, participants were routed to the preference elicitation page to collect information
to train the collaborative recommenders. In this phase, 24 movies were displayed to the users, asking
them to select movies they previously watched and liked. The displayed movies were sampled from
the dataset using the following procedure. We calculated estimated relevance (w.r.t. average user
profile), novelty (w.r.t. item’s mean popularity complement [ 13]), and diversity (w.r.t. CF-ILD [14])
characteristics for each item in the dataset. For each characteristic, we divided items into "low" and
"high" buckets and sampled four items from each bucket.2</p>
        <p>The displayed items were organized in a grid, and each item was represented by its poster image,
title, and genres. Users could express their preferences by simply clicking on the ones they watched
and liked before. We recommended selecting at least 5-10 items, but users were allowed to continue
even if fewer items were selected. To make the elicitation more thorough, users could dynamically load
additional items (repeating the procedure above) or search for specific ones via a text prompt.</p>
        <p>After the elicitation phase, we estimated the initial user’s propensities of users toward individual
objectives based on the normalized marginal gains calculated for each selected vs. each displayed
movie - i.e., as compared to all displayed movies, how much were the ones selected by the user
relevant/novel/diverse. Please see [3, 15] for more details on the procedure.
2Note that we first sampled from the relevance and novelty buckets and only then calculated the diversity w.r.t. already
selected items.
3.2.3. Recommendation comparison
The main part of the study comprised six iterations, where users received two lists of top-10
recommendations side-by-side. One list of recommendations was supplied by the relevance-only baseline
algorithm, while the other was provided by the multi-objective RS. In particular, we employed
generalized matrix factorization as a relevance-only baseline.3 Note the relevance-only baseline will also
be referred to as single-objective RS or simply SORS in the results. For the multi-objective RS, we
employed the RLprop algorithm [15] aiming to maintain the proportionality between user-defined
propensities and the fraction of individual objectives in the results. The considered relevance, novelty,
and diversity objectives were used as defined in Section 2.2 while noting that internally, the RLprop
algorithm normalizes the objectives using empirical cumulative distribution function (CDF) to make
them comparable against each other.</p>
        <p>Note that each participant received their own copy of the recommending algorithms (and objectives’
weights) that were updated after each step. That is, for each participant, the algorithms were gradually
ifne-tuned w.r.t. recommended items the user selected in previous iterations. 4 However, these “sandbox"
updates did not afect the recommendations given to the other users. Also note that if the item was
recommended in one iteration, it was removed from the set of candidates in the subsequent ones so
that the user was not overwhelmed with repeating recommendations.</p>
        <p>Regarding the number of iterations, we opted for a rather lower number to maintain a reasonable
study duration. Otherwise, too much of users’ attention could be lost, which would compromise the
results. Following the findings of [ 16], recommendation lists were organized into columns, and the
placement (i.e., left or right) of RS variants was randomized. Each item was represented with its poster
image, title, genres, a short plot summary, and a link to its trailer to allow users to thoroughly inspect
the previously unknown ones.5 See Figure 2 for a screenshot of the layout.</p>
        <p>At each iteration, users were asked for both low-level and high-level feedback. Similarly to the
preference elicitation phase, the low-level feedback was a simple click on relevant items. However, we
used a diferent prompt: “Select items that you would consider watching tonight.” As for the high-level
feedback, users were required to assign 1-5 stars to both recommendation lists so as to compare their
overall quality. Once the users finished the feedback provision, they were directed to the customization
UI, where they could modify their propensities towards individual objectives. These were then supplied
to the MORS algorithm to generate the next list of recommendations. We evaluated in total four variants
of the customization UIs (details in Section 3.3), while one variant was assigned to the user for the whole
duration of the study (i.e., a between-user variable). We opted for this due to supposedly substantial
carry-over efects that could otherwise compromise our results.
3.2.4. Post-study questionnaire
During the final phase, users were asked to fill out a post-study questionnaire containing 19 questions
together with 6 attention checks (instruction manipulation, nonsensical questions, and memory-based
questions). The questionnaire was inspired by the ResQue framework [17], but we altered it to primarily
cover the efect of customization UIs (see Figure 3 for the exact wording). Users were allowed to reply
in the form of a 5-point Likert scale. The exact prompts were “Strongly Disagree”, “Disagree”, “Neutral”,
“Agree”, “Strongly Agree”, and we also allowed users to answer “I don’t understand”. In the subsequent
analysis, we assigned the numeric values of -2, -1, 0, 1, and 2 to these prompts, while discarding the “I
don’t understand” answers.</p>
        <p>In the evaluation, we grouped the question into individual evaluated aspects of users’ attitudes
towards the system: through perceived relevance, novelty, and diversity, we aim to observe to what extent
these correspond to the users’ feedback and measurable characteristics of the resulting recommendations.
Several questions aim to determine whether users receive suficient information to participate in the study
3Based on the implementation from https://www.tensorflow.org/recommenders/examples/basic_retrieval.
4Unlike [3], algorithms were only updated by selections originating from that particular algorithm.
5Only the movie’s poster was initially visible; other information was displayed on mouse hover.
and whether the study interface was easy to use. Then, a series of questions focused on the customization
UIs: whether the initial state (i.e., estimated propensities) already provided good recommendations,
whether the efect of changing propensities was both positive and substantial, and whether the UI was
understandable, easy to use and gave the users suficient control to express their preferences. Finally, we
also enquired about the overall perceived satisfaction of the users.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Customization UIs</title>
        <p>The user study evaluated four diferent UI variants: sliders, sliders_shifted, options, and buttons (see
Figure 4). The sliders layout comprised three sliders, one for each objective, that automatically normalize
values to unit sum, i.e., when one value was being increased, others decreased proportionally. Note
that the sliders were initialized with the previous values of each objective. The sliders_shifted layout
appeared the same from the user’s point of view, but in line with the findings of [ 3], the relative weight
of relevance was increased. In particular, upon receiving the user-defined weights, we silently increased
the relevance’s weight by the factor of  = 0.5 and then re-normalized the weights again. As such,
both sliders variants provide users an interface with a well-perceivable tradeof between individual
objectives and a chance to provide fine-grained preferences.</p>
        <p>The options layout provided five radio buttons for each objective that allowed users to manipulate the
objectives relative to their previous weights. At -th iteration, objective weights [] were calculated as
[− 1] *  , where the factor  was derived from selected options (less: 1/2, slightly less: 2/3, same: 1/1,
slightly more: 3/2, and more: 2/1). As such, the options UI gives users a chance to relate their feedback
to the previous recommendations while allowing them to supply coarse-grained responses only.</p>
        <p>Finally, the plus-minus buttons layout utilized “virtual coins” to allow users to increase/decrease the
objective’s importance. At the very beginning, ten coins were assigned w.r.t. preference elicitation, and
at each iteration, the user received four additional coins to assign. Users could also transfer the coins
already assigned to other objectives (via a minus button). This is a very similar setting to many RPG
games, where players define their avatar’s statistics when the game beginnings and then incrementally
update them after some level-ups are accumulated. Therefore, we believe users may be quite familiar
with such a UI as well. Similarly as sliders, buttons UI is tuned to visualize the tradeof between individual
objectives. However, it only allows for a coarse-grained response and nudges users towards smaller,
incremental changes.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>The study was conducted in June 2023. In total, 142 participants were recruited using the Prolific.com
service. Participants were pre-screened for fluent English, no less than 10 previous submissions, and a
99% approval rate. Twelve users did not finish the study, and, in addition, we rejected 9 participants due
to failed attention checks, which resulted in 121 participants uniformly distributed along individual UIs
(i.e., at least 30 participants evaluated each UI). The study sample size was constrained by the funds
allocated for participants’ compensations. Nonetheless, we also conducted a sensitivity analysis in
G*Power software [18] using ANOVA with four groups,  = 0.05, and 1 −  = 0.8, concluding that the
study should be capable of discovering medium efects (Cohen’s  = 0.305) with reasonable probability.</p>
      <p>As for the participant’s demographics, the sample was rather well-balanced regarding gender (50%
female, 49% male, 1% unspecified). Participants were rather younger in general (mean age = 27, standard
deviation = 7.7, median age = 24), mostly white (64%) or black (21%), and mostly from South Africa (21%)
or several European countries (55% in total). The average time to complete the study was 15 minutes.</p>
      <p>In the analysis of the results, we focused on three main aspects: (i) whether diferent UIs afected
the received implicit and explicit user feedback, (ii) whether the UIs afected perceived RS qualities
as expressed in the questionnaire, and (iii) how individual questionnaire answers correlate with each
other.</p>
      <sec id="sec-4-1">
        <title>4.1. Users Feedback Analysis</title>
        <p>4.1.1. Comparison of single-objective and multi-objective RS
Let us first focus on the overall diference between the results of single and multi-objective RS. We
ifrst analyzed, whether the single- and multi-objective RS actually supplied users with diferent lists of
recommendations. To do so, we compared the corresponding pairs of lists given to the user at each
iteration w.r.t. the size of their intersection. Depending on the customization UI, the mean intersection
ranged from 12% (buttons UI) to 28% (sliders_shifted). Therefore we can conclude that the lists were
suficiently diferent to perform the subsequent analyses.</p>
        <p>Next, Table 1 contains estimated relevance and beyond-accuracy metrics evaluated on the resulting
lists. Similarly as in [3], we observed that MORS provided recommendations of higher diversity
(CFILD) and novelty. MORS also provided more diverse recommendations w.r.t. content-based ILD (cosine
similarity of associated genres; denoted as CB-ILD), more recent movies (mean year of release), and had
higher coverage of topics (w.r.t. associated genres). We evaluated these metrics both w.r.t. individual lists
and w.r.t. all items the algorithm recommended to a particular user throughout the six recommendation
iterations - yet the conclusions were the same. Also, all considered customization UIs exhibited the
same trend. However, the improvements in beyond-accuracy metrics were achieved at the expense of a
significant drop in the estimated relevance of recommended items (1.537 vs. 1.136). It seemed that the
deficiency w.r.t. relevance was perceived also by the users, who selected items with significantly higher
estimated relevance than the average values (1.282 vs. 1.136, T-test p-value: 1.8e-50). While a similar
trend was also observed for SORS selections, its magnitude was much smaller.</p>
        <p>Overall, single-objective RS obtained more user selections (3096 vs. 2200) and also received a higher
average rating from participants (3.36 vs. 2.78). On the other hand, significantly higher diversity, novelty,
and recency were also maintained within the selected items recommended by MORS as compared to
those recommended by SORS. Furthermore, the ratio of selected items recommended by SORS tends to
drop with subsequent iterations, while the volume of selected items recommended by MORS remained
roughly the same throughout all iterations. As such, we can conclude that despite not beating the
single-objective RS w.r.t. short-term utility, MORS brings favorable features that might pay of in the
long run.
4.1.2. Comparison of diferent customization UIs
6
s
in4
o
t
p
o2
0 0.0</p>
        <p>0.5
initial values
q1(relevaqn2ce(n)oqvq3e4l(t(dyin)ivfoe.rqss5iutyf(fR)icSieeqna6csy(eR)-Soqf-e7uas(seien)-fooqf.-8usus(efinf)ifcoie.nsucqyff9)ic(iUenIcinyi)t.qs1t0at(eU)Iqe1ff1ec(Ut)Iqe1ff2ec(Ut)Iqe1fqf3e1c(4Ut)I(UeIfqfse1uc5ftf)i(cUieIqnqs1uc16yf7f)i(c(UUieIInsuucnyffd)iceireqsn1tac8ny(d)UaIbeilaiqtsy1e)-9o(f-suasteis)faction)
Figure 6: Results of the questionnaire analysis. The mean of the numeric values corresponding to individual
answers is displayed.</p>
        <p>Relevance
Novelty
Diversity
1.0
0.5
0.0
0.5
diferences in the initial weights (i.e., after preference elicitation), so we trust this was a deliberate act
of the users. Furthermore, Figure 5 depicts the distribution of propensity scores in each iteration and
for each UI. It can be seen that the change is rather gradual for all UIs, but the average vector of the
change in buttons UI is opposite to those of other UIs (i.e., demoting rather than promoting relevance).
Also, while the other UIs tend to disperse the propensities more, these remain fairly compact in the
case of buttons UI.</p>
        <p>As an additional observation, we can see that while the feedback manipulation introduced by the
sliders_shifted UI had a visible efect on the propensity scores, it did not fully translate to the users’
feedback. While the fraction of selections and the mean rating were slightly higher for sliders_shifted
than for sliders, the diference was not significant, and also sliders achieved a slightly higher hit ratio. We
hypothesize that the measured diference in the resulting recommendations was simply not substantial
enough to trigger a significantly diferent response from the users. This is in line with the observations
of [19] regarding perceived diversity, where users often perceived the diversity of the presented lists
inversely or indiferently, despite relatively large diferences in the measured diversity levels.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Questionnaire Analysis</title>
        <p>The feedback analysis revealed that the buttons UI leads to inferior results w.r.t. short-term relevance,
while the other three UIs perform comparably with a slight preference towards sliders and sliders_shifted.
However, it is not yet clear whether the users perceived the results alike. So, in the questionnaire
analysis, we focused on evaluating additional axes of RS’s and customization UI’s quality.</p>
        <p>Figure 6 depicts the results of the questionnaire analysis. Let us start with overall remarks. Generally,
users were able to understand and answer required questions; We received less than 1% of “I don’t
understand” answers in total. The only question with more (9) of such answers was Q16 targeting
UI satisfaction. This might be partially caused by the fact that it was the only question with negative
phrasing. Therefore, we approach Q16 cautiously here and plan to rephrase it in future studies. Overall,
users answered that recommendations were suficiently relevant (Q1) and diverse (Q3) 6, but not quite as
novel (Q2). This might be an efect of using a bit older dataset or not taking movie recency directly into
account. The experiment environment’s validity is supported by overly positive answers to RS
ease-ofuse and information suficiency (Q4-Q8), but the suficiency and efect of the customization UIs may be
questioned to some extent, given slightly less positive answers for Q11, Q12, and Q14-Q16. Nevertheless,
answers on all questions except Q16 were above the neutral point (p-vals &lt; 0.0002). We plan to explore
this issue in the future by providing users with more options to tune the recommendations. Finally, users
of all UI variants agreed that tweaking the values had a visible efect on resulting recommendations
(Q13) and that they were generally satisfied with recommendations (Q19).</p>
        <p>Moving to compare diferent UIs, one of the main results was the superiority of the options and
buttons UI w.r.t. information suficiency. In particular, participants perceived the description of relevance,
novelty, and diversity as clearer (Q7) and better understood the purpose of tweaking their values (Q8).
This seemingly afected the perceived usefulness of the UI usage (Q10) and the understandability of the
UIs’ mechanisms (Q17).7 These findings can be, to some extent, backed by the work of Funke [ 6] - if we
accept that users internally perceive the task as one with a limited option space.</p>
        <p>Let us now briefly mention some more speculative observations. Despite its inferior efectivity, the
buttons UI surpassed both sliders and sliders_shifted in perceived ease of setting proper weights for
objectives (Q18). This might be an artifact of the finer-grained slider’s scale [ 20, 21], but the same
was not suficiently corroborated for the options UI. The perceived diversity (Q3) of buttons-based
recommendations was significantly higher than for sliders_shifted – in accordance with the diferences
of the user-defined diversity weights. In contrast, although the average weights of relevance criterion
were much higher for sliders_shifted than for sliders, users perceived sliders-based recommendations
as significantly more matching to their interests (Q1). This supports our previous hypothesis on the
somewhat inconsistent perception of individual objectives. However, a dedicated future study is needed
to quantify the magnitude of such inconsistencies.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Questionnaire correlations</title>
        <p>Finally, let us focus on the interdependence of the questionnaire answers. Figure 7 depicts the correlation
matrix of the responses to individual questions in the post-study questionnaire. We derive several
interesting observations from the results.</p>
        <p>First, considering overall satisfaction (Q19) as a target variable, we can see that no other evaluated
quality criteria exhibited a substantial negative correlation with satisfaction.8 Also, while the perceived
relevance (Q1) was strongly correlated with the overall satisfaction ( = 0.5), several questions related
to UI’s efect and suficiency had an even larger impact (Q11, Q12, Q14, Q16). This also translates to
6Means significantly above the neutral point; one-sample t-test p-vals &lt; 2.6e-19.
7In Q7, options improved over sliders (one-sided T-test p-value: 0.002) and sliders_shifted (p-val: 0.03). Also, buttons improved
over sliders (p-val: 0.02). In Q8, options and buttons improved over sliders (p-vals: 0.002 and 0.02 resp.). In Q10, options
improved over sliders (p-val: 0.027). In Q17, options improved over sliders (p-val: 0.041). Also, if all information suficiency
answers are merged together, options UI significantly outperforms sliders and sliders_shifted, while buttons UI outperforms
sliders.
8Note that Q16 was negatively formulated itself, so negative values actually indicate a positive efect.
compound statistics,9 where the mean UI efect and mean UI suficiency answers are strongly correlated
with the overall satisfaction ( = 0.57 and  = 0.62, respectively). Furthermore, UI’s understandability
(Q17) and ease of use (Q18) also exhibited a non-negligible correlation with overall satisfaction. To sum
up, these findings indicate a possible strong influence of UI controls and their function on the overall
user experience with the recommender systems. In this study, we only evaluated limited graphical
user interfaces. However, in light of emerging conversational RS powered by large-language models,
it may be crucial to focus on this aspect of user experience also in connection with additional UI and
interactional designs.</p>
        <p>Second, some of the considered quality axes (information suficiency , RS ease-of-use, UI efect , and UI
suficiency ) were targetted by multiple questions. However, while the questions targetting UI’s efect and
UI’s suficiency were highly correlated in most cases, this was not true for the information suficiency
and RS ease-of-use. As such, a finer-grained division of these objectives may be considered in future
work.</p>
        <p>In addition, we also focused on whether the diference in the feedback users provided on MORS and
SORS recommendations can be explained by some of the questionnaire answers. To do so, we also
calculated the correlations for the per-user diferences in mean ratings to MORS and SORS. In most
cases, we obtained close-to-zero results, with the exception of two moderate correlations: Q8 and Q17,
both targeting possible understandability issues (Q8: “I understood the purpose of tweaking relevance,
diversity, and novelty.”; Q17: “The mechanism (slider) for tweaking the objectives was understandable and
intuitive.”). Therefore, we can preliminarily conclude that the main driving force behind the adoption of
customizable individual MORS is actually whether we did a good job of explaining why &amp; how should
9I.e., using the mean of all answers targeting the same evaluated aspect.
users tune their propensities.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions and Limitations</title>
      <p>We tackled the problem of allowing users to indicate their propensity towards diferent recommendation
objectives, so as to shape more efective and better-tailored MORS. To this end, we conducted a user
study that allowed users to customize MORS via sliders, sliders_shifted, buttons, and options UIs. Results
show that while multiple UIs can lead to similarly efective recommendations (w.r.t. user feedback),
they can significantly vary in some of the user-perceived quality criteria. In particular, buttons UI
resulted in the lowest consumption-related statistics as well as lowest user ratings, while the other
three UIs performed without significant diferences from each other. The main driving force behind this
inferiority was a diferent distribution of propensities the users set through this UI. Further research is
needed to focus on the causes of this behavior and possible ways to support users in setting the best
possible values for their current needs.</p>
      <p>A subsequent questionnaire analysis revealed certain advantages of the options UI variant as compared
to more standard slider-based UIs. In particular, options UI dominated over sliders and sliders_shifted in
information suficiency , perceived usefulness, and UI’s ease-of-use aspects. As such, we can tentatively
recommend options UI with prompts relative to the previous criteria values as a good variant for
customizing local MORS.</p>
      <p>As an initial work on a rather complex topic, the study has numerous limitations, which we plan
to address in the future. First, when designing the evaluated UIs, we primarily aimed at the most
commonly used UI components. Even though, there was a plethora of parameters and design options
that we could not test due to the limits imposed on the number of participants. In particular, the current
study could not disentangle whether the diferences between sliders and options UIs were mainly caused
by the diferent “grounding” of the choices (i.e., relative to other criteria vs. relative to previous choices),
diferent response granularity, or diferent appearance. So, although we can conclude that there are
viable alternatives to the most common choice (i.e., sliders UI), the selection of the best such alternative
is a matter for future work.</p>
      <p>Similarly as in some related works [3], the study revealed several features that should favor MORS
over single-objective RS in the long term. However, a truly long-term study should be conducted to
verify these assumptions. As indicated by not-so-positive scores for Q14-Q16, there is some space for
revisiting the objectives by incorporating additional criteria or re-defining the current ones. Finally,
while the pool of participants was suficient to reveal the diferences in user feedback and corroborate
medium-sized efects in the questionnaire analysis, subtle efects might have been overlooked, which
could be remedied by contracting more users.</p>
      <p>Overall, we plan a series of larger follow-up studies that will focus on a more detailed long-term
analysis of user interaction and perception of MORS. This should also include studying the impact of
diferent domains, dataset properties, recommending algorithms, and study designs.</p>
      <sec id="sec-5-1">
        <title>Acknowledgments</title>
        <p>This paper has been supported by Czech Science Foundation (GAČR) project 22-21696S, Charles
University grant SVV-260698/2023, and Charles University Grant Agency (GA UK) project number
188322.
[2] D. Jannach, Multi-objective recommendation: Overview and challenges, in: H. Abdollahpouri,
S. Sahebi, M. Elahi, M. Mansoury, B. Loni, Z. Nazari, M. Dimakopoulou (Eds.), Proceedings of the
2nd Workshop on Multi-Objective Recommender Systems co-located with 16th ACM Conference
on Recommender Systems (RecSys 2022), Seattle, WA, USA, 18th-23rd September 2022, volume 3268
of CEUR Workshop Proceedings, CEUR-WS.org, 2022. URL: https://ceur-ws.org/Vol-3268/paper1.pdf.
[3] P. Dokoupil, L. Peska, L. Boratto, Looks can be deceiving: Linking user-item interactions and
user’s propensity towards multi-objective recommendations, in: Proceedings of the Seventeenth
ACM Conference on Recommender Systems, RecSys ’23, Association for Computing Machinery,
New York, NY, USA, 2023. URL: https://doi.org/10.1145/3604915.3608848. doi:10.1145/3604915.
3608848.
[4] F. M. Harper, F. Xu, H. Kaur, K. Condif, S. Chang, L. Terveen, Putting users in control of
their recommendations, in: Proceedings of the 9th ACM Conference on Recommender Systems,
RecSys ’15, Association for Computing Machinery, New York, NY, USA, 2015, p. 3–10. URL:
https://doi.org/10.1145/2792838.2800179. doi:10.1145/2792838.2800179.
[5] Y. Liang, M. C. Willemsen, Personalized recommendations for music genre exploration, in:
Proceedings of the 27th ACM Conference on User Modeling, Adaptation and Personalization,
UMAP ’19, Association for Computing Machinery, New York, NY, USA, 2019, p. 276–284. URL:
https://doi.org/10.1145/3320435.3320455. doi:10.1145/3320435.3320455.
[6] F. Funke, A web experiment showing negative efects of slider scales compared to visual analogue
scales and radio button scales, Social Science Computer Review 34 (2016) 244–254. doi:10.1177/
0894439315575477.
[7] B. Shneiderman, Designing the User Interface: Strategies for Efective Human-Computer
Interaction, 3rd ed., Addison-Wesley Longman Publishing Co., Inc., USA, 1997.
[8] Q. Zhao, The superior psychological impact of absolute (vs. relative) standing feedback does
not depend on the reward criterion, Social Psychology of Education 26 (2023) 473–484. URL:
https://doi.org/10.1007/s11218-023-09758-2. doi:10.1007/s11218-023-09758-2.
[9] L. Peska, S. Balcar, The efect of feedback granularity on recommender systems performance,
in: Proceedings of the 16th ACM Conference on Recommender Systems, RecSys ’22, Association
for Computing Machinery, New York, NY, USA, 2022, p. 586–591. URL: https://doi.org/10.1145/
3523227.3551479. doi:10.1145/3523227.3551479.
[10] D. Jannach, H. Abdollahpouri, A survey on multi-objective recommender systems, Frontiers in
Big Data 6 (2023). URL: https://www.frontiersin.org/articles/10.3389/fdata.2023.1157899. doi:10.
3389/fdata.2023.1157899.
[11] P. Dokoupil, L. Peska, Easystudy: Framework for easy deployment of user studies on recommender
systems, in: Proceedings of the 17th ACM Conference on Recommender Systems, RecSys ’23,
Association for Computing Machinery, New York, NY, USA, 2023, p. 1196–1199. URL: https:
//doi.org/10.1145/3604915.3610640. doi:10.1145/3604915.3610640.
[12] F. M. Harper, J. A. Konstan, The movielens datasets: History and context, ACM Trans. Interact.</p>
        <p>Intell. Syst. 5 (2015). URL: https://doi.org/10.1145/2827872. doi:10.1145/2827872.
[13] S. Vargas, P. Castells, Rank and relevance in novelty and diversity metrics for recommender
systems, in: Proceedings of the Fifth ACM Conference on Recommender Systems, RecSys ’11,
Association for Computing Machinery, New York, NY, USA, 2011, p. 109–116. URL: https://doi.org/
10.1145/2043932.2043955. doi:10.1145/2043932.2043955.
[14] K. Bradley, B. Smyth, Improving recommendation diversity, in: Proceedings of the twelfth Irish
conference on artificial intelligence and cognitive science, Maynooth, Ireland, volume 85, Citeseer,
2001, pp. 141–152.
[15] L. Peska, P. Dokoupil, Towards results-level proportionality for multi-objective recommender
systems, in: Proceedings of the 45th International ACM SIGIR Conference on Research and
Development in Information Retrieval, SIGIR ’22, Association for Computing Machinery, New
York, NY, USA, 2022, p. 1963–1968. URL: https://doi.org/10.1145/3477495.3531787. doi:10.1145/
3477495.3531787.
[16] P. Dokoupil, L. Peska, L. Boratto, Rows or columns? minimizing presentation bias when comparing
multiple recommender systems, in: Proceedings of the 46th International ACM SIGIR Conference
on Research and Development in Information Retrieval, SIGIR ’23, Association for Computing
Machinery, New York, NY, USA, 2023, p. 2354–2358. URL: https://doi.org/10.1145/3539618.3592056.
doi:10.1145/3539618.3592056.
[17] P. Pu, L. Chen, R. Hu, A user-centric evaluation framework for recommender systems, in:
Proceedings of the Fifth ACM Conference on Recommender Systems, RecSys ’11, Association
for Computing Machinery, New York, NY, USA, 2011, p. 157–164. URL: https://doi.org/10.1145/
2043932.2043962. doi:10.1145/2043932.2043962.
[18] F. Faul, E. Erdfelder, A. Buchner, A.-G. Lang, Statistical power analyses using g*power 3.1: Tests
for correlation and regression analyses, Behavior Research Methods 41 (2009) 1149–1160. URL:
https://doi.org/10.3758/BRM.41.4.1149. doi:10.3758/BRM.41.4.1149.
[19] P. Dokoupil, L. Boratto, L. Peska, User perceptions of diversity in recommender systems, in:
Proceedings of the 32nd ACM Conference on User Modeling, Adaptation and Personalization,
UMAP ’24, Association for Computing Machinery, New York, NY, USA, 2024, p. 212–222. URL:
https://doi.org/10.1145/3627043.3659555. doi:10.1145/3627043.3659555.
[20] C. C. Preston, A. M. Colman, Optimal number of response categories in rating scales: reliability,
validity, discriminating power, and respondent preferences, Acta Psychologica 104 (2000) 1–15.
URL: https://www.sciencedirect.com/science/article/pii/S0001691899000505. doi:https://doi.
org/10.1016/S0001-6918(99)00050-5.
[21] E. I. Sparling, S. Sen, Rating: How dificult is it?, in: Proceedings of the Fifth ACM Conference
on Recommender Systems, RecSys ’11, Association for Computing Machinery, New York, NY,
USA, 2011, p. 149–156. URL: https://doi.org/10.1145/2043932.2043961. doi:10.1145/2043932.
2043961.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>A survey of recommender systems with multi-objective optimization</article-title>
          ,
          <source>Neurocomputing</source>
          <volume>474</volume>
          (
          <year>2022</year>
          )
          <fpage>141</fpage>
          -
          <lpage>153</lpage>
          . URL: https://doi.org/10.1016/j.neucom.
          <year>2021</year>
          .
          <volume>11</volume>
          .041. doi:
          <volume>10</volume>
          . 1016/j.neucom.
          <year>2021</year>
          .
          <volume>11</volume>
          .041.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>