<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>DOLAP</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Explanations for Sequential Group Recom mendations</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Md. Mahade Hasan</string-name>
          <email>mdmahade.hasan@tuni.fi</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Soha Pervez</string-name>
          <email>soha.pervez@tuni.fi</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maria Stratigi</string-name>
          <email>maria.stratigi@tuni.fi</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kostas Stefanidis</string-name>
          <email>konstantinos.stefanidis@tuni.fi</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Languages and Analytical Processing of Big Data</institution>
          ,
          <addr-line>co-located with</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Tampere University</institution>
          ,
          <addr-line>Tampere</addr-line>
          ,
          <country country="FI">Finland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>26</volume>
      <abstract>
        <p>A growing number of applications enable users to form groups for activities, like visiting a restaurant or watching a movie, making group recommenders more prevalent than ever. SQUIRREL is a framework for sequential group recommendations, providing a diferent recommendation in each round. It relies on Reinforcement Learning to select appropriate group recommendation algorithms based on the current state of the group. At each round of recommendations, it calculates the satisfaction of each group member and selects a recommendation method that will produce the maximum reward. In this paper, we incorporate two new reward functions, utilizing the m-proportionality measure to produce recommendations that are fairer to the group by promoting at least  items in the group recommendation list that each member prefers. Moreover, we study a user case explaining the SQUIRREL recommendations.</p>
      </abstract>
      <kwd-group>
        <kwd>Group recommendations</kwd>
        <kwd>Sequential recommendations</kwd>
        <kwd>Fairness</kwd>
        <kwd>Explanations</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR</p>
      <p>ceur-ws.org
1. Introduction
munication have encouraged more people to socialize
and engage in group activities. This has led to the group
recommendation research area becoming one of the more
popular ones for recommender systems. Group
recommendations should be able to recommend items relevant
to the group members without any bias, meaning that a
mendation list that is relevant to them. In addition, the
system needs to take into account previous interactions
quential group recommender, which ensures that, after a
sequence of recommendations, all members are satisfied</p>
      <p>Group recommendations can be approached in many
ways, but it is not always clear which approach is best
for each situation. Due to the wide range of group
recommendation methods and their widely difering
approaches, evaluating all of them and deciding which one
works best for each test scenario is dificult and
timeconsuming. Not all group recommendation methods
have the same performance when utilized in diferent
domains. For example, a group recommender system
build for movie recommendations has a vastly diferent
performance when it is used in diferent domains [ 1].</p>
      <p>Because of this, transferring a recommender system to</p>
      <p>CEUR
Workshop
Proce dings
htp:/ceur-ws.org
ISN1613-073</p>
      <p>CEUR</p>
      <p>Workshop Proceedings (CEUR-WS.org)
another domain is not an easy task.</p>
      <p>The SQUIRREL framework [2] provides a solution to
which is an ideal model since it directly mirrors the
sequential nature of recommendations. The model is
composed of three main elements: state, actions, and reward.</p>
      <p>The state of the model is how satisfied the group
members are with all the previous recommendation rounds.</p>
      <p>Actions are the diferent group recommendation methods
cessful the action chosen was. At each recommendation
round, SQUIRREL automatically selects the best group
satisfied the group members are (state).</p>
      <p>In this work, we mainly focus on achieving fair
seutilizing reward functions that promote fairness among
the group. We aim to maximize the satisfaction of the
group members, while trying to minimize their
disagreement. We define a user’s satisfaction by how relevant
the recommended items are to each user and their
disagreement as the diference between their satisfaction
and the highest satisfied member. In addition, we
examine the performance of the model when utilizing the
m-proportionality measure [3] that considers the
recommended list fair when users like at least  items on it.</p>
      <p>Finally, we not only want our model to produce fair
recommendations but also to be able to explain why these
suggestions were made. Since the group
recommendation process is primarily treated as a black box,
explanations can increase users’ trust in the system. In this
work, we demonstrate how to explain the SQUIRREL
recommendations throughout multiple recommendation
user should have at least one item in the group recom- that the model has available. Rewards represent how
sucwith the group. This type of system is known as a se- recommendation method to apply (action) based on how
with the suggestions and no one is discriminated against. quential group recommendations. We facilitate this by
© 2024 Copyright for this paper by its authors. Use permitted under Creative Commons License rounds, using the historical data of the group.</p>
    </sec>
    <sec id="sec-2">
      <title>2. SQUIRREL for Group</title>
    </sec>
    <sec id="sec-3">
      <title>Recommendations</title>
      <p>Let I be a set of data items and U a set of users. G denotes
a group of users where  ⊆  . For each user  in G,   is
the list of recommended items, as a single recommender
system has generated them at a recommendation round
 for  . At round  , SQUIRREL chooses an appropriate
aggregation function based on the current state of the
group, to combine the individual members’
recommendation lists   into one group recommendation list</p>
      <p>SQUIRREL can be described as a Markov decision
pro

.
cess, where an agent interacts with an environment 
to maximize the accumulative reward after each
recommendation round. The Markov decision process can be
described by a tuple of (, A,  , )</p>
      <p>. The goal of the model
is to find a policy  (|)
that takes action  ∈</p>
      <p>during
state  ∈  to maximize the expected discounted
cumulative reward after  recommendation rounds:  [()]</p>
      <p>∑=0    (,  ′), with 0 ≤  ≤ 1 .
where () =
∑
∑
∈
∈ ,

   (,)
   (,)
dation list (, 
(, 

 ) =
relevant the items recommended to the group were
with the most relevant items for that user. That is:

 ) is calculated by comparing how
, where   (, )</p>
      <p>returns the
S is a continuous space that describes the environ- 
∈
(, ℝ) − 
∈
(, ℝ)
. This measure is
mental state, i.e., the group state, that is, how
satisfying the recommended items were for each group
member. A user  ’s satisfaction with the group recommen- functions needed by the F-score.
often referred to as the F-score. We simulate the group
agreement using 1 −</p>
      <p>, considering the input
  (ℝ  ) = 2
 (ℝ
 (ℝ
 ) ∗ (1 −  (ℝ
 ) + (1 −  (ℝ
and
 ))
 ))
(1)
 (
as:  (ℝ) =
ifne the overall group satisfaction of a group  for a
recommendation sequence ℝ</p>
      <p>of  group recommendations
(,
||
∑∈
(,ℝ)
||</p>
      <p>.
isfaction concerning a group recommendation list  
to be the average of the satisfaction scores in the group:</p>
      <p>) =
∑∈

 ) . Subsequently, we
de</p>
      <p>The overall group satisfaction can be used as an
expression of the reward achieved by action  in
recommendation round  , that is:   (ℝ</p>
      <p>) =  (ℝ
where ℝ  refers to all the rounds up to the jℎ one.</p>
      <p>),</p>
      <p>However, this reward may lead us to somehow ignore
the dissatisfaction of a user, since it considers only the
average of the group members’ individual satisfaction
scores. This observation leads us to a new utility score
that considers the user disagreement. We utilize the
harmonic mean of overall group satisfaction  
overall group disagreement</p>
      <p>, which is the
difisfaction among the group members:  (ℝ) =
, ference between the maximum and minimum overall
satpredicted score of item  for user  at recommendation
round  , as it was produced by a single recommender.</p>
      <p>The group state is calculated based on how satisfied
each of its members is in overall, that is, their average</p>
      <p>∑=1 (,


 ) .</p>
      <p>Formally: (, ℝ) =</p>
      <p>In turn, A is a set of distinct actions consisting of
SQUIRREL aggregation functions. We employ 6 diferent
methods as presented in [2], namely: Average, RP80 [4],
Par [5], SDAA [6], SIAA [6] and Avg+ [6].   (,  ′) defines
the probability to transition from state  to state  ′ during
round  under the action  . Formally,   (,  ′) =   (
′
 |</p>
      <p>= ,   = ) . Finally,   (,  ′) is the reward gained
from transitioning from state  to state  ′. The reward</p>
      <p>+1 =
describes the quality of recommendations given by the
model. The model determines if an action is appropriate
based on the reward that it receives by taking the action.
satisfaction up to the current recommendation round. To form a group recommendation list,   , that is fair</p>
      <sec id="sec-3-1">
        <title>2.1. Satisfaction &amp; Disagreement Rewards</title>
        <p>One option to define the reward is via the group
satisfaction score. This score will indicate how well the
system is able to balance the individual needs of the
group members. Specifically, we define the group
 ’s
sat</p>
        <p>This reward function reflects the degree of preference
for an item among the group members as well as the level
of disagreement or agreement between members.
2.2. Fairness Reward


to every member of the group, we consider one fairness
criterion called proportionality [3]. The fairness concept
of fair division of resources inspires this function. The
main idea behind proportionality is to count for each
group member, the items that are present in the group
recommendation list that they prefer. A user likes an
item, if it is in the top Δ% of the user’s individual
recommendation list,   , as it was produced by a single
recommender system. When a recommendation list
contains at least m ( ≥ 1 ) items that a user likes, then we
call it m-proportional. If  = 1 , we call this list
singleproportional, otherwise referred to as multi-proportional.</p>
        <p>The m-proportional approach creates mutual
acceptance of the recommended items among the group
members. When a list contains at least m items that a user
strongly prefers, that user is likely to be more accepting
of other items in the list that they may not like. This
tolerance is based on the understanding that other group
members may prefer those items. Then, m-proportionality
can be defined as follows.
[m-PROPORTIONALITY] Given a group G, and
group recommendation list  
proportionality as:  
(, 


 ) = ||
, We express the
m|  | , where   ⊆ 
represents the set of users within the group for which
  is m-proportional.</p>
        <p>We utilize the m-proportionality fairness measure, to
introduce another reward function, which is the
harmonic mean of the overall group satisfaction and
mproportionality. Formally:
 
(, 

 ) = 2
 (, 
 (, 


 ) ∗  
 ) +  
(, 
(, 

 )

 )
(2)</p>
        <p>This measure combines the overall satisfaction of a
group along with m-proportionality fairness. In this
reward function, the goal is to ensure all members are
recommended some of their preferred items while
maintaining the group’s overall satisfaction.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. Evaluation</title>
      <p>To evaluate our work, we use the 20M MovieLens dataset
and divide it into two parts [2]. The system initiates
with the first part which contains 60% of the movies
to avoid the cold start problem. The latter portion is
divided into 14 chunks and used for the experiments. We
use (4+1) groups that contain 4 similar and 1 dissimilar
user, counting similarity using the Pearson Correlation
function. At each round, we suggest a set of 10 items
to the group, without suggesting those items again. We
generated 100 diferent groups: 80 groups were used for
chose  = 2 and Δ = 20%.
training and 20 for testing. For</p>
      <sec id="sec-4-1">
        <title>3.1. Experimental Results</title>
        <p>and</p>
        <p>we
When evaluating the outcomes for the 4+1 training set
across the four reward functions, we observe in Figure
1 that the   reward function yields the best group
satisfaction scores (Figure 1(a)). On the other hand,  
comes up with the lowest value. This is because the  
reward function aims to maximize overall group
satisfaction and</p>
        <p>tries to recommend at least a fixed
number of preferred items per user. In our current
settings, it is hard for a group of five people to retain two
items from a list of ten items because there is at least one
(a)
(b)
Group Satisfaction and Disagreement for
 ,   ,  
,</p>
        <p>in 4+1 groups.
 
to   . This outcome matches our expectations since</p>
        <p>takes into account both group satisfaction and
m-proportionality in its calculation which results in a
balanced set of recommended items.</p>
        <p>We also calculate the Normalized Discounted
Cumulative Gain (NDCG) value to analyze the quality of the
methods. NDCG shows how many items in the group
recommended list are also present in the user preferences
list (Table 1). The table showcases the average scores
calculated after all 15 rounds of recommendations. From the
table, we can see that the new fairness-based SQUIRREL
models</p>
        <p>and  
to others. This means it can identify more relevant items
for the groups. In Figure 2, we plotted the NDCG scores
have higher values compared
member who stands out from the rest. That means it is dif- for the two new models. There, it becomes evident that
ifcult to satisfy every group member. When satisfaction
in the initial round of recommendations, these methods
is included with m-proportionality in  
better satisfaction. However, note that   and  
result in higher group disagreement compared to the  
and  
reward functions (Figure 1(b)). Also,  
it shows
show a higher performance, which varies in the
subsequent rounds. The reason for this is two-fold. First, due
to the varied sparsity of the dataset as each chunk is
added to the system. Second, after each recommendation
produces slightly lower disagreement scores compared
round, the top 10 items for each group are excluded from
visualizations showing:
1. Group recommendations with a satisfaction score.
2. Group recommendations with disagreement score.
3. Single-user recommendations for all the users of a group.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Related Work</title>
      <p>For producing group recommendations (e.g., [4, 7, 8]), we
employ a standard single-user recommender, apply it to
(a) each individual group member, and aggregate the group
Figure 2: NDCG Values per recommendation round for members lists into one single group recommendation list.
   −   ,   in the 4+1 test scenario. Ipnro[b9l]e, mfa.irFnoersas gisivpernesseenttoefdraasnkaicnognss,ttrhaeinmeedthoopdtimpriozvatidioens
the most similar ranking to the provided sets that satisfy
a specific fairness requirement. For assessing fairness,
consideration, contributing to the performance decline. [5] measures the degree of satisfaction for each group
member with the group recommendation list, based on
4. Explanations for SQUIRREL the relevance of the recommended items for each member.
Diferently, [ 10] uses the position of the items in the
Typically, explanations results in the users becoming group recommendation list, exploiting the concept of
more trusting in the system, which makes the system Pareto optimality. More recently, [3] counts fairness
more persuasive and efective, resulting in a higher level using proportionality: when the user  likes at least 
of satisfaction for the users. For our work, we have de- items in the recommended list, the user considers the
veloped why questions and their respective explanations. list fair for them. In our work, we take inspiration from
We explore genres of movies and how many times they various fairness definitions and apply a comparable one
have been recommended to the group as a whole as well using the users’ satisfaction and disagreement scores.
as to individual users of the group. We format our ques- Additionally, we use directly the proportionality fairness
tion as follows: Why ‘selected genre’ occur ‘selected fre- measures as a reward function for our model.
quency’? In recent years, Reinforcement Learning (RL) has
be</p>
      <p>There are 19 genres and 3 frequencies to choose from: come increasingly popular in recommendation systems
selected genre = [ ‘Action’, ‘Adventure’, ‘Animation’, ‘Chil- [11]. For example, [12] proposes an online
personaldren’, ‘Comedy’, ‘Crime’, ‘Documentary’, ‘Drama’, ‘Fan- ized news recommendation framework based on DQNs,
tasy’, ‘Film-Noir’, ‘Horror’, ‘Musical’, ‘Mystery’, ‘Romance’, while [13] optimizes recommendation models for
long‘Sci-Fi’, ‘Thriller’, ‘War’, ‘Western’, ‘IMAX’ ], term accuracy using RL techniques. [14] incorporates
selected frequency = [ ‘a few times’, ‘many times’, ‘not at randomness for fairness throughout, using Variational
all’ ]. Autoencoders (VAEs), and penalizes items based on their</p>
      <p>Each group will have the freedom of choice about historical popularity to promote diversity and minimize
which genre to inquire about. For example, if a group bias [15, 16, 17]. In our work, we aim to be more
verwants to ask a question related to the ‘Musical’ genre and satile regarding the domains in which our solution can
why it appears ‘a few times’, the question becomes: Why be utilized and can incorporate a variety of strategies to
‘Musical’ occurs ‘a few times’?. compensate for the limitations inherent to every
recom</p>
      <p>We explain the ‘Why’ question in terms of a general mendation method.
explanation that is based on single-user recommendation
lists for the users of the group, along with a model-based 6. Summary
explanation that incorporates the summarized version
of diferent aggregation methods. An example of such In this work, we present a new reward function for
an explanation is the following: The genre Musical is less the SQUIRREL model, aiming to increase the fairness
likely to be enjoyed by 4 members of this group, therefore of the sequential group recommendations. Using  −
it occurs less frequently.    produces better results in terms of group
This explanation uses the SDAA action that balances the members’ satisfaction and disagreement. In addition, we
average predicted score of an item for the group with the showcase how to provide explanations for
recommendapredicted score of the least satisfied member . tions produced by SQUIRREL.</p>
      <p>Finally, to facilitate easy comprehension of
recommendations by the users of the group, we have produced
mitigating user unfairness in recommendation
systems, in: SAC, 2022.
[1] R. D. Burke, M. Ramezani, Matching recommen- [17] R. Borges, K. Stefanidis, On measuring popularity
dation technologies and domains, in: F. Ricci, bias in collaborative filtering data, in: EDBT/ICDT
L. Rokach, B. Shapira, P. B. Kantor (Eds.), Rec- Workshops, 2020.
ommender Systems Handbook, Springer, 2011, pp.</p>
      <p>367–386.
[2] M. Stratigi, E. Pitoura, K. Stefanidis, Squirrel: A
framework for sequential group recommendations
through reinforcement learning, Inf. Syst. 112
(2023). URL: https://doi.org/10.1016/j.is.2022.102128.</p>
      <p>doi:10.1016/j.is.2022.102128.
[3] D. Serbos, S. Qi, N. Mamoulis, E. Pitoura,</p>
      <p>P. Tsaparas, Fairness in package-to-group
recommendations, in: WWW, 2017.
[4] S. Amer-Yahia, S. B. Roy, A. Chawlat, G. Das, C. Yu,</p>
      <p>Group recommendation: Semantics and eficiency,</p>
      <p>PVLDB 2 (2009) 754–765.
[5] L. Xiao, Z. Min, Z. Yongfeng, G. Zhaoquan, L. Yiqun,</p>
      <p>M. Shaoping, Fairness-aware group
recommendation with pareto-eficiency, in: RecSys, 2017.
[6] M. Stratigi, E. Pitoura, J. Nummenmaa, K.
Stefanidis, Sequential group recommendations based on
satisfaction and disagreement scores, Journal of</p>
      <p>Intelligent Information Systems (2021).
[7] L. Baltrunas, T. Makcinskas, F. Ricci, Group
recommendations with rank aggregation and
collaborative filtering, in: RecSys, 2010.
[8] E. Ntoutsi, K. Stefanidis, K. Nørvåg, H.-P. Kriegel,</p>
      <p>Fast group recommendations by applying user
clustering, in: International conference on conceptual
modeling, Springer, 2012, pp. 126–140.
[9] C. Kuhlman, E. Rundensteiner, Rank aggregation
algorithms for fair consensus, Proceedings of the</p>
      <p>VLDB Endowment 13 (2020).
[10] D. Sacharidis, Top-n group recommendations with</p>
      <p>fairness, in: SAC, 2019.
[11] M. M. Afsar, T. Crump, B. Far, Reinforcement
learning based recommender systems: A survey, 2021.</p>
      <p>arXiv:2101.06286.
[12] G. Zheng, F. Zhang, Z. Zheng, Y. Xiang, N. J. Yuan,</p>
      <p>X. Xie, Z. Li, Drn: A deep reinforcement learning
framework for news recommendation, in: WWW,
2018.
[13] L. Huang, M. Fu, F. Li, H. Qu, Y. Liu, W. Chen,</p>
      <p>A deep reinforcement learning based long-term
recommender system, Knowledge-Based Systems
(2021).
[14] R. Borges, K. Stefanidis, Enhancing long term
fairness in recommendations with variational
autoencoders, in: MEDES, 2019.
[15] R. Borges, K. Stefanidis, On mitigating popularity
bias in recommendations via variational
autoencoders, in: SAC, 2021.
[16] R. Borges, K. Stefanidis, F2VAE: a framework for</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>