<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Data Sparsity via Neuro-Symbolic Knowledge Transfer</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Tommaso Carraro</string-name>
          <email>tcarraro@fbk.eu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessandro Daniele</string-name>
          <email>daniele@fbk.eu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabio Aiolli</string-name>
          <email>aiolli@math.unipd.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luciano Serafini</string-name>
          <email>serafini@fbk.eu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Data and Knowledge Management, Fondazione Bruno Kessler</institution>
          ,
          <addr-line>Via Sommarive, 18, 38123 Povo</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Mathematics, University of Padova</institution>
          ,
          <addr-line>Via Trieste, 63, 35131 Padova</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>ings via Collaborative Filtering. Then</institution>
          ,
          <addr-line>this knowledge is</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>tial for e-services (e.g.</institution>
          ,
          <addr-line>Amazon, Netflix, Spotify). Given</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Data sparsity is a well-known historical limitation of recommender systems that still impacts the performance of state-of-theart approaches. One practical technique to mitigate this issue involves transferring information from other domains or tasks to compensate for scarcity in the target domain, where the recommendations must be performed. Following this idea, we propose a novel approach based on Neuro-Symbolic computing designed for the knowledge transfer task in recommender systems. In particular, we use a Logic Tensor Network (LTN) to train a vanilla Matrix Factorization (MF) model for rating prediction. We show how the LTN can be used to regularize the MF model using axiomatic knowledge that permits injecting pre-trained information learned by Collaborative Filtering on a diferent task or domain. Extensive experiments comparing our model with a baseline MF on two versions of a novel real-world dataset prove our proposal's potential in the knowledge transfer task. In particular, our model consistently outperforms the MF, suggesting that the knowledge is efectively transferred to the target domain via logical reasoning. Moreover, an experiment that drastically decreases the density of user-item ratings shows that the benefits of the acquired knowledge increase with the sparsity of the dataset, showing the importance of exploiting knowledge from a denser source of information when training data is scarce in the target domain. knowledge transfer, matrix factorization, neuro-symbolic integration, logic tensor networks, rating prediction, explicit Recommender systems (RSs) have recently become essen- the context of recommender systems, knowledge transthe user's historical data, these tools mitigate informa- categories: feature-based models and fine-tuning ∗Corresponding author.</p>
      </abstract>
      <kwd-group>
        <kwd>feedback</kwd>
        <kwd>data sparsity</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        longs to the research field of transfer learning [ 14]. In
fer techniques [
        <xref ref-type="bibr" rid="ref2">15, 16, 17, 18</xref>
        ] can be subdivided into two
tion overload by suggesting novel items (e.g., products,
movies, songs) that match the user’s preferences [1].
      </p>
      <sec id="sec-1-1">
        <title>Since the beginning of the RSs literature, Collaborative</title>
      </sec>
      <sec id="sec-1-2">
        <title>Filtering (CF) [2, 3, 4, 5] has been one of the most success</title>
        <p>
          ful recommendation approaches. Latent Factor models,
in particular Matrix Factorization (MF), have dominated
the CF scene [6, 7, 8] for years, and this has been further
emphasized with the deep learning rise [
          <xref ref-type="bibr" rid="ref11 ref5 ref7">9, 10, 11, 12</xref>
          ].
        </p>
      </sec>
      <sec id="sec-1-3">
        <title>Despite their success in improving recommendation</title>
        <p>performance, state-of-the-art models still sufer from a
plicability in real-world scenarios. One way to address
data sparsity consists of leveraging models pre-trained
on other sources of information (i.e., source domains)
to make the final model more accurate in the target
do</p>
        <p>Despite Neuro-Symbolic (NeSy) [19] approaches
having been successfully applied in many AI fields [ 20, 21,
22], including RSs [23, 24, 25, 26], they have not yet been
investigated in the task of knowledge transfer for
recommender systems, where we believe their application
is particularly suited. In particular, NeSy aims at
integrating knowledge, usually expressed using logical
axhistorical issue, i.e., data sparsity, that limits their ap- recommendation task, namely the target domain. The
main, where the recommendations must be performed. transferred to the target domain via logical reasoning.
to be particularly beneficial in tasks with poor training 31, 32, 33]. Among them, some directly use logical
fordata [27], giving insights that this paradigm can help deal mulas to learn a recommendation model [23, 26, 31]. The
with data sparsity in recommendation datasets. seminal approach that applied NeSy to RSs has been</p>
        <p>Following this intuition, we propose using a Logic Ten- Neural Collaborative Reasoning (NCR) [26]. In NCR, the
sor Network (LTN) [28] to encode axiomatic knowledge sequential recommendation task is formalized as a logical
to enable efective knowledge transfer and injection for reasoning problem. In particular, the user’s ratings are
a vanilla MF model1 trained on movie ratings. LTN is a represented using propositional variables, and logical
opNeSy framework that efectively integrates logical rea- erators (e.g., ∧, ⟹ ) are used to construct propositional
soning and neural networks. Our approach uses it as the formulas that express sequential patterns between them.
interface between the model pre-trained on the source Then, NCR maps the variables to logical embeddings and
domain and the final model trained on the target domain. the operators to neural networks (NNs) that act on those
We called it an interface as it allows the explicit transfer embeddings. The NNs are forced to behave as classical
of information between domains via logical reasoning. logic operators through logical regularization. By doing
To perform our experiments, we use MindReader [30], so, the formulas can be organized as a neural network to
a novel recommendation dataset containing explicit rat- conduct logical reasoning and prediction in a continuous
ings from real users on movies and non-recommendable space. They compared NCR with many linear and deep
entities, such as movie genres, actors, and producers. In baselines, showing it can reach state-of-the-art
perforparticular, we use ratings on movie genres to learn our mance. However, this approach is not properly NeSy as
pre-trained model via Collaborative Filtering and ratings neural networks implement the symbolic part (i.e.,
logion movies to train the final model to provide accurate cal connectives). LTN uses fuzzy logic semantics instead,
recommendations in the target domain. The pipeline of making the framework theoretically sound. Moreover,
the proposed approach consists of two steps. In the first NCR uses propositional logic, which makes it impossible
step, we train a genre classifier using an MF model to to encode complex and expressive knowledge due to the
learn which genres the users like and dislike in the source simplicity of the language syntax.
domain. In the second step, we use LTN to transfer the In Graph Collaborative Reasoning (GCR) [31], NCR
pre-trained knowledge to an MF model trained on the is extended to work with knowledge graphs. In
partictarget domain for the movie rating prediction task. ular, they provide a simple approach for translating the</p>
        <p>We compare our approach with a baseline MF model to graph structure into logical expressions to convert the
understand if knowledge transfer is successfully2 reached link prediction task into a logical reasoning problem. As
thanks to Neuro-Symbolic integration. The results show in NCR, they use logically constrained neural modules
that our model consistently outperforms the MF, prov- to build the network architecture according to the
logiing its ability to transfer knowledge across domains. In cal expression. They conducted experiments similar to
addition, an experiment that drastically reduces the den- NCR, showing that GCR can improve NCR on knowledge
sity of user-item ratings shows that the benefits of the graphs.
knowledge increase with the sparsity of the dataset. This In Counterfactual Collaborative Reasoning (CCR) [32],
gives insight that our model can successfully deal with NCR is used to perform data augmentation based on
logdata sparsity thanks to Neuro-Symbolic reasoning. To ical reasoning for sequential recommendation.
Specifithe best of our knowledge, this is the first time a NeSy cally, counterfactual logic reasoning is exploited to
generapproach has been successfully applied to the knowledge ate counterfactual examples for data augmentation. The
transfer task for recommender systems. examples are generated by discovering slight changes in
users’ explicit feedback (i.e., the sequence of purchases)
by solving a counterfactual optimization problem. They
2. Related works showed that these new examples, together with the
original examples, can alleviate scarcity and enhance the
In the last two years, the RS community has seen the performance of sequential recommendations.
emergence of some Neuro-Symbolic approaches [25, 24, Another approach that successfully integrated logical
reasoning and learning has been HYbrid Probabilistic
1Note we selected Matrix Factorization for our experiments because, Extensible Recommender (HyPER) [25], which is based
tdhees-pairtet aitpspsriomapchliecsity[2,9it].isTshtiilsl oisnneootfttohebemionsttenpdoewderafsualsltiamteit-ooff- on Probabilistic Soft Logic [ 34]. In particular, HyPER
our approach, as any other state-of-the-art model could be used in exploits the expressiveness of First-Order Logic (FOL)
principle. to encode knowledge from a wide range of information
2Note the objective of the experiment is not to obtain state-of-the-art sources, such as multiple user and item similarity
meaperformance. Instead, our goal is to show that a NeSy approach sures, content, and social information. Then, Hinge-Loss
tchains benedu,stehde foonrlykndoifewrelendcgeebetrtawnesefnert hine rmeocodmelsminentdheercsoymstpeamriss.onTo Markov Random Fields are used to learn how to balance
is the addition of knowledge via LTN. the diferent information types. HyPER is highly related
considered more theoretically sound compared to previ- is defined as a set of  − couples  − = {(, ) () |(, ,  )
() ∈
to our work with LTN since the logical formulas we use
resemble the ones used in HyPER.</p>
      </sec>
      <sec id="sec-1-4">
        <title>In [24], they propose a NeSy approach to encode FOL</title>
        <p>formulas to enhance knowledge graph embeddings [35]
and provide accurate knowledge-aware
recommendations. Their approach consists of three steps: () the
FOL formulas are automatically extracted from a
recommendation knowledge graph, then () the knowledge
graph embeddings are learned jointly with the extracted
formulas using a NeSy approach. Finally, ()
the
useritem embeddings are fed to a neural architecture to get
predictions. Specifically, in the second step, they used
KALE [36], a NeSy approach that allows learning
knowledge graph embeddings jointly with FOL formulas. In
particular, the learning can be unified as the graph triples
can be interpreted as FOL atoms. As KALE uses fuzzy
semantics to learn graph embeddings, this approach can be
ous deep approaches, as the symbolic component remains
symbolic during learning, as it happens with LTN and</p>
      </sec>
      <sec id="sec-1-5">
        <title>HyPER.</title>
      </sec>
      <sec id="sec-1-6">
        <title>One [23] of the last published approaches is highly</title>
        <p>related to ours. Specifically, they try to mitigate data
sparsity by using LTN to inject content information into
an MF model. In particular, they encode FOL formulas
to use side information as a regularizer for the latent
factors of the MF model. They show that the proposed
NeSy approach can outperform the MF. However, the
improvement is poor, and the model has some
scalability issues due to the number of times the formulas have
to be evaluated during training. Our approach difers
from [23] on how axiomatic knowledge is used. In [23],
the knowledge extends the MF model using content
information, while our model uses it to enable the transfer of
Bold notation diferentiates vectors, e.g., x = [3.2, 2.1],
and scalars, e.g.,  = 5 . Matrices and tensors are denoted
with upper case bold notation, e.g., X. X is used to denote
the  -th row of X, while X, to denote the item at row 
and column  . We refer to the set of users of an RS with
 , where | | =</p>
        <p>. Similarly, the set of items is denoted
as ℐ, such that |ℐ | =  . We use  to denote a dataset. 
is defined as a set of  triples  = {(, ,  )
R ∈ ℕ× , such that R, =  if (, ,  ) ∈ 
 ∈</p>
        <p>,  ∈ ℐ , and  ∈ {0, 1} is a binary explicit rating.
 can be reorganized in the so-called user-item matrix
() }</p>
        <p>=1 , where
, 0 otherwise.</p>
        <p>Then, since we work with binary feedback, we refer to
 + (resp.  −) as the dataset of positive (resp. negative)
user-item pairs.  + is defined as a set of  + couples
 + = {(, )
() |(, ,  )
() ∈  ,</p>
        <p>() = 1}=1 . Similarly,  −
 ,</p>
        <p>() = 0}=1 . Finally,  ? denotes the dataset of
useritem pairs for which the rating is unknown, i.e.,  ? =
{(, ) ∈  × ℐ |(, ,  ) ∉  }
. Clearly,  ? =  ⋅  − 
.</p>
        <sec id="sec-1-6-1">
          <title>3.2. Matrix Factorization</title>
          <p>Matrix Factorization (MF) is a Latent Factor Model that
aims at factorizing the user-item matrix R into the
product of two lower-dimensional rectangular matrices, U ∈
ℝ× and I ∈ ℝ× , containing the users’ and items’ latent
factors, respectively.  represents the number of latent
factors. More formally, the objective of MF is to find U
and I such that R ≈ U ⋅ I⊤. An efective way to learn the
latent factors is by using gradient-descent optimization.</p>
          <p>Given the dataset  , an MF model seeks to minimize
the Mean Squared Error (MSE) between predicted and

1
∑(,,)∈</p>
          <p>|| −̃  || 2 + ||  ||2. In
U ⋅ I⊤ + u
 + i , where u and i
pre-trained knowledge learned via Collaborative Filter- target ratings, defined as
ing on another task (i.e., movie genre rating prediction). the formulation,  ̃ =
Moreover, our model provides a logical formalization for
are bias terms for user  and item  , respectively, and
the rating prediction task3, while [23] proposes a ranking-  = {U, I, u, i}.  is a hyper-parameter to set the strength
based method. Finally, note all the presented models use
some form of knowledge (e.g., knowledge graphs, logic)
to improve the recommendation performance. However,
none of them have designed experiments to understand
if the advantages of a NeSy system can help mitigate one
or some of the historical issues of recommender systems
(e.g., data sparsity, cold-start, explainability). Hence, it
is unclear if the properties (e.g., few-shot learning,
interpretability) of a NeSy approach are totally exploited.
of the 2</p>
          <p>regularization.</p>
          <p>In our setting, we use a diferent implementation of MF
since we treat the recommendation problem as a binary
classification task. Specifically, we need to recommend if
a user likes (1) or dislikes (0) an item. Hence, the focal
loss is used in place of MSE for the training, and the
logistic function is applied to the prediction of MF to
restrict the output between 0 and 1. Focal loss is defined
as</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>3. Background</title>
      <p>This section provides useful notation and terminology
used in the remainder of the paper.
3Note our approach can be easily extended to the top-n
recommendation task by changing the formalization of the knowledge base.
1
 (,,)∈
∑</p>
      <p>−  (1 −   ) log   + ||  ||2
  = {
  = {


1 − 
main of real numbers. We refer to the term grounding4, tor, generally defined using
allows mapping every symbolic expression into the do- SatAgg ∶ [0, 1]∗ ↦ [0, 1] is a formula aggregating
operaME .
items of the dataset), with   ∈ ℕ+,   &gt; 0. As a conse- ternal source of information (e.g., content information,
quence, a term ()</p>
      <p>or a formula P() , will be mapped
to a sequence of   values too. Afterward, connectives
additional ratings), a pre-trained model  
learned to generate features5 for users and items in the
(, |  1) is
are grounded using fuzzy semantics (i.e., operators deal- source domain. In the notation,  is a user index,  is an
for gradient-based optimization [37]. In the notation, source domain via Matrix Factorization; hence, the
outclassify, and  =  (
logistic function.</p>
      <p>U ⋅ I⊤ + u + i ), where  is the</p>
      <sec id="sec-2-1">
        <title>3.3. Logic Tensor Networks</title>
        <sec id="sec-2-1-1">
          <title>Logic Tensor Networks (LTN) [28] is a Neuro-Symbolic framework that allows using a knowledge base composed</title>
        </sec>
        <sec id="sec-2-1-2">
          <title>LTN uses a specific first-order language, called Real Logic,</title>
          <p>that is fully diferentiable and has concrete semantics that
formally denoted by  , as the function that defines this
mapping. Real Logic allows LTN to ground logical
formulas into computational graphs, enabling gradient-based
optimization.</p>
          <p>In particular,  maps individuals (e.g., users) to tensors
of real features (e.g., users’ demographic information),
functions (e.g., Score( , )
) as real functions (e.g.,
inner product), and predicates (e.g., Likes( , )
real functions with output in [0, 1]. Then, a variable
 is mapped to a sequence of   individuals (e.g., some
) as</p>
          <p>Given a Real Logic knowledge base  = {
1, … ,   },
where  1, … ,   are closed formulas, LTN allows learning
the grounding of constants, functions, and predicates
appearing in them. In particular, if constants are grounded
as embeddings and functions/predicates onto neural
networks, their grounding  depends on some learnable
parameters  . We denote a parametric grounding as  (⋅|  ).
by finding parameters 
of  , namely 
∗ = argmax SatAgg∈
∗ that maximize the satisfaction
 (|  ), where</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>3.4. Neuro-Symbolic knowledge transfer</title>
        <p>Our model is inspired by feature-based knowledge
transfer models [13]. In these approaches, a pre-trained model
is learned on some additional source of information,
usually side information. Then, the acquired knowledge
is transferred to the target domain during the training
of the final model to help deal with the data sparsity
of user-item interactions. More formally, given the
exwhere  is a hyper-parameter to give diferent weights
to the two classes,  is a hyper-parameter that represents
the penalty assigned to the examples that are hard to
individual of each variable.
quantifying over specific tuples of individuals of
variables  1, … ,   , such that the  -th tuple contains the  -th
item index, and  1 are the parameters of the pre-trained
model. Instead of generating features, our approach
learns</p>
        <p>to predict user-genre preferences in the
put of this model is a score for the given user-genre pair
in the source domain. Then, we assume  
a model that learns to predict user-movie preferences in
the target domain. In the notation,  is a user index, 
is an item index, and  2 are the parameters of the final
(, |  2) is
recommendation model. Specifically,  
score for the given user-movie pair in the target domain.
outputs a
In our approach,  
knowledge from  
is to maximize function
Neuro-Symbolic computing. The objective of our model
is learned while transferring6
via logical reasoning thanks to
ℱ ( 1,  2) = ( 
(, |  1),  
(, |  2), (, ))
where  is a logic-based aggregation function,  is a user
Note</p>
        <p>is a function that relates movies and genres</p>
        <sec id="sec-2-2-1">
          <title>5In feature-based models, the pre-trained model is usually used</title>
          <p>to obtain features for users and items from the content or side
information. In our scenario, it learns to predict user-item ratings
ing with fuzzy values), while quantifiers are grounded as
special aggregation functions (e.g., generalized means).</p>
        </sec>
        <sec id="sec-2-2-2">
          <title>This paper uses the product configuration , best suited</title>
          <p>, ,  1, … ,   ∈ [0, 1] and  ≥ 1 .</p>
          <p>(∧) =</p>
          <p>T</p>
          <p>( , ) =  ∗ 
 ( ⟹ ) = I ( , ) = 1 −  +  ∗ 
 (¬) =</p>
          <p>N ( ) = 1 − 
 (∀) =</p>
          <p>ME ( 1, … ,   ) = 1 − (
 (∃) =</p>
          <p>M ( 1, … ,   ) = (


1
 =1
1</p>
          <p>∑ 
 =1
∑(1 −   ) )</p>
          <p>1
  ) 1 )
The intuition behind the choice of hyper-parameter 
is that the higher that  is, the more weight that M</p>
          <p>converging to the max (resp. min) operator. Real Logic
also provides a special type of quantification, called</p>
          <p>diagonal quantification , denoted as Diag( 1, … ,   ). It allows
(resp. ME ) will give to true (resp. false) truth-values, index,  is a movie genre index, and  is a movie index.
4Notice that this is diferent from the common use of the term</p>
          <p>grounding in logic, which indicates the process of replacing the variables of
a term or formula with constants or terms containing no variables.
via Collaborative Filtering.
6Note that parameters  1 (of  
parameters  2 (of  
).</p>
          <p>) are frozen during the training of
using content information. It can be seen as a lookup 4.2. Grounding of symbols
table denoting which movies belong to a specific genre
in the dataset. In our approach, this function is used by  The grounding  defines how logical symbols are mapped
iItnno crpeoagmrutbilcaiunrilazateri,ottnhhewetiartaghigntrhieneggapotriefodnicftuionncvstiioaonfl o gciacanl braeenadsseoenninags. okLnnTtoNowmtlheodedgreeel.balaIsfienelddtheaifinnsedwthhoeernkcc,oe m(hpouwtat∗tih)oen=aalx⟨gi or(am)p|sh(, io)nf (tt)hhee∈
the logical knowledge base that LTN seeks to maximally  ∗⟩=∗1 and  (  ∗) = ⟨ () |(, ) () ∈  ∗⟩=∗1 , namely
satisfy during training. Specifically,  is a composition of  ∗ and   ∗ are grounded as a sequence of the  ∗
logical formulas that defines how the source and target user and movie indexes in  ∗, with ∗ ∈ {+, −, ?}.
Indomain interact during training. In other words,  makes stead,  ( ) = ⟨1, … ,   ⟩, namely   is grounded
it possible to transfer knowledge between domains via as a sequence of   genre indexes, where   is the
logical reasoning and is carefully formalized in Section 4. number of movie genres in the dataset. Afterward,
 ( Likes |U, I, u, i) ∶ ,  ↦  ( U ⋅ I⊤ + u + i ), namely
Likes is grounded onto a function that takes as input a
4. Method user index  and a movie index  and returns the
prediction in [0, 1] of the MF7 model for the given
useritem pair. U ∈ ℝ× , I ∈ ℝ× , u ∈ ℝ , and i ∈ ℝ are
the matrices of the users’ and items’ latent factors, and
the vectors of users’ and items’ biases, respectively.
Notice Likes is the predicate that implements   . Then,
Our approach uses an LTN to enable domain adaptation
for efective knowledge transfer. Specifically, the LTN
is trained using a Real Logic knowledge base containing
facts designed to intuitively transfer information about
movie genre preferences (i.e., the source domain) to a Ma-  ( LikesGenre) ∶ ,  ↦ G, , where G ∈ {0, 1}×  ,
trix Factorization model trained on movie ratings (i.e., the namely LikesGenre is grounded onto a function that
target domain). In the next subsections, we will present takes as input a user index  and a genre index  and
our knowledge base, how  is used to convert it into returns the prediction contained in matrix G for user 
a computational graph suitable for gradient-based opti- and genre  . In particular, G can be seen as a lookup table
mization, and how the learning of the LTN takes place. containing the binarized8 predictions of a pre-trained
genre classifier. LTN has shown to work better with
4.1. Real Logic knowledge base binarized outputs as the classifier was returning
preThe objective of our LTN model is the satisfaction of the dictions too near the decision boundary for LTN to
unfollowing Real Logic knowledge base. derstand9 the diference between like and dislike. Note
LikesGenre is the predicate that implements   .
Finally,  (  ) ∶ ,  ↦ {0, 1} , namely HasGenre is
∀ Diag( +,   +) Likes( +,   +) (2) grounded onto a function that takes as input a movie
index  and a genre index  and returns one if the movie
∀ Diag( −,   −)¬ Likes( −,   −) (3)  belongs to genre  , zero otherwise. Note HasGenre
∀ Diag( ?,   ?)(∃ ¬ LikesGenre( ?,  ) is the predicate that implements  . Intuitively,
∧ HasGenre(  ?,  )) ⟹ ¬ Likes( ?,   ?()4) t(heLsikoeusrGceendroem)caoinn.taIinnscothnetrkasnto, w(
leLdigkeesp|rUe,-tIr,aui,nie)dreopnresents the MF model we need to train on the target</p>
          <p>Specifically,  + and   + are variable symbols de- domain.
noting positive user-item pairs,  − and   − are vari- Intuitively, Axiom (2) forces Likes to be true for each
able symbols denoting negative user-item pairs,  ? and positive user-item pair in  +, while Axiom (3) forces
  ? are variable symbols denoting user-item pairs for Likes to be false for each negative user-item pair in  −.
which the rating is unknown, and   is a variable sym- In other words, by maximizing the satisfaction of
Axbol denoting the genres of the movies. Then, Likes(, ) iom (2) and Axiom (3), the model learns to factorize the
is a predicate symbol denoting whether a user  likes a user-item matrix using the ground truth. In contrast,
movie  , LikesGenre(, ) is a predicate symbol denoting Axiom (4) is designed to transfer knowledge from the
whether a user  likes a movie genre  , and HasGenre(, )
is a predicate symbol denoting whether a movie  belongs 7iNteomticpeatihr.atInLitkheiss cwaonrbke,
waneyufsuenacntioMnFremtuordneiln.gTahsiscohraesfonrotatuosebreto genre  . intended as a limit of our approach as any other state-of-the-art</p>
          <p>Intuitively, Axiom (2), Axiom (3), and Axiom (4) are ap- model could be used in principle.
plied to user-item pairs in  +,  −, and  ?, respectively. 8A binarized prediction is obtained by using the decision boundary
Diag is used to quantify over the desired user-item pairs 9Foofrthaebcilnaassriyfiercloansstihfieer, o0u.4t5puatnodf 0th.5e5maroedpelretodigcetitovnaslubeesloinngin{0g, 1to}.
rather than quantifying over all possible combinations diferent classes. For a logical framework, those values represent
of user and item indexes in the dataset. similar truth values.
source domain to the target domain through logical rea- 5.1. Datasets
soning. Specifically, it forces Likes to be false whenever
a user  does not10 like at least one genre  of a movie
 . Note this axiom is applied only to unknown user-item
pairs in  ?. In fact, when no movie ratings are available
on the target domain, knowing something about movie
genre preferences is better than knowing nothing. In
other words, we believe transferring knowledge from
the source domain is crucial when data is missing in the
target domain.</p>
          <p>To perform our experiments, we selected MindReader13
(MR) [30], a novel dataset containing ratings from real
users for movies and non-recommendable entities, such
as movie genres, actors, and producers. We performed
experiments on both MR-100k and MR-200k, the two
available versions of the dataset. We used the ratings on movie
genres as the source domain14 and the movie ratings as
the target domain. To guarantee the users in the source
and target domains totally overlapped, we removed the
users that only rated movie genres or movies. After this
4.3. Learning of the LTN pre-processing, MR-100k (resp. MR-200k) comprised 962
The objective of our model is to learn  ( Likes |U, I, u, i) (resp. 2,182) users, 3,034 (resp. 3,806) movies, and 140
by maximizing the satisfaction of the knowledge base. In (resp. 159) movie genres. The density of user-movie
other words, LTN seeks to minimize the following loss ratings was 0.62% (resp. 0.58%), while for user-genre
ratfunction: ings was 8.09% (resp. 6.37%). Selecting ratings on movie
genres as the source domain allowed us to use
particularly dense information15 for knowledge transfer. When
L( ) = (1 − SatAgg∈  ( +, +)←ℬ+(|  )) + ||  ||2  ( ) ≫  ( ) , knowledge transfer is
( −, −)←ℬ− more likely to be efective [ 39].</p>
          <p>( ?, ?)←ℬ? (5) MR provides three types of ratings: likes (1), unknown
where ℬ∗ denotes a batch of training examples randomly (0), and dislikes (-1). As in [30], we removed the unknown
sdaemnoptleesdthfraot mva ria∗b.leTs he n∗oatnatdi o n ( ∗ a∗r,e  groun∗)de←d wℬith∗ irnagtisngfrso.mAfte-1r tthoa0t,. wAes cthheandgaetadstehteplraobveildfeosrbnineagraytiveexprlaitc-it
actual user-movie pairs coming from the corresponding feedback, we treated the recommendation problem as
batch ℬ∗, where ∗ ∈ {+, −, ?}. Notice the loss does not a binary classification task 16, where one has to predict
specify how the variable   is grounded. At each whether a user likes or dislikes an item. This choice
altraining step, we ground it with the sequence of all the lowed us to use the focal loss (Equation (1)) to train the
movie genre indexes in the dataset. Note ℬ? is created MF models and the F-measure as an evaluation metric.
by uniformly sampling user-item pairs from  ? at each This helped in dealing with class imbalance. The class
imtraining step. While all the user-item pairs in  + and  − balance ratio in MR-100k (resp. MR-200k) is 21%(-)/79%(+)
are iterated at each epoch, going through all the possible (resp. 20%(-)/80%(+)) for movie genres, and 38%(-)/62%(+)
unknown pairs is unnecessary and would be unfeasible. (resp. 36%(-)/64%(+)) for movies. In both cases, the
negaIn this sense, ℬ? has not to be considered a mini-batch tive class is the minority one. Hence, we used it as the
in the usual sense. positive one to compute evaluation metrics in Table 2.</p>
          <p>As the splitting strategy for the target domain, we
randomly sampled 20% of the movie ratings from each user
5. Experiments to construct the test set. Then, we randomly sampled 10%
of the remaining movie ratings from each user to
construct the validation set. Instead, for the source domain,
we only created the validation set by randomly sampling
20% of the movie genre ratings from each user. The test</p>
        </sec>
        <sec id="sec-2-2-3">
          <title>This section presents the experiments we performed with</title>
          <p>our method. They have been executed on an Apple
MacBook Pro (2019) with a 2,6 GHz 6-Core Intel Core i7. The
models have been implemented in Python using PyTorch.</p>
          <p>In particular, we used the LTNtorch11 library [38].
Moreover, we used Weights and Biases (WandB) for
hyperparameter optimization. Our source code is freely
available12.
13https://mindreader.tech/dataset/
14Notice that our approach is flexible on the type of knowledge that
has to be transferred. In this work, we use movie genre ratings,
but every type of rating (e.g., ratings on books, actors) or side
information can be used in principle. One has just to change the
10Note the negated formula is used on purpose, as it is likely that knowledge base formalization accordingly.
a user dislikes the majority of movies belonging to a genre she 15Notice ratings on movie genres are dense as they are easily
obtaindislikes. For example, if  does not like horror, likely, she will not able. It is more likely a user will provide a rating about some genre
like all the horror movies. Moreover, it is more critical to avoid over hundreds rather than some movie (or actor) over thousands.
recommending something users do not like than not recommend- 16MindReader provides binary explicit ratings rather than usual
ing something they like. This will restrict the recommendation to 1-5 star ratings. For this reason, the recommendation task can
a few positive movies. be interpreted as a binary classification problem rather than a
11https://github.com/logictensornetworks/LTNtorch regression one. Finally, we work in a rating prediction task rather
12https://github.com/tommasocarraro/NESYKnowledgeTransfer than ranking as the feedback is clearly explicit.
set is not needed in the source domain, as we only need
to computational time, the searches have been conducted
a validation set to find the optimal hyper-parameters for
only for the first seed of the experiment and just for the
the pre-trained model.</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>5.2. Experimental setting</title>
        <p>complete dataset (i.e., 100% ratings). The best
hyperparameters found for the models have been then used
in the rest of the experiment. For all the models, we
tried a number of latent factors  ∈ {5, 10, 25, 50} ,
regularization coeficient  ∈ {0.01, 0.001, 0.0001, 0.00005} ,
learnOur experiment compares the proposed Neuro-Symbolic
approach, denoted as NESYMF, with a baseline MF model, ing rate  ∈ {0.01, 0.001, 0.0001} , training batch size  ∈
denoted as MF, to check if NESYMF can efectively
trans{64, 128, 256}. For the MF models, we additionally tried
fer knowledge from source to target domain and im- focal loss hyper-parameters  ∈ {0.05, 0.1, 0.2, 0.3, 0.4, 0.5}
prove the performance when training data becomes
scarce. Specifically, the experiment consists of the
following pipeline: (1) additional training sets are
generated by randomly sampling the 50%, 20%, 10%, and 5%
and  ∈ {0, 1, 2, 3} . For the MF model trained on the
source domain, we also tried diferent thresholds for the
decision boundary  ∈ {0.3, 0.4, 0.5, 0.6, 0.7} . Due to the
huge class imbalance in the source domain, we preferred
of the movie ratings17 from the entire training set, re- finding a threshold to maximize
ferred to as 100%. Notice ratings are sampled indepen- recall19. Finally, for NESYMF, we additionally tried
difprecision rather than
dently from the user, diferently from the splitting
strategy explained previously. Then, (2) for each training
set   ∈ {100%, 50%, 20%, 10%, 5%}
and for each model
ferent values for hyper-parameter  ∈ {2, 4, 6, 8, 10} of
ME and M . Table 1 presents the best hyper-parameters</p>
        <p>found for the MF model trained on the source domain
 ∈ { MF, NESYMF}: (2) 
is trained on</p>
        <p>using hyper- (pre-trained model) and the models trained on the target
parameters found through a bayesian search. Finally, (2)
 is evaluated on the test set. Note that for NESYMF, step
(2) consists of two steps: () a standard MF model is
pretrained on the source domain to populate matrix G, then
() NESYMF is trained on the target domain, namely   .</p>
        <sec id="sec-2-3-1">
          <title>We repeated the entire procedure 30 times using seeds from 0 to 29. The test metrics have been averaged across these runs and reported in Table 2.</title>
        </sec>
      </sec>
      <sec id="sec-2-4">
        <title>5.3. Training details</title>
        <sec id="sec-2-4-1">
          <title>All the models have been trained for 500 epochs using</title>
          <p>the Adam optimizer. Early stopping has been used to stop
the training if no improvements were found on the
validation set for ten epochs. For all the models, the user
and item latent factors, U and I, and the user and item
biases, u and i, have been randomly initialized using the</p>
        </sec>
        <sec id="sec-2-4-2">
          <title>Glorot initialization. The MF models have been trained</title>
          <p>using Equation (1), while NESYMF using Equation (5).
For NESYMF, Axiom (4) has been added to the loss from
epoch five 18, allowing LTN to learn something about the
latent factors before starting reasoning on the acquired
knowledge.</p>
        </sec>
        <sec id="sec-2-4-3">
          <title>We used Bayesian optimization to find the optimal</title>
          <p>hyper-parameters for our models. We executed every
hyper-parameter search for 150 runs and selected the
configuration that led to the best validation score. Due
domain (i.e., MF and NESYMF).</p>
          <p>For the MF model trained on the source domain, we
used F0.5-measure20 as the validation metric, while in
all the other cases, we used F1-measure. Clearly, the
performance of NESYMF depends on the quality of G (i.e.,
the pre-trained model). Using F0.5-measure allowed us
to obtain more precise predictions for G. In particular,
we reduced the number of false positives, namely cases in
which G erroneously predicts that a user dislikes a genre.</p>
        </sec>
        <sec id="sec-2-4-4">
          <title>In such cases, Axiom (4) would have been unreliably applied.</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>6. Results</title>
      <sec id="sec-3-1">
        <title>The results obtained with our experiments are summa</title>
        <p>rized in Table 2. The discussion is limited to MR-200k as
for MR-100k we obtained similar results. Note the results
for MR-100k are better since the dataset is slightly denser.</p>
      </sec>
      <sec id="sec-3-2">
        <title>By looking at the F1-measure, it is possible to observe that</title>
        <p>NESYMF outperforms MF on all five folds. In particular,
the performance gap increases with the sparsity of the
user-item ratings, starting from a 1.04% improvement on
the full dataset (i.e., 100% fold) and ending with a 6.69%
improvement on the most sparse dataset (i.e., 5% fold).</p>
        <p>This shows the benefits of transferring knowledge from
a denser domain when training data is poor. Moreover, it
suggests our proposal can be efectively used in the task
17Notice the ratings on movie genres are kept untouched as they are
of knowledge transfer for recommendation.</p>
        <p>used for accurate pre-training.
18Notice this is an arbitrary choice. The idea of using knowledge
transfer is to correct the misclassifications made by the MF model
when training data are sparse. The more the data sparsity, the more
likely the MF will erroneously classify user-item pairs. During the
ifrst steps of learning, the MF has not learned enough information
to predict accurately. For this reason, it is likely knowledge transfer
is applied to random predictions, hence useless.</p>
        <p>Interestingly, by looking at recall, it is possible to
observe that the addition of knowledge helps NESYMF in
19By maximizing precision, we obtained a more accurate pre-trained
model for predicting user-genre preferences. This helped in
transferring knowledge more efectively.
20F0.5-measure gives more weight to precision than recall. It is used
when avoiding false positives is particularly important.
Best hyper-parameters (referred to as h-p) found through Bayesian optimization for the pre-trained model (left) and the
models learned on the target domain (right). The hyper-parameters are subdivided by dataset.
Comparison of MF and NESYMF on the selected datasets. The test metrics are averaged across 30 runs.
confidence) that a user really dislikes an item, only
know</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>7. Conclusions and future works</title>
      <sec id="sec-4-1">
        <title>In this paper, we presented a Neuro-Symbolic approach to</title>
        <p>knowledge transfer for RSs. Specifically, we used a Logic
Tensor Network to encode axiomatic knowledge suitable
for transferring information from source to target
domain via logical reasoning. We showed that our model
outperforms a standard MF model on all the presented
tasks, proving its potential in the knowledge transfer
task. Moreover, an experiment that drastically reduced
the user-item ratings in the dataset showed the ability
of our proposal to deal with data sparsity. In the
future, we would like to extend our model to the
crossdomain recommendation task. In particular, instead of
using movie-genre preferences, one could use more usual
source domains (e.g., books, songs).
ArXiv abs/2105.05330 (2021). we really making much progress? a worrying
anal[22] T. Campari, L. Lamanna, P. Traverso, L. Serafini, ysis of recent neural recommendation approaches,
L. Ballan, Online learning of reusable abstract in: Proceedings of the 13th ACM Conference on
models for object goal navigation, in: 2022 Recommender Systems, RecSys ’19, Association for
IEEE/CVF Conference on Computer Vision and Computing Machinery, New York, NY, USA, 2019,
Pattern Recognition (CVPR), IEEE Computer Soci- p. 101–109. URL: https://doi.org/10.1145/3298689.
ety, Los Alamitos, CA, USA, 2022, pp. 14850–14859. 3347058. doi:10.1145/3298689.3347058.
URL: https://doi.ieeecomputersociety.org/10.1109/ [30] A. H. Brams, A. L. Jakobsen, T. E. Jendal, M.
LissanCVPR52688.2022.01445. doi:10.1109/CVPR52688. drini, P. Dolog, K. Hose, Mindreader:
Recommen2022.01445. dation over knowledge graph entities with explicit
[23] T. Carraro, A. Daniele, F. Aiolli, L. Serafini, user ratings, in: Proceedings of the 29th ACM
Logic tensor networks for top-n recommenda- International Conference on Information &amp;
Knowltion, in: AIxIA 2022 – Advances in Arti- edge Management, CIKM ’20, Association for
Comifcial Intelligence: XXIst International Confer- puting Machinery, New York, NY, USA, 2020, p.
ence of the Italian Association for Artificial In- 2975–2982. URL: https://doi.org/10.1145/3340531.
telligence, AIxIA 2022, Udine, Italy, November 3412759. doi:10.1145/3340531.3412759.
28 – December 2, 2022, Proceedings, Springer- [31] H. Chen, Y. Li, S. Shi, S. Liu, H. Zhu, Y. Zhang,
Verlag, Berlin, Heidelberg, 2023, p. 110–123. Graph collaborative reasoning, in: Proceedings
URL: https://doi.org/10.1007/978-3-031-27181-6_8. of the Fifteenth ACM International Conference on
doi:10.1007/978-3-031-27181-6_8. Web Search and Data Mining, WSDM ’22,
Asso[24] G. Spillo, C. Musto, M. De Gemmis, P. Lops, G. Se- ciation for Computing Machinery, New York, NY,
meraro, Knowledge-aware recommendations based USA, 2022, p. 75–84. URL: https://doi.org/10.1145/
on neuro-symbolic graph embeddings and first- 3488560.3498410. doi:10.1145/3488560.3498410.
order logical rules, in: Proceedings of the 16th [32] J. Ji, Z. Li, S. Xu, M. Xiong, J. Tan, Y. Ge,
ACM Conference on Recommender Systems, Rec- H. Wang, Y. Zhang, Counterfactual
collaboraSys ’22, Association for Computing Machinery, tive reasoning, in: Proceedings of the Sixteenth
New York, NY, USA, 2022, p. 616–621. URL: https: ACM International Conference on Web Search
//doi.org/10.1145/3523227.3551484. doi:10.1145/ and Data Mining, WSDM ’23, Association for
3523227.3551484. Computing Machinery, New York, NY, USA, 2023,
[25] P. Kouki, S. Fakhraei, J. Foulds, M. Eirinaki, p. 249–257. URL: https://doi.org/10.1145/3539597.</p>
        <p>L. Getoor, Hyper: A flexible and extensible prob- 3570464. doi:10.1145/3539597.3570464.
abilistic framework for hybrid recommender sys- [33] Y. Xian, Z. Fu, H. Zhao, Y. Ge, X. Chen, Q. Huang,
tems, in: Proceedings of the 9th ACM Confer- S. Geng, Z. Qin, G. de Melo, S.
Muthukrishence on Recommender Systems, RecSys ’15, As- nan, Y. Zhang, Cafe: Coarse-to-fine neural
sociation for Computing Machinery, New York, NY, symbolic reasoning for explainable
recommendaUSA, 2015, p. 99–106. URL: https://doi.org/10.1145/ tion, in: Proceedings of the 29th ACM
Interna2792838.2800175. doi:10.1145/2792838.2800175. tional Conference on Information &amp; Knowledge
[26] H. Chen, S. Shi, Y. Li, Y. Zhang, Neural collab- Management, CIKM ’20, Association for
Comorative reasoning, in: Proceedings of the Web puting Machinery, New York, NY, USA, 2020, p.
Conference 2021, WWW ’21, Association for Com- 1645–1654. URL: https://doi.org/10.1145/3340531.
puting Machinery, New York, NY, USA, 2021, p. 3412038. doi:10.1145/3340531.3412038.
1516–1527. URL: https://doi.org/10.1145/3442381. [34] A. Kimmig, S. Bach, M. Broecheler, B. Huang,
3449973. doi:10.1145/3442381.3449973. L. Getoor, A short introduction to probabilistic
[27] A. Daniele, L. Serafini, Knowledge enhanced neural soft logic, Mansinghka, Vikash, 2012, pp. 1–4. URL:
networks, in: A. C. Nayak, A. Sharma (Eds.), PRI- https://lirias.kuleuven.be/retrieve/204697.
CAI 2019: Trends in Artificial Intelligence, Springer [35] Z. Wang, J. Zhang, J. Feng, Z. Chen, Knowledge
International Publishing, Cham, 2019, pp. 542–554. graph embedding by translating on hyperplanes,
doi:10.1007/978-3-030-29908-8_43. Proceedings of the AAAI Conference on
Artifi[28] S. Badreddine, A. d’Avila Garcez, L. Ser- cial Intelligence 28 (2014). URL: https://ojs.aaai.org/
afini, M. Spranger, Logic tensor networks, index.php/AAAI/article/view/8870. doi:10.1609/
Artificial Intelligence 303 (2022) 103649. aaai.v28i1.8870.</p>
        <p>URL: https://www.sciencedirect.com/science/ [36] S. Guo, Q. Wang, L. Wang, B. Wang, L. Guo, Jointly
article/pii/S0004370221002009. doi:https: embedding knowledge graphs and logical rules, in:
//doi.org/10.1016/j.artint.2021.103649. Proceedings of the 2016 Conference on Empirical
[29] M. Ferrari Dacrema, P. Cremonesi, D. Jannach, Are Methods in Natural Language Processing,
Associa</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <year>2016</year>
          , pp.
          <fpage>192</fpage>
          -
          <lpage>202</lpage>
          . URL: https://aclanthology.org/
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <fpage>D16</fpage>
          -
          <lpage>1019</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>D16</fpage>
          -1019. [37]
          <string-name>
            <surname>E. van Krieken</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Acar</surname>
          </string-name>
          ,
          <string-name>
            <surname>F. van Harmelen</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>tors</surname>
          </string-name>
          ,
          <source>Artificial Intelligence</source>
          <volume>302</volume>
          (
          <year>2022</year>
          )
          <fpage>103602</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>article/pii/S0004370221001533. doi:https:</mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          //doi.org/10.1016/j.artint.
          <year>2021</year>
          .
          <volume>103602</volume>
          . [38]
          <string-name>
            <given-names>T.</given-names>
            <surname>Carraro</surname>
          </string-name>
          , LTNtorch: PyTorch implementation
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <source>of Logic Tensor Networks</source>
          ,
          <year>2023</year>
          . URL: https://doi.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <source>org/10</source>
          .5281/zenodo.7778157. doi:
          <volume>10</volume>
          .5281/zenodo.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          7778157. [39]
          <string-name>
            <given-names>F.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>ference on Artificial Intelligence, IJCAI-21</source>
          , Interna-
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Organization</surname>
          </string-name>
          ,
          <year>2021</year>
          , pp.
          <fpage>4721</fpage>
          -
          <lpage>4728</lpage>
          . URL: https://doi.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <source>org/10</source>
          .24963/ijcai.
          <year>2021</year>
          /639. doi:
          <volume>10</volume>
          .24963/ijcai.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <year>2021</year>
          /639, survey Track.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>