<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Using Ontology-based Data Summarization to Develop Semantics-aware Recommender Systems?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vito Walter Anelli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tommaso Di Noia</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrea Maurino</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matteo Palmonari</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anisa Rula</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Polytechnic University of Bari</institution>
          ,
          <addr-line>Via Orabona, 4, 70125 Bari</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Milano-Bicocca</institution>
          ,
          <addr-line>Piazza dell'Ateneo Nuovo, 1, 20126 Milano</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <fpage>24</fpage>
      <lpage>27</lpage>
      <abstract>
        <p>In the current information-centric era, recommender systems are gaining momentum as tools able to assist users in daily decision-making tasks. Within the recommendation process, Linked Data have been already proposed as a valuable source of information to enhance the predictive power of recommender systems but an open issues is still related to feature selection of the most relevant subset of data in the whole semantic web. In this paper, we show how ontology-based (linked) data summarization can drive the selection of properties/features useful to a recommender system. In particular, we compare a fully automated feature selection method based on ontology-based data summaries with more classical ones, and we evaluate the performance of these methods in terms of accuracy and aggregate diversity of a recommender system exploiting the top-k selected features.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Semantics-aware Recommender Systems (RSs) exploiting information held in knowledge
graphs, as the ones available as Linked Data (LD), represent one of the most interesting
and challenging application scenarios for LD [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. A high number of solutions and tools
have been proposed in the last years showing the effectiveness of adopting LD as
knowledge sources to feed a recommendation engine (see [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and references therein for an
overview). Nevertheless, how to automatically select the “best” subset of a LD dataset to
feed a LD-based RS without affecting the performance of the recommendation algorithm
is still an open issue. Notice that the selection of the top-k features to use in a RSs means
to discover which properties in a LD-dataset (e.g., DBpedia) encode the knowledge useful
in the recommendation task and which ones are just noise [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. In most of the approaches
proposed so far, usually, the FS process is performed by human experts that manually
choose properties resulting more “suitable” for a given scenario. Over the years, many
algorithms and techniques for feature selection , e.g., Information Gain, Information Gain
Ratio, Chi Squared Test and Principal Component Analysis, have been proposed with
reference to machine learning tasks but they do not consider a characteristic which makes
unique LD: they come with semantics attached.
      </p>
      <p>
        The main objective of this paper is to investigate how ontology-based data
summarization [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ] can be used as a new and semantic-oriented feature selection technique for
LD-based RSs. We define a feature selection method that automatically extracts the top-k
properties that are deemed to be more important to evaluate similarity between instances
of a given class on top of data summaries built with the help of an ontology. We
perform an experimental evaluation on three well-known datasets in the RS domain
(Movielens, LastFM, LibraryThing) in order to analyze how the choice of a particular FS
technique may influence the performance of recommendation algorithms in terms of typical
accuracy and diversity metrics. Experimental results show that information provided in
ontology-based data summaries selects features that achieve comparable, or, in most of
the cases, better performance than state-of-the-art, semantic-agnostic analytical methods
such as Information Gain [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. We believe that these results are interesting also because
of practical reasons. LD summaries are published on-line and summary-based FS can be
performed even without acquiring the entire dataset and efficiently (on top of summary
information). The use of frequency associated with schema patterns in a FS approach was
initially tested in a previous work[
        <xref ref-type="bibr" rid="ref24">24</xref>
        ].
      </p>
      <p>The paper is organized as follows: in Section 2, we introduce the ontology-based data
summarization approach used in this work, while in Section 3, we describe the feature
selection and recommendation methods. Section 4 is devoted to the explanation and
discussion of the experimental results. Section 5 briefly reviews related literature for schema
and data summarization as well as on recommender systems while Section 6 discuss
conclusions and future work.</p>
    </sec>
    <sec id="sec-2">
      <title>2 Ontology-driven Linked Data Summarization</title>
      <p>
        While relevance-oriented data summarization approaches are aimed at finding subsets of a
dataset or an ontology that are estimated to be more relevant for the users [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ],
vocabularyoriented approaches are aimed at profiling a dataset, by describing the usage of
vocabularies/ontologies used in the dataset. The summaries returned by these approaches are
complete, i.e., they provide statistics about every element of the vocabulary/ontology used in
the dataset [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ]. Statistics captured by these summaries that can be useful for the feature
selection process are the ones concerning the usage of properties for a certain class of
items to recommend.
      </p>
      <p>
        Patterns and frequency. In our approach we use pattern-based summaries
extracted using the ABSTAT framework. Pattern-based summaries describe the content
of a dataset using schema patterns having the form hC; P; Di, where C and D, are
types (either classes or datatypes) and P is a property. For example, the pattern
hdbo:Film; dbo:starring; dbo:Actori tells that films exist in the dataset, in
which star some actors. Differently from similar pattern-based summaries [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ],
ABSTAT uses the subclass relations in the data ontology, represented in a Type Graph, to
extract only minimal type patterns from relational assertions, i.e, the patterns that are
more type-wise specific according to the ontology. A pattern hC; P; Di is a minimal type
pattern for a relational assertion ha; P; bi according to a type graph G iff C and D are
the types of a and b respectively, which are minimal in G. In a pattern hC; P; Di, C
and D are referred to as source and target types respectively. A minimal type pattern
hdbo:Film; dbo:starring; dbo:Actori (simply referred to as pattern in the
following) tells that there exist entities that have dbo:Film and dbo:Actor as minimal
ABSTAT summaries for several datasets can be explored at http://abstat.disco.
unimib.it
If no ontology is specified, all types are minimal and patterns are extracted like in frameworks
that do not adopt minimalization
types which are connected through the property P . Non minimal patterns can be inferred
from minimal patterns and the type graph. Therefore, they can be excluded as redundant
without information loss, making summaries more compact [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ]. Each pattern hC; P; Di
is associated with a frequency, which reports the number of relational assertions ha; P; bi
from which the pattern has been extracted.
      </p>
      <p>Local cardinality descriptors. For this work, we have extended ABSTAT to extract
local cardinality descriptors, i.e., cardinality descriptors of RDF properties, which are
specific to the patterns in which the properties occur. To define these descriptors, we first
introduce the concept of restricted property extensions. The extension of a property P
restricted to a pattern hC; P; Di is the set of pairs hx; yi such that the relational assertion
hx; P; yi is part of the dataset and hC; P; Di is a minimal-type pattern for hx; P; yi. Given
a pattern with a property P , we can define the functions (that return the closest integer):
minS( ), maxS( ), avgS( ): denoting respectively the minimum, maximum and
average number of distinct subjects associated to unique objects in the extension of P
restricted to ;</p>
      <p>minO( ), maxO( ), avgO( ): denoting respectively the minimum, maximum and
average number of distinct objects associated to unique subjects in the extension of P
restricted to .</p>
      <p>ABSTAT can also compute global cardinality descriptors by adjusting the above
mentioned definition so as to consider unrestricted property extensions. Local cardinality
descriptors carry information about the semantics of properties as used with specific types
of resources (in specific patterns) and can be helpful for selecting features used to
compute the similarity between resources. For example, to compute similarity for movies, one
would like to discard properties that occur in patterns with dbo:Film as source type
and avgS( ) = 1. We remark that the values of local cardinality descriptors for patterns
with a property P may differ from values of global cardinality descriptors for P . Some
examples of local cardinality descriptors can be found in the faceted-search interface
(ABSTATBrowse). In conclusion, ABSTAT takes a linked dataset and - if specified - one or
more ontologies as input, and returns a summary that consists of: a type graph, a set of
patterns, their frequency, local an global cardinality descriptors.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Semantics-aware Feature Selection</title>
      <p>
        Feature selection is the process of selecting a subset of relevant attributes in a dataset.
Thanks to the feature selection process it is possible to improve the prediction
performance, and to give a better understanding of the process that generates the data [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
There are three typical measures of feature selection (i) ”filters”, statistical measures to
assign a score to each feature (here the feature selection process is a preprocessing step
and can be independent from learning[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]); (ii) ”wrapper” where the learning system is
used as a black box to score subsets of features [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]; (iii) embedded methods that
perform the selection within the process of training [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. In the following, we discuss two
approaches used for the feature selection task: the first operates on the summarization of
the datasets and the second operates on the instances of the datasets.
      </p>
      <p>Feature Selection with Ontology-based Summaries. As described in Section 2, the
ABSTAT framework provides two useful statistics: the pattern frequency and the
cardinality descriptors that are used in the feature selection process as described in Figure 1. The
http://abstat.disco.unimib.it/browse</p>
      <p>P D # avgS
1 dbo:director foaf:Person 93k 4
2 dbo:director dbo:Person 39k 5
3 dbo:director dbo:Producer 16k 7
4 dbo:wikiPageExternalLink owl:Thing 110k 1
5 dbo:starring dbo:Person 49k 4
6 dbo:starring dbo:Actor 218k 7
7 dbo:starring foaf:Person 306k 4
8 owl:sameAs owl:Thing 757k 1
9 dcterms:subject skos:Concept 934k 18</p>
      <p>P
dcterms:subject
dbo:starring
( = )
process starts by considering all patterns = f 1; 2; : : : ; ng of a given class C
occurring as a source type. The example in Figure 1 shows a subset of with dbo:Film as
source type. The first step of our approach (FILTERBY) filters out properties based on the
local cardinality descriptors. In particular, it filters only properties for which the average
number of distinct subjects associated with unique objects is more than one (avgS &gt; 1).
The second step of the process (SELECTDISTINCTP) selects all properties of the
patterns in by applying the maximum of the pattern frequency (# in the Figure). Then,
the properties are ranked (ORDERBY) in a descending order on pattern frequency and
then k properties (TOPK) are selected (k=2).</p>
      <p>In some datasets, such as DBpedia, properties may use redundant information
by using same properties with different namespaces, e.g., dbo:starring and
dbp:starring. For this reason, in such case, a pre-processing step for removing
replicated properties to avoid redundant ones is requested (see Section 4).</p>
      <p>. 6
1 dbo:direPctor foaf:PeDrson #93k
2 dbo:director dbo:Person 39k 
3 dbo:director dbo:Producer 16k ( # )
5 dbo:starring dbo:Person 49k
6 dbo:starring dbo:Actor 218k
7 dbo:starring foaf:Person 306k
9 dcterms:subject skos:Concept 934k
dbo:direPctor #93k
dbo:starring 306k
dcterms:subject 934k</p>
      <p>P
(#) dcterms:subject
dbo:starring
dbo:director</p>
      <p>
        Feature Selection with State-of-the-art Techniques. In this work we consider RDF
properties as features, so among the different feature selection techniques available in the
literature, we initially selected Information Gain, Information Gain Ratio, Chi-squared
test and Principal Component Analysis as their computation can be adapted to categorical
features as LD and we then evaluated their effect over the recommendation results. The
features selected from each technique have been used as an input of the recommendation
algorithm that uses the Jaccard index as similarity measure. In order to identify the best
technique among the one we selected, they have been evaluated by using Information
Gain (IG), Gain Ration and Chi Squared Test. At the end, IG[
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] resulted as the best
performing one. Then features are ranked according to their IG value and the top-k ones
are returned.
      </p>
      <p>Feature pre-processing. LD datasets usually have a quite large feature set that can
be very sparse depending on the knowledge domain. For instance, taking into
account the movies available in Movielens, properties as dbp:artDirection or
dbp:precededBy are very specific and have a lot of missing values. On the other hand,
properties as dbo:wikiPageExternalLink or owl:sameAs always have different
and unique values, so they are not informative for a recommendation task.</p>
      <p>For the sake of conciseness we do not report all the results here. Results obtained with other FS
techniques can be found at http://ow.ly/zAA530d0wu0</p>
      <p>
        For this reason, before starting the feature selection process with IG, we reduced
redundant or irrelevant features. The pre-processing step has been done following [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]: we
fixed a threshold tm = td = 97% both for missing values and for distinct values and,
then, we discarded features for which we had more than tm of missing values and more
than td of distinct values.
      </p>
      <p>
        Recommendation Method. We implemented a content-based recommender system
using an item-based nearest neighbors algorithm as in [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], where the similarity is
computed by means of Jaccard’s index (widely adopted for categorical features). In this work,
the neighborhood of a resource includes all the nodes in the graph reachable starting from
the resource i (respectively j) following the properties selected by the feature selection
phase. The neighbors are thus one-hop features. The similarity values are then used to
recommend to each user the items which result most similar to the ones she has liked
in the past. Ratings are predicted as a normalized sum of neighbors ratings, weighted by
their similarity values [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ].
      </p>
    </sec>
    <sec id="sec-4">
      <title>4 Experimental Evaluation</title>
      <p>
        Datasets. The evaluation has been carried out on the three well-known datasets
belonging to different domains, i.e. movies(Movielens 1M), books (LibraryThing), and music
(Last.fm). The datasets contains, respectively, 1,000,209, 626,000 and 92,834 ratings.
Movielens 1M and LibraryThing provide explicit ratings over 1-5 and 1-10 scales whereas
Last.fm [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] provides users listening counts.
      </p>
      <p>
        Measures. For evaluating the quality of our recommendation algorithm we are interested
in measuring its performances in terms of accuracy of the predicted results and diversity.
To evaluate recommendation accuracy, we used Precision (Precision@N) and Mean
Reciprocal Rank (MRR). Precision@N is a metric denoting the fraction of relevant items in
the Top-N recommendations. MRR computes the average reciprocal rank of the first
relevant recommended item [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]. A good recommender system should provide
recommendations equally distributed among the items, otherwise, even if accurate, they indicate a low
degree of personalization [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. To evaluate aggregate diversity, we considered catalog
coverage (the percentage of recommended items in the catalog) and aggregate entropy
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Please note that here we are considering a global diversity rather than a personalized
one [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>Implementation. We tried different ranking and filtering functions to study their
effect on feature selection. For lack of space we include the best combination from those
proposed in section 3, AbsOccAvgS, that considers as input of FILTERBY the avgS and
SELECTDISTINCTP the maximum of the pattern frequency. Both for ABSTAT and IG
we considered both the Onlydbp configuration in which, if among the first N features
selected there are both dbo: and dbp: feature we consider only the dbp: one. Conversely
in Onlydbo we take into account only the dbo: one.</p>
      <p>ABSTAT Baseline. As a baseline for ABSTAT-based feature selection we use TF-IDF
as is a well-known measure to identify most relevant terms (properties in this case) for
a document (a class in this case). We adopt TF-IDF in our context where by document
we refer to a set of patterns having the same subject-type and by term we refer to a
property. TF-IDF is based on the number of properties occurring in a document (TF) and
the logarithm of the ratio between the total number of documents and the ones containing
the property (IDF).</p>
      <p>Results Tables 1, 2, 3 show the experimental results obtained on, respectively,
Movielens, Last.FM and LibraryThing datasets in terms of Precision, MRR,
catalogCoverage, and aggrEntropy. Results are computed over lists of top-10 items recommended
by the RS. We conducted experiments using top-k selected features for different k, i.e.,
k = 5; 10; 15; 20, and all configurations but we report only results for k = 5; 20 and
best configurations . We highlight in bold only the values for which there is a statistical
significant difference. For Lastfm dataset the differences are not statistical significant so
the two methods are equivalent in selecting features.</p>
      <p>Discussion. As an overall result, ABSTAT-based FS leads to the best results in terms
of accuracy and diversity for both the movie and books domains while IG leads to better
results (although not statistically significant) for music.</p>
      <p>Specifically, considering the results on Movielens (Table 1), ABSTAT produces
better accuracy with respect to IG in all the configurations both with 5 and 20 features.
In terms of aggregate diversity, i.e. itemCoverage and aggrEntropy, ABSTAT is still the
best choice, overcoming IG in almost all the situations. On Lastfm (Table 2) there are
no particular differences, and hence the choice of the method seems irrelevant: both
summarization-based and statistical methods are comparable. Eventually, on
LibraryThing (Table 3), ABSTAT strongly beats IG in almost all the configurations. In particular,
it gets more than twice of the precision and MRR respect to IG in top-5 features
scenario. Summing up, ABSTAT beats IG in almost all the configurations on the two datasets
Movielens and LibraryThing, while they act in the same way on the Lastfm dataset.</p>
      <p>In order to investigate the reasons behind the different behaviors depending on the
selected knowledge domain, we measured: (i) the number of minimal patterns and (ii) the
average number of triples per resource and the corresponding variance. Regarding the
former we may say that a higher number of minimal patterns means a richer and more diverse
The interested reader can find results for all values of k and configurations on GitHub: http:
//ow.ly/zAA530d0wu0
ontological representation of the knowledge domain. As for the latter, a high variance in
the number of triples associated to resources is a clue of an unbalanced representation of
the items to recommend. Hence, items with a higher number of triples associated result
“more popular” in the knowledge graph compared to those with only a few. This may
reflect in the rising of a stronger content popularity bias while computing the
recommendation results. If we look at the values represented in Table 4 we may assert that a higher
sparsity in the knowledge graph data may give chance to statistical methods to beat
ontological ones. In other words, it seems that the higher the sparsity of the knowledge graph
at the data level, the lower the influence of the ontological schema in the selection of the
most informative features to build a pure content-based recommendation engine.</p>
    </sec>
    <sec id="sec-5">
      <title>5 Related Work</title>
      <p>
        Summarization. Different approaches have been proposed for schema and data
summarization[
        <xref ref-type="bibr" rid="ref28">28</xref>
        ]. Several data profiling approaches are aimed to describe linked data by
reporting statistics about the usage of the vocabularies, types and properties. SchemeEx
extracts interesting measures , by considering the co-occurrence of types and properties [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
Linked Open Vocabularies, RDFStats [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] and LODStats [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] provide such statistics. In
contrast, ABSTAT represents connections between types using schema patterns, for which
it also provides cardinality descriptors. does not include cardinality descriptors for
properties or patterns. TermPicker extracts [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] patterns consisting in triples hS; E; Oi, where
S and O are sets of types and E is a set of predicates. Instead, ABSTAT and Loupe
extract patterns each consisting in a triple hC; P; Di where C and D are types and P a
property. TermPicker summaries do not describe cardinality and are extracted from RDF
data without considering relationships between types. According to [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], which proposes
a method to define and discover classes of cardinality constraints with some preliminary
results, current approaches focus only on mining keys or pseudo-keys (e.g., [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]). We
discover richer statistics about property cardinality like the above mentioned work, and in
addition we compute cardinality descriptors for properties occurring in specific schema
patterns.
      </p>
      <p>
        Recommender Systems. One of the first approaches for using LD in a recommender
system was proposed by Heitmann and Hayes[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. A system for recommending artists
and music using DBpedia was presented in [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. The task of cross-domain
recommendation leveraging DBpedia as a knowledge-based framework was addressed in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], while in
[
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] the authors present a semantics-aware approach to deal with cold-start situations. In
[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] the authors use a a hybrid graph-based algorithm built upon DBpedia and collaborative
information. To the best of our knowledge, the only approaches proposing an automatic
selection of LD features are [
        <xref ref-type="bibr" rid="ref19 ref24">19, 24</xref>
        ]. Finally, we observe that even approaches that do not
perform automatic FS like [
        <xref ref-type="bibr" rid="ref19 ref4 ref9">19, 4, 9</xref>
        ] used (hand-crafted) FS to improve their performance.
      </p>
    </sec>
    <sec id="sec-6">
      <title>6 Conclusions</title>
      <p>In this work we investigated the role of ontology-based data summarization for feature
selection in recommendation tasks. Here we compare results coming from ABSTAT, a
http://lov.okfn.org/
schema summarization tool, with classical methods for feature selection and we show that
the former are allowed to compute better predictions not just in terms of precision of the
recommended items but also considering other dimensions such as diversity. Experiments
have been carried out in three different knowledge domains thus showing the effectiveness
of a feature selection based on schema summarization over classical techniques such as
Information Gain.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>G.</given-names>
            <surname>Adomavicius</surname>
          </string-name>
          and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kwon</surname>
          </string-name>
          .
          <article-title>Improving aggregate recommendation diversity using ranking-based techniques</article-title>
          .
          <source>IEEE Trans. on Knowl. and Data Eng</source>
          .,
          <volume>24</volume>
          (
          <issue>5</issue>
          ),
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>S.</given-names>
            <surname>Auer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Demter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Martin</surname>
          </string-name>
          ,
          <string-name>
            <surname>and J. Lehmann.</surname>
          </string-name>
          <article-title>LODStats - An Extensible Framework for High-Performance Dataset Analytics</article-title>
          .
          <source>In EKAW (2)</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>I.</given-names>
            <surname>Cantador</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Brusilovsky</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Kuflik</surname>
          </string-name>
          . 2nd workshop
          <article-title>on information heterogeneity and fusion in recommender systems (hetrec 2011)</article-title>
          .
          <source>In RecSys. ACM</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>M. de Gemmis</surname>
            , P. Lops,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Musto</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Narducci</surname>
            , and
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Semeraro</surname>
          </string-name>
          .
          <article-title>Semantics-aware content-based recommender systems</article-title>
          .
          <source>In Recommender Systems Handbook</source>
          .
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>T. Di</given-names>
            <surname>Noia</surname>
          </string-name>
          .
          <article-title>Knowledge-enabled recommender systems: Models, challenges, solutions</article-title>
          .
          <source>In Proceedings of the 3rd International Workshop on Knowledge Discovery on the WEB, Cagliari, Italy, September 11-12</source>
          ,
          <year>2017</year>
          .,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>T. Di</given-names>
            <surname>Noia</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Cantador</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V. C.</given-names>
            <surname>Ostuni</surname>
          </string-name>
          .
          <article-title>Linked open data-enabled recommender systems: ESWC 2014 challenge on book recommendation</article-title>
          .
          <source>In Semantic Web Evaluation Challenge - SemWebEval 2014 at ESWC</source>
          <year>2014</year>
          , Anissaras, Crete, Greece, May
          <volume>25</volume>
          -29,
          <year>2014</year>
          , Revised Selected Papers, pages
          <fpage>129</fpage>
          -
          <lpage>143</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>T. Di</given-names>
            <surname>Noia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Magarelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Maurino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Palmonari</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Rula</surname>
          </string-name>
          .
          <article-title>Using ontology-based data summarization to develop semantics-aware recommender systems</article-title>
          .
          <source>In The Semantic Web - 15th International Conference, ESWC</source>
          <year>2018</year>
          , Heraklion, Crete, Greece, June 3-7,
          <year>2018</year>
          , Proceedings, pages
          <fpage>128</fpage>
          -
          <lpage>144</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>T. Di</given-names>
            <surname>Noia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Ostuni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rosati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Tomeo</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E. Di</given-names>
            <surname>Sciascio</surname>
          </string-name>
          .
          <article-title>An analysis of users' propensity toward diversity in recommendations</article-title>
          .
          <source>In RecSys 2014 - Proceedings of the 8th ACM Conference on Recommender Systems</source>
          , pages
          <fpage>285</fpage>
          -
          <lpage>288</lpage>
          . Association for Computing Machinery, Inc,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>T. Di</given-names>
            <surname>Noia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. C.</given-names>
            <surname>Ostuni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Tomeo</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E. Di</given-names>
            <surname>Sciascio</surname>
          </string-name>
          . Sprank:
          <article-title>Semantic path-based ranking for top-N recommendations using linked open data</article-title>
          .
          <source>ACM TIST</source>
          ,
          <volume>8</volume>
          (
          <issue>1</issue>
          ):9:
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          :
          <fpage>34</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10. I.
          <article-title>Ferna´ndez-Tob´ıas, I</article-title>
          . Cantador,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kaminskas</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Ricci</surname>
          </string-name>
          .
          <article-title>A generic semantic-based framework for cross-domain recommendation</article-title>
          .
          <source>In 2nd Workshop on Information Heterogeneity and Fusion in Recommender Systems</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>X.</given-names>
            <surname>Geng</surname>
          </string-name>
          , T.-Y. Liu,
          <string-name>
            <given-names>T.</given-names>
            <surname>Qin</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <article-title>Feature selection for ranking</article-title>
          .
          <source>In SIGIR. ACM</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>I.</given-names>
            <surname>Guyon</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Elisseeff</surname>
          </string-name>
          .
          <article-title>An introduction to variable and feature selection</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          ,
          <volume>3</volume>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>B.</given-names>
            <surname>Heitmann</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Hayes</surname>
          </string-name>
          .
          <article-title>Using linked data to build open, collaborative recommender systems</article-title>
          .
          <source>In AAAI Spring Symposium: Linked Data Meets Artificial Intelligence</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <given-names>R.</given-names>
            <surname>Kohavi</surname>
          </string-name>
          and
          <string-name>
            <given-names>G. H.</given-names>
            <surname>John</surname>
          </string-name>
          .
          <article-title>Wrappers for feature subset selection</article-title>
          .
          <source>Artificial Intelligence</source>
          ,
          <volume>97</volume>
          (
          <issue>1-2</issue>
          ),
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>M. Konrath</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Gottron</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Staab</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Scherp</surname>
          </string-name>
          .
          <article-title>Schemex - efficient construction of a data catalogue by stream-based indexing of linked data</article-title>
          .
          <source>Web Semant., 16</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <given-names>A.</given-names>
            <surname>Langegger</surname>
          </string-name>
          and
          <string-name>
            <given-names>W.</given-names>
            <surname>Wo</surname>
          </string-name>
          <article-title>¨ß. RDFStats - an extensible RDF statistics generator and library</article-title>
          .
          <source>In DEXA Workshops</source>
          , pages
          <fpage>79</fpage>
          -
          <lpage>83</lpage>
          . IEEE,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>N.</given-names>
            <surname>Mihindukulasooriya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. Poveda</given-names>
            <surname>Villalon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Garcia-Castro</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Gomez-Perez</surname>
          </string-name>
          .
          <article-title>Loupe - An Online Tool for Inspecting Datasets in the Linked Data Cloud</article-title>
          .
          <source>In ISWC Posters &amp; Demonstrations</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18. E. Mun˜oz.
          <article-title>On learnability of constraints from RDF data</article-title>
          .
          <source>In ESWC</source>
          . Springer,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>C. Musto</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Lops</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Basile</surname>
            , M. de Gemmis, and
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Semeraro</surname>
          </string-name>
          .
          <article-title>Semantics-aware graph-based recommender systems exploiting linked open data</article-title>
          .
          <source>In UMAP</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20. P. Nguyen,
          <string-name>
            <given-names>P.</given-names>
            <surname>Tomeo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. Di</given-names>
            <surname>Noia</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E. Di</given-names>
            <surname>Sciascio</surname>
          </string-name>
          .
          <article-title>An evaluation of SimRank and Personalized PageRank to build a recommender system for the web of data</article-title>
          .
          <source>In WWW '15 Companion</source>
          , pages
          <fpage>1477</fpage>
          -
          <lpage>1482</lpage>
          . ACM,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <given-names>X.</given-names>
            <surname>Ning</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Desrosiers</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Karypis</surname>
          </string-name>
          .
          <article-title>A comprehensive survey of neighborhood-based recommendation methods</article-title>
          .
          <source>In Recommender Systems Handbook</source>
          , pages
          <fpage>37</fpage>
          -
          <lpage>76</lpage>
          .
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <given-names>A.</given-names>
            <surname>Passant</surname>
          </string-name>
          . Dbrec:
          <article-title>Music Recommendations Using DBpedia</article-title>
          .
          <source>In 9th ISWC</source>
          , pages
          <fpage>209</fpage>
          -
          <lpage>224</lpage>
          . Springer-Verlag,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <given-names>H.</given-names>
            <surname>Paulheim</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Fu</surname>
          </string-name>
          <article-title>¨mkranz. Unsupervised generation of data mining features from linked open data</article-title>
          .
          <source>In WIMS</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <given-names>A.</given-names>
            <surname>Ragone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Tomeo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Magarelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. Di</given-names>
            <surname>Noia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Palmonari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Maurino</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E. Di</given-names>
            <surname>Sciascio</surname>
          </string-name>
          .
          <article-title>Schema-summarization in linked-data-based feature selection for recommender systems</article-title>
          .
          <source>In Proceedings of the Symposium on Applied Computing, SAC</source>
          <year>2017</year>
          , Marrakech, Morocco, April 3-
          <issue>7</issue>
          ,
          <year>2017</year>
          , pages
          <fpage>330</fpage>
          -
          <lpage>335</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>J. Schaible</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Gottron</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Scherp</surname>
          </string-name>
          .
          <article-title>Termpicker: Enabling the reuse of vocabulary terms by exploiting data from the linked open data cloud</article-title>
          .
          <source>In The Semantic Web. Latest Advances and New Domains - 13th International Conference, ESWC</source>
          <year>2016</year>
          , Heraklion, Crete, Greece, May 29 - June 2,
          <year>2016</year>
          , Proceedings, pages
          <fpage>101</fpage>
          -
          <lpage>117</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Karatzoglou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Baltrunas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Larson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Oliver</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <surname>A. Hanjalic.</surname>
          </string-name>
          <article-title>CLiMF: learning to maximize reciprocal rank with collaborative less-is-more filtering</article-title>
          .
          <source>In RecSys. ACM</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <given-names>T.</given-names>
            <surname>Soru</surname>
          </string-name>
          , E. Marx,
          <article-title>and</article-title>
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Ngomo</surname>
          </string-name>
          . ROCKER:
          <article-title>A refinement operator for key discovery</article-title>
          .
          <source>In WWW. ACM</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <given-names>B.</given-names>
            <surname>Spahiu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Porrini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Palmonari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rula</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Maurino</surname>
          </string-name>
          . ABSTAT:
          <article-title>ontology-driven linked data summaries with pattern minimalization</article-title>
          .
          <source>In 2nd Workshop on Summarizing and Presenting Entities and Ontologies</source>
          , co
          <article-title>-located with ESWC</article-title>
          .,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <given-names>P.</given-names>
            <surname>Tomeo</surname>
          </string-name>
          ,
          <string-name>
            <surname>I.</surname>
          </string-name>
          <article-title>Ferna´ndez-Tob´ıas, I. Cantador, and</article-title>
          <string-name>
            <given-names>T. Di</given-names>
            <surname>Noia</surname>
          </string-name>
          .
          <article-title>Addressing the cold start with positive-only feedback through semantic-based recommendations</article-title>
          .
          <source>International Journal of Uncertainty, Fuzziness and Knowledge-Based Systems</source>
          ,
          <volume>25</volume>
          (Supplement-2):
          <fpage>57</fpage>
          -
          <lpage>78</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30. G. Troullinou,
          <string-name>
            <given-names>H.</given-names>
            <surname>Kondylakis</surname>
          </string-name>
          , E. Daskalaki, and
          <string-name>
            <given-names>D.</given-names>
            <surname>Plexousakis</surname>
          </string-name>
          . RDF Digest:
          <article-title>Efficient Summarization of RDF/S KBs</article-title>
          . In ESWC,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>