<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Semantic Network-driven News Recommender Systems: a Celebrity Gossip Use Case</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marco Fossati</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Claudio Giuliano</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giovanni Tummarello</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>fossati</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>giuliano</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>tummarellog@fbk.eu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Fondazione Bruno Kessler</institution>
          ,
          <addr-line>via Sommarive 18, 38123 Trento</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Information overload on the Internet motivates the need for ltering tools. Recommender systems play a signi cant role in such a scenario, as they provide automatically generated suggestions. In this paper, we propose a novel recommendation approach, based on semantic networks exploration. Given a set of celebrity gossip news articles, our systems leverage both natural language processing text annotation techniques and knowledge bases. Hence, real-world entities detection and cross-document entity relations discovery are enabled. The recommendations are enhanced by detailed explanations to attract end users' attention. An online evaluation with paid workers from crowdsourcing services proves the e ectiveness of our approach.</p>
      </abstract>
      <kwd-group>
        <kwd>Data Integration</kwd>
        <kwd>Natural Language Processing</kwd>
        <kwd>Information Retrieval</kwd>
        <kwd>Information Filtering</kwd>
        <kwd>Entity Linking</kwd>
        <kwd>Recommendation Strategy</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The amount of publicly available data on the World Wide Web nowadays has
dramatically increased and has led to the problem of information overload.
Recommender systems try to tackle this issue by o ering personalized suggestions.
News recommendation is a real-world application of such systems and is growing
as fast as the online news reading practice: it is estimated that, in May 2010,
57% of U.S. Internet users consumed online news by visiting news portals [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
Recently, online news consumers seem to have changed the way they access news
portals: \just a few years ago, most people arrived at our site by typing in the
website address. (...) Today the picture is very di erent. Fewer than 50% of the
8 million+ visitors to the News website every day see our front page and the
rest arrive directly at a story", a product manager of the BBC News website
a rms,1 indicating the need for news information ltering tools.
      </p>
      <p>The online reading practice leads to the so-called post-click news
recommendation problem: when a user has clicked on a news link and is reading an article,
he or she is likely to be interested in other related articles. This is still a
typical editor's task, namely an expert who manually looks for relevant content and</p>
    </sec>
    <sec id="sec-2">
      <title>1 http://www.bbc.co.uk/blogs/bbcinternet/2012/03/bbc_news_facebook_app.</title>
      <p>
        html
builds a recommendation set of links, which will be displayed below or next to the
current article. The primary aim is to keep users navigating on the visited portal.
News recommender systems attempt to automate such task. Current strategies
can be clustered into 3 main categories [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], namely (a) collaborative ltering, (b)
content-based recommendation, and (c) knowledge-based recommendation. (a)
focuses on the similarities between users of a service, thus relying on user
proles data. (b) leverages term-driven information retrieval techniques to compute
similarities between items. (c) mines external data to enrich item descriptions.
      </p>
      <p>In this paper, we propose a novel news recommendation strategy, which
leverages both natural language processing techniques and semantically structured
data. We show that entity linking tools can be coupled to existing knowledge
bases in order to compute unexpected suggestions. Such knowledge bases are
used to discover meaningful relations between entities. As a preliminary work to
assess the validity of our approach, we focus on a celebrity gossip use case and
consume data from the TMZ news portal and the Freebase graph database.2 For
instance, given a TMZ article on Michael Jackson, our strategy is able to detect
from Freebase that Michael Jackson (a) is a dead celebrity who had drug
problems and (b) dated with Brooke Shields, thus suggesting other TMZ articles on
Amy Winehouse, Kurt Cobain (other dead celebrities who had drug problems)
and Brooke Shields. We investigate if user attention can be attracted via
speci c explanations, which clarify why a given recommendation set is proposed.
Such explanations are built on top of the entity relations. Finally, we conducted
an online evaluation with real users. We outsourced a set of experiments to the
community of paid workers from Amazon's Mechanical Turk (AMT)
crowdsourcing service.3 The collected results con rm the e ectiveness of our approach.</p>
      <p>
        Our primary aim is to attract the attention of a generic user, since
postclick news recommendation generally relies on a single click user pro le data.
Therefore, we are set apart from most traditional recommender systems with
respect to three main features:
1. User agnosticity : user interests are deduced from user pro le data and
contribute to the quality of recommendations. Collecting explicit feedback is a
costly task, as it requires motivated users. Our approach gives low priority
to user pro les.
2. Unexpectedness : similarity, novelty and coherence are key components for
satisfactory news recommendations [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Content-based strategies tend to
propose too similar items and create an 'already seen' sensation. We believe
entity relations discovery can augment both novelty and coherence, thus
leading to unexpected suggestions.
3. Speci c explanation: in news web portals, generic sentences such as Related
stories or See also are typically shown together with the recommendation
set. We expect that more speci c sentences can improve the trustworthiness
of the system.
      </p>
    </sec>
    <sec id="sec-3">
      <title>2 http://www.tmz.com, http://www.freebase.com/</title>
    </sec>
    <sec id="sec-4">
      <title>3 https://www.mturk.com/mturk/welcome</title>
      <sec id="sec-4-1">
        <title>Related Work</title>
        <p>
          Content-based recommendation applies to unstructured text, such as news
articles. Document representation with bag-of-words vector space models and the
cosine similarity function still represent a valid starting point to suggest
topicrelated documents [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. Knowledge extraction from structured data is an attested
knowledge-based strategy. Linked Open Data (LOD) datasets, e.g., DBpedia4
and Freebase are queried to enrich with properties the entities extracted from
news articles [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], to collect movie information for movie schedules
recommendations [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], or to suggest music for photo albums [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. Structured data may be
also mined in order to compute similarities between items, then between user
and items [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. Content-based and knowledge-based approaches must be
combined into hybrid systems in order to achieve better results. Lasek [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] proposes a
hybrid news articles recommendation system, which merges content processing
techniques and data enrichment via LOD.
        </p>
        <p>
          Recommender systems evaluation frameworks boil down to two main
approaches [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], namely (a) o ine and (b) online. (a) leverages gold-standard datasets
and aims at estimating the performance of a recommendation algorithm via
statistical measures. (b) relies on real user studies. Ziegler et al. [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] adopt both
approaches. Hayes et al. [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] argue that user satisfaction corresponds to the
actual use of a system and can be e ectively measured only via online evaluation.
The interest in exploiting crowdsourcing services for dataset building and
online evaluation has recently grown, especially with respect to natural language
processing tasks [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] and behavioral research [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
3
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Approach</title>
        <p>Our strategy merges content-based and knowledge-based approaches and is
dened as a hybrid entity-oriented recommendation strategy enhanced by
humanreadable explanations. Given a source article from a news portal, we recommend
other articles from the portal archive, namely the corpus, by leveraging both
entity linking techniques and knowledge extraction from semantically structured
knowledge bases. Speci cally, we gathered a celebrity gossip corpus from TMZ
and chose Freebase as the knowledge base.</p>
        <p>We consider both the corpus and the knowledge base as a unique object,
namely a dataspace, which results from heterogeneous data sources integration.
Each data source is converted into an RDF graph and becomes an element of
the dataspace. Such dataspace can then be queried in order to retrieve sets of
recommendations. A semantic recommender exploits SPARQL graph navigation
capabilities to output recommendation sets. Each recommender is built on top
of a concept, e.g., substance abuse.</p>
        <p>The entity linking step in the corpus processing phase enables the
detection of both real-world entities and encyclopedic concepts. We compute concept
statistics on the whole corpus and assume that the most frequent ones are likely
to generate interesting recommendations. A mapping between corpus concepts
and meaningful relations of the knowledge base allows the creation of
recom</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4 http://dbpedia.org/</title>
      <p>menders. Table 1 shows the TMZ-to-Freebase n-ary concept mapping we
manually built. Each Freebase value represents the starting point for the construction
of a recommender, while the string after the last dot becomes the name of the
recommender, e.g., parents.</p>
      <p>Given an entity of the source article, a name of a recommender and an
entity contained in the recommendation sets, we are able to construct a speci c
explanation. Ultimately, a ranking of all the recommendation sets produces the
nal top-N suggestions output.</p>
      <p>TMZ</p>
      <p>Family
Intimate relationship</p>
      <p>Dating
Ex (relationship)</p>
      <p>Net worth
Substance abuse</p>
      <p>Conviction</p>
      <p>Court</p>
      <p>Arrest</p>
      <p>Legal case
Criminal charge</p>
      <p>Judge</p>
      <p>Death
Television program</p>
      <p>
        TMZ Processing Pipeline. Given as input a set of TMZ articles, we output an
RDF graph and load it into the dataspace. Corpus documents are harvested via
a subscription to the TMZ RSS feed. The RSS feed returns semi-structured XML
documents. A cleansing script extracts raw text from each XML document. The
entity linking step exploits The Wiki Machine,5 a state-of-the-art [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] machine
learning system designed for linking text to Wikipedia, based on a word sense
disambiguation algorithm [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. For each raw text document, real-world entities
such as persons, locations and organizations are recognized, as well as
encyclopedic concepts. This enables (a) the assignment of a unique identi er, namely
a DBpedia URI to each annotation and (b) the choice of top corpus concepts
for recommenders building purposes. The Wiki Machine takes a plain text as
input and produces an RDFa document.6 The extracted terms are assigned an
rdf:type, namely NAM for real-world entities or NOM for encyclopedic concepts.
The hasLink property connects the terms to the article URL they belong, thus
enabling the computation of the recommendation set. Other metadata, such as
      </p>
    </sec>
    <sec id="sec-6">
      <title>5 http://thewikimachine.fbk.eu</title>
      <p>6 The full corpus of TMZ RDFa documents is available at http://bit.ly/QLph9B
Knowledge base
(Freebase)</p>
      <p>Corpus (TMZ)
the link to the corresponding Wikipedia page and the annotation con dence
score are also expressed. RDFa documents are converted into RDF data via the
Any23 library.7 RDF data is loaded into a Virtuoso8 triple store instance, which
serves the dataspace for querying.</p>
      <p>Freebase Processing Pipeline. Freebase provides exhaustive granularity for
several domains, especially for celebrities. Given that such knowledge base is
large, we avoid loading its complete version, because of severe performance issues
we encountered. Consequently, meaningful slices corresponding to the corpus
domains, e.g., celebrities, people, are selected. A domain-dependent subset is
then produced via a lter written in Java. The dataset is converted into RDF
data with logic implemented in Java. Finally, RDF data is loaded into a Virtuoso
triple store instance.
4.1</p>
      <sec id="sec-6-1">
        <title>Querying the Dataspace</title>
        <p>A recommender performs a join between an entity belonging to the TMZ graph
and the corresponding entity belonging to the Freebase graph. TMZ entities are
identi ed by a DBpedia URI, which di ers from the Freebase one. Therefore,
we exploit sameAs links between DBpedia and Freebase URIs. Recommenders
are divided in two categories, namely (a) entity-driven and (b) property-driven.9
For each detected entity of the source article, we run Freebase schema inspection
queries10 and retrieve its types and properties. Thus, we are able to recognize
which recommenders can be triggered for a given entity. Building a recommender
7 http://incubator.apache.org/any23/
8 http://virtuoso.openlinksw.com/
9 The full sets are available at http://bit.ly/MWGu06 and http://bit.ly/MWGsW3
10 Available at http://bit.ly/MVGVtE
requires (a) knowledge of relevant Freebase schema parts in order to properly
browse its graph and (b) a su ciently expressive RDFa model for named
entities and link retrieval. The NAM type and the hasLink property provide such
expressivity.</p>
        <p>Entity-Driven Recommenders. The queries behind entity-driven
recommenders contain an %entity% parameter that must be programmatically lled
by an entity belonging to the source article. For instance, given an article in
which Jessica Simpson is detected and triggers the sexual relationships
recommender, we are able to return all the corpus articles (if any) that mention
entities who had sexual relationships with her, e.g., John Mayer. To avoid
running empty-result recommenders, we built a set of ASK queries,11 which check
if recommendation data exists for a given entity. The sexual relationships query
follows:
PREFIX fb: &lt;http://rdf.freebase.com/ns/&gt;
PREFIX twm: &lt;http://thewikimachine.fbk.eu#&gt;
SELECT DISTINCT ?had_relationship_with ?link
WHERE &lt;http://dbpedia.org/resource/%entity%&gt; owl:sameAs ?fb_entity .
?fb_entity fb:celebrities.celebrity.sexual_relationships ?fb_sexual_rel .
?fb_sexual_rel fb:celebrities.romantic_relationship.celebrity ?fb_celeb .
?fb_celeb fb:type.object.name ?had_relationship_with .
?dbp_celeb owl:sameAs ?fb_celeb ; a twm:NAM ; twm:hasLink ?link ; twm:hasConfidence ?conf .
FILTER (?fb_entity != ?fb_celeb) . FILTER (lang(?had_relationship_with)='en') .
ORDER BY DESC (?conf)
Property-Driven Recommenders. After the schema inspection step, an
entity of the source article can directly trigger one of these recommenders if it
contains the corresponding property. Property-driven queries return articles that
mention entities who share the same property. Hence, they do not require a
parameter to be lled. For instance, given an article in which Lindsay Lohan is
detected and the property legal entanglements is identi ed during the schema
inspection step, we can suggest other articles on people who had legal
entanglements, e.g., Britney Spears.</p>
        <p>Building Explanations. Speci c explanations are handcrafted from &lt;s, r, o&gt;
triples, where s is a subject entity that was extracted from the source article, r is
the relation expressed by the triggered recommender and o is an object entity for
which the recommendation set is computed. Therefore, we are able to construct
di erent explanations depending on the elements we use. For instance, (a) s,r,o
yields: Jessica Simpson had sexual relationships with John Mayer. Read
more about him. (b) s,r yields: Read more about Jessica Simpson's sexual
relationships. (c) r,o yields: Read more about her sexual relationships
with John Mayer.
4.2</p>
      </sec>
      <sec id="sec-6-2">
        <title>Ranking the Recommendation Sets</title>
        <p>
          Since recommendations originate from database queries, they are unranked and
in some cases too many. To overcome the problem, we implemented an
information retrieval ranking algorithm and are able to provide top-N recommendations.
11 Available at http://bit.ly/NDNORH
The bag-of-words (BOW) cosine similarity function is known to perform e
ectively for topic-related suggestions [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. However, it does not take into account
language variability. Consequently, we also leverage a latent semantic analysis
(LSA) algorithm.12 The nal score of each corpus article is the sum of BOW
and LSA scores and is assigned to the article URL. Afterwards, we run all the
recommenders and intersect their result sets with the BOW+LSA ranking of the
whole corpus, thus producing a so-called semantic ranking. This represents our
nal output, which consists of a ranked set of article URLs associated to the
corresponding recommenders names.
5
        </p>
        <sec id="sec-6-2-1">
          <title>Evaluation</title>
          <p>
            The assessment of end user satisfaction has high priority in our work.
According to Hayes et al. [
            <xref ref-type="bibr" rid="ref4">4</xref>
            ], we consequently decided to adopt an online evaluation
approach with real users. In this scenario, the major issue consists of
gathering a su ciently large group of people who are willing to evaluate our systems.
Crowdsourcing services provide a solution to the problem, as they allow us to
outsource the evaluation task to an already available massive community of paid
workers. To the best of our knowledge, no news recommender systems have been
evaluated with crowdsourcing services so far. We set up an experimental
evaluation framework for AMT, via the CrowdFlower platform.13 A description of
the mechanisms that regulate AMT is beyond the scope of the present paper:
the reader may refer to [
            <xref ref-type="bibr" rid="ref8">8</xref>
            ] for a detailed analysis.
          </p>
          <p>Our primary aim is to demonstrate that evaluators generally prefer our
recommendations. Thus, we need to put our strategy in competition with a baseline.
We leveraged the already implemented BOW+LSA information retrieval ranking
algorithm. In addition, we set two speci c objectives, related to the speci c
explanation and unexpectedness assumptions, as outlined in Section 1: (a) con rm
that a speci c explanation better attracts user attention rather than a generic
one; (b) check if the recommended items are interesting, although they may
appear unrelated and no matter what kind of explanation is provided.</p>
          <p>Quality control of the collected judgements is a key factor for the success
of the experiments. The essential drawback of crowdsourcing services relies on
the cheating risk: workers (from now on called turkers) are generally paid a few
cents for tasks which may only need a single click to be completed. Hence, it
is highly probable to collect data coming from random choices that can heavily
pollute the results. The issue is resolved by adding gold units, namely data for
which the requester already knows the answer. If a turker misses too many gold
answers within a given threshold, he or she will be agged as untrusted and his
or her judgments will be automatically discarded.
5.1</p>
        </sec>
      </sec>
      <sec id="sec-6-3">
        <title>General Setting</title>
        <p>Our evaluation framework is designed as follows: (a) the turker is invited to
read a complete news article. (b) A set of recommender systems are displayed
12 http://hlt.fbk.eu/en/technology/jlsi
13 http://crowdflower.com/
below the article. Each system consists of a natural language explanation and a
news title recommendation. (c) The turker is asked to give a preference on the
most attracting recommendation, namely the one he or she would click on in
order to read the suggested article. A single experiment (or job) is composed
of multiple data units. A unit contains the text of the article and the set of
explanation-recommendation pairs. Figure 2 shows a unit fragment of the actual
web page that is given to a turker who accepted one of our evaluation jobs. Both
instructions and question texts need to be carefully modeled, as they must mirror
the main objective of the task and should not bias turkers' reaction. Since we
aim at evaluating user attention attraction, we formulated them as per Figure 2.
Table 2 provides an overview of our experimental environment. The parameters
we have isolated for a single experiment are presented in Table 2a. On top of the
possible variations, we built a set of nine experiments, which are described in
Table 2b. We modeled two Q values, namely direct (as per Figure 2) and
indirect (Which recommendation do you consider to be more trustworthy?),
to monitor a possible alteration of turkers' reaction. Experiments having A = 5
aim at decreasing the probability a turker gets trusted by chance, because he or
she accidentally selected correct gold answers. They have an additional F value
in the Rec parameter, as we randomly extracted 3 fake recommendations per
unit from a le with more than 2 million news titles. However, such an
architectural choice generated noisy results, since it occurred that some fake titles
were selected.14 Exp is a key parameter, which allows us to check whether the
presence or the absence of a speci c explanation represents a discriminating
factor. SExp is intended to measure the e ectiveness of a speci c explanation while
reducing its complexity.</p>
        <p>Each job contains 8 regular + 2 gold units, namely 5 articles proposed
twice, in combination with 2 signi cant (and eventually 3 fake)
explanationrecommendation pairs. The recommendation titles of the regular units are
extracted from the top-2 links of the baseline and the semantic rankings. Gold is
created by extracting the title from the last, i.e., less related link of the baseline
ranking, the top link of the semantic ranking and assigning the correct answer
to the latter. We collected a minimum of 10 valid judgments per unit and set
the number of units per page to 3.</p>
        <p>Once the results obtained, it frequently occurred that the expected number
of judgments was higher: depending on their accuracy in providing answers to
gold units, turkers switched from untrusted to trusted, thus adding free extra
judgments. The proposed articles come from the TMZ website, which is well
known in the United States. Therefore, we decided to gather evaluation data
only from American turkers. The total cost of each experiment was 3.66$.</p>
        <p>After visiting some news web portals, we chose the following generic
explanations and randomly assigned them to both the baseline and the fake
recommendations: (a) The most related story selected for you; (b) If you liked
14 See Table 3 for further details.
this article, you may also like; (c) Here for you the hottest story
from a similar topic; (d) More on this story; (e) People who read this
article, also read. 2 regular units were removed from the relation only and
the object + relation experiments: it was impossible to build speci c
explanations with an implicit subject or object, since the entities that triggered the
recommendations di ered from the main entity of the source article.
5.3</p>
      </sec>
      <sec id="sec-6-4">
        <title>Results</title>
        <p>Table 3 provides an aggregated view of the results obtained from the Crowd ower
platform.15 With respect to the absolute percentage values, we rst observe
that our approach always outperformed the baseline. Furthermore, statistical
signi cance di erences emerge when a complete &lt;s; r; o&gt; speci c explanation is
given. We ran twice, i.e., in two separate days the pilot experiment and noticed an
improvement. The indirect experiment only di ers from the pilot in the question
parameter and yielded similar results. The 4 generic + 1 speci c experiment
has the highest semantic percentage: this translates into an expected behavior,
since the presence of a single speci c explanation against four generic ones is
likely to bias turkers' reaction towards our approach. As the complexity of the
speci c explanation decreases, i.e., in the subject + relation, object + relation
and relation only experiments or when only generic explanations are presented,
namely in the 5 generic and same explanation experiments, judgments towards
our approach tend to decrease too. Hence, we evince the importance of providing
speci c explanations in order to attract user attention.
5.4</p>
      </sec>
      <sec id="sec-6-5">
        <title>Discussion</title>
        <p>Experiments containing a speci c explanation aim at assessing its attractive
power (assumption 3). If we compare experiments which only di er in the Exp
parameter, namely 4 generic + 1 speci c and 5 generic, pilot 1-2 and Same
explanation, in the formers turkers prefer our strategy with a statistically
sig15 The complete set of full reports is available at http://bit.ly/MOrN30
ni cant di erence. Therefore, speci c explanations are proven to enhance the
trustworthiness of the system.</p>
        <p>The evaluation of the unexpectedness factor (assumption 2) boils down to
check whether turkers privilege the novelty of a recommendation or its similarity
to the source article. In experiments including only generic explanations, namely
Same explanation and 5 generic, we noticed the following: (a) no statistically
signi cant di erences exist between the strategies; (b) when the baseline returns
articles that are unrelated to the topic or the entity of the source article, turkers
prefer our strategy and vice versa. Hence, we argue that users tend to privilege
similarity if they are given a generic explanation. On the other hand, when the
baseline strategy suggests a clearly related article and when a speci c explanation
is provided, turkers tend to choose our strategy even if it suggests an apparently
unrelated article. This is a rst proof of the unexpectedness factor: users are
attracted by the speci c explanation and are eager to read an unexpected article
rather than another article on the same topic/entity.
6</p>
        <sec id="sec-6-5-1">
          <title>Conclusion</title>
          <p>
            In this paper, we presented a novel recommendation strategy leveraging entity
linking techniques in unstructured text and knowledge extraction from
structured knowledge bases. On top of it, we build hybrid entity-oriented
recommender systems for news ltering and post-click news recommendation. We
argued that entity relations discovery leads to unexpected suggestions and speci c
explanations, thus attracting user attention. The adopted online evaluation
approach via crowdsourcing services assessed the validity of our systems. A demo
prototype consumes Freebase data to recommend TMZ celebrity gossip articles
and can be viewed at http://spaziodati.eu/widget_recommendation/. For
our future work, we have set the following milestones:
1. Ecological evaluation. AMT allowed us to build fast and cheap online
evaluation experiments. However, the collected judgments may be biased by
the politeness e ect of the economical reward and the turkers' awareness
of performing a question-answering task. Therefore, we intend to set up an
ecological evaluation scenario, which simulates a real-world usage of our
recommender systems and enables natural user reactions. We will adopt the
Google AdWords16 approach proposed by Guerini et al. [
            <xref ref-type="bibr" rid="ref3">3</xref>
            ].
2. Methodology for building recommenders. Currently, we have manually
implemented a domain-speci c list of recommenders, based on the most frequent
corpus concepts. We plan to automate this process by extracting generic
relations from Freebase via data analytics techniques.
3. Methodology for building speci c explanations. Explanations are naively mapped
to the relations and the corresponding subject/object entities. How to
automatically build linguistically correct sentences remains an open problem.
4. User pro le construction. Explicit and implicit user preferences acquisition
can improve the quality of the recommendations. Our demo page may serve
16 http://adwords.google.com/
as a platform for gathering such data. Otherwise, we may adapt our systems
to datasets containing user ratings.
          </p>
          <p>Acknowledgements. This work was supported by the EU project
Eurosentiment, contract number 296277.</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Chao</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , Zhang,
          <string-name>
            <given-names>W.</given-names>
            ,
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y.</surname>
          </string-name>
          :
          <article-title>Tunesensor: A semantic-driven music recommendation service for digital photo albums</article-title>
          .
          <source>In: Proceedings of the 10th International Semantic Web Conference. ISWC2011 (October</source>
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Giuliano</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gliozzo</surname>
            ,
            <given-names>A.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Strapparava</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Kernel methods for minimally supervised wsd</article-title>
          .
          <source>Computational Linguistics</source>
          <volume>35</volume>
          (
          <issue>4</issue>
          ),
          <volume>513</volume>
          {
          <fpage>528</fpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Guerini</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Strapparava</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stock</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Ecological evaluation of persuasive messages using google adwords</article-title>
          .
          <source>In: Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics. ACL2012</source>
          , vol.
          <source>abs/1204.5369 (July</source>
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Hayes</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cunningham</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Massa</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>An on-line evaluation framework for recommender systems</article-title>
          .
          <source>Tech. Rep. TCD-CS-2002-19</source>
          , Trinity College Dublin, Department of Computer Science (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Jannach</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zanker</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Felfernig</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Friedrich</surname>
          </string-name>
          , G.:
          <source>Recommender Systems: An Introduction</source>
          . Cambridge University Press (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Lasek</surname>
          </string-name>
          , I.:
          <article-title>Dc proposal: Model for news ltering with named entities</article-title>
          . In: Aroyo,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Welty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Alani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Taylor</surname>
          </string-name>
          , J.,
          <string-name>
            <surname>Bernstein</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kagal</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Noy</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blomqvist</surname>
          </string-name>
          , E. (eds.)
          <source>The Semantic Web { ISWC 2011, Lecture Notes in Computer Science</source>
          , vol.
          <volume>7032</volume>
          , pp.
          <volume>309</volume>
          {
          <fpage>316</fpage>
          . Springer Berlin / Heidelberg (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Lv</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moon</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kolari</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zheng</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Learning to model relatedness for news recommendation</article-title>
          .
          <source>In: Proceedings of the 20th international conference on World wide web</source>
          . pp.
          <volume>57</volume>
          {
          <fpage>66</fpage>
          . WWW '11,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Mason</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suri</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Conducting behavioral research on amazon's mechanical turk</article-title>
          .
          <source>Behavior Research Methods</source>
          <volume>44</volume>
          ,
          <issue>1</issue>
          {
          <fpage>23</fpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Mendes</surname>
            ,
            <given-names>P.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jakob</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garc</surname>
            a-Silva,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Dbpedia spotlight: shedding light on the web of documents</article-title>
          .
          <source>In: Proceedings of the 7th International Conference on Semantic Systems</source>
          . pp.
          <volume>1</volume>
          {
          <issue>8</issue>
          . I-Semantics '
          <fpage>11</fpage>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Negri</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bentivogli</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mehdad</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giampiccolo</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marchetti</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Divide and conquer: crowdsourcing the creation of cross-lingual textual entailment corpora</article-title>
          .
          <source>In: Proceedings of the Conference on Empirical Methods in Natural Language Processing</source>
          . pp.
          <volume>670</volume>
          {
          <fpage>679</fpage>
          . EMNLP '
          <volume>11</volume>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computational Linguistics, Stroudsburg, PA, USA (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Pazzani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Billsus</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Content-based recommendation systems</article-title>
          . In: Brusilovsky,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Kobsa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Nejdl</surname>
          </string-name>
          , W. (eds.)
          <source>The Adaptive Web, Lecture Notes in Computer Science</source>
          , vol.
          <volume>4321</volume>
          , pp.
          <volume>325</volume>
          {
          <fpage>341</fpage>
          . Springer Berlin / Heidelberg (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Thalhammer</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ermilov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nyberg</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Santoso</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Domingue</surname>
          </string-name>
          , J.:
          <article-title>Moviegoer - semantic social recommendations and personalized location-based o ers</article-title>
          .
          <source>In: Proceedings of the 10th International Semantic Web Conference. ISWC2011 (October</source>
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Ziegler</surname>
            ,
            <given-names>C.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lausen</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmidt-Thieme</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Taxonomy-driven computation of product recommendations</article-title>
          .
          <source>In: Proceedings of the thirteenth ACM international conference on Information and knowledge management</source>
          . pp.
          <volume>406</volume>
          {
          <fpage>415</fpage>
          . CIKM '04,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>