<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>PhD Workshop, August</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Zainab Zolaktaf Supervised by Rachel Pottinger University of British Columbia Vancouver</institution>
          ,
          <addr-line>B.C</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <volume>28</volume>
      <issue>2017</issue>
      <abstract>
        <p>In many domains, such as scienti c computing, users can directly access and query data that is stored in large, and often structured, data sources. Discovering interesting patterns and e ciently locating relevant information, however, can be challenging. Users must be aware of the data content and its structure, before they can query it. Furthermore, they have to interpret the retrieved results and possibly re ne their query. Essentially, to nd information, the user has to engage in a repeated cycle of data exploration, query composition, and query answer analysis. The focus of my PhD research is on designing techniques that facilitate this interaction. Speci cally, I examine the utility of recommender systems for the data exploration and query composition phases, and propose techniques that assist users in the query answer analysis phase. Overall, the solutions developed in my thesis aim to increase the e ciency and decision quality of users.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        With the advent of technology and the web, large volumes
of data are generated and stored in data sources that evolve
and grow over time. Often, these sources are structured as
relational databases that users can directly query and
explore. For instance, astronomical measurements are stored
in a large relational database, called the Sloan Digital Sky
Survey (SDSS) [
        <xref ref-type="bibr" rid="ref19 ref21">19, 21</xref>
        ]. Climate data collected from various
sources is integrated in relational databases and o ered for
analysis by users [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>At a high level, user interaction with data involves two
phases: a query composition phase, where the user
composes and submits a query, and a query answer analysis
phase, where the user analyses query answers produced by
the system. During both phases, however, users can face
problems in understanding the data.</p>
      <p>
        Consider, for example, the scienti c computing domain.
The SDSS schema has over 88 tables, 51 views, 204
userde ned functions, and 3440 columns [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. A variety of users,
ranging from high school students to professional astronomers,
with varied levels of skills and knowledge, interact with this
database. Furthermore, scienti c databases are typically
used for Interactive Data Exploration (IDE), where users
pose exploratory queries to understand the content and nd
patterns [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. E ciently composing queries over this data to
discover interesting patterns, is one of their main challenges.
      </p>
      <p>After successfully composing the query, the next step is
to interpret query answers. However, the retrieved results
can often be di cult to understand. For example, consider
an aggregate query SELECT AVG(TEMPERATURE) over climate
data. In the weather domain, observational data
regarding atmospheric conditions is collected by several weather
stations, satellites, and ships. For the same data point, e.g.,
temperature on a given day, there can be con icting and
duplicate values. Consequently, the aggregate query can have
an overwhelming number of correct and con icting answers.
Here, mechanisms that aid the user in understanding the
query answers are required.</p>
      <p>In my thesis, I develop techniques that assist user
interaction with data. I consider the data exploration and query
composition phase, and examine the utility of
recommendation systems for this phase. Furthermore, I consider the
query answer analysis phase and devise e cient techniques
that provide insights about query answers. More precisely,
I study three problems: 1. how do classical
recommendation systems perform with regards to exploration tasks in
standard recommendation domains, and how can we
modify them to facilitate data exploration more rigorously
(Section 2)? 2. what are the challenges of recommendation in the
relational database context and which algorithms are
appropriate for helping users explore data and compose queries
(Section 3)? 3. how can we assist users in the query
answer analysis phase (Section 4)? Overall, I aim to develop
techniques that help users explore data and increase their
decision quality.
2.</p>
    </sec>
    <sec id="sec-2">
      <title>FACILITATING DATA EXPLORATION WITH</title>
    </sec>
    <sec id="sec-3">
      <title>RECOMMENDER SYSTEMS</title>
      <p>
        One way to facilitate data navigation and exploration is to
nd and suggest items of interest to users by deploying a
recommendation system [
        <xref ref-type="bibr" rid="ref16 ref29 ref6">6, 16, 29</xref>
        ]. Classical recommendation
systems are categorized into content-based and collaborative
ltering methods.
      </p>
      <p>Content-based methods use descriptive features such as
genre of movies, or user demographics, to construct
informative user and item pro les, and measure similarity between
them. But descriptive features might not be available.
Collaborative ltering methods instead infer user interests from
user interaction data. The main intuition is that users with
similar interaction patterns have similar interests.</p>
      <p>
        The interaction data may include explicit user feedback on
items, such as user ratings on movies, or implicit feedback,
such as purchasing history, browsing and click logs, or query
logs [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. An important property of the interaction data is
that the majority of items (users) receive (provide) little
feedback and are infrequent, while a few receive (provide)
lots of feedback and are frequent. But many models only
work well when there is a lot of data available, i.e., they make
good recommendations for frequent users, and are biased
toward recommending frequent items [
        <xref ref-type="bibr" rid="ref15 ref17 ref6">6, 15, 17</xref>
        ].
      </p>
      <p>However, recommending popular items is not su cient for
exploratory tasks. Users are likely already aware of
popular items or can nd them on their own. Concentrating on
popular items also means the system has low overall
coverage of the item space in its recommendations. It is essential
to develop methods that help users discover new items that
may be less common but more interesting. Therefore, we
investigate the following research question:</p>
      <p>
        How do existing recommendation models perform with
regard to data exploration tasks in standard
recommendation domains, and how can they be modi ed to
facilitate data exploration more rigorously?
To answer this question, we focus on top-N item
recommendation, where the goal is to recommend the most appealing
set of N items to each user [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Informally, the problem
setting is as follows: we are given a log of explicit user feedback,
e.g., ratings, for di erent items. We want to assign a set of
N unseen items to each user.
2.1
      </p>
    </sec>
    <sec id="sec-4">
      <title>Solution</title>
      <p>
        In our solution [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], we focused on promoting less
frequent items, or long-tail items, in top-N sets to facilitate
exploration. Recommending these items introduces novelty
and serendipity into top-N sets, and allows users to discover
new items. It also increases the item-space coverage, which
increases pro ts for providers of the items [
        <xref ref-type="bibr" rid="ref22 ref26 ref3 ref6">3, 6, 26, 22</xref>
        ].
Our main challenge was in promoting long-tail items in a
targeted manner, and in designing responsive and scalable
models. We used historical rating data to learn user
preference for discovering new items. The main intuition was
that the long-tail preference of user u, captured by u,
depends on the types of long-tail items she rates. Moreover,
the long-tail type or weight of item i, captured by wi,
depends on the long-tail preference of users who rate that item.
Based on this, we formulated a joint optimization objective
for learning both unknown variables, and w.
      </p>
      <p>Next, we integrated the learned user preference estimates,
, into a generic re-ranking framework to provide customized
balance between accuracy and coverage. Speci cally, we
dened a re-ranking framework that required three
components: 1. an accuracy recommender that was responsible for
recommending accurate top-N sets. 2. a coverage
recommender that was responsible for suggesting top-N sets that
maximized coverage across the item space, and consequently
promoted long-tail items. 3. the user long-tail preference.</p>
      <p>
        In contrast to prior related work [
        <xref ref-type="bibr" rid="ref1 ref10 ref27">1, 10, 27</xref>
        ], our
framework learned the personalization rather than optimizing
using cross-validation or parameter tuning; in other words, our
0K PMoFp [[268]]
-20 5D ACC [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]
T Co R [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]
M PureSVD [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
      </p>
      <p>
        Dyn900 [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]
Pop
personalization method was independent of the underlying
recommendation model.
      </p>
      <p>
        We evaluated our framework on several standard datasets
from the movie domain. Table 1 shows the top-5
recommendation performance for the MovieTweetings 200K
(MT200K) dataset [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] which contains voluntary movie rating
tweets from users. For accuracy, we computed precision
(P@5) and recall (R@5) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] wrt the test items of users.
Longtail accuracy (L@5) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], is the normalized number of
longtail items in top-5 sets per user. Long-tail items are those
that generate the lower 20% of the total ratings in the train
set, based on the Pareto principle or the 80=20 rule [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ].
Coverage (C@5) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] is the ratio of the number of distinct
items recommended to all users, to the number of items.
      </p>
      <p>
        We compared with non-personalized baselines: Random
that has high coverage but low accuracy, and most
popular recommendation (Pop) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], that provides accurate top-N
sets but has low coverage and long-tail accuracy. We also
compared with personalized algorithms: matrix
factorization (MF) with 40 factors, L2-regularization, and stochastic
gradient descent optimization [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ], a resource allocation
approach that re-ranks MF (5D ACC) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], Co Rank with
regression loss (Co R) [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], and PureSVD with 300 factors [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
On MT-200K, we chose the non-personalized Pop algorithm
as our accuracy recommender, and combined it with a
dynamic coverage recommender (Dyn900) introduced in [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
Our personalized algorithm is denoted PDoypn900. Table 1
shows that while most baselines achieve best performance in
either coverage or accuracy metrics, PDoypn900 has high
coverage, while maintaining reasonable accuracy levels.
Furthermore, it outperforms the personalized algorithms, PureSVD
and Co R, in both accuracy and coverage metrics.
3.
      </p>
    </sec>
    <sec id="sec-5">
      <title>FACILITATING DATA EXPLORATION AND</title>
    </sec>
    <sec id="sec-6">
      <title>QUERY COMPOSITION</title>
      <p>
        Getting information out of database systems is a major
challenge [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Users must be familiar with the schema to be
able to compose queries. Some relational database systems,
e.g, SkyServer, provide a sample of example queries to aid
users with this task. However, compared to the size of the
database and complexity of potential queries, this sample
set is small and static. The problem is exacerbated as the
volume of data increases, particularly for IDE. A mechanism
that helps users navigate the schema and data space, and
exposes relevant data regions based on their query context,
is required. We consider using recommendation systems in
this setting and focus on the following research question:
What are the challenges of recommendation in the
database context, and which algorithms are suitable
for facilitating interactive exploration and navigation
of relational databases?
      </p>
      <p>To answer this question, we address top-N aspect
recommendation, where the goal is to suggest a set of N aspects
to the user that facilitate query composition and database
exploration. Similar to the collaborative ltering setting in
Section 2, we analyse user interaction data, available in a
query log. Informally, the problem setting is as follows: we
are given a query log that is partitioned into sessions, sets
of queries submitted by the same user. Furthermore, we
also have a relational database synopsis with information
about the schema of the database (#relations, #attributes,
and foreign key constraints) and the range of numerical
attributes. Given a new partial session, the objective is to
recommend potential query extensions, or aspects.
3.1</p>
    </sec>
    <sec id="sec-7">
      <title>Proposed Work</title>
      <p>To formulate an adequate solution, the following
challenges must be addressed:
1. Aspect De nition. There is no clear notion of \item"
or aspect in this setting. Instead, we need to nd an
adequate set of aspects that can be used to to
capture user intent and characterize queries. Given the
exploratory nature of queries in the scienti c domains,
the aspects should enable both schema navigation and
data space exploration.
2. Sequential Aspects and Domain-Speci c Constraints.</p>
      <p>
        Individual elements in a SQL query are sequential and
there is dependency between them. For instance, in
SELECT T.A FROM T WHERE X &gt; 10, the domain of
variable X is attributes in table T. Thus, given partial
query, only a subset of the aspects are syntactically
valid. Queries in the same session, are also submitted
sequentially.
3. Session and Aspect Sparsity. In SDSS, the typical
session has six SQL queries and lasts thirty minutes [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]
which indicates aspect sparsity in queries and sessions.
The relational database setting exhibits some similarities
to standard recommendation domains (e.g., movie): Some
aspects, e.g., tables, attributes, data regions, are popular
while the majority of them are unpopular. Some sessions
are frequent, i.e., many queries are submitted, while the
majority are infrequent. Scalability and responsiveness is
important in both domains.
      </p>
      <p>Analogous to our work in Section 2, our main hypothesis
is that merely recommending popular aspects is not su
cient for exploratory tasks. Although popular aspects can
help familiarize novice users with concepts like the
important tables and attributes, given the exploratory nature of
queries in IDE, recommendations are deemed more useful if
they can help users narrow down their queries and expose
relevant data regions. For example, recommending a
speci c interval like b1 &lt; BRIGHTNESS &lt; b2 is more useful than
just suggesting the attribute BRIGHTNESS.</p>
      <p>
        Based on these intuitions, we will focus on recommending
interesting aspects that enable data exploration and schema
navigation for users of a relational database, and in
particular, in IDE settings. Using the query log and the database
synopsis, we will devise a set of aspects that include not just
the relations, attributes, and user-de ned functions, but also
intervals of numeric attributes, e.g., b1 &lt; BRIGHTNESS &lt; b2.
Subsequently, we can use a vector-based query
representation model where each element denotes the presence of a
certain aspect. Alternatively, a graph-based representation [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]
might be more suitable. After formulating similarity
measures between queries (or sessions) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], we can use a nearest
neighbour model to suggest relevant aspects to the user.
      </p>
      <p>
        In contrast to prior work that focuses on supervised
learning and query rewriting [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], we focus on aspect de nition
and extraction. In contrast to [
        <xref ref-type="bibr" rid="ref4 ref7 ref8">4, 7, 8</xref>
        ], we rely on the
database synopsis only. Accessing a large scienti c database
like SDSS to retrieve the entire set of tuples is expensive. In
contrast to [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] our recommendations include intervals not
just tables and attributes. The intermediate query format
in [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] is complementary to our work.
4.
      </p>
    </sec>
    <sec id="sec-8">
      <title>FACILITATING QUERY ANSWER ANAL</title>
    </sec>
    <sec id="sec-9">
      <title>YSIS</title>
      <p>After users have successfully submitted a query, their next
challenge is to analyse and understand the query answers.
When the answer set is small, this task is attainable. The
challenge is in examining and interpreting large, or even
con icting, answer sets.</p>
      <p>To illustrate the problem, consider again climate data
that is reported by various sources and integrated in
relational databases. Because the sources were independently
created and maintained, a given data point can have
multiple, inconsistent values across the sources. For example,
one source may have the high temperature for Vancouver
on 06/11/2006 as 17C, while another may list it as 19C.
As a result of this value-level heterogeneity, an aggregate
query such as SELECT AVG(TEMPERATURE) does not have a
single true answer. Instead, depending on the choice of data
source combinations that are used to answer the query,
different answers can be generated. Reporting the entire set
of answers can overwhelm the user. Here, mechanisms that
summarize the results and help the user understand query
answers are required. Therefore, we study the following
research question:</p>
      <p>
        After a query has been submitted to the system, how
can we help the user understand and interpret the
query answers?
Speci cally, we address the problem of helping users
understand aggregate query answers in integration contexts
where data is segmented across several sources. We assume
meta-information that describes the mappings and bindings
between data sources is available [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ].Our main concern is
how to handle the value-level heterogeneity that exists in
the data, to enable the user to better understand the range
of possible query answers.
4.1
      </p>
    </sec>
    <sec id="sec-10">
      <title>Solution</title>
      <p>
        In our solution [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ], we represented the answer to the
aggregate query as an answer distribution instead of a single
scalar value. We then proposed a suite of methods for
extracting statistics that convey meaningful information about
the query answers. We focused on the following challenges
1. determining which statistics best represent and answer's
distribution 2. e ciently computing the desired statistics.
In deriving our algorithms, we assumed prior knowledge
regarding the sources is unavailable and all sources are equal.
      </p>
      <p>
        A high coverage interval is one of the statistics we
extract to convey the shape of the answer distribution and
00.6 0.7 0.8 0.9 1 1.1 1.2 1.3 1.4 1.5 00.6 0.8
x104
1
1.2 1.4 1.6 1.8
x104
(a) S1
(b) S4
the intervals where the majority of viable answers can be
found. Figure 1 shows the multi-modal answer distributions
of the aggregate query AVG(TEMP), on Canadian climate data
(S1) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and synthetic data (S4) [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ], and their corresponding
high coverage intervals.
      </p>
    </sec>
    <sec id="sec-11">
      <title>5. SUMMARY AND OUTLOOK</title>
      <p>The goal of my thesis is to devise techniques that facilitate
user interaction with data. I address three aspects:
(Accomplished) Facilitating data exploration with
recommender systems in standard domains (Section 2).
(In progress) Facilitating data exploration and query
composition in the relational database context
(Section 3). I am currently working on extracting a dataset,
and narrowing down the problem statement.
(Accomplished) Facilitating query answer analysis by
extracting statistics and semantics about the range of
query answers (Section 4).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Gediminas</given-names>
            <surname>Adomavicius and YoungOk Kwon</surname>
          </string-name>
          .
          <article-title>Improving aggregate recommendation diversity using ranking-based techniques</article-title>
          .
          <source>TKDE</source>
          ,
          <volume>24</volume>
          (
          <issue>5</issue>
          ):
          <volume>896</volume>
          {
          <fpage>911</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Julien</given-names>
            <surname>Aligon</surname>
          </string-name>
          , Matteo Golfarelli, Patrick Marcel, Stefano Rizzi, and
          <string-name>
            <given-names>Elisa</given-names>
            <surname>Turricchia</surname>
          </string-name>
          .
          <article-title>Similarity measures for olap sessions</article-title>
          .
          <source>Knowledge and information systems</source>
          ,
          <volume>39</volume>
          (
          <issue>2</issue>
          ):
          <volume>463</volume>
          {
          <fpage>489</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Pablo</given-names>
            <surname>Castells</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Neil J.</given-names>
            <surname>Hurley</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Saul</given-names>
            <surname>Vargas</surname>
          </string-name>
          .
          <source>Recommender Systems Handbook, chapter Novelty and Diversity in Recommender Systems</source>
          . Springer US,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Gloria</given-names>
            <surname>Chatzopoulou</surname>
          </string-name>
          , Magdalini Eirinaki, and
          <string-name>
            <given-names>Neoklis</given-names>
            <surname>Polyzotis</surname>
          </string-name>
          .
          <article-title>Query recommendations for interactive database exploration</article-title>
          .
          <source>In International Conference on Scienti c and Statistical Database Management</source>
          , pages
          <fpage>3</fpage>
          <lpage>{</lpage>
          18. Springer,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Climate</given-names>
            <surname>Canada</surname>
          </string-name>
          .
          <article-title>Canada climate data</article-title>
          . http://climate. weatheroffice.gc.ca/climateData/canada_e.html,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Cremonesi</surname>
          </string-name>
          , Yehuda Koren, and
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Turrin</surname>
          </string-name>
          .
          <article-title>Performance of recommender algorithms on top-n recommendation tasks</article-title>
          .
          <source>In RecSys</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Julien</given-names>
            <surname>Cumin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Jean-Marc</surname>
            <given-names>Petit</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vasile-Marian Scuturici</surname>
            , and
            <given-names>Sabina</given-names>
          </string-name>
          <string-name>
            <surname>Surdu</surname>
          </string-name>
          .
          <article-title>Data exploration with sql using machine learning techniques</article-title>
          .
          <source>In EDBT</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Kyriaki</given-names>
            <surname>Dimitriadou</surname>
          </string-name>
          , Olga Papaemmanouil, and
          <string-name>
            <given-names>Yanlei</given-names>
            <surname>Diao</surname>
          </string-name>
          .
          <article-title>Explore-by-example: An automatic query steering framework for interactive data exploration</article-title>
          .
          <source>In Proceedings of the 2014 ACM SIGMOD international conference on Management of data</source>
          , pages
          <volume>517</volume>
          {
          <fpage>528</fpage>
          . ACM,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Simon</given-names>
            <surname>Dooms</surname>
          </string-name>
          , Toon De Pessemier, and
          <string-name>
            <given-names>Luc</given-names>
            <surname>Martens</surname>
          </string-name>
          .
          <article-title>Movietweetings: a movie rating dataset collected from twitter</article-title>
          .
          <source>In CrowdRec at RecSys</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Yu-Chieh</surname>
            <given-names>Ho</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yi-Ting Chiang</surname>
          </string-name>
          , and
          <string-name>
            <surname>Jane</surname>
            <given-names>Yung-Jen</given-names>
          </string-name>
          <string-name>
            <surname>Hsu</surname>
          </string-name>
          .
          <article-title>Who likes it more?: mining worth-recommending items from long tails by modeling relative preference</article-title>
          .
          <source>In WSDM</source>
          , pages
          <volume>253</volume>
          {
          <fpage>262</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Yifan</surname>
            <given-names>Hu</given-names>
          </string-name>
          , Yehuda Koren, and
          <string-name>
            <given-names>Chris</given-names>
            <surname>Volinsky</surname>
          </string-name>
          .
          <article-title>Collaborative ltering for implicit feedback datasets</article-title>
          .
          <source>In 2008 Eighth IEEE International Conference on Data Mining</source>
          , pages
          <volume>263</volume>
          {
          <fpage>272</fpage>
          . IEEE,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>HV</given-names>
            <surname>Jagadish</surname>
          </string-name>
          , Adriane Chapman, Aaron Elkiss, Magesh Jayapandian,
          <string-name>
            <given-names>Yunyao</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Arnab</given-names>
            <surname>Nandi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Cong</given-names>
            <surname>Yu</surname>
          </string-name>
          .
          <article-title>Making database systems usable</article-title>
          .
          <source>In Proceedings of the 2007 ACM SIGMOD international conference on Management of data</source>
          , pages
          <volume>13</volume>
          {
          <fpage>24</fpage>
          . ACM,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Martin</surname>
            <given-names>L Kersten</given-names>
          </string-name>
          , Stratos Idreos, Stefan Manegold,
          <string-name>
            <given-names>Erietta</given-names>
            <surname>Liarou</surname>
          </string-name>
          , et al.
          <article-title>The researchers guide to the data deluge: Querying a scienti c database in just a few seconds</article-title>
          .
          <source>PVLDB Challenges and Visions</source>
          ,
          <volume>3</volume>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Nodira</surname>
            <given-names>Khoussainova</given-names>
          </string-name>
          , YongChul Kwon, Magdalena Balazinska, and
          <string-name>
            <given-names>Dan</given-names>
            <surname>Suciu</surname>
          </string-name>
          .
          <article-title>Snipsuggest: context-aware autocompletion for sql</article-title>
          .
          <source>Proceedings of the VLDB Endowment</source>
          ,
          <volume>4</volume>
          (
          <issue>1</issue>
          ):
          <volume>22</volume>
          {
          <fpage>33</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Joonseok</given-names>
            <surname>Lee</surname>
          </string-name>
          , Samy Bengio, Seungyeon Kim, Guy Lebanon, and
          <string-name>
            <given-names>Yoram</given-names>
            <surname>Singer</surname>
          </string-name>
          .
          <article-title>Local collaborative ranking</article-title>
          .
          <source>In WWW</source>
          , pages
          <volume>85</volume>
          {
          <fpage>96</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Joonseok</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Mingxuan</given-names>
            <surname>Sun</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Guy</given-names>
            <surname>Lebanon</surname>
          </string-name>
          .
          <article-title>A comparative study of collaborative ltering algorithms</article-title>
          .
          <source>arXiv preprint arXiv:1205.3193</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>Andriy</given-names>
            <surname>Mnih</surname>
          </string-name>
          and
          <string-name>
            <given-names>Ruslan</given-names>
            <surname>Salakhutdinov</surname>
          </string-name>
          .
          <article-title>Probabilistic matrix factorization</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          , pages
          <volume>1257</volume>
          {
          <fpage>1264</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Hoang</given-names>
            <surname>Vu</surname>
          </string-name>
          <string-name>
            <surname>Nguyen</surname>
          </string-name>
          , Klemens Bohm, Florian Becker, Bertrand Goldman, Georg Hinkel, and
          <article-title>Emmanuel Muller. Identifying user interests within the data space-a case study with skyserver</article-title>
          .
          <source>In EDBT</source>
          , pages
          <volume>641</volume>
          {
          <fpage>652</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>M</given-names>
            <surname>Jordan Raddick</surname>
          </string-name>
          , Ani R Thakar, Alexander S Szalay, and Rafael DC Santos.
          <article-title>Ten years of skyserver i: Tracking web and sql e-science usage</article-title>
          .
          <source>Computing in Science &amp; Engineering</source>
          ,
          <volume>16</volume>
          (
          <issue>4</issue>
          ):
          <volume>22</volume>
          {
          <fpage>31</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <article-title>Information removed for double-blind review</article-title>
          . Submitted paper,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Vik</surname>
            <given-names>Singh</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jim Gray</surname>
          </string-name>
          , Ani Thakar, Alexander S Szalay, Jordan Raddick, Bill Boroski, Svetlana Lebedeva, and
          <string-name>
            <given-names>Brian</given-names>
            <surname>Yanny</surname>
          </string-name>
          .
          <article-title>Skyserver tra c report-the rst ve years</article-title>
          .
          <source>arXiv preprint cs/0701173</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>Saul</given-names>
            <surname>Vargas</surname>
          </string-name>
          and
          <string-name>
            <given-names>Pablo</given-names>
            <surname>Castells</surname>
          </string-name>
          .
          <article-title>Improving sales diversity by recommending users to items</article-title>
          . In RecSys,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Roy</surname>
            <given-names>Villafane</given-names>
          </string-name>
          , Kien A Hua,
          <string-name>
            <given-names>Duc</given-names>
            <surname>Tran</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Basab</given-names>
            <surname>Maulik</surname>
          </string-name>
          .
          <article-title>Mining interval time series</article-title>
          .
          <source>In International Conference on Data Warehousing and Knowledge Discovery</source>
          , pages
          <volume>318</volume>
          {
          <fpage>330</fpage>
          . Springer,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Markus</surname>
            <given-names>Weimer</given-names>
          </string-name>
          , Alexandros Karatzoglou, Quoc Viet Le, and
          <string-name>
            <given-names>Alex</given-names>
            <surname>Smola</surname>
          </string-name>
          .
          <article-title>Maximum margin matrix factorization for collaborative ranking</article-title>
          .
          <source>Advances in neural information processing systems</source>
          , pages
          <fpage>1</fpage>
          <issue>{8</issue>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>Jian</given-names>
            <surname>Xu</surname>
          </string-name>
          and
          <string-name>
            <given-names>Rachel</given-names>
            <surname>Pottinger</surname>
          </string-name>
          .
          <article-title>Integrating domain heterogeneous data sources using decomposition aggregation queries</article-title>
          .
          <source>Information Systems</source>
          ,
          <volume>39</volume>
          (
          <issue>0</issue>
          ),
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Hongzhi</surname>
            <given-names>Yin</given-names>
          </string-name>
          , Bin Cui,
          <string-name>
            <given-names>Jing</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Junjie</given-names>
            <surname>Yao</surname>
          </string-name>
          , and
          <string-name>
            <surname>Chen</surname>
            <given-names>Chen</given-names>
          </string-name>
          .
          <article-title>Challenging the long tail recommendation</article-title>
          .
          <source>PVLDB</source>
          ,
          <volume>5</volume>
          (
          <issue>9</issue>
          ):
          <volume>896</volume>
          {
          <fpage>907</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Weinan</surname>
            <given-names>Zhang</given-names>
          </string-name>
          , Jun Wang,
          <string-name>
            <surname>Bowei Chen</surname>
            , and
            <given-names>Xiaoxue</given-names>
          </string-name>
          <string-name>
            <surname>Zhao</surname>
          </string-name>
          .
          <article-title>To personalize or not: a risk management perspective</article-title>
          .
          <source>In RecSys</source>
          , pages
          <volume>229</volume>
          {
          <fpage>236</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Yong</surname>
            <given-names>Zhuang</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wei-Sheng</surname>
            <given-names>Chin</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu-Chin Juan</surname>
          </string-name>
          , and
          <string-name>
            <surname>Chih-Jen Lin</surname>
          </string-name>
          .
          <article-title>A fast parallel sgd for matrix factorization in shared memory systems</article-title>
          . In RecSys, pages
          <volume>249</volume>
          {
          <fpage>256</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>Sedigheh</given-names>
            <surname>Zolaktaf and Gail C Murphy</surname>
          </string-name>
          .
          <article-title>What to learn next: recommending commands in a feature-rich environment</article-title>
          .
          <source>In ICMLA</source>
          , pages
          <volume>1038</volume>
          {
          <fpage>1044</fpage>
          . IEEE,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <surname>Zainab</surname>
            <given-names>Zolaktaf</given-names>
          </string-name>
          , Jian Xu,
          <string-name>
            <given-names>and Rachel</given-names>
            <surname>Pottinger</surname>
          </string-name>
          .
          <article-title>Extracting aggregate answer statistics for integration</article-title>
          .
          <source>EDBT</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>