<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Exploring the “Double-Edged Sword” Efect of Auto-Insight Recom mendation in Exploratory Data Analysis</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Quan Li</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Huanbin Lin</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chunfeng Tang</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xiguang Wei</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zhenhui Peng</string-name>
          <email>zpengab@connect.ust.hk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xiaojuan Ma</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tianjian Cheng</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>AI Group</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>WeBank</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>AI Group</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>WeBank</string-name>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>College Station, USA</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>AI Department, Shenzhen Semacare Medical Technology Co</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Computer Science and Engineering, The Hong Kong University of Science and Technology</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>School of Information Science and Technology, ShanghaiTech University</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Modern data analytics tools often provide visualizations as an accessible data window to users in exploratory data analysis (EDA). Still, many analysts feel lost in this process due to issues such as the high complexity of data. Auto-insight recommendations ofer a promising alternative by suggesting possible interpretations of the data to users during EDA but might impose undesirable efects on users. In this study, we systematically explore the “double-edged sword” efect of auto-insight recommendations on EDA in terms of exploration assistance, message reliability, and interference. Particularly, we design and develop two versions of a Tableau-like visualization system termed TurboVis: one supports auto-insight recommendations while the other does not. We first demonstrate how typical visualization specification tools can be augmented by incorporating auto-insight recommendations and then conduct a within-subjects user study with 18 participants during which they experience both versions in EDA tasks. We find that although auto-insight recommendations encourage more visualization inspections, they also introduce biases to data exploration. The perceived level of message reliability and interference of auto-insight recommendations depend on data familiarity and task structures. Our work elicits design implications for embedding auto-insight recommendations into the EDA process.</p>
      </abstract>
      <kwd-group>
        <kwd>Analysis</kwd>
        <kwd>Visualization recommendations</kwd>
        <kwd>exploratory data analysis</kwd>
        <kwd>auto-insight</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Exploratory Data Analysis (EDA) refers to the critical
process of performing initial investigations on data to
discover patterns, spot anomalies, test hypotheses, and
check assumptions with the help of summary statistics
and graphical representations [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. For example, financial
terprise business metrics, discover data outliers to locate
potential problems that need attention or even action, and
form further hypotheses for more in-depth data
explorations. The outcomes of EDA processes are data insights
      </p>
      <sec id="sec-1-1">
        <title>In a sense, EDA is kind of a creative process [4], during which users leverage their knowledge and intuitions to</title>
        <p>
          mendation techniques such as Voyager 2 [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] and
Foresight [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] primarily focus on the design of insight
discovery algorithm or perceptually efective insight
presentations without deep considerations on the
representation of the intended goal of the recommendations,
ease of understanding and contexts, and user
preferthat are often integrated into a visual dashboard [
          <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
          ]. In other words, as users visualize data, these tools could
siences. Consequently, the resulting auto-insight recom- simpler version of TurboVis that only has the later three
menders may introduce the following side efects to the features but no suggestions from the system. Our
withinEDA process. ⃝ 1 Bias. Prior studies in the field of rec- subjects (TurboVis with vs. without auto-insight
recomommender systems indicate that without a clear repre- mendation) study with 18 industrial business analysts
sentation of the intended goal of the recommendations, shows that the auto-insight design makes them inspect
elaborately designed recommendation algorithms have more visualizations but introduces bias to the direction
the potential to limit exploration breadth as users may of exploration. Auto-insight recommendations ofer new
unconsciously confine their explorations to the items perspectives when analysts are not familiar with the data
recommended [
          <xref ref-type="bibr" rid="ref12 ref13 ref14">12, 13, 14</xref>
          ]. Systems such as Foresight or have a vague idea about how to proceed, and may
and Voyager 2 uncover visual insights independent of the distract analysts when they are facing a familiar dataset
EDA pipeline and represent visual insights in an implicit or usage scenario. Meanwhile, auto-insight
recommenway [
          <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
          ]. An anecdotal evidence shows that, if such dations would interrupt analysts more when they have
auxiliary findings generated by these systems and the specific target for exploration in mind than when they
intended goal of these insights are not explicitly repre- are completely open-minded. Based on these results, we
sented and mentioned, users still have little clue as to further elicit design implications for embedding
autohow the augmented information can be pieced into the insight recommendations into the EDA process.
ifnal story and where it leads them in EDA [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ].
Consequently, it brings bias to analysts in interpreting and
exploiting the recommended visual insights for EDA. ⃝ 2 2. Related Work
Reliability. Prior research in recommender systems
suggested that recommendation service is context-specific Literature that overlaps with this work can be classified
and should improve its readability in order to make the re- into three categories: exploratory data analysis,
visualsults more reliable to users [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. For one thing, previous ization recommendations, and recommender system.
studies illustrate that the perceived utility of recommen- Exploratory Data Analysis. EDA is a term coined
dations is context-specific, i.e., with a limited knowledge by John W. Tukey for describing the act of “looking at
of users, intelligent systems will be less competent in of- data to see what it seems to say” [21]. In EDA, attempts
fering recommendation services [
          <xref ref-type="bibr" rid="ref12 ref16 ref17">12, 16, 17</xref>
          ]. For another, are made to identify the major features of a dataset of
as we may not know the system users at all, “some auto- interest and to generate ideas for further investigation.
Particularly, Tukey drew an analogy between EDA and a
mmyatsitcearyllytorelacoymmmene”n[d1i8n]g. iInfstihgehtasuatore-instsiilglhlitkereacobmaflimngen- series of detective work, during which analysts form a
dation service fails to present its suggestions in an easy- set of hypotheses by asking questions, and integrate their
to-understand manner, the insights’ reliability drops. ⃝ 3 domain knowledge to obtain rich data insights [22, 21].
Interruption. Given that the structure of an EDA pro- Data visualization is perhaps the most widely used tool
cess could be anywhere from fully open exploration to in the EDA process. With the rise of interest in data
target-oriented inspection, automatically recommending science and the need to derive value from data, analysts
insights of the data may bring interruption to the EDA increasingly leverage visualization tools to conduct
exprocess when users prefer open exploration or exami- ploratory data analysis, spot data anomalies, and
correlanation on the non-suggested data aspects [
          <xref ref-type="bibr" rid="ref14">14, 19, 20</xref>
          ]. tions, and identify patterns and trends [23, 24]. The state
Auto-insight recommendation service could thus be a of the art in data visualization involves a lot of manual
“double-edged sword” in EDA, and it would impair the generation of visualizations through tools such as Excel,
analysis experience and results if it is designed inap- Tableau [25] and Qlik [26] to facilitate the EDA process
propriately. Hence, to design an efective auto-insight for non-expert analysts. However, it is still challenging
recommendation service to support EDA, we need to first for data visualization novices to rapidly construct
visuidentify what are required and concerned in a smooth alizations during the EDA process. Grammel et al. [27]
EDA pipeline, and then we need to systematically explore conducted an exploratory laboratory study in which data
the potential “double-edged sword” efect of auto-insight visualization novices explored fictitious sales data by
recommendations on the EDA process and outcome. communicating visualizations to a human mediator, who
        </p>
        <p>To this end, we design TurboVis, a Tableau-like visual- rapidly constructed the visualizations using commercial
ization system that supports analysts in EDA activities visualization software. Apart from identifying activities
with features including auto-insight recommendation that are central to the iterative visualization construction
based on an extensible repository of statistic metrics, process, they also found that the major barriers faced
graphics matching, manual visualization specification, by the participants are translating questions into data
and dashboard editing and interaction. To evaluate how attributes, designing visual mappings, and interpreting
auto-insight recommendation could positively or nega- the visualizations. In this study, we explore the role of
tively afect the EDA process and outcome, we create a auto-insight recommendation in the EDA process.</p>
        <p>
          Visualization Recommendations. As the demand Inspired by the previous findings and to obtain a
systemfor rapid analysis for visualization grows, there is an atic understanding of how the auto-insight
recommendaincreasing requirement to design visualization tools al- tion systems might pose a hindrance to the exploratory
lowing users to eficiently generate visualizations. Prior data analysis process, we design and develop two
verstudies show that the relevant authoring tools are in- sions of a Tableau-like visualization tool and attempt to
creasingly towards automatic [28], which can be classi- explore the “double-edged sword” efect of auto-insight
ifed into four categories. Initially, users have to manually recommendations on EDA in terms of exploration
assiswrite codes for visualizing data by using imperative lan- tance, message reliability, and interference.
guages and libraries such as D3 [29], Vega-Lite [30], and Recommender System. Recommender systems
colECharts [31], which are designed for users who are famil- lect their target users’ preferences for a set of items
iar with coding and visualizations. Later, researchers con- such as movies, songs, books, and travel destinations.
tributed visual building frameworks for easy visualiza- They leverage diferent sources of information for
protions including template editing [32, 33], shelf configura- viding users with predictions and recommendations of
tion [34], and visual building [35, 36]. These tools are de- items [
          <xref ref-type="bibr" rid="ref19">46</xref>
          ]. With the ever-growing volume of online
insigned for users who can write codes but not familiar with formation, recommender systems have been an efective
visualizations. Particularly, users need to “pre-conceive strategy to overcome such information overload [
          <xref ref-type="bibr" rid="ref20">47</xref>
          ],
blueprints, then interact with the system” [28] to obtain particularly useful when users do not have suficient
more expressive, appropriate and aesthetic visualizations. experience to make a choice from a large number of
Then, semi-automatic methods involved with few inter- alternatives [
          <xref ref-type="bibr" rid="ref21">48</xref>
          ]. Existing research in the field of
recomactions were proposed for eficiently obtaining visualiza- mender systems mainly focus on the recommendation
tions like SAGE [37] and Tableau [25]. Fully automatic accuracy and the explainability of recommendation
algomethods are designed for no-human-involved tools for rithms [
          <xref ref-type="bibr" rid="ref19 ref20 ref22 ref23 ref24">46, 47, 49, 50, 51</xref>
          ], which inevitably result in that
eficiently obtaining visualization recommendations such people are increasingly relying on recommender systems
as Text-to-Viz [38], Click2Annotate [39], Data2Vis [40], that employ algorithmic content curation to organize,
seand DeepEye [41]. These tools resolve the issues when lect and present information [
          <xref ref-type="bibr" rid="ref12 ref25 ref26">12, 52, 53</xref>
          ]. Despite its wide
users are not familiar with either visualizations or cod- utility, researchers have indicated that its potential
iming. On the other hand, to resolve the issues that analysts pact to improve problems related to over-choice should
often have no idea what they are looking for especially be concerned [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. Therefore, inspired by the studies
at the initial stage of data exploration, researchers have on the potential harm of such recommendations, we
bedeveloped various algorithms and systems to recommend lieve that the potential “double-edged sword” efect of
insightful visualizations that can depict data trends and auto-insight recommendations warrant a separate study.
patterns [41, 42, 43]. In this study, we combine a
semiautomatic method that involves user interactions along
with algorithms that can recommend interesting insights 3. Research Questions
on the basis of extensible metric repository. Therefore,
analysts can benefit from a quick launch of data explo- The literature suggested that visualization
recommendaration from automated recommendations of potentially tions encourage users to explore more visualizations [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ],
interesting data patterns. while prior studies in the field of recommender systems
        </p>
        <p>
          The existing studies mainly focus on employing either indicated that recommendations would introduce
explomachine learning algorithms or user-defined rules and ration bias, i.e., users might confine their exploration to
visual embellishments into the creation of infographics the recommended items [
          <xref ref-type="bibr" rid="ref13 ref27">13, 54</xref>
          ]. In other words,
recomto lower the barrier for data exploration by automatically mendations can encourage exploration breath but may
generating visualizations. However, indicators from pre- introduce a lack of breadth diversity. We posit that with
vious studies also point out that recommendations may the auto-insight recommendations, target users may
expotentially hamper users during data exploration [44, 45]. plore more visualizations but will be biased to the
visualizations that are only supported by the recommended eratively refine the system through a series of informal
items. Therefore, we have our first research question: usability testings. To be specific, TurboVis consists of
RQ⃝ 1 How do analysts utilize auto-insight recom- four main modules, namely, the data processing,
automendations and manual specification of visualiza- insight recommendation, interactive visual analysis, and
tions collectively as a latent impetus in their EDA dashboard editing and export modules. Particularly, the
process? data processing module (Figure 2(A)) handles common
        </p>
        <p>
          Previous studies in the field of recommendations indi- data formats, obtains and analyzes data fields, and
procate that the perceived reliability of recommendations is vides necessary processing functions such as sorting,
context-specific [
          <xref ref-type="bibr" rid="ref12 ref15 ref16 ref17 ref27">12, 15, 16, 17, 54</xref>
          ]. For example, when ifltering, and editing. The auto-insight recommendation
users are exploring an unfamiliar dataset, recommen- module serves as assistance to inspect potentially
indation services should improve its readability to make teresting data patterns including trending, correlation,
the results more reliable to users. Therefore, we posit distribution, clustering, and outlier detection based on an
a similar efect of auto-insight recommendations need extensible metric repository (covered later). Particularly,
to verify it in the second research question: RQ⃝ 2 How data auto-insight recommendation is embedded in a
typidoes the perceived reliability of auto-insight recom- cal EDA process, recommending interesting visualization
mendations depend on the context of the EDA pro- and enables automatically modifying users’ visualization
cess such as dataset familiarity? specifications to achieve the desired visualizations. The
        </p>
        <p>
          Analysts may hold diferent purposes when conduct- interactive visual analysis module supports analysts to
ing EDA. Previous studies indicate that the degree of perform a drag-and-drop interaction on-demand to
manapplying auto-insight recommendation may vary in dif- ually specify visualizations (Figure 2(B)) . The dashboard
ferent EDA scenarios [
          <xref ref-type="bibr" rid="ref28">55</xref>
          ]. Given that the structure of an editing and export module (Figure 2(C)) allows analysts
EDA process could be anywhere between a fully open to interactively edit each data insight visualization by
exploration and a specific target-oriented inspection, we “linking + view” technique. The finalized dashboard can
want to know RQ⃝ 3 Would auto-insight recommen- be exported on demand.
dations act diferently due to diferent exploration TurboVis without Auto-Insight Recommendation.
purposes in EDA? TurboVis without auto-insight recommendation version
only supports manual visualization specification on the
basis of graph matching. To be specific, as shown in
4. TurboVis Figure 2(B), after loading the data ( ⃝ 1 ), TurboVis
automatically splits data fields into dimensions ( ⃝ 2 ) and metrics
To understand how auto-insight recommendations could (⃝ 3 ). Analysts can perform a drag-and-drop interaction to
be leveraged to assist analysts’ EDA processes and ex- drag any attribute(s) onto the x- or y-axis (⃝ 4 ) and the
displore its potential “double-edged sword” efect, we design play area (⃝ 5 ) would simultaneously present the default
and develop TurboVis, an auto-insight recommendation- chart ranking the first in the recommendation list ( ⃝ 6 ).
powered exploratory data analysis system. To enhance We do not explicitly provide other visual encoding
chanthe generalizability of our findings to common data ana- nels such as size and color due to the observation that
lytics tools, we design TurboVis by reference to existing participants often get lost to determining where each
secommercial exploratory data analytics software and tools lected attribute goes. Therefore, TurboVis automatically
such as Tableau. TurboVis serves as instruments “for de- determines an appropriate visual encoding channel, e.g.,
sign and understand” and are not intended to suggest color and size after analysts’ specification of the  - and
new interaction techniques [
          <xref ref-type="bibr" rid="ref29">56</xref>
          ]. Through a prolonged  -axis. Quantitative attributes can be aggregated in seven
collaborative design process with data analysts, we
it
        </p>
        <p>ways, i.e., sum, mean, maximum, minimum, median,
variance, and standard deviation. For instance, analysts can ommendation serves as assistance to help achieve
visualdrag a quantitative attribute to the  -axis shelf and an- izations with interesting patterns. Asides from manual
other quantitative attribute to the  -axis shelf to generate visualization specification, this version supports
proaca scatterplot; to create a bar chart, analysts can drag a tive unsolicited recommendations and reactive
querynominal attribute to the  -axis shelf and a quantitative based recommendations. Particularly, proactive
unsoattribute by mean, and the chart can be replaced to a line licited recommendations list all potential auto-insights,
chart if the  -axis is filled with a temporal attribute. and reactive query-based recommendations list all the</p>
        <p>
          With respect to the recommendation list, in the version recommended charts on the right side of the interface.
without auto-insight recommendation, the list only has This design considers the selected attributes a query and
the results from graph matching. Specifically, we classify generates recommendations relevant to the selected
atbasic visualizations based on input data format, which tributes.
comes in the form of a decision tree [
          <xref ref-type="bibr" rid="ref30">57</xref>
          ] that leads to a Regarding the proactive unsolicited recommendations,
set of potentially appropriate visualizations to represent we put the entry to the auto-insight recommendation
the current data configuration. For example, considering above the data table (Figure 2(A)⃝ 2 ), maximizing the
utilquantitative attributes, if with only one numeric variable, ity of auto-insight recommendation service. Particularly,
graphs that are appropriate in this case are histogram as shown in Figure 3, diferent types of auto-insight
recand density plot. We adopt the graph matching on the ommendations are classified into pull-down list options
basis of two underlying philosophies [
          <xref ref-type="bibr" rid="ref30">57</xref>
          ]. First, most (⃝ 1 ), i.e., patterns measured by statistical metrics and
data analysis can be summarized in about twenty difer- clustering and outlier detection algorithms. Currently,
ent dataset formats. Second, both data and context can auto-insight recommendation supports four insight types:
determine the appropriate chart. Therefore, our graph trend detection shows line charts with obvious increasing
matching scheme consists of identifying and trying all or decreasing temporal pattern between temporal and
feasible chart types to find out which one(s) suit(s) the quantitative dimensions; the correlation between two
data and idea best. In ⃝ 7 , analysts can rename the current highly correlated attributes between two quantitative
divisualization or adjust the color and size encoding if no mensions; pairwise distribution comparison concerns two
additional attributes are encoded by color or size. groups where the distributions are significantly
difer
        </p>
        <p>
          TurboVis with Auto-Insight Recommendation. We ent, and clustering and outlier detection shows potential
design the second version based on prior studies in proac- clusters and outliers. To facilitate quick browsing and
intive, e.g., Voder [
          <xref ref-type="bibr" rid="ref31">58</xref>
          ] and reactive, e.g., DIVE [
          <xref ref-type="bibr" rid="ref32">59</xref>
          ] insight- spection and easy-to-understand, each recommendation
based recommendations. TurboVis with auto-insight rec- contains a concise natural language description which
is generated by templates (⃝ 2 ), such as “Miles per Gallon insight configuration are automatically filled. For
exand Weight have a strong correlation”. With respect to ample, given a match of both “type_x” and “type_y” is
clustering and outlier detection, we select t-SNE as the di- a quantitative attribute, we generate a candidate
automensionality reduction technique because it shows supe- insight template with the following fields: ‘ mask’ : {‘type’
riority in generating 2D projection that “can reveal mean- : supported_graphs[0], ‘tooltip’ : True}, ‘encoding’ : {‘x’
ingful insights about data, e.g., clusters and outliers” [
          <xref ref-type="bibr" rid="ref33">60</xref>
          ]. : {‘field ’ : name_x, ‘type’ : “quantitative”},‘y’ : {‘field ’ :
In addition, advanced parameter settings are also pro- name_y, ‘type’ : “quantitative”}}, ‘priority’ : priority.
Furvided such as quantitative attributes for projection and thermore, according to diferent auto-insight measures,
t-SNE parameters in terms of perplexity, learning rate, we compute diferent metrics for each candidate
visualmaximum iterations, and distance metric (⃝ 3 ). By pre- ization specification. For example,  with a quantitative
viewing all the auto-insight recommendations, analysts attribute and  with a quantitative attribute are
evalucan select any of them by clicking on + to submit to the ated by a Spearman correlation coeficient with associated
target dashboard. We employ exhaustion, match, gener- p-value while  with a temporal attribute and  with a
ate for mining interesting visualizations. We show how quantitative attribute are evaluated by the trend
detecto generate data auto-insights through the following four tion measure. Candidates with a metric value higher
steps. than a predefined threshold parameter are recommended
        </p>
        <p>Step ⃝ 1 Determining attribute types. After loading to analysts. Based on the results, we fill the ‘   ’ in
a dataset, TurboVis first gathers metadata of data types the template with {‘correlation’ : correlation, ‘p-value’ :
by iterating on all data records. Particularly, we main- p-value}; ‘message’ : “name_x and name_y has a strong
tain several metadata to determine whether the value correlation depending the value of p-value”.
on a certain attribute is e.g., numeric, date, or coordi- Regarding reactive query-based recommendations,
Turnate. We also maintain the number of unique values and boVis with auto-insight recommendations merges
autothe maximum of replication corresponding to a certain insight recommendations into the recommendation list
data attribute. Data types can be thus determined for that appear in the right panel (Figure 2(B)⃝ 6 ), which
subsequent processing. tailors the auto-insight recommendations into the EDA</p>
        <p>Step ⃝ 2 Maintaining recommendation configura- pipeline. In Figure 4, the left subfigure shows the graph
tion. Each data auto-insight corresponds to a recommen- matching based on a particular data attribute
configuradation configuration, which consists of six dimensions in tion and the right subfigure displays the auto-insights in
terms of “type_x”, “type_y”, “position_exchange”, “mea- the version with auto-insight recommendations. In other
sure”, “supported_graphs” and “priority”. For example, words, we only display graph matching in the version
trend corresponds to a bar recommendation configura- without auto-insight recommendations and display both
tion: {type_x : temporal, type_y : quantitative, position_ex- in the version with auto-insight recommendations.
change : 0, measure : trend, supported_graphs : [‘bar’],
priority : 0 }, which means that when a temporal attribute
meets a quantitative attribute, we can use bar to visualize 5. Experiment
the relationship with exchangeable axes. “Priority”
indicates the recommendation priority when demonstrating To investigate the “double-edged sword” efect of
autoall the data auto-insight patterns to audiences. Similarly, insight recommendation design on EDA, we conduct a
correlation corresponds to a scatterplot recommenda- within-subjects study with 18 data analysts in EDA tasks
tion configuration with both “ type_x” and “type_y” is a on two datasets.
quantitative attribute and the “supported_graphs” can be Participants. We recruit 18 industrial data analysts
[‘scSattetper⃝’3]. Preparing and matching all feasible com- (n9etfebmanalke,sm,9osmt aolfesw, haogme: h2a8v±e 32.0to3)5fryoemarsa
olofcwalorInkitnergbinations to recommendation configuration. We ini- experiences. We invite participants who need to
contialize feasible combinations to mine potential patterns duct EDA almost every day according to their self report.
hidden in the combination of any two diferent attributes. Particularly, participants had used tools for EDA,
includTo optimize the exhaustion process, we allow users to ing Tableau (10/18), Excel (18/18), Python (8/10), and
specify mask attributes thus the calculation will not con- R (12/18). They are representatives of our target users
sider those masked attributes. We then match each feasi- and could provide us more comprehensive insights. We
ble combination against the targets of the auto-insight compensate participants with a $20 gift card.
meSatseupr e⃝4s oGnetnheerbaatisnisgocfa“ n did_a t”eaanudto“ -Ins_igh”.t recom- to DevaatlausaettestahnedefeDcatstaoPfrthoeceasustion-gin. sWigehcthroeocosemtmwoenddaata-sets
mendation. Upon a match between a feasible combi- tions. The first one is the happiness ranking dataset that
nation and an auto-insight recommendation configura- our participants analyze less in their daily work, and it
tion, parameters in the corresponding candidate auto- consists of 1093 records with attributes of date, country,
region, happiness ranking, happiness, GDP per capita, GDP TurboVis to ensure that they have no problems
conductper family, healthy, freedom, trustness, generosity, and res- ing the subsequent tasks on their own. In the main study,
idence1. The second dataset is the Chinese bank’s annual for participants who start with the one TurboVis version,
report, which is closer to the type of data that our partici- we ask them to conduct a 15-minute data exploration of
pants use everyday, and it comprises 1590 records with 18 the happiness ranking sub dataset (session 1). Then, they
ifnancial attributes. As a demo to introduce our system proceed to another 15-minute data exploration of another
to the participants, we also include a car dataset which happiness ranking sub dataset using another version of
consists of 403 records with attributes of name, miles TurboVis (session 2). Participants are also asked to think
per gallon, cylinders, displacement, horsepower, weight, aloud their ideas when performing all the tasks. Their
acceleration, year, and origin. exploration processes are automatically recorded as
sys</p>
        <p>To prepare datasets for the two versions, we split the tem logs for the subsequent quantitative analysis. Then,
happiness ranking dataset and bank dataset into two we repeat the above process by using another dataset,
parts. During the study, the first dataset is used in the i.e., the bank annual report dataset for session 3 and 4.
ifrst data exploration session while the second one is used After finishing all the tasks, participants are required
in the second session. This mitigates potential learning to complete a questionnaire with 7-point Likert scale
efects across the two sessions while ensuring that the questions, followed by a semi-structured interview with
data collected from the two sessions could be compared. each participant to make sense of their ratings and
col</p>
        <p>Procedure. After obtaining the participants’ consent, lect their opinions about auto-insight recommendations.
we conduct the experiment in four sessions, each with a The whole experiment lasts around 90 minutes for each
subset of a dataset and one version of TurboVis. In other participant.
words, every participant gets to explore both datasets Measures. In the above-mentioned experiment, we
in diferent tasks using both versions of our tool. We collect 72 log files ( 18 participants × 4 sessions). The
counterbalance the order of TurboVis’s version in each log data details every action and the associated entities
dataset to minimize the learning efects. We arrange the conducted by the participant. The actions include but are
experimental procedure in the following steps. First, we not limited to click, select, delete, and drag, and objects
give a tutorial on how to use both versions of TurboVis are like the name of auto-insight recommendation and
(each for 10 minutes; the system is running on ThinkPad attributes. We then derive our dependent measures from
X270 notebook with a 12.5-inch display) to explore data these data in relation to the previously mentioned three
with the car dataset and then allow the participants to research questions.
freely explore the tool for another 10 minutes. During M⃝ 1 Number of inspected visualizations. We count
this process, we encourage the participants to raise any the number of visualization if a specific visualization is
question about the usability, functions, and features of loaded, selected, edited, or added in a dashboard. This
measure is a conservative estimate of the inspected
visualizations for the exploration with recommendations.</p>
        <p>1https://www.kaggle.com/mathurinache/world-happinessreport</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>6. Results and Analysis</title>
      <p>M⃝ 2 Proportion of supported visualizations. To
understand whether the potential exploration bias, i.e.,
confining to the scope of auto-insight recommendations, We report the quantitative analysis of participants’
opwe compare the proportion of visualizations in terms of eration logs and quantitative ratings and feedback on
correlation, trend, pairwise distribution comparison, and the three research questions, as shown in Figure 5.
Parclusters and outliers among those visualizations logged as ticularly, we analyze the first five measures ( M⃝ 1 - ⃝ 5 )
manual or auto-insight recommendations between with using Wilcoxon signed-rank tests (very appropriate for
and without auto-insight recommendation versions. a repeated measure design where the same subjects are</p>
      <p>M⃝ 3 Number of manually specified visualizations. evaluated under two diferent conditions) with a
signifiWe count the number of manually created visualizations cance level of 0.05 and we report the median values for
for each session. This measure indicates the adoption or the subjective measures collected from the questionnaires
potential reliance on the auto-insight recommendations on each item (M⃝ 6 - ⃝ 9 ).
for EDA, since participants might utilize the auto-insights RQ⃝ 1 Utilizing auto-insights and manual
visualand thus create fewer visualizations on their own. ization specification collectively . As shown in
Fig</p>
      <p>M⃝ 4 Time duration between opening and closing ure 5(a), we find a significant diference in the M⃝ 1
numauto-insight recommendations. To investigate the ef- ber of inspected visualizations at a significance level of
fect of how auto-insights may advance their EDA process, 0.05 (both  = −3.726,  = 0.00019 ) by using happiness
we record and calculate the time intervals between partic- ranking and bank datasets, respectively. On average,
paripant opening and closing (actions recorded in the logs) ticipants inspect more visualizations when using
Turbothe auto-insight recommendation panel when given the Vis with auto-insight recommendations ( = 45.33 with
version of our tool with this function, to explore diferent happiness ranking and  = 44.11 with bank dataset)
datasets. than in the without auto-insight recommendation
condi</p>
      <p>M⃝ 5 Number of modifications to the recommended tion ( = 12.61 with happiness ranking and  = 15.44
visualizations. When participants are exploring the with bank dataset), indicating a wider coverage of
vidata using the auto-insight recommendation version, sualizations with auto-insight recommendations. M⃝ 2
they can select any visualization from both the chan- Proportion of visualizations that are supported by
recomnels of graph matching and the tool’s recommendations. mendation is higher with auto-insight recommendations
When participants drag “asset size” to one axis and the ( = .465 with happiness ranking and  = .36 with
display area would immediately present the auto-insight bank dataset) than without this feature (Figure 5(b)). The
recommendation relevant to this attribute and fill the diference is significant for both the happiness
rankother axis with information e.g., asset size that has a ing ( = −3.725,  = .000196 ) and the bank dataset
high correlation relationship. However, one issue we fre- ( = −3.483,  = .000499 ), indicating that the
autoquently observe is that if this “automatic completion” is insight recommendations bias participants towards
cerinconsistent with participants’ intent, they would modify tain types of visualizations during the EDA process.
the recommendation result. Therefore, to investigate to Generally, participants specify more visualization in
what extend participants would directly accept the auto- the absence of auto-insight recommendations ( = 6.61
insights when they are exploring dataset with diferent with happiness ranking and  = 9.94 with bank dataset)
familiarity, we calculate the number of modification ac- than in the presence of this service ( = 4.5 with
haptions immediately occur after a recommendation result piness ranking and  = 8.67 with bank dataset), as
is populated in the display area. shown in Figure 5(c) (M⃝ 3 ). However, we observe
difer</p>
      <p>To complement the quantitative data and provide in- ent results of the manual specification of visualization
depth understanding of users’ perceptions towards the on our two datasets. Specifically, the diference is
sigauto-insight recommendations, we also collect the par- nificant when participants explore the happiness
rankticipants’ responses to an end-of-study questionnaires, ing dataset ( = −2.166,  = .03 ), but not significant
in which we ask them about their preference of tool ( = −1.742,  = .081 ) when using a bank dataset.
Conversions when conducing EDA with a clear exploration sidering that our participants frequently analyze bank
task, and whether the auto-insights ofer new knowl- data and rarely inspect happiness data in their daily work,
edge. Particularly, we have: M⃝ 6 usefulness of auto- this implies that our participants have a tendency to
insights in an open exploration, M⃝ 7 usefulness of resort to auto-insight recommendations for inspecting
auto-insights with a target-oriented inspection (1 - visual patterns rather than constructing visualizations
Extremely unuseful, 7 - Extremely useful), M⃝ 8 version manually when this service is available for exploring
preference when exploring in an open exploration, an unfamiliar dataset, but not as much when facing a
and M⃝ 9 version preference when exploring with a familiar dataset. “I think I can do more inspection
withtarget-oriented inspection (1 - Prefer non-auto-insight out auto-insight recommendations” (P12, male, age: 31).
version a lot, 7 - Prefer auto-insight version a lot). “With my intuition and knowledge, I just want to see the
data this way” (P10, female, age: 29). “I feel like I can ploration or with a target-oriented inspection) in terms of
reason more on my own than the recommendations” (P6, perceived usefulness and version preference. We ask
parmale, age: 26). ticipants in the questionnaire about the perceived
useful</p>
      <p>RQ⃝ 2 Auto-insight reliability in EDA. To evaluate ness and preference of auto-insights when they conduct
the acceptance of participants towards the auto-insights an open exploration or a target-oriented inspection in our
conveyed by the recommendations, we measure the M⃝ 4 EDA tasks. During the experiment, we often find that
partime duration between opening and closing the auto-insight ticipants modify the recommended results by deleting the
recommendation panel given the version of TurboVis with attribute that has been automatically populated on one
this feature. As shown in Figure 5(d), we find a signif- axis. Therefore, we obtain the M⃝ 5 number of
modificaicant diference in this measure (  = −3.301,  = .001 ). tions to the recommended visualizations and compare the
An average duration of 2.88 minutes is spent on brows- counts between the two dataset. As shown in Figure 5(e),
ing the results of auto-insight recommendations of the we find that although the mean value of the number of
happiness ranking dataset, compared with an average modification difers, i.e.,  = 3.39 with the happiness
duration of 1.97 minutes on the bank dataset’s recom- ranking dataset and  = 4.56 with the bank dataset, the
mended auto-insights. “I am not familiar with happiness diference is not significant (  = −1.579,  = .114 ). The
ranking dataset so I have a lower expectation with auto- questionnaire item of M⃝ 6 ⃝ 7 usefulness of auto-insight
insight recommendations” (P2, male, age: 28). “I would recommendations in an open exploration or with a
targettry to find why these auto-insights are recommended when oriented inspection also shows that participants
appreI am exploring the happiness ranking dataset” (P14, fe- ciate the usefulness of auto-insight recommendations
male, age: 25). “I probably would not have figured out the regardless of the tasks (open:  = 6.11,  = .96 and
taroutliers by intuition and I am happy that it has been recom- get:  = 6.06,  = .87 ), suggesting that the acceptance
mended” (P6, male, age: 26). With respect to a relatively of auto-insights in diferent EDA tasks does not change
more familiar dataset (e.g., bank dataset in our case), the significantly.
auto-insight recommendation serves as assistance for However, in participants’ response to the question
quick verification, “ I can quickly identify the interesting M⃝ 8 ⃝ 9 version preference in an open exploration or with
auto-insights since I am familiar with them” (P8, female, a target-oriented inspection, the median rating was 5.89
age: 30). with an SD of 0.83 for open exploration on a scale from</p>
      <p>RQ⃝ 3 Auto-insights in open exploration and tar- preferring without auto-insight recommendation much
get oriented inspection. We investigate the acceptance more (1) to preferring without auto-insight
recommendaof auto-insights in diferent EDA tasks (i.e., an open ex- tions much more (7), suggesting that participants prefer
having the service much more when they only have a Message Reliability. When analysts are quite
familvague idea about what they are looking for. “When you iar with the dataset and exploration scenarios, they have
introduce a new dataset that I haven’t see before, I don’t a higher expectation of the auto-insight
recommendaknow where to start” (P2, male, age: 28). “It is hard for tions. They would try to draw conclusions by observing
me to figure out where to go first and auto-insights help the auto-insight recommendation results, e.g.,
determinme with the first step ” (P14, female, age: 25). However, ing whether these insights make sense or not, i.e., they
when they have specific questions to investigate, they may question the message reliability. Otherwise, when
prefer TurboVis without auto-insight recommendations they have a vague idea about what they are looking for,
( = 3.8 ,  = 1.1 ). “I have a very clear target in my they have a lower expectation on the auto-insight
recommind so I directly turn to the manually specifying visual- mendations; they appreciate the interestingness of the
izations interface to see what I can get”, “since I am quite recommended patterns, instead of identifying whether
familiar with the data and I know how to select attributes these insights are right or wrong. A design implication
that have relationships” (P8, female, age: 30). “When I is that it is necessary to note next to the auto-insight
was exploring on my own, I feel like I am creating what I recommendations what methods are adopted to generate
want” (P12, male, age: 31). these recommendations and what the system has done
in order to make analysts clear about the underlying
recommendation mechanisms.
7. Discussion Exploration Interruption. When we inferring user
intention from interaction log data, we observed that
In this section, we first discuss the identified “double- auto-insight recommendations sometimes interrupt
anaedged sword” efect of auto-insight recommendations on lysts. When they were exploring a familiar dataset and
the EDA process. Then, we elicit the design implications trying to construct a desired visualization for
inspecregarding the observed findings. In the end, we reflect tion, they commented that they prefer the without
autoon the limitations of this study. insight recommendation version in the drag-and-drop</p>
      <p>Exploration Bias. We further our awareness of the process. For example, when analysts drag a data attribute
side efects of auto-insight recommendations on EDA by that has been involved in an auto-insight
recommendaifrst highlighting potential exploration bias and excessive tion to one axis, TurboVis automatically refreshes the
reliance. Analysts tend to adjust their degree of reliance recommendation view and lists all the auto-insight
recon auto-insight recommendations or manual visualiza- ommendations related to the specific data attribute and
ttiaosnk ssptreuccitficuarteio. nFsoronontheethbiansgi,s tohfeddaotmafaainmeilxiapreirtytsawnidth even populates the other fields, “ I was intending to put
a high degree of familiarity with data and analytic tasks ‘acceleration’ and ‘cylinders’ together to see what kind of
are more likely to explore more visualizations on their visualization results could appear, but the recommended
own. For another, they believe that their domain knowl- saiustios-itnhsaitgthhtes
glorwabcboedgnmityivaettceonsttioonf.g”leAanpilnaugsiinbsleighhytspofrtohmeedge and intuitions can help them achieve a smooth EDA the recommendations makes them too tempting to
conprocess when facing familiar data and scenarios. How- sume, thereby inducing undesirable interruption efects.
ever, if they encounter a new dataset, they might heavily To mitigate the interruption efects of auto-insight
recrely on the auto-insight recommendations by immers- ommendations, one alternative is to hide the auto-insight
ing themselves in browsing the recommendation results. recommendation services in a toolbar by following Show
When auto-insight recommendations were present, par- Me or split the recommended results from the existing
ticipants demonstrated less desire to explore data on their panel and display them in a separate panel.
own, e.g., the number of manually-specified visualiza- Limitations. First, although we conducted a
protions drops significantly. Auto-insight recommendation longed collaborative design with a limited number of
service can produce a large number of recommendations industrial domain experts, TurboVis is still limited in its
to implicitly impel users to explore more visualizations, current form with respect to the raised requirements
thus, leading to biased data exploration. A potential de- collected from their feedback. Second, we derived
design implication is that auto-insight recommendations sign alternatives by surveying prior systems and
conshould be designed diferently based on how people can ducting iterative design with our collaboration experts.
tolerant auto-insights. In scenarios that welcome diverse Admittedly, we only tapped into a limited design space of
exploration results, recommendations should be hidden recommendations, i.e., by providing a button to see the
or at least receive less concerns. Also, tooltips can be pro- auto-insight recommendation on-demand and linking
vided to explicitly inform analysts that how many auto- auto-insight recommendations to manual user
interacinsight recommendations or the ratio of auto-insights to tion. With a diferent design of auto-insight
recommenthe overall visualizations have been added to the dash- dation service, users may perceive diferently. Third, we
board. design the auto-insight recommendation service only
based on a limited number of an extensible repository
of statistic metrics, which quantify interesting
visualization patterns in basic charts. Meanwhile, only
experienced data analysts working on a particular set of
problems were included in the user study. The results
therefore might generalize only to this kind of users.</p>
      <p>Furthermore, auto-insight recommendations fail to
recommend any patterns if involving multiple attributes.</p>
      <p>Our collaboration experts also commented that there
should be more types of recommendations. Future work
will systematically conduct more investigation into more
real-world business scenarios to identify more preferable
auto-insight types.</p>
    </sec>
    <sec id="sec-3">
      <title>8. Conclusion</title>
      <sec id="sec-3-1">
        <title>In this study, we explore the potential “double-edged</title>
        <p>sword” efects of auto-insight recommendations on the
EDA process. We demonstrate how auto-insight
recommendations could be incorporated into a self-developed
Tableau-like visualization tool termed TurboVis. By
comparing two versions of TurboVis, we find that auto-insight
recommendations not only encourage more visualization
inspections but also introduce biases to data exploration.
Meanwhile, the perceived level of message reliability and
interruption of auto-insight recommendation service
depend on data familiarity and task structures. Our work
ofers initial implications for embedding auto-insight
recommendations into the EDA process.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <sec id="sec-4-1">
        <title>We are grateful for the valuable feedback and comments provided by the anonymous reviewers.</title>
        <p>[19] D. Booth, The human-computer interaction hand- knowledge management and sensemaking tools for
book: fundamentals, evolving technologies and intelligence analysts, in: Proceedings of the 15th
emerging applications, L Erlbaum Associates Inc ACM international conference on Information and
15 (2012) 85–87. knowledge management, 2006, pp. 513–521.
[20] M. Vartak, S. Huang, T. Siddiqui, S. Madden, [35] Z. Liu, J. Thompson, A. Wilson, M. Dontcheva, J.
DeA. Parameswaran, Towards visualization recom- lorey, S. Grigg, B. Kerr, J. Stasko, Data illustrator:
mendation systems, Acm Sigmod Record 45 (2017) Augmenting vector design tools with lazy data
bind34–39. ing for expressive visualization authoring, in:
Pro[21] J. Tukey, Exploratory Data Analysis, Addison- ceedings of the 2018 CHI Conference on Human</p>
        <p>Wesley Pub. Co.„ ???? Factors in Computing Systems, 2018, pp. 1–13.
[22] J. T. Behrens, Principles and procedures of ex- [36] D. Ren, T. Höllerer, X. Yuan, ivisdesigner:
Expresploratory data analysis., Psychological Methods sive interactive design of information visualizations,
2 (1997) 131. IEEE transactions on visualization and computer
[23] P. Hanrahan, Analytic database technologies for a graphics 20 (2014) 2092–2101.
new kind of user: the data enthusiast, in: Proceed- [37] S. F. Roth, J. Kolojejchick, J. Mattis, M. C. Chuah,
ings of the 2012 ACM SIGMOD International Con- Sagetools: An intelligent environment for
sketchference on Management of Data, 2012, pp. 577–578. ing, browsing, and customizing data-graphics, in:
[24] K. Morton, M. Balazinska, D. Grossman, J. Mackin- Conference companion on Human factors in
comlay, Support the data enthusiast: Challenges for puting systems, 1995, pp. 409–410.
next-generation data-analysis systems, Proceed- [38] W. Cui, X. Zhang, Y. Wang, H. Huang, B. Chen,
ings of the VLDB Endowment 7 (2014) 453–456. L. Fang, H. Zhang, J.-G. Lou, D. Zhang,
Text-to[25] M. D’Agostino, D. M. Gabbay, R. Hähnle, J. Posegga, viz: Automatic generation of infographics from
Handbook of tableau methods, Springer Science &amp; proportion-related natural language statements,
Business Media, 2013. IEEE transactions on visualization and computer
[26] O. Troyansky, T. Gibson, C. Leichtweis, QlikView graphics 26 (2019) 906–916.</p>
        <p>Your Business: An Expert Guide to Business Dis- [39] Y. Chen, S. Barlowe, J. Yang, Click2annotate:
Autocovery with QlikView and Qlik Sense, John Wiley mated insight externalization with rich semantics,
&amp; Sons, 2015. in: 2010 IEEE Symposium on Visual Analytics
Sci[27] L. Grammel, M. Tory, M.-A. Storey, How informa- ence and Technology, IEEE, 2010, pp. 155–162.
tion visualization novices construct visualizations, [40] V. Dibia, Ç. Demiralp, Data2vis: Automatic
genIEEE transactions on visualization and computer eration of data visualizations using
sequence-tographics 16 (2010) 943–952. sequence recurrent neural networks, IEEE
com[28] S. Zhu, G. Sun, Q. Jiang, M. Zha, R. Liang, A sur- puter graphics and applications 39 (2019) 33–46.
vey on automatic infographics and visualization [41] Y. Luo, X. Qin, N. Tang, G. Li, Deepeye: Towards
aurecommendations, Visual Informatics (2020). tomatic data visualization, in: 2018 IEEE 34th
Inter[29] M. Bostock, V. Ogievetsky, J. Heer, D3 data-driven national Conference on Data Engineering (ICDE),
documents, IEEE transactions on visualization and IEEE, 2018, pp. 101–112.</p>
        <p>computer graphics 17 (2011) 2301–2309. [42] R. Ding, S. Han, Y. Xu, H. Zhang, D. Zhang,
Quick[30] A. Satyanarayan, D. Moritz, K. Wongsuphasawat, insights: Quick and automatic discovery of insights
J. Heer, Vega-lite: A grammar of interactive graph- from multi-dimensional data, in: Proceedings of
ics, IEEE transactions on visualization and com- the 2019 International Conference on Management
puter graphics 23 (2016) 341–350. of Data, 2019, pp. 317–332.
[31] D. Li, H. Mei, Y. Shen, S. Su, W. Zhang, J. Wang, [43] B. Tang, S. Han, M. L. Yiu, R. Ding, D. Zhang,
M. Zu, W. Chen, Echarts: A declarative framework Extracting top-k insights from multi-dimensional
for rapid construction of web-based visualization, data, in: Proceedings of the 2017 ACM
InternaVisual Informatics 2 (2018) 136–146. tional Conference on Management of Data, 2017,
[32] R. López-Cortijo, J. G. Guzmán, A. A. Seco, icharts: pp. 1509–1524.</p>
        <p>charts for software process improvement value [44] J. Heer, Agency plus automation: Designing
artifimanagement, in: European Conference on Software cial intelligence into interactive systems,
ProceedProcess Improvement, Springer, 2007, pp. 124–135. ings of the National Academy of Sciences 116 (2019)
[33] M. Mauri, T. Elli, G. Caviglia, G. Uboldi, M. Azzi, 1844–1850.</p>
        <p>Rawgraphs: a visualisation platform to create open [45] K. Wongsuphasawat, Z. Qu, D. Moritz, R. Chang,
outputs, in: Proceedings of the 12th biannual con- J. Heer, Voyager 2: Augmenting visual analysis
ference on Italian SIGCHI chapter, 2017, pp. 1–5. with partial view specifications, in: the 2017 CHI
[34] N. J. Pioch, J. O. Everett, Polestar: collaborative Conference, 2017.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tamizharasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Singhal</surname>
          </string-name>
          ,
          <article-title>Exploratory data analysis (</article-title>
          <year>1999</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Diamond</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mattia</surname>
          </string-name>
          ,
          <article-title>Data visualization: An exploratory study into the software tools used by businesses</article-title>
          .,
          <source>Journal of Instructional Pedagogies</source>
          <volume>18</volume>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Gustafson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. H.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Pape-Haugaard</surname>
          </string-name>
          ,
          <article-title>Designing a dashboard to visualize patient information</article-title>
          ,
          <source>in: Proceedings from The 16th Scandinavian Conference on Health Informatics</source>
          <year>2018</year>
          , Aalborg,
          <source>Denmark August 28-29</source>
          ,
          <year>2018</year>
          ,
          <volume>151</volume>
          , Linköping University Electronic Press,
          <year>2018</year>
          , pp.
          <fpage>23</fpage>
          -
          <lpage>28</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R. K.</given-names>
            <surname>Sawyer</surname>
          </string-name>
          ,
          <article-title>Explaining creativity: The science of human innovation</article-title>
          , Oxford university press,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Tukey</surname>
          </string-name>
          ,
          <article-title>We need both exploratory and confirmatory</article-title>
          ,
          <source>The American Statistician</source>
          <volume>34</volume>
          (
          <year>1980</year>
          )
          <fpage>23</fpage>
          -
          <lpage>25</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Vartak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rahman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Madden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Parameswaran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Polyzotis</surname>
          </string-name>
          , Seedb:
          <article-title>Eficient data-driven visualization recommendations to support visual analytics</article-title>
          ,
          <source>in: Proceedings of the VLDB Endowment International Conference on Very Large Data Bases</source>
          , volume
          <volume>8</volume>
          ,
          <string-name>
            <given-names>NIH</given-names>
            <surname>Public Access</surname>
          </string-name>
          ,
          <year>2015</year>
          , p.
          <fpage>2182</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>K.</given-names>
            <surname>Wongsuphasawat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Moritz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Anand</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mackinlay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Howe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Heer</surname>
          </string-name>
          , Voyager:
          <article-title>Exploratory analysis via faceted browsing of visualization recommendations</article-title>
          ,
          <source>IEEE transactions on visualization and computer graphics 22</source>
          (
          <year>2015</year>
          )
          <fpage>649</fpage>
          -
          <lpage>658</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Cui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. K.</given-names>
            <surname>Badam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Yalçin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Elmqvist</surname>
          </string-name>
          , Datasite:
          <article-title>Proactive visual data exploration with computation of insight-based recommendations</article-title>
          ,
          <source>Information Visualization</source>
          <volume>18</volume>
          (
          <year>2019</year>
          )
          <fpage>251</fpage>
          -
          <lpage>267</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Mackinlay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Hanrahan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Stolte</surname>
          </string-name>
          ,
          <article-title>Show me: Automatic presentation for visual analysis</article-title>
          ,
          <source>IEEE Transactions on Visualization &amp; Computer Graphics</source>
          <volume>13</volume>
          (
          <year>2007</year>
          )
          <fpage>1137</fpage>
          -
          <lpage>1144</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>K.</given-names>
            <surname>Wongsuphasawat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Qu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Moritz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Heer</surname>
          </string-name>
          ,
          <article-title>Voyager 2: Augmenting visual analysis with partial view specifications</article-title>
          ,
          <source>in: the 2017 CHI Conference</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Ç. Demiralp</surname>
            ,
            <given-names>P. J.</given-names>
          </string-name>
          <string-name>
            <surname>Haas</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Parthasarathy</surname>
          </string-name>
          , T. Pedapati, Foresight: Recommending visual insights,
          <source>arXiv preprint arXiv:1707.03877</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zeng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Yi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Ma</surname>
          </string-name>
          , T. Chen,
          <article-title>Friend network as gatekeeper: A study of wechat users' consumption of friend-curated contents</article-title>
          , in: The eighth International Workshop of Chinese CHI,
          <year>2020</year>
          , pp.
          <fpage>21</fpage>
          -
          <lpage>31</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Jameson</surname>
          </string-name>
          ,
          <article-title>Adaptive interfaces and agents, in: The human-computer interaction handbook</article-title>
          , CRC Press,
          <year>2007</year>
          , pp.
          <fpage>459</fpage>
          -
          <lpage>484</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>A. D. Jameson</surname>
          </string-name>
          ,
          <article-title>Understanding and dealing with usability side efects of intelligent processing</article-title>
          ,
          <source>AI</source>
          Magazine
          <volume>30</volume>
          (
          <year>2009</year>
          )
          <fpage>23</fpage>
          -
          <lpage>23</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>F.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Plaisant</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Spring</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Crowley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Shneiderman</surname>
          </string-name>
          ,
          <article-title>Eventaction: A visual analytics approach to explainable recommendation for event sequences</article-title>
          ,
          <source>ACM Transactions on Interactive Intelligent Systems (TiiS) 9</source>
          (
          <issue>2019</issue>
          )
          <fpage>1</fpage>
          -
          <lpage>31</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>D.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Teevan</surname>
          </string-name>
          ,
          <article-title>Implicit feedback for inferring user preference: a bibliography</article-title>
          ,
          <source>in: Acm Sigir Forum</source>
          , volume
          <volume>37</volume>
          , ACM New York, NY, USA,
          <year>2003</year>
          , pp.
          <fpage>18</fpage>
          -
          <lpage>28</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>J.</given-names>
            <surname>Xiao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Stasko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Catrambone</surname>
          </string-name>
          , et al.,
          <article-title>An empirical study of the efect of agent competence on user performance and perception</article-title>
          ,
          <source>in: AAMAS</source>
          , volume
          <volume>4</volume>
          ,
          <string-name>
            <surname>Citeseer</surname>
          </string-name>
          ,
          <year>2004</year>
          , pp.
          <fpage>178</fpage>
          -
          <lpage>185</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>G.</given-names>
            <surname>Dove</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Halskov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Forlizzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zimmerman</surname>
          </string-name>
          ,
          <article-title>Ux design innovation: Challenges for working with machine learning as a design material</article-title>
          ,
          <source>in: Chi Conference on Human Factors in Computing Systems</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [46]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bobadilla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ortega</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hernando</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gutiérrez</surname>
          </string-name>
          ,
          <article-title>Recommender systems survey</article-title>
          ,
          <source>Knowledge-Based Systems</source>
          <volume>46</volume>
          (
          <year>2013</year>
          )
          <fpage>109</fpage>
          -
          <lpage>132</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [47]
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tay</surname>
          </string-name>
          ,
          <article-title>Deep learning based recommender system: A survey and new perspectives, ACM Computing Surveys (CSUR) 52 (</article-title>
          <year>2019</year>
          )
          <fpage>1</fpage>
          -
          <lpage>38</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [48]
          <string-name>
            <given-names>D.</given-names>
            <surname>Parra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sahebi</surname>
          </string-name>
          ,
          <article-title>Recommender systems: Sources of knowledge and evaluation metrics</article-title>
          ,
          <source>in: Advanced techniques in web intelligence-2</source>
          , Springer,
          <year>2013</year>
          , pp.
          <fpage>149</fpage>
          -
          <lpage>175</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [49]
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Herlocker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Konstan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. G.</given-names>
            <surname>Terveen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. T.</given-names>
            <surname>Riedl</surname>
          </string-name>
          ,
          <article-title>Evaluating collaborative filtering recommender systems</article-title>
          ,
          <source>ACM Transactions on Information Systems (TOIS) 22</source>
          (
          <year>2004</year>
          )
          <fpage>5</fpage>
          -
          <lpage>53</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [50]
          <string-name>
            <given-names>P.</given-names>
            <surname>Pu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <article-title>A user-centric evaluation framework for recommender systems</article-title>
          ,
          <source>in: Proceedings of the fith ACM conference on Recommender systems</source>
          ,
          <year>2011</year>
          , pp.
          <fpage>157</fpage>
          -
          <lpage>164</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [51]
          <string-name>
            <given-names>R.</given-names>
            <surname>Sinha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Swearingen</surname>
          </string-name>
          ,
          <article-title>The role of transparency in recommender systems</article-title>
          , in: CHI'
          <article-title>02 extended abstracts on Human factors in computing systems</article-title>
          ,
          <year>2002</year>
          , pp.
          <fpage>830</fpage>
          -
          <lpage>831</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [52]
          <string-name>
            <given-names>M.</given-names>
            <surname>Eslami</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rickman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Vaccaro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Aleyasen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vuong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Karahalios</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Hamilton</surname>
          </string-name>
          , C. Sandvig, ”
          <article-title>i always assumed that i wasn't really that close to [her]” reasoning about invisible algorithms in news feeds</article-title>
          ,
          <source>in: Proceedings of the 33rd annual ACM conference on human factors in computing systems</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>153</fpage>
          -
          <lpage>162</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [53]
          <string-name>
            <given-names>E.</given-names>
            <surname>Rader</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Gray</surname>
          </string-name>
          ,
          <article-title>Understanding user beliefs about algorithmic curation in the facebook news feed</article-title>
          ,
          <source>in: Proceedings of the 33rd annual ACM conference on human factors in computing systems</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>173</fpage>
          -
          <lpage>182</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [54]
          <string-name>
            <given-names>A.</given-names>
            <surname>Jameson</surname>
          </string-name>
          ,
          <article-title>Understanding and dealing with usability side efects of intelligent processing</article-title>
          ,
          <source>Ai Magazine</source>
          <volume>30</volume>
          (
          <year>2009</year>
          )
          <fpage>23</fpage>
          -
          <lpage>40</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [55]
          <string-name>
            <given-names>P. D.</given-names>
            <surname>Adamczyk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. P.</given-names>
            <surname>Bailey</surname>
          </string-name>
          ,
          <article-title>If not now, when? the efects of interruption at diferent moments within task execution</article-title>
          ,
          <source>in: Proceedings of the SIGCHI conference on Human factors in computing systems</source>
          ,
          <year>2004</year>
          , pp.
          <fpage>271</fpage>
          -
          <lpage>278</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [56]
          <string-name>
            <given-names>J.</given-names>
            <surname>Wallace</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mccarthy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. C.</given-names>
            <surname>Wright</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Olivier</surname>
          </string-name>
          ,
          <article-title>Making design probes work</article-title>
          ,
          <source>in: Sigchi Conference on Human Factors in Computing Systems</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [57]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Holtz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Healy</surname>
          </string-name>
          , From data to viz,
          <year>2018</year>
          , URL https://www. data-to-viz.
          <source>com</source>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [58]
          <string-name>
            <given-names>A.</given-names>
            <surname>Srinivasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Drucker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Endert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Stasko</surname>
          </string-name>
          ,
          <article-title>Augmenting visualizations with interactive data facts to facilitate interpretation and communication</article-title>
          ,
          <source>IEEE transactions on visualization and computer graphics 25</source>
          (
          <year>2018</year>
          )
          <fpage>672</fpage>
          -
          <lpage>681</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [59]
          <string-name>
            <given-names>K.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Orghian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hidalgo</surname>
          </string-name>
          ,
          <article-title>Dive: A mixedinitiative system supporting integrated data exploration workflows</article-title>
          ,
          <source>in: Proceedings of the Workshop on Human-In-the-Loop Data Analytics</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>7</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [60]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. S.</given-names>
            <surname>Njotoprawiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Haleem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Yi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Ma</surname>
          </string-name>
          ,
          <article-title>Embeddingvis: A visual analytics approach to comparative network embedding inspection</article-title>
          ,
          <source>in: 2018 IEEE Conference on Visual Analytics Science and Technology (VAST)</source>
          , IEEE,
          <year>2018</year>
          , pp.
          <fpage>48</fpage>
          -
          <lpage>59</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>