<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>VisualFacts: A Platform for In-Situ Visual Exploration and Real-time Entity Resolution</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>George Papastefanatos</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giorgos Alexiou</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nikos Bikakis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stavros Maroulis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vasilis Stamatopoulos</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>ATHENA Research Center</institution>
          ,
          <country country="GR">Greece</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Nat. Techn. Univ. of Athens</institution>
          ,
          <country country="GR">Greece</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>VisualFacts is an open-source data visualization platform for big geo-located data. VisualFacts combines in-situ visualization with real-time entity resolution capabilities to address scenarios where users wish to visual explore and eficiently perform analytic operations directly on raw data files, which are aggregated from multiple overlapping data sources. VisualFacts is based on an adaptive index for eficiently scaling up to big volumes; it combines a grid with a tree and a blocking-based structure for eficiently processing spatial, categorical and overlapping data, respectively. It is progressively created based on the user interaction, adapting to the areas and details of the user exploration. This paper provides the architecture of the platform and presents a demonstration of its main features.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Data Visualization</kwd>
        <kwd>Visual Analytics</kwd>
        <kwd>Adaptive Indexing</kwd>
        <kwd>In-situ Processing</kwd>
        <kwd>Entity Resolution</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>ift in memory or when complex analytic tasks, such as
real-time ER, are employed.</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        In-situ visual exploration and analysis has become a
common practise due to the availability of big datasets in Contribution. In this demo paper, we present VisualFacts,
raw formats (e.g., csv/json). In-situ techniques attempt to a visual analytics platform for big geo-located data that
avoid the overhead of fully loading and indexing the data assists users perform analysis of raw data files of varying
in a DBMS, and improve performance by progressively quality (with duplicates or missing data) in rich visual
building an index during data exploration. In this work, ways. VisualFacts combines in-situ visualization with
we focus on in-situ visual analysis of data that is collected real-time ER capabilities. It allows users to use their own
and aggregated from multiple sources (e.g., scholarly data, data file(s) and start visually interacting with the data on
POIs, etc) and thus they contain dirty/duplicate entities, a map without loading or indexing the data in a database.
i.e., multiple records in the file may refer to the same The backbone of the platform is a visual aware in-memory
real-world entity. Furthermore, the data can be visual- index, which is constructed on-the-fly and adjusted to
ized and explored on a 2D layout (e.g., map/scatter plot) user interaction, as well as an engine which ofers
on-theand the types of analysis include faceted filtering over lfy visual ER and clustering of duplicate data. Specifically,
categorical attributes, spatial clustering of datapoints, VisualFacts combines and integrates technologies from
computation of aggregate uni/bivariate statistics which RawVis [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and QueryER [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The first is a visual tool for
are visualized in charts and, finally real-time entity res- in-situ visual exploration and the second ofers the ER
olution (ER) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The latter refers to the analysis-aware query engine. The novelty of this work lies in the
detection of matching entities between the data points, extensions we made in the underlying in-situ index
included in the visualized area and the analysis of the to accommodate new structures that can speed up
properties, which characterize them as duplicates. Many real-time ER in visual exploration scenarios, the
commercial and research tools ofer the functionality to integration of the ER query engine in the in-situ
visualize data from raw files [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]; however many of them query engine, and finally the novel UI modules that
sufer from large initialization times and poor interac- enhanced VisualFacts to combine visual analytics
tivity performance, when the files become too large to on geo-located data with ER analysis tasks. With
these enhancements, the platform can scale up the
visualization, interactive exploration and ER analysis to
millions of data points on a map, using commodity
hardware.
      </p>
      <p>VisualFacts as well as the individual technologies are
open source and can be used independently or integrated
into existing data management or visual analysis systems.</p>
      <p>Published in the Workshop Proceedings of the EDBT/ICDT 2022 Joint
Conference (March 29-April 1, 2022), Edinburgh, UK</p>
      <p>0000-0002-9273-9843 (G. Papastefanatos); 0000-0002-6307-4053
(G. Alexiou); 0000-0001-6859-1941 (N. Bikakis);
0000-0003-2816-4368 (S. Maroulis); 0000-0002-9044-796X
(V. Stamatopoulos)</p>
      <p>© 2022 Copyright for this paper by its authors. Use permitted under Creative</p>
      <p>CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g CCoEmmUoRns LWiceonsrekAstthribouptionP4r.0oIncteerenadtiionnagl s(CC(CBYE4U.0)R.-WS.org)
tributes, filtering conditions on the categorical attributes,
aggregate functions, etc) for evaluating a visual user
action. It also denotes whether the user has requested to
perform an ER task on the query results.</p>
      <sec id="sec-2-1">
        <title>Related Work. Techniques for progressive loading and</title>
        <p>
          indexing exist for generic in-situ analysis [
          <xref ref-type="bibr" rid="ref5 ref6 ref7 ref8">5, 6, 7, 8</xref>
          ],
and visualization [
          <xref ref-type="bibr" rid="ref10 ref11 ref9">9, 10, 11</xref>
          ]. Also, query-driven ER tech- Architecture Overview. The web-based frontend (Fig.
niques have been proposed [
          <xref ref-type="bibr" rid="ref1 ref12">1, 12</xref>
          ] in the context of the 1) communicates with the backend via REST API calls.
generic problem of analysis-aware data processing [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. ➊ The user selects a data file A and an initial operation
To the best of our knowledge, this is the first work and on its contents (e.g., select and render the points within a
system that combines these two problems, and considers specific area on the map) to start the analysis. VisualFacts
real-time ER in in-situ visual analytics settings. parses the file and creates an initial version of the index,
called Tile-Tree-Block Index (Index Initialization). During
2. Platform Overview the index construction, the results corresponding to the
ifrst user request are also fetched. ➋ The user follows a
sequence of visual exploration and analysis operations
Exploration Scenario and Operations. Fig. 1 provides (i.e., interactions), which are sent to the backend and
an overview of our platform. The exploration scenario translated to exploratory queries (Operation Translation).
considers that a user visually interacts with data stored ➌ Queries are evaluated over the index structure (Query
in a single data file A on disk using a 2D visualization Evaluation) to compute and fetch the results (data objects
technique (e.g., map, scatter plot) B , and analyzes it us- evaluated by the query). ➍ The query parameters are
ing visual (e.g., bar, line, pie charts, and heatmaps), and used by the Index Adaptation component to adjust the
statistical methods C . The file contains records that cor- index, i.e., re-organize its tile structure and update its
respond to duplicate entities and there is no attribute
(e.g., id) that explicitly identifies duplicate records in the stored statistics. ➎ The results are further processed
ifle. Attribute values are numerical, categorical, or tex- (Data Reduction component) and reduced (e.g., via spatial
tual. Two of the numerical ones (e.g., longitude, latitude) clustering) such that over-plotting issues are properly
are selected explicitly by the user or implicitly by the addressed at the frontend visualization. Simultaneously,
system (based on the datatype) for the X and Y axis of the ER component considers the query results to find
the 2D visualization. and group duplicates into clusters. ➏ The output of
        </p>
        <p>
          A visual operation performed by the user is character- these two components (spatial and duplicate clusters)
ized by a window area and a set of parameters (filters, are combined at the Result Aggregator component and
analytic operations, requested statistics in charts, etc.). sent to the frontend as the final query results. Note that,
At the backend, a visual operation is mapped to the query during the index construction or the query evaluation,
semantics of our framework. An exploratory query [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] is the index structure may not fit in main memory, in such
an SQL-like expression which includes the data access cases, the Eviction Handler component stores parts of the
operations (e.g., a window selection on the X and Y at- index structure on the disk.
        </p>
        <p>At the frontend, the results are visualized on a map B ,
statistics are rendered in various charts C and duplicate
1https://github.com/VisualFacts
2https://visualfacts.imsi.athenarc.gr/software.html</p>
      </sec>
      <sec id="sec-2-2">
        <title>Tile-Tree-Block Index. The Tile-Tree-Block (TTB) is an</title>
        <p>
          in-memory index. The TTB index is an extension of the
VETI index [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] with structures that allow real-time ER
analysis. VETI is a grid index, which organizes the data
objects in non-overlapping rectangle tiles. The grid is
deifned over the domains of the X and Y numeric attributes,
and each tile in the grid is defined from intervals in the
two domains, respectively. Each tile encloses the objects,
whose values for the two axis attributes fall within the
tile intervals. Every tile is associated with a tree, which
organizes the data objects based on values from the
categorical attributes. Each level in the tree corresponds
to a categorical attribute (e.g, assume two categorical
attributes for Country, Gender) and the nodes of a level
correspond to the distinct values found in the objects
contained in this tile (e.g., Greece, Italy are nodes in
Country level). Note that, diferent values capturing
the same concept (e.g., Greece vs. GR) in a categorical
attribute are not sanitized or resolved in the tree; thus, a
level may contain multiple nodes for the same concept.
The reference to a data object is kept in the leaves and it
is the ofset of the record in the data file.
        </p>
        <p>
          The TTB extends the VETI index with an additional
blocking index per tile which groups the ofsets of objects
in a tile in blocks. It considers all the possible attribute
values of the objects in a tile (thus increasing the recall of
the ER) and associates each value (blocking key) with the
list of objects it appears (e.g., Greece: {1, 3, ..., } is
the block for ’Greece’ value with reference to the objects
containing this value in the ’Country’ attribute).
Blocking is a common index structure used in ER Papadakis
et al. [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] for reducing the number of pairwise
comparisons that need to be performed between entities.
        </p>
        <p>Finally, a global link index is used to store the duplicate
relations between the objects, which are detected within
the exploration session of the user. It is used to avoid
performing comparisons and speed up ER tasks for areas
previously visited by the user.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. User Interface Functionality</title>
      <sec id="sec-3-1">
        <title>This section outlines VisualFacts’s user interface (Fig. 2) visualizing data about hotels in US (see Section 4). The basic features include:</title>
      </sec>
      <sec id="sec-3-2">
        <title>Dataset Selection. The user can select to explore one</title>
        <p>of the already added datasets or upload a new CSV file
A . For the exploration and analysis of a new CSV file,
the schema (e.g., latitude and longitude fields, categorical
attributes) is automatically detected by the UI and can be
further customized by the user.</p>
        <p>
          Map-based Visual Exploration. The user is able to
explore and analyze data over diferent geographical areas
using operations like panning and zooming, or focus on a
specific area (e.g., neighborhood) by drawing a rectangle
over the map B . The data objects are clustered on the
map to avoid overplotting issues. By clicking on a green
cluster, the cluster breaks down into an arachnoid of
Index initialization and Query Evaluation. The raw points, where each point represents a single data object.
ifle is parsed and the index is initialized with an initial By clicking on a single point the user can see the details
set of tiles. Details on the initialization policy are pre- about that entity C .
sented in [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. The index progressively adjusts itself to
the user interactions, by splitting visited tiles into more Faceted Filtering. Faceted filtering enables users to
deifne-grained ones. Thus, the grid becomes more dense ifne multiple filters over the categorical attributes via the
in the area that the user explores. The tree and block Filtering dropdown D . For example, in Fig. 2, the user
indexes are spatially bounded within a tile (which keeps has selected to view and analyze the 4-star hotels that
them small in size) and they are rebuilt when a tile is provide fitness facilities. Active filters are summarized
split into new ones, during index adaptation. below the filtering dropdown.
        </p>
        <p>In brief, an exploratory query is evaluated over the Statistics. In the statistics panel E , the user can examine
TTB index as follows: Based on the requested window, univariate (e.g., mean, variance, standard deviation) or
it first identifies the tiles overlapping with the query</p>
        <p>A
D</p>
        <p>G</p>
        <p>C
bivariate statistics (e.g., the Pearson correlation, covari- that appear in its duplicate objects I . Additionally, a pie
ance). The statistics are computed for the data points chart presents the percentage of each attribute value over
being visible in the entire window or the selected rectan- the objects of the cluster K .
gle area and refreshed following every user action on the
map (e.g., pan, zoom). In Figure 2, univariate statistics
are presented for the hotel rating field, and bivariate for 4. Demonstration Outline
the rating and price fields.</p>
      </sec>
      <sec id="sec-3-3">
        <title>In this section, we outline our demonstration scenario.</title>
        <p>Visual Analysis. The user is able to visually analyze the The tool is available at:
data by selecting the most suitable visualization type and https://visualfacts.imsi.athenarc.gr/platform/visualize/hotels.
metrics for their analysis F . For example, in Fig. 2, the The attendees will be able to interact with the tool and
user has selected a heatmap to visualize the average hotel analyze two real-world datasets, regarding Hotels and
price w.r.t. the type and stars of the hotels. Following the Restaurants, respectively. The Hotel dataset contains
a user action, charts are refreshed following the visible records for hotels in NY (about 180K hotels). Each hotel
data points on the map. is described by several attributes, such as name, address,
price, type. We assume that the dataset contains data
Visual-based ER Analysis. The ER analysis is activated retrieved from diferent booking platforms (e.g., Booking,
via a toggle button (merge duplicates) G . When enabled, Trivago), so, multiple records for a hotel may be included
the ER operation is performed for the visible points on the in the data. In brief, we generate this dataset by using as
map and the duplicate records are clustered. A duplicate "base" collection, hotel entries retrieved from the public
cluster is visualized by a red sign annotated with the factual API at 20153. Based on these entries, we generate
number of duplicates H . duplicate records by using diferent values for various
ER Statistics. With data deduplication enabled, the attributes (e.g., name of the hotel, price, location),
repreStatistics E and Analysis F panels on the right are up- senting records coming from diferent booking platforms
dated to present statistics and metrics evaluated over the (details are omitted due to lack of space). The Restaurant
deduplicated data. dataset contains about 180K restaurants from Europe and
is provided by the TripAdvisor4.</p>
        <p>Attribute-based ER Analysis. The user can select and
analyze a specific duplicate cluster on the map, examining
its details and the diferent and common attribute values</p>
      </sec>
      <sec id="sec-3-4">
        <title>3www.factual.com</title>
        <p>4www.kaggle.com/stefanoleone992/tripadvisor-europeanrestaurants</p>
        <p>Users will be able to interact with the prototype and
perform several operations such as:</p>
      </sec>
      <sec id="sec-3-5">
        <title>Innovation (ELIDEK) and by the General Secretariat for</title>
        <p>Research and Technology (GSRT).
• Interact with the map to pan, zoom in/out to find
areas of interest and filter the visualized objects.
• Focus on and analyze specific areas by using the
rectangle selection functionality.
• Select the statistics to examine during the
exploration.
• Select the visualization type, the attributes and
the metrics that will be generated in order to
support their analysis tasks
• Use the ER functionalities to: (1) on-the-fly detect
and visually present clusters of duplicate entries;
(2) examine aggregated statistics w.r.t. the
duplicate entities; and (3) analyze the (common and
diferent) attribute’s values that appear in the
duplicate entities.</p>
        <p>Users will be also presented with specific use case
scenarios and analytic tasks that will provide better insights
into VisualFacts capabilities. Using the Hotel dataset,
we assume a scenario where a data analyst, working for
a consulting company, assists hotels to advertise their
business and oferings across booking platforms. The
analyst uses data from diferent booking platforms (e.g.,
Booking, Trivago) and compares them against hotel’s
location, amenities, rating, and prices to match clients’
presence and oferings (e.g., price per night) with the
most appropriate platform. Assuming that the analyst’s
company specializes in 4-stars hotels, the following tasks
are considered. Initially, the analyst gains an overview
of the booking platforms, which have many 4-star
hotels. She navigates to the location of interest, filters out
the 4-stars hotels and generates a chart with the number
of the hotels per data source. Then, she inspects in a
heatmap the platforms’ coverage for diferent types (e.g.,
business, motel) of hotels and amenities and decides the
one(s) which covers most of her clients. Finally, using
the ER functionalities, she detects and analyzes how the
same hotels are presented in diferent platforms, selects
the hotels she has as clients and investigates on the
differences and similarities of their ofers in the diferent
booking platforms.</p>
        <p>VisualFacts has been assessed by an evaluation study
of 40 users, who were requested to perform visual tasks
similar to the aforementioned and provide their feedback
on the usability and interactivity of the tool. Participants
reported that the tool provides real-time interaction, the
exploratory tasks were easy to perform, and the online
ER task could saved time from manually cleaning and
preparing the data for visualization. Due to lack of space
we omit the details of the study and its findings.</p>
      </sec>
      <sec id="sec-3-6">
        <title>Acknowledgements. The VisualFacts project (1614) has</title>
        <p>been funded by the Hellenic Foundation for Research and</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>H.</given-names>
            <surname>Altwaijry</surname>
          </string-name>
          , S. Mehrotra, ,
          <string-name>
            <given-names>D. V.</given-names>
            <surname>Kalashnikov</surname>
          </string-name>
          ,
          <article-title>Query: A framework for integrating entity resolution with query processing</article-title>
          ,
          <source>in: Proceedings of the VLDB Endowment</source>
          , volume
          <volume>9</volume>
          (
          <issue>3</issue>
          ),
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>N.</given-names>
            <surname>Bikakis</surname>
          </string-name>
          ,
          <article-title>Big Data Visualization Tools, in: Encyclopedia of Big Data Technologies</article-title>
          , Elsevier,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Maroulis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Bikakis</surname>
          </string-name>
          , G. Papastefanatos, P. Vassiliadis,
          <article-title>RawVis: A System for Eficient In-situ Visual Analytics</article-title>
          , in: ACM SIGMOD,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>G.</given-names>
            <surname>Alexiou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Papastefanatos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Stamatopoulos</surname>
          </string-name>
          , G. Koutrika,
          <string-name>
            <given-names>N.</given-names>
            <surname>Koziris</surname>
          </string-name>
          ,
          <article-title>Queryer: A framework for fast analysis-aware deduplication over dirty data (</article-title>
          <year>2022</year>
          ). arXiv:
          <volume>2202</volume>
          .
          <fpage>01546</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>I. Alagiannis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Borovica</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Branco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Idreos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ailamaki</surname>
          </string-name>
          ,
          <source>Nodb: Eficient Query Execution on Raw Data Files, in: ACM ACM SIGMOD</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Karpathiotakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Branco</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Alagiannis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ailamaki</surname>
          </string-name>
          ,
          <source>Adaptive Query Processing on Raw Data, PVLDB</source>
          <volume>7</volume>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Olma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Karpathiotakis</surname>
          </string-name>
          , I. Alagiannis,
          <string-name>
            <given-names>M.</given-names>
            <surname>Athanassoulis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ailamaki</surname>
          </string-name>
          , Slalom:
          <article-title>Coasting through Raw Data Via Adaptive Partitioning and Indexing</article-title>
          , PVLDB
          <volume>10</volume>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Olma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Karpathiotakis</surname>
          </string-name>
          , I. Alagiannis,
          <string-name>
            <given-names>M.</given-names>
            <surname>Athanassoulis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ailamaki</surname>
          </string-name>
          ,
          <article-title>Adaptive partitioning and indexing for in situ query processing</article-title>
          ,
          <source>VLDBJ</source>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Maroulis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Bikakis</surname>
          </string-name>
          , G. Papastefanatos,
          <string-name>
            <given-names>P.</given-names>
            <surname>Vassiliadis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Vassiliou</surname>
          </string-name>
          ,
          <article-title>Resource-Aware Adaptive Indexing for In-situ Visual Exploration and Analytics</article-title>
          ,
          <source>VLDB Journal</source>
          (
          <year>2022</year>
          - to appear).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>N.</given-names>
            <surname>Bikakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Maroulis</surname>
          </string-name>
          , G. Papastefanatos, P. Vassiliadis, In-situ
          <source>Visual Exploration over Big Raw Data, Information Systems</source>
          <volume>95</volume>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>S.</given-names>
            <surname>Maroulis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Bikakis</surname>
          </string-name>
          , G. Papastefanatos,
          <string-name>
            <given-names>P.</given-names>
            <surname>Vassiliadis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Vassiliou</surname>
          </string-name>
          ,
          <article-title>Adaptive indexing for in-situ visual exploration and analytics</article-title>
          , in: DOLAP,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>G.</given-names>
            <surname>Alexiou</surname>
          </string-name>
          , G. Papastefanatos,
          <article-title>Query driven entity resolution in data lakes</article-title>
          , in: Springer, Cham, volume International Workshop on Information Search, Integration, and
          <string-name>
            <surname>Personalization</surname>
          </string-name>
          ,
          <year>2019</year>
          , pp.
          <fpage>117</fpage>
          -
          <lpage>130</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S.</given-names>
            <surname>Giannakopoulou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Karpathiotakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ailamaki</surname>
          </string-name>
          ,
          <article-title>Cleaning denial constraint violations through relaxation</article-title>
          ,
          <source>in: ACM SIGMOD</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>G.</given-names>
            <surname>Papadakis</surname>
          </string-name>
          , G. Alexiou, G. Papastefanatos, , G. Koutrika,
          <article-title>Schema-agnostic vs schema-based conifgurations for blocking methods on homogeneous data</article-title>
          ,
          <source>in: PVLDB</source>
          ,
          <volume>9</volume>
          (
          <issue>4</issue>
          ),
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>