<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Automated Visualization Support for Linked Research Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Belgin Mutlu</string-name>
          <email>bmutlu@know-center.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Patrick Hoefler</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vedran Sabol</string-name>
          <email>vsabol@know-center.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gerwald Tschinkel</string-name>
          <email>gtschinkel@know-center.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michael Granitzer</string-name>
          <email>michael.granitzer@uni-passau.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Know-Center</institution>
          ,
          <addr-line>Graz</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Passau</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <fpage>40</fpage>
      <lpage>44</lpage>
      <abstract>
        <p>Finding, organizing and analyzing research data (i.e. publications) published in various digital libraries are often tedious tasks. Each digital library deploys their own meta-model and technology to query and analyze the knowledge (in further text, scientific facts) contained in research publications. The goal of the EU-funded research project CODE is to provide methods for federated querying and analysis of such data. To achieve this, the CODE project offers a platform, that extracts scientific facts from research data and integrates them within the Linked Data Cloud using a common vocabulary (i.e. meta-model). To support users in analyzing scientific facts, the project provides means for easy-to-use visual analysis. In this paper, we present the web-based CODE Visualization Wizard, which aims to analyze research data visually with an emphasis on automating the visualization process. The main focus of the paper lies on a mapping strategy, which integrates various vocabularies to facilitate the automated visualization process.</p>
      </abstract>
      <kwd-group>
        <kwd>Linked Data</kwd>
        <kwd>Visualization</kwd>
        <kwd>Research Data</kwd>
        <kwd>RDF Data Cube</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Digital libraries, which control the lifecycle of research publications (i.e.
publishing and making them accessible for certain communities) mainly expose the
research knowledge using domain-specific meta-models and technologies.
Moreover, they only focus on some structural attributes and often don’t consider the
content of the publications. This domain-specificity and weakness in specifying
querying attributes limit the ability to effectively find desired information, since
the number of published content is continuously growing. The goal of the CODE1
[4] [5] project is to offer a solution for this issue by providing a platform that
structures (heterogeneous) research data using the RDF Data Cube Vocabulary2
and releases them as Linked Data.</p>
      <p>The RDF Data Cube Vocabulary is a generic vocabulary used to describe
quantitative data (e.g. research results from tables). To simplify the analysis
of this data, the web-based CODE Visualization Wizard3 has been developed,
which integrates several visualizations. To achieve a Linked Data-based
visualization, these visualizations (e.g. charts) should also be described semantically.
For this purpose, we defined the Visual Analytics (VA) Vocabulary4 in the form
of an OWL ontology. This vocabulary is an interface between the RDF Data
Cube and visualization-specific technologies, and together with the RDF Data
Cube Vocabulary it forms the basis for automating the visualization process.</p>
      <p>In this paper, we summarize the current status of the CODE Visualization
Wizard and its ongoing research.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>Semantic description of visualizations using RDF is a new research topic and the
literature, up to now, offers just a few related publications. The most significant
research, the Statistical Graph Ontology [3], comes from the biomedical domain
and presents a new approach to annotate visualizations semantically.</p>
      <p>While the Statistical Graph Ontology provides a sophisticated ground for
describing statistical graphs, some key issues (e.g. the description of size and color
as visualization component or the datatype of a visualization component etc.)
for our applications were missing. This is why we have extended this vocabulary
for our Visualization Wizard.</p>
      <p>At Stanford University, an interactive Web-based visualization system, the
Vispedia [1], has been developed to visualize heterogeneous datasets. The
visualization process of Vispedia is based on the integration of the selected data
into an iterative and interactive data exploration and analysis process enabling
non-experts to more effectively visualize the semi-structured data available.
Vispedia was an inspiration for the Visualization Wizard, but being a Wikipedia
plugin it only supports visualization of Wikipedia data. Also, it does not provide
automatic binding of heterogeneous data onto visualizations.
3</p>
      <p>
        Approach for Automated Visualization Support
In contrast to other available solutions for visualizing Linked Data [2], the CODE
Visualization Wizard automatically suggests suitable visualizations based on
(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) the content and structure of the provided research data and (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) semantic
description of the visualizations. The following parts of our wizard contribute to
these features:
Vocabularies: The RDF Data Cube is a W3C Standard and has been developed
to represent statistical data as RDF. In the CODE project we use this standard
to define the meta-model for the basic research data in order to capture the
3CODE Visualization Wizard: http://code.know-center.tugraz.at/vis
4VA Vocabulary: http://code-research.eu/ontology/visual-analytics
evaluation results from publications. The results are represented as a collection of
observations consisting of a set of dimensions and measures, which represent the
structure of the data. Dimensions identify the observation, measures are related
to concrete values and attributes add semantics to them. For example: when we
have a dataset representing the result of a scientific challenge (such as PAN5)
for several teams, there will be a collection of observations with dimensions
describing the teams with concrete values for the challenge result and with an
attribute percent to identify the unit of the value it is measured in.
      </p>
      <p>Our VA Vocabulary is used to represent the information about visualizations.
It describes the visualization axes and other visual channels, such as color or
size of visual symbols, used to visually represent the data. The vocabulary also
describes suitable datatypes that can be represented by the axes and visual
channels, including the allowed occurrence of the axes and visual channels. The
definition of the occurrence is important to identify whether the axes or the visual
channel can be instantiated only once (e.g. bar chart x-axis) or multiple times
(e.g. parallel coordinates x-axis). In fact, this model is technology-independent
and used by the Visualization Wizard to generate the specific visualization code.
We use in our Wizard the D36 visualization library and Google Charts7 to create
our visualizations but as mentioned above, it is possible to use other technologies.</p>
      <p>Currently, the Visualization Wizard supports nine different charts and a
table. For the integration of each new visualization, a generator needs to be
implemented, which has well-defined interfaces and can be plugged-in to the
Visualization Wizard easily.</p>
      <p>Mapping Vocabularies: The mapping between both mentioned vocabularies,
the RDF Data Cube and the VA Vocabulary, is a relation from dimensions and
measures of the RDF Data Cube (i.e. cube components) to the corresponding
axes and visual channels of the visualization. The mapping combinations will be
found based on the structural compatibility and on the datatype compatibility
between a RDF Data Cube and visualizations.</p>
      <p>
        The number of the dimension and measures in a RDF Data Cube is
unbounded. The possible combinations (i.e. in the format dimension: measure) for
each RDF Data Cube are: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) 1:1, (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) 1:n, (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) n:1 and (
        <xref ref-type="bibr" rid="ref4">4</xref>
        ) n:n. The
structural definition of a visualization represents, how many axes/visual channels
the visualization has. To find a valid mapping, the VA Vocabulary has to
suggest visualizations with the same structural definition like the structural
definition of the corresponding RDF Data Cube. To clarify this, let us analyze the
bar chart from the Figure 1: The bar chart has two axis, x-axis and y-axis,
and can only visualize RDF Data Cubes with one dimension and one measure
(1:1). The structural compatibility is not sufficient for a valid mapping, but also
the datatype compatibility. The datatype compatibility is based on the
primitive datatypes8 (string, integer, float etc.) supported by the both vocabularies
5PAN: http://pan.webis.de/
6D3: http://d3js.org/
7Google Charts: https://developers.google.com/chart/
8Datatypes: http://www.w3.org/TR/2001/REC-xmlschema-2-20010502/
(see Visualization Process). Since the RDF Data Cube may expose composite
datatypes, these must be mapped to supported primitive datatypes.
Visualization Process: Based on the provided RDF Cube model, the
Visualization Wizard proposes (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) visualizations and (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) possible variants of the
mapping (see Fig. 2). The mapping is done by depicting dimensions and measures on
the provided axes or on the visual channels of the visualizations. For instance, a
bar chart consists of two axes: x-axis with a string and y-axis with a decimal
datatype. Here, a dimension of datatype string will be mapped onto the x-axis
and a measure of datatype decimal onto the y-axis (see Figure 1). However, if
there are more dimensions or measures with the same datatype, we have various
mapping variations for a visualization with axes which have the same datatype
like these cube components. In this case (the option 2), the wizard creates a
candidate table including all possible combinations between both models. The
user can choose between different combinations, and for each combination, a
specific visualization will be created and the provided data will be automatically
visualized.
4
      </p>
    </sec>
    <sec id="sec-3">
      <title>Conclusion and Future Work</title>
      <p>The challenge of the first iteration in developing the CODE Visualization
Wizard was to show that pitfalls of traditional visualization principles, such as the
need for the manual work and high maintenance while visualizing datasets, can
be effectively overcome by describing data and visualizations in dedicated
vocabularies and by mapping these vocabularies. From the technical viewpoint,
the main challenge was to automatically determine the right mapping between
instances of the RDF Data Cube and the existing visualizations. Another, and</p>
      <p>G
N
I
P
P
A
M
S
E
I
R
A
LBU RDF Data Cube Vocabulary
A
C
O
V</p>
      <sec id="sec-3-1">
        <title>SS RDF Cube Datasets</title>
        <p>
          E
C
O
R
IILZTPANO Dat1aset
A
U
S
I
V (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) Dataset Query
        </p>
        <sec id="sec-3-1-1">
          <title>Mapping Vocabularies by common datatypes</title>
        </sec>
        <sec id="sec-3-1-2">
          <title>Mapping by code generators</title>
          <p>Visual Analytics Vocabulary
Visual Analytics Vocabulary
Visualization Technology</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Generated Visualizations</title>
      </sec>
      <sec id="sec-3-3">
        <title>Suggested Visualizations</title>
      </sec>
      <sec id="sec-3-4">
        <title>Suggested Mapping</title>
        <p>x-Axis
y-Axis
more serious challenge was to determine only valid suggestions among the
provided visualizations.</p>
        <p>
          The ongoing topics, which are parts of the project’s next iterations, are (
          <xref ref-type="bibr" rid="ref1">1</xref>
          )
the investigation and the implementation of methods on how to use the previous
user’s knowledge (i.e. stored mappings) in order to effectively suggest mappings,
(
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) the extension of the automated visualization model for RDF Data Cubes with
no explicit datatypes and (
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) the implementation of refinement functionalities,
like zooming, filtering etc.
        </p>
        <p>The development of the prototype will continue throughout the rest of the
year, leading to a final evaluation at the beginning of 2014.</p>
        <p>Acknowledgement This work is being developed at the Know Center within the
CODE project funded by the EU Seventh Framework Programme, grant agreement
number 296150. The Know-Center is funded within the Austrian COMET
ProgramCompetence Centers for Excellent Technologiesunder the auspices of the Austrian
Federal Ministry of Transport, Innovation and Technology, the Austrian Federal Ministry
of Economy, Family and Youth and by the State of Styria. COMET is managed by the
Austrian Research Promotion Agency (FFG).</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Chan</surname>
          </string-name>
          et al.
          <article-title>Vispedia: Interactive Visual Exploration of Wikipedia Data via SearchBased Integration</article-title>
          .
          <source>IEEE Trans. Vis. Comput. Graphics</source>
          ,
          <volume>14</volume>
          (
          <issue>6</issue>
          ),
          <year>2008</year>
          ,
          <fpage>1213</fpage>
          -
          <lpage>1220</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Dadzie</surname>
          </string-name>
          et al.
          <article-title>Approaches to visualising linked data: A survey</article-title>
          .
          <source>Semant. web 2</source>
          (
          <issue>2</issue>
          ),
          <year>2011</year>
          ,
          <fpage>89</fpage>
          -
          <lpage>124</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Dumontier</surname>
          </string-name>
          et al.
          <article-title>Modeling and querying graphical representations of statistical data</article-title>
          .
          <source>Web Semant</source>
          .
          <volume>8</volume>
          (
          <issue>2-3</issue>
          ),
          <year>2010</year>
          ,
          <fpage>241</fpage>
          -
          <lpage>254</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Seifert</surname>
          </string-name>
          et al.
          <article-title>Crowdsourcing Fact Extraction from Scientific Literature</article-title>
          .
          <source>Proc. of HCI-KDD 2013 Workshop</source>
          , pp.
          <fpage>160</fpage>
          -
          <lpage>172</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Stegmaier</surname>
          </string-name>
          et al.
          <source>Unleashing Semantics of Research Data. 2nd Workshop on Big Data Benchmarking</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>