<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>and Journal of
Educational Technology &amp; Society (JETS) Special Issue on LAK.
The data are represented in the RDF form</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Paperista: Visual Exploration of Semantically Annotated Research Papers</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nikola Milikic</string-name>
          <email>nikola.milikic@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Uros Krcadinac</string-name>
          <email>sk@uzrok.com</email>
          <email>uros@krcadinac.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jelena Jovanovic</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bojan Brankov</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Srdjan Keca</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Learning Analytics, Visualization, Research Papers</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Faculty of Organizational Sciences, University of Belgrade</institution>
          ,
          <addr-line>Jove Ilića 154, Belgrade 11000, Serbia, +381-11-3950853</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Faculty of Organizational Sciences, University of Belgrade</institution>
          ,
          <addr-line>Jove Ilića 154, Belgrade 11000, Serbia, +381-11-3950853</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Faculty of Organizational Sciences, University of Belgrade</institution>
          ,
          <addr-line>Jove Ilića 154, Belgrade 11000, Serbia, +381-11-3950853</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>UZROK Labs</institution>
          ,
          <addr-line>107 Nehruova, Belgrade 10070, Serbia, +381-61-3115661</addr-line>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>UZROK Labs</institution>
          ,
          <addr-line>107 Nehruova, Belgrade 10070, Serbia, +381-63-581879</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>We consider the problem of visualizing and exploring a dataset about research publications from the fields of Learning Analytics (LA) and Educational Data Mining (EDM). Our approach is based on semantic annotation that associates publications from the dataset with Wikipedia topics. We present a visualization and exploration tool, called Paperista (www.uzrok.com/paperista), which presents these topics in the form of bubble and line charts. The tool provides multiple views, thus allowing users to observe and interact with topics, understand their evolution and relationships over time, and compare data originating from different research fields (i.e., LA and EDM). Moreover, user can explore papers to which the presented topics are related to, and make related Web searches to access the papers themselves.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Categories and Subject Descriptors</title>
      <p>D.2.2 [Software Engineering]: Design Tools and Techniques
user interfaces</p>
    </sec>
    <sec id="sec-2">
      <title>MOTIVATION</title>
      <p>The field of Learning Analytics is emerging in the past few years
and attracting more and more researchers from other areas of
Technology Enhanced Learning (TEL). It aims to address the
current needs in the broad area of education by making use of the
latest trends in information technologies where everything is
moving towards Big Data and real-time analytics.</p>
      <p>
        Learning Analytics (LA) is defined as “the measurement,
collection, analysis and reporting of data about learners and their
contexts, for purposes of understanding and optimising learning
and the environments in which it occurs” [
        <xref ref-type="bibr" rid="ref15">16</xref>
        ]. It is often equated
with other similar fields in the TEL area, such as Academic
Analytics or Educational Data Mining (EDM) [
        <xref ref-type="bibr" rid="ref13">14</xref>
        ]. EDM is a
research field that focuses on using computational approaches,
namely data mining and machine learning, to analyze educational
data in order to facilitate and enhance educational process, and
contribute to the overall improvement of students’ learning
experience [
        <xref ref-type="bibr" rid="ref16">17</xref>
        ]. Even though both LA and EDM are
selfcontained research fields, they are intertwined and overlap in
topics they cover. They share many similarities, but also have
some distinct differences as discussed by Siemens and Baker [
        <xref ref-type="bibr" rid="ref17">18</xref>
        ].
One of the similarities emphasized by these authors is that both
fields reflect the emergence of data-intensive approaches to
education, where both communities have the goal of analyzing
large-scale educational data in order to support research and
practice in education. They differ in the level of automation they
aim to achieve. In particular, EDM has a greater focus on
automating support for educational processes, such as adaptation
and personalization of learning environments and learning
processes. On the other hand, LA has a considerably greater focus
on leveraging human judgment, on informing and empowering
instructors and learners to reflect over and improve learning
processes.
      </p>
      <p>In this paper, we propose an approach to visualizing and exploring
the LAK dataset. It is centered around the topics covered by the
papers from the dataset, and is intended to give an overall view of
the topics that LA and the EDM fields cover. As the focus of
researchers and the degree of relevance of particular topics have
been changing over years, our approach tries to show a trend of
those changes through the whole period the dataset covers,
namely from 2008 to 2012. It also allows for topic-based
exploration of research papers and easy navigation to them.</p>
    </sec>
    <sec id="sec-3">
      <title>2. RELATED WORK</title>
      <p>
        In [
        <xref ref-type="bibr" rid="ref10">11</xref>
        ], authors present an interesting work aimed at automating
the creation of relations between research areas by using
semantically annotated data about research papers in a particular
      </p>
      <sec id="sec-3-1">
        <title>1 www.solaresearch.org/resources/lak-dataset</title>
        <p>
          area. As a continuation of this work, the same authors have
created a tool, called Rexplore2, which, among other things,
visualizes authors migration patterns across research areas [
          <xref ref-type="bibr" rid="ref14">15</xref>
          ].
In terms of visual representation, we find interesting an approach
to visualization of tags (topics) and categories of tags over time.
For example, Dubinko et al. [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] consider the problem of
visualizing the evolution of Flickr tags. The authors present a new
slider-based approach based on a characterization of the most
interesting tags. A Flash-based animation in a web browser allows
the user to observe and interact with the tags. Zhang et al. [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]
present an approach to classification and visualization of temporal
and geographic tag distributions. The authors argue that their
approach can help humans recognize semantic relationships
between tags. Lemma [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] presents the Ebony system, an
application for browsing, navigation, and visualization of the
DBLP database. Wattenberg [7] introduces arc diagrams for
representing complex patterns of repetition in string data.
Watteberg application, the Shape of Song, visualizes music files,
creating a static representation of repetition throughout a time
series. However, to our knowledge, there has been no (published)
research work on the visualization of research topics and
publications in the areas of LA and EDM.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. THE PAPERISTA SYSTEM</title>
      <p>Our approach is illustrated through a Web application called
Paperista. The application visualizes topics associated with
research publications from the LAK dataset, allowing users to
browse through papers, compare LA and EDM research fields,
and make related Web searches. Visualizations are created for
each individual year in order to display relevant topics in the LA
and EDM fields for a specific year, but also for all years
combined in order to give an overall depiction of the topic
distribution in these research areas.</p>
    </sec>
    <sec id="sec-5">
      <title>3.1 Data Preparation and Analysis</title>
      <p>LAK dataset consists of data about conferences and journal papers
published in the LA and EDM research fields in the 2008-2012
period. For each paper, the following elements are available: title,
author(s), abstract, keyword(s) and full text. Also, basic
information about authors is available, such as name and
affiliation.</p>
      <sec id="sec-5-1">
        <title>3.1.1 Topic Extraction</title>
        <p>Since one of the main features of Paperista is visualization of
research topics relevant for the given corpus, the first step in the
data preparation process was to extract main topics of the papers
encompassed by the LAK Dataset. A straightforward approach
was to use keywords associated with the papers. This is because
the authors themselves have compiled those keywords, and it is
them who know the best which topics describe their work in the
most appropriate way. However, the downside of this approach is
that those keywords are given as free form text and are not
consistent with any existing formal vocabulary. This makes them
inconsistent throughout the corpus. Furthermore, the dataset is
incomplete in regard to keywords as for conferences EDM 2008,
2009 and 2010 no keywords are provided.</p>
        <p>Thus, we decided to employ a service for semantic annotation in
order to detect paper topics. We took into consideration two
Wikipedia based semantic annotators: TagMe3 and DBpedia</p>
        <sec id="sec-5-1-1">
          <title>2 http://technologies.kmi.open.ac.uk/rexplore</title>
          <p>
            3 http://tagme.di.unipi.it
Spotlight4. The decision to use Wikipedia based annotator was
motivated by the fact that Wikipedia is the largest corpus of open
encyclopedic knowledge and is often used as a well established
large-scale taxonomy [
            <xref ref-type="bibr" rid="ref7">8</xref>
            ]. Both annotator services are designed to
look for and retrieve recognized Wikipedia concepts from the
given text. They can be configured to the specific needs of any
particular usage scenario (i.e., corpus). TagMe is designed to
identify Wikipedia concepts specifically in short texts. Its REST
API5 allows for configuration of two parameters: i) the rho
parameter which refers to the "goodness" of an annotation with
respect to the topics of the input text, and ii) the epsilon parameter
which is used for fine-tuning the disambiguation process and
indicates whether to favor the most-common topics or to take the
context more into account [
            <xref ref-type="bibr" rid="ref8">9</xref>
            ]. DBpedia Spotlight annotates a
given text with concepts from DBpedia, a structured
representation of Wikipedia [
            <xref ref-type="bibr" rid="ref11">12</xref>
            ]. DBpedia Spotlight REST API6
exposes two parameters: confidence of the annotation process that
takes into account factors such as the topical pertinence and the
contextual ambiguity; support parameter specifies the minimum
number of inlinks7 [
            <xref ref-type="bibr" rid="ref9">10</xref>
            ]. We used only paper title and abstract for
topic extraction, based on an assumption that these two elements
contain mentions of the most important and interesting topics a
paper is related to. In order to decide which service for semantic
annotation to use, the two services were tested with a random
sample comprising 5% of all papers and with different parameter
settings. The best results were achieved by the TagMe service
(rho=0.15; epsilon=0.5). For this reason, TagMe service was
employed to annotate all papers in the corpus.
          </p>
        </sec>
      </sec>
      <sec id="sec-5-2">
        <title>3.1.2 Identifying Popular Topics</title>
        <p>Once having all papers associated with topics, we calculated the
significance of each topic. Numerical statistic called TF-IDF
(Term Frequency – Inverse Document Frequency)8 was used as it
calculates how important a word is to a document in a corpus of
documents. This metric was adapted to our case and used to
calculate the importance of a topic in a paper. Instead of
calculating the frequency of a word, we calculate the frequency of
a topic.</p>
        <p>Since Paperista allows for visualizing topics in a specific year and
overall (in all years, 2008-2012), the significance was calculated
for corpora containing papers from each of these different time
periods. Accordingly, we had six different corpora and calculated
the significance of a topic for each corpus. In order to present only
the most significant topics, we have filtered the topic set to only
those whose significance for a particular period was over 0.01.
This threshold was empirically chosen and presents the best
balance between the relevance of topics and their presentation in
the Paperista’s visualizations (i.e., assuring easy comprehension
by users).</p>
      </sec>
      <sec id="sec-5-3">
        <title>3.1.3 Topic Cleaning</title>
        <p>Even though the output of TagMe service consisted of topics that
are relevant to the papers’ content, some of them can hardly be
considered as relevant research topics in the LA and EDM fields
as they are too general. For instance, topics like Methodology,
4 http://spotlight.dbpedia.org
5 http://tagme.di.unipi.it/tagme_help.html
6http://github.com/dbpedia-spotlight/dbpedia-spotlight/wiki/Webservice
7 Inlinik, or inline link, are incoming links from other DBpedia
concepts to the observed DBpedia concept
8 http://en.wikipedia.org/wiki/Tf-idf
Research, and Experiment can be associated with almost every
paper in this corpus. Actually, these topics can be related to
research papers from almost any other research area. Similarly,
some of the retrieved topics were not relevant to research papers
from the LAK dataset. Such topics resulted from imperfection of
the TagMe tool (and semantic annotation tools, in general). Some
examples of these alien topics include The T.O. Show, Ade Easily,
Henry Snapp, etc. For instance, Henry Snapp topic was apparently
mistaken with the SNAPP tool9, a popular learning analytics and
visualization tool. Hence, it was important to detect and exclude
all these generic and alien topics from the final visualization in
order to reduce the noise.</p>
        <p>
          We applied topic cleaning approach similar to [
          <xref ref-type="bibr" rid="ref10">11</xref>
          ]. The idea is to
identify topics that have little or no relationships with other topics
in the corpus. This can be an indicator that a topic is too specific
or alien to our set of identified topics and thus can be considered
as an exclusion candidate. On the other hand, if a topic has
relationships with too many other topics, this can be an indicator
that a topic is too generic and again should be considered as an
exclusion candidate. In order to detect these outlier topics, we
needed a measure of relatedness between topics. To that end, we
used the Wikipedia Miner10 service that calculates semantic
relatedness of two topics by finding the corresponding Wikipedia
articles, and calculating similarity of those articles by comparing
their incoming and outgoing links [
          <xref ref-type="bibr" rid="ref12">13</xref>
          ]. Wikipedia Miner has a
REST API11 that allows for retrieving this information
programmatically.
        </p>
        <p>Once having relatedness calculated for all the topics in our corpus,
we compiled two lists to help us detect removal candidates. In the
first list, each topic was associated with the number of other topics
that topic is related to. This gave us an insight into which topics
can be considered too general/specific (the higher the number of
related topics, the more generic the topic is, and vice versa). In the
second list, each topic was associated with a sum of its relatedness
with all the other topics. This list was meant to complement the
first one. The rationale here is that there might be a topic with fair
number of relations to other topics, but those relatedness values
are weak. This behavior also qualifies a topic to be considered as
too specific or alien.</p>
        <p>The initial idea with compiling these two lists was that topics to
be removed will be at the beginning and the end of the lists (top
and bottom 10%), and that they could be removed automatically.
However, by examining the lists, among the obvious exclusion
candidates, there were also several topics that should not have
been excluded. For instance, topics like Online tutoring, Process
mining, Educational data mining etc. were at the end of both lists
making them removal candidates, even though these topics are
obviously highly relevant for LA and EDM fields. The reason for
this lays in the nature of Wikipedia itself and the fact that not
many other articles in Wikipedia link to these topics. Thus, the
topic removal process could not be done completely automatically
and an expert in the area was consulted to mark the topics that
should not be excluded.</p>
        <sec id="sec-5-3-1">
          <title>9 http://www.snappvis.org 10 http://wikipedia-miner.cms.waikato.ac.nz 11 http://wikipedia-miner.cms.waikato.ac.nz/services</title>
          <p>Figure 1 - Paperista Interface</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>3.2 Data Visualization and Exploration</title>
      <p>
        The topic visualization applied in Paperista is inspired by the New
York Times visualizations Four Ways to Slice Obama’s 2013
Budget Proposal [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and At the National Conventions, the Words
They Used [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>The Paperista visualization includes bubble and line charts,
allowing users to gain insights into topic trends within the LA and
EDM fields. Bubble charts show the importance of a certain topic
for the entire dataset, each year, and/or each field. By changing
different views, users can watch the changes within the dataset
and compare the two fields. Animated transitions between charts
help users understand these processes. In addition, since the
animation does not show precise changes in topic’s relevancy
(calculated using TF-IDF metric, see Sect. 3.1), users are also
presented with a relevancy line chart for each topic.</p>
      <p>The user interface (Figure 1) consists of an animated bubble
cloud, two button sliders, a sidebar, and an optional timeline. The
first slider button (All Years / By Year) allows users to choose
between the “All Years” and “By Year” views. “All Years” view
presents relevant topics for the entire corpus of publications. “By
Year” view activates a timeline, showing relevant topics for each
year. By using the slider, users can follow the change in topic
relevancy through the years the data is available for (2008-2012).
The second slider button (All Topics / Group Topics) allows for
grouping and regrouping of topics. “All Topics” view shows one
circle-shaped bubble chart. “Group Topics” view divides the chart
into two groups of bubbles. The first group presents topics that
appear only in the EDM field. The third one shows topics related
only to the LA field (i.e., LAK and JETS publications). The group
in the middle shows “mixed” topics, i.e., those that appear at least
once in both EDM and LAK/JETS. Different views of the bubble
chart are presented on Figure 2.</p>
      <p>The size of a bubble represents the topic’s relevancy (i.e., TF–IDF
value). Two research fields, EDM and LA, are color-coded. Each
bubble is divided into two slices the size of which corresponds to
the frequency of that topic within publications of each of the two
sources. For the years 2008-2010, the dataset contains data only
for the EDM research field, so the bubbles are one-colored.
The order of topic bubbles is intended to help users compare the
two fields. The leftmost bubbles represent mostly EDM-related
topics, while the rightmost bubbles mostly belong to the LA field.
Moreover, clicking on a bubble creates a line chart in a sidebar.
The line chart shows the growth and decline of a certain topic.
In addition to the visualization, the Paperista application allows
users to browse papers by topic. When a user clicks on a
particular bubble (topic), a list of papers related to that topic
appears in the right sidebar (represented by a title and a list of
authors). Clicking on the particular paper opens a link to Google
scholar with a name of the article as a search query. Thus, if a
paper is available online, a user could easily obtain the paper
using the Paperista system.
Furthermore, when a user hovers over the paper title, all topics
related to that paper become highlighted. By hovering over
papers, users can gain quick insight about topic connections
between publications. Users can also distinguish papers annotated
with highly relevant topics from those marked with insignificant
ones. This can show which papers are more related to the fields of
EDM and LA, and which can be viewed as “outliers”.</p>
    </sec>
    <sec id="sec-7">
      <title>3.3 Paperista Architecture and Dataset API</title>
      <p>
        The Paperista system consists of a Web application and a server
application that provides RESTful API for communicating with
the dataset. The Web-based visualization is written in D3, a
JavaScript library for manipulating documents based on data [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
We have chosen D3 because of its good performance for
animation and interaction within the Web environment. The
visualization is available at the following address:
www.uzrok.com/paperista.
      </p>
      <p>All data about conference topics and their significance (explained
in Section 3.1) is available as a part of Paperista Dataset API. This
API supports a REST model for accessing the data and it is
available at: http://147.91.128.71:9090/LAKChallenge2013. The
Paperista’s Web application calls these operations in order to
access data from the dataset (for example, a click on a topic
triggers a call to the API, which returns a list of papers).</p>
    </sec>
    <sec id="sec-8">
      <title>4. DISCUSSION</title>
      <p>When looking at the view displaying topic distribution in all years
(Figure 2.1), one can observe that EDM conference dominates in
almost all topics. This is due to the fact that EDM conference is
being organized longer than the LAK conference (3 years longer),
and thus the LAK dataset contains overall more papers coming
from the EDM conference.</p>
      <p>Filtering topics by years allows for observing the popularity of
topics in a particular year and a particular field (LA or EDM).
This further enables one to observe the shift in interest for a
particular topic by researchers in the LA and EDM fields
throughout the years. For instance, one can observe that before
2011, the topic of Learning Analytics was not much popular in the
papers from the EDM field; thus this topic is not displayed at all
in visualizations for years 2008-2010. In 2011, it boomed in
popularity as indicated by the significant rise in the number of
papers covering it. In fact, this was the first year the LAK
conference was organized, and it immediately occupied the
attention of researchers interested in the topic of Learning
Analytics. Interestingly, this topic also gained some traction
among the researchers publishing in the EDM field. In 2012, the
topic’s popularity grew even bigger and the researchers covering
it directed their effort toward the LA field. This resulted in papers
published within the LA field to almost exclusively cover the
topic of Learning Analytics. Similarly, we can observe topics that
have kept high popularity in both areas over years. For instance,
this is the case with the Data topic, obviously as a consequence of
research in both areas concentrating on the analysis of large
amounts of data coming from various learning systems and other
sources.</p>
      <p>The application also allows us to observe that topics such as
Intelligent Tutoring System, Prediction and Accuracy and
Precision mostly kept their popularity throughout the years and
stayed exclusively within the EDM field. On the other hand, one
can observe that the large majority of topics have been covered by
both fields. This suggests that the similarities between the two
fields are significant as they share many research topics.</p>
    </sec>
    <sec id="sec-9">
      <title>5. CONCLUSION</title>
      <p>In this paper we have presented our approach to visualizing topics
and their trends in the LA and EDM fields. Our application allows
for easy identification of the main topics researchers in these
fields have been focusing on, and also exploration of papers
related to those topics.</p>
      <p>When compared to other similar tools that provide visualization of
research topics, our tool is the most similar to the previously
mentioned Rexplore tool. However, while Rexplore is more
focused on relations between authors and topics in research areas,
Paperista’s focus is on research topics and their trends over time.
Also, Paperista allows for exploring papers related to different
topics.</p>
      <p>Future work for Paperista will be primarily directed towards
extending the system to support other datasets, related to other
research areas. Since the LAK dataset is RDF-based, Paperista
can easily be expanded to support other RDF-based datasets
expressed using the same or related vocabulary, such as the
Semantic Web Dog Food corpus12. Regarding the interface, we
plan to introduce keyword-based search functionality for
searching a topic by its name. This would allow for easy
navigation to a desired topic and filtering papers related to it. The
final goal for Paperista is to become a universal visualization tool
for research papers.
12 http://data.semanticweb.org</p>
      <p>Wattenberg, M. Arc Diagrams: Visualizing Structure in
Strings. InfoVis 2002. Available online:
http://hint.fm/papers/arc-diagrams.pdf</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Carter</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <article-title>Four Ways to Slice Obama's 2013 Budget Proposal</article-title>
          . New York Times,
          <year>2012</year>
          . Available online: http://www.nytimes.com/interactive/2012/02/13/us/politics/2 013-budget-proposal-graphic.html
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Bostok</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carter</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Ericson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>At the National Conventions, the Words They Used</article-title>
          . New York Times,
          <year>2012</year>
          . Available online: http://www.nytimes.com/interactive/2012/09/06/us/politics/c onvention-word-counts.html
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Bostok</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ogievetsky</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Heer</surname>
            <given-names>J. D3</given-names>
          </string-name>
          :
          <article-title>Data-Driven Documents</article-title>
          .
          <source>IEEE Trans. Visualization &amp; Comp. Graphics (Proc. InfoVis)</source>
          ,
          <year>2011</year>
          . Available online: http://vis.stanford.edu/papers/d3
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Dubinko</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          et. al.
          <article-title>Visualizing Tags over Time</article-title>
          .
          <source>WWW</source>
          <year>2006</year>
          ,
          <article-title>Edinbourgh</article-title>
          . Available online: http://labs.rightnow.com/colloquium/papers/visualizing_tags. pdf
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Zhang</surname>
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Korayem</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>You</surname>
            <given-names>E.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Crandall D. J. Beyond</surname>
          </string-name>
          Co-occurrence:
          <article-title>Discovering and Visualizing Tag Relationships from Geo-spatial and Temporal Similarities</article-title>
          . Available online: http://www.cs.indiana.edu/~zhanhaip/wsdm2012- clustering.pdf
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Lemma</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <article-title>Visualizing the DBLP Database</article-title>
          .
          <source>Bachelor Thesis</source>
          ,
          <year>2010</year>
          . Available online: http://www.inf.usi.ch/faculty/lanza/Downloads/Lemm2010a. pdf
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Ponzetto</surname>
            ,
            <given-names>S. P.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Strube</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2007</year>
          ,
          <article-title>July)</article-title>
          .
          <article-title>Deriving a large scale taxonomy from Wikipedia</article-title>
          .
          <source>In Proceedings of the national conference on artificial intelligence(</source>
          Vol.
          <volume>22</volume>
          , No.
          <volume>2</volume>
          , p.
          <fpage>1440</fpage>
          ). Menlo Park, CA; Cambridge, MA; London; AAAI Press; MIT Press;
          <year>1999</year>
          . Available online: http://www.hits.org/english/research/nlp/papers/ponzetto07b.pdf
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Ferragina</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Scaiella</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          (
          <year>2010</year>
          ,
          <article-title>October)</article-title>
          .
          <article-title>TAGME: onthe-fly annotation of short text fragments (by wikipedia entities)</article-title>
          .
          <source>In Proceedings of the 19th ACM international conference on Information and knowledge management</source>
          (pp.
          <fpage>1625</fpage>
          -
          <lpage>1628</lpage>
          ). ACM.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Mendes</surname>
            ,
            <given-names>P. N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jakob</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>García-Silva</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2011</year>
          ,
          <article-title>September)</article-title>
          .
          <article-title>Dbpedia spotlight: Shedding light on the web of documents</article-title>
          .
          <source>In Proceedings of the 7th International Conference on Semantic Systems</source>
          (pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          ). ACM.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Osborne</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Motta</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>Mining semantic relations between research areas</article-title>
          .
          <source>The Semantic Web-ISWC</source>
          <year>2012</year>
          ,
          <volume>410</volume>
          -
          <fpage>426</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Auer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bizer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kobilarov</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cyganiak</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Ives</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          (
          <year>2007</year>
          ).
          <article-title>Dbpedia: A nucleus for a web of open data</article-title>
          .
          <source>The Semantic Web</source>
          ,
          <fpage>722</fpage>
          -
          <lpage>735</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Milne</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Witten</surname>
            ,
            <given-names>I. H.</given-names>
          </string-name>
          (
          <year>2008</year>
          ,
          <article-title>October)</article-title>
          .
          <article-title>Learning to link with wikipedia</article-title>
          .
          <source>InProceedings of the 17th ACM conference on Information and knowledge management</source>
          (pp.
          <fpage>509</fpage>
          -
          <lpage>518</lpage>
          ). ACM.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Siemens</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Long</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>Penetrating the Fog: Analytics in Learning and Education</article-title>
          .
          <source>Educause Review</source>
          ,
          <volume>46</volume>
          (
          <issue>5</issue>
          ),
          <fpage>30</fpage>
          -
          <lpage>32</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Osborne</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Motta</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>Making Sense of Research with Rexplore. The Semantic Web-ISWC 2012</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <source>[16] 1st International Conference on Learning Analytics and Knowledge</source>
          , Banff, Alberta,
          <source>February 27-March 1</source>
          ,
          <year>2011</year>
          , link https://tekri.athabascau.ca/analytics/
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Romero</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Ventura</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>Educational data mining: a review of the state of the art</article-title>
          .
          <source>Systems, Man, and Cybernetics</source>
          , Part C:
          <article-title>Applications</article-title>
          and Reviews, IEEE Transactions on,
          <volume>40</volume>
          (
          <issue>6</issue>
          ),
          <fpage>601</fpage>
          -
          <lpage>618</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Siemens</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Baker</surname>
            ,
            <given-names>R. S. D.</given-names>
          </string-name>
          (
          <year>2012</year>
          , April).
          <article-title>Learning analytics and educational data mining: towards communication and collaboration</article-title>
          .
          <source>In Proceedings of the 2nd International Conference on Learning Analytics and Knowledge</source>
          (pp.
          <fpage>252</fpage>
          -
          <lpage>254</lpage>
          ). ACM.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>