<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards Automatic Topical Classification of LOD Datasets</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Robert Meusel</string-name>
          <email>robert@dwslab.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Blerina Spahiu</string-name>
          <email>spahiu@disco.unimib.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Heiko Paulheim</string-name>
          <email>heiko@dwslab.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christian Bizer</string-name>
          <email>chris@dwslab.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Data and Web Science Group, University of Mannheim</institution>
          ,
          <addr-line>B6 26, Mannheim</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Computer</institution>
          ,
          <addr-line>Science, Systems and, Communication</addr-line>
          ,
          <institution>University of Milan Bicocca</institution>
          ,
          <addr-line>Viale Sarca, 336 20126 Milano</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>The datasets that are part of the Linking Open Data cloud diagramm (LOD cloud) are classi ed into the following topical categories: media, government, publications, life sciences, geographic, social networking, user-generated content, and cross-domain. The topical categories were manually assigned to the datasets. In this paper, we investigate to which extent the topical classi cation of new LOD datasets can be automated using machine learning techniques and the existing annotations as supervision. We conducted experiments with di erent classi cation techniques and di erent feature sets. The best classi cation technique/feature set combination reaches an accuracy of 81:62% on the task of assigning one out of the eight classes to a given LOD dataset. A deeper inspection of the classi cation errors reveals problems with the manual classi cation of datasets in the current LOD cloud.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        The Web of Linked Data o ers a rich collection of
structured data provided by hundreds of di erent data sources
that use common standards such as dereferencable URIs
and RDF. The central idea of Linked Data is that data
sources set RDF links pointing at other data sources { e.g.,
owl:sameAs links { so that all data is connected into a global
data space [
        <xref ref-type="bibr" rid="ref3 ref8">3, 8</xref>
        ]. In this data space, agents can navigate
from one data source to another by following RDF links,
thereby discovering new data sources on the y.
      </p>
      <p>
        Since the proposal of the Linked Data best practices in
2006, the Linked Open Data cloud (LOD cloud) has grown to
roughly 1 000 datasets (as of April 2014) [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. The datasets
cover various topical domains, with social media,
government data, and metadata about publications being the most
prominent areas [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>The most well-known categorization of LOD datasets by
topical domain is the coloring of the LOD cloud diagram.1
Up till now, the topical categories were manually assigned
to the datasets in the cloud either by the publishers of the
datasets themselves via the datahub.io dataset catalog or
by the authors of the LOD cloud diagram. In this paper, we
investigate to which extent the topical classi cation of new
LOD datasets can be automated for upcoming versions of
the LOD cloud diagram using machine learning techniques
and the existing annotations as supervision.</p>
      <p>
        Beside creating upcoming versions of the LOD cloud
diagram, the automatic topical classi cation of LOD datasets
can be interesting for other purposes as well: Agents
navigating on the Web of Linked Data should know the topical
domain of datasets that they discover by following links in
order to judge whether the datasets might be useful for their
use case at hand or not. Furthermore, as shown in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], it is
interesting to analyze characteristics of datasets grouped by
topical domain, so that trends and best practices that exist
only in a particular topical domain can be identi ed.
      </p>
      <p>In this paper, we present { to the best of our knowledge {
the rst automatic approach to classify LOD datasets into
the topical categories that are used by the LOD cloud
diagram. Using the data catalog underlying the recent LOD
cloud, we train machine learning classi ers with di erent
sets of features. Our best classi cation technique/feature
set combination reaches an accuracy of 82%.</p>
      <p>The rest of this paper is structured as follows. Section 2
introduces the methodology of our experiments, followed by
a presentation of the results in Section 3 and a discussion of
remaining classi cation errors in Section 4. Section 5 gives
an overview of related work. We conclude with a summary
and an outlook on future work.</p>
      <p>Copyright is held by the author/owner(s).</p>
      <p>WWW2015 Workshop: Linked Data on the Web (LDOW2015).</p>
    </sec>
    <sec id="sec-2">
      <title>METHODOLOGY</title>
      <p>In this section, we rst brie y describe the data corpus
that we use for our experiments and the di erent feature
sets we derive from the data. We than brie y introduce the
classi cation techniques that we considered and sketch the
nal experimental setup that was used for the evaluation.
2.1</p>
    </sec>
    <sec id="sec-3">
      <title>Data Corpus</title>
      <p>
        In order to extract our features for the di erent datasets
which are contained in the LOD cloud, we used the data
corpus that was crawled by Schmachtenberg et al. [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] and
which was used to draw the most recent LOD cloud diagram.
Schmachtenberg et al. used the LD-Spider framework [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]
to gather Linked Data from the Web in April 2014. The
crawler was seeded with URIs from three di erent sources:
(1) dataset descriptions in lod-cloud group of the datahub.io
dataset catalog, as well as other datasets marked with Linked
Data related tags within the catalog; (2) a sample of the
Billion Triple Challenge 2012 dataset2; and (3) datasets
advertised on the public-lodw3.org mailing list since 2011. The
nal crawl contains data from 1 014 di erent LOD datasets.3
Altogether 188 million RDF triples were extracted from 900 129
documents describing 8 038 396 resources. Figure 1 shows
the distribution of the number of resources and documents
per dataset contained in the crawl.
      </p>
      <p>
        In order to create the 2014 version of the LOD cloud
diagram, newly discovered datasets were manually classi ed
into one of the following categories: media, government,
publications, life sciences, geographic, social networking,
usergenerated content, and cross-domain. A detailed de nition
of each category is available in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>Figure 2 shows the number of datasets per category
contained in the 2014 version of the LOD cloud. As we can
see, the LOD cloud is dominated by datasets belonging to
the category social networking (48%), followed by
government (18%) and publications (13%) datasets. The categories
media and geographic are only represented by less than 25
datasets within the whole corpus.
2.2</p>
    </sec>
    <sec id="sec-4">
      <title>Feature Sets</title>
      <p>For each of the datasets, we created the following eight
feature sets based on the crawled data.</p>
      <p>Vocabulary Usage (VOC): As many vocabularies target
a speci c topical domain, e.g. bibo bibliographic
information, we assume that the vocabularies that are
2http://km.aifb.kit.edu/projects/btc-2012/
3The crawled data is publicly available: http://data.dws.
informatik.uni-mannheim.de/lodcloud/2014/ISWC-RDB/</p>
      <p>
        used by a dataset form a helpful indicator for
determining the topical category of the dataset. Thus, we
determine the vocabulary of all terms that are used as
predicates or as the object of a type statement within
each dataset. Altogether we identi ed 1 439 di erent
vocabularies being used by the datasets (see [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] for
details about the most widely used vocabularies).
Class URIs (CUri): As a more ne-grained feature, the
rdfs: and owl:classes which are used to describe
entities within a dataset might provide useful
information to determine the topical category of the dataset.
Thus, we extracted all the classes that are used by at
least two di erent datasets, resulting in 914 attributes
for this feature set.
      </p>
      <p>Property URIs (PUri): Beside the class information of
an entity, information about which properties are used
to describe the entity can be helpful. For example it
might make a di erence, if a person is described with
foaf:knows statements or if her professional a liation
is provided. To leverage this information, we collected
all properties that are used within the crawled data by
at least two datasets. This feature set consists of 2 333
attributes.</p>
      <p>Local Class Names (LCN): Di erent vocabularies might
contain synonymous (or at least closely related) terms
that share the same local name and only di er in their
namespace, e.g. foaf:Person and dbpedia:Person.
Creating correspondences between similar classes from
di erent vocabularies reduces the diversity of features,
but on the other side might increase the number of
attributes which are used by more than one dataset.
As we lack correspondences between all the
vocabularies, we bypass this, by using only the local names
of the type URIs, meaning vocab1:Country and
vocab2:Country are mapped to the same attribute. We
used a simple regular expression to determine the
local class name checking for #, : and / within the type
object. By focusing only on the local part of a class
name, we increase the number of classes that are used
by more than one dataset in comparison to CUri and
thus generate 1 041 attributes for the LCN feature set.
Local Property Names (LPN): Using the same
assumption as for the LCN feature set, we also extracted the
local name of each property that is used by a dataset.
This results in treating vocab1:name and vocab2:name
as a single property. We used the same heuristic for
the extraction as for the LCN feature set and generated
3 493 di erent local property names which are used by
more than one dataset, resulting in an increase of the
number of attributes in comparison to the PUri feature
set.</p>
      <p>Text from rdfs:label (LAB): Beside the vocabulary-level
features, the names of the described entities might
also indicate the topical domain of a dataset. We
thus extracted all values of rdfs:label properties,
lower-cased them, and tokenized the values at
spacecharacters. We further excluded tokens shorter than
three and longer than 25 characters. Afterward, we
calculated the TF-IDF value for each token while
excluding tokens that appeared in less than 10 and more
than 200 datasets, in order to reduce the in uence of
noise. This resulted in a feature set consisting of 1 440
attributes.</p>
      <p>Top-Level Domains (TLD): Another feature which might
help to assign datasets to topical categories is the
toplevel domain of the dataset. For instance,
government data is often hosted in the gov top-level domain,
whereas library data might be found more likely on
edu or org top-level domains.4
In &amp; Outdegree (DEG): In addition to vocabulary-based
and textual features, the number of outgoing RDF
links to other datasets and incoming RDF links from
other datasets could provide useful information for
classifying the datasets. This feature could give a hint
about the density of the linkage of a dataset, as well
as the way the dataset is interconnected within the
whole LOD cloud ecosystem.</p>
      <p>We were able to create all features (except LAB) for 1 001
datasets. As only 470 datasets provide rdfs:labels, we
only use these datasets for evaluating the utility of the LAB
feature set.</p>
      <p>As the total number of occurrences of vocabularies and
terms is heavily in uenced by the distribution of entities
within the crawl for each dataset, we apply two di erent
normalization strategies to the values of the vocabulary-level
features VOC, CUri, PUri, LCN, and LPN: On the one hand
side, we create a binary version (bin) where the feature
vectors of each feature set consist of 0 and 1 indicating presence
and absence of the vocabulary or term. The second version,
the relative term occurrence (rto), captures the fraction of
vocabulary or term usage for each dataset.</p>
      <p>The following table shows an example of the two di erent
feature set versions for the terms ti:</p>
      <p>Feature Set Version
Term Occurrence
Binary (bin)
Relative Term Occurrence (rto)
Feature Vector
t1 t2 t3
10 0 2
1 0 1
0:5 0 0:1
4We restrict ourselves to top-level domains, and not public
su xes.
k-Nearest Neighbor: k-Nearest Neighbor (k-NN) classi
cation models make use of the similarity between new
cases and known cases to predict the class for the new
case. A case is classi ed by its majority vote of its
neighbors, with the case being assigned to the class
most common among its k nearest neighbors measured
by the distance function. In our experiments we used a
k equal to 5 with Euclidean-similarity for non-binary
term vectors and J accard-similarity for binary term
vectors.</p>
      <p>
        J48 Decision Tree: A decision tree is a owchart-like tree
structure which is built top-down from a root node and
involves some partitioning steps to divide data into
subsets that contain instances with similar values. For
our experiments we use the Weka implementation of
the C4.5 decision tree [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. We learn a pruned tree,
using a con dence threshold of 0:25 with a minimum
number of 2 instances per leaf.
      </p>
      <p>
        Naive Bayes: As a last classi cation method, we used Naive
Bayes (NB). NB uses joint probabilities of some
evidence to estimate the probability of some event.
Although this classi er is based on the assumption that
all features are independent, which is violated in many
use cases, NB has shown to work well in practice [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
2.4
      </p>
    </sec>
    <sec id="sec-5">
      <title>Experimental Setup</title>
      <p>In order to evaluate the performance of the three classi
cation methods, we use 10-fold cross-validation and report
the average accuracy in the end.</p>
      <p>As the number of datasets per category is not equally
distributed within the LOD cloud, which might in uence the
performance of the classi cation models, we also explore the
e ect of balancing the training data. We used two di erent
balancing approaches: (1) we down sample the number of
datasets used for training until each category is represented
by the same number of datasets; this number is equal to
the number of datasets within the smallest category; and
(2) we up sample the datasets for each category until each
category is at least represented by the number of datasets
equal to the number of datasets of the largest category. The
rst approach, reduces the chance to over t a model into the
direction of the larger represented classes, but it might also
remove valuable information from the training set, as
examples are removed and not taken into account for learning
the model. The second approach, ensures that all possible
examples are taken into account and no information is lost
for training, but by creating the same entity many times
can result in emphasizing those particular data points. For
example a neighborhood based classi er might look at the
5 nearest neighbors, which than could be one and the same
data point, which would result into looking only at the
nearest neighbor.
3.</p>
    </sec>
    <sec id="sec-6">
      <title>RESULTS</title>
      <p>In the following, we rst report the results of our
experiments using the di erent feature sets in separation.
Afterward, we report the results of experiments combining
attributes from multiple feature sets.</p>
      <p>Table 1 shows the accuracy that is reached using the three
di erent classi cation algorithms with and without
balancing the training data. Majority Class is the performance
of a default baseline classi er always predicting the largest
class: social networking.</p>
      <p>As a general observation, the vocabulary-based feature
sets (VOC, LCN, LPN, CUri, PUri) perform on a similar
level, where DEG and TLD alone show a relatively poor
performance and in some cases are not at all able to beat
the majority class baseline. Classi cation models based on
the attributes of the LAB feature set perform on average
(without sampling) around 20% above the majority
baseline, but predict still in half of all cases the wrong category.
Algorithm-wise, the best results are achieved using the
decision tree (J48) without balancing (maximal accuracy 80:59%
for LCNrto) and the k-NN algorithm, also without
balancing for the PUribin and LPNbin feature sets. Comparing
the two balancing approaches, we see better results using
the up sampling approach for almost all feature sets (except
VOCrto and DEG). In most cases, the category-speci c
accuracy of the smaller categories is higher when using up
sampling. Using down sampling the learned models make
more errors for predicting the larger categories.
Furthermore, when comparing the results of the models trained on
unbalanced data with the best model trained on balanced
data, the models on the unbalanced data are more accurate
except for the VOCbin feature set. Having a closer look at
the confusion matrices, we see that the balanced approaches
are in general making more errors when trying to predict
datasets for the larger categories, like social networking and
government.
3.2</p>
    </sec>
    <sec id="sec-7">
      <title>Results for Combined Feature Sets</title>
      <p>For our second set of experiments, we combine the
available attributes from the di erent feature sets and train again
our classi cation models using the three described algorithms.
As before, we generate a binary and relative term occurrence
version of the vocabulary-based features. In addition, we
create a second set (binary and relative term occurrence ),
where we omit the attributes from the LAB feature set, as
we wanted to measure the in uence of this particular set
of attributes, which is only available for less than half of
the datasets. Furthermore we created a combined set of
attributes consisting of the three best performing feature sets
from the previous section.
ALLrto: Combination of the attributes from all eight
feature sets, using the rto version of the vocabulary-based
features.</p>
      <p>ALLbin: Combination of the attributes from all eight
feature sets, using the bin version of the vocabulary-based
features.</p>
      <p>NoLabrto: Combination of the attributes from all feature,
without the attributes of the LAB feature set, using
the rto version of the vocabulary-based features.
NoLabbin: Combination of the attributes from all feature,
without the attributes of the LAB feature set, using
the bin version of the vocabulary-based features.
Best3: Includes the attributes from the three best
performing feature sets from the previous section based on
their average accuracy: PUribin, LCNbin, and LPNbin.</p>
      <p>We can observe that when selecting a larger set of
attributes, our model is able to reach a slightly higher accuracy
of 81:62% than using just the attributes from one feature set
(80:59%, LCNbin). Still the trained model is unsure for
certain decisions and has a stronger bias towards the categories
publications and social networking.
4.</p>
    </sec>
    <sec id="sec-8">
      <title>DISCUSSION</title>
      <p>In the following, we look at the best performing approach
(Naive Bayes trained on the attributes of the NoLabbin
feature set using up sampling). Table 3 shows the confusion
matrix of this experiment, where on the left side we list the
predictions by the learned model, while the head names the
actual category of the dataset. As observed in the table,
there are three kinds of errors which occur more frequently
than 10 times.</p>
      <p>The most common confusion occurs for the publication
domain, where a larger number of datasets are predicted to
belong to the government domain. A reason for this is that
government datasets often contain metadata about
government statistics which are represented using the same
vocabularies and terms (e.g. skos:Concept ) that are also used
in the publication domain. This makes it challenging for
a vocabulary-based classi er to distinguish those two
categories apart. In addition, for example the http://mcu.es
dataset { the Ministry of Culture in Spain { was
manually labeled as publication within the LOD cloud, whereas
the model predicts government which turns out to be a
borderline case in the gold standard. A similar frequent
problem is the prediction of life sciences for datasets in the
publications category. This can be observed, e.g., for the
http://ns.nature.com/publications/, which describe the
publications in Nature. Those publications, however, are
often in the life sciences eld, which makes the labeling in the
gold standard a borderline case.</p>
      <p>
        The third most common confusion occurs between the
usergenerated content and the social networking domain.
Here, the problem is in the shared use of similar
vocabularies, such as foaf. At the same time, labeling a dataset as
either one of the two is often not so simple. In [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], it has been
de ned that social networking datasets should focus on the
presentation of people and their interrelations, while
usergenerated content should have a stronger focus on the
content. Datasets from personal blogs, such as www.wordpress.com,
however, can convey both aspects. Due to the labeling rule,
these datasets are labeled as usergenerated content, but our
approach frequently classi es them as social networking.
      </p>
      <p>In summary, while we observe some true classi cation
errors, many of the mistakes made by our approach actually
point at datasets which are di cult to classify, and which
are rather borderline cases between two categories.</p>
    </sec>
    <sec id="sec-9">
      <title>RELATED WORK</title>
      <p>
        Topical pro ling has been studied in the data mining,
database, and information retrieval communities. The
resulting methods nd application in domains such as
documents classi cation, contextual search, content management
and review analysis [
        <xref ref-type="bibr" rid="ref1 ref11 ref16 ref17 ref2">1, 11, 2, 16, 17</xref>
        ].
      </p>
      <p>Although topical pro ling has been studied in other
settings before, only a small number of methods exist for pro
ling LOD datasets. These methods can be categorized based
on the general learning approach that is employed into the
categories unsupervised and supervised. Where the rst
category does not rely on labeled input data, the latter is only
applicable for labeled data.</p>
      <p>
        Elle et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] try to de ne the pro le of datasets using
semantic and statistical characteristics. They use statistics
about vocabulary, property, and datatype usage, as well as
statistics on property values, like string lengths, for
characterizing datasets. For classi cation, they propose a
feature/characteristic generation process, starting from the top
discovered types of a dataset and generating property/value
pairs. In order to integrate the property/value pairs they
consider the problem of vocabulary heterogeneity of the datasets
by de ning correspondences between features in di erent
vocabularies. The authors have pointed out that it is
essential to automate the feature generations and proposed the
framework to do so, but do not evaluate their approach on
real-world datasets. In our work, we draw from their ideas of
using schema-usage characteristics as features for the topical
classi cation, but focus on LOD datasets.
      </p>
      <p>
        An approach to detect latent topics in
entity-relationship graphs is introduced by Bohm et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Their
approach works in two phases: (1) A number of subgraphs
having strong relations between classes are discovered from
the whole graph, and (2) the subgraphs are combined to
generate a larger subgraph, which is assumed to represent
a latent topic. Their approach explicitly omits any kind of
features based on textual representations and solely relies on
the exploitation of the underlying graph. Bohm et al. used
the DBpedia dataset to evaluate their approach.
      </p>
      <p>
        Fetahu et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] propose an approach for creating dataset
pro les represented by a weighted dataset-topic graph which
is generated using the category graph and instances from
DBpedia. In order to create such pro les, a processing
pipeline that combines tailored techniques for dataset
sampling, topic extraction from reference datasets, and relevance
ranking is used. Topics are extracted using
named-entityrecognition techniques, where the ranking of the topics is
based on their normalized relevance score for a dataset.
      </p>
      <p>While the mentioned approaches are unsupervised, we
employ supervised learning techniques as we want to exploit the
existing topical annotation of the datasets in the LOD cloud.</p>
    </sec>
    <sec id="sec-10">
      <title>CONCLUSION AND FUTURE WORK</title>
      <p>In this paper, we investigate to which extent the topical
classi cation of new LOD datasets can be automated using
machine learning techniques. Our experiments indicate that
vocabulary-level features are a good indicator for the topical
domain, yielding an accuracy of around 82%.</p>
      <p>
        The analysis of the limitations of our approach, i.e., the
cases where the automatic classi cation deviates from the
manually labeled one, points to a problem of the
categorization approach that is currently used for the LOD cloud: All
datasets are labeled with exactly one topical category,
although sometimes two or more categories would be equally
appropriate. One such example are datasets describing life
science publications, which can be either labeled as
publications or as life sciences. Thus, the LOD dataset classi
cation task might be more suitably formulated as a multi-label
classi cation problem [
        <xref ref-type="bibr" rid="ref10 ref18">18, 10</xref>
        ].
      </p>
      <p>
        A particular challenge of the classi cation is the heavy
imbalance of the dataset categories, with roughly half of the
datasets belonging to the social networking domain. Here, a
two-stage approach might help, in which a rst classi er tries
to separate the largest category from the rest, while a second
classi er then tries to make a prediction for the remaining
classes. When regarding the problem as a multi-label
problem, the corresponding approach would be classi er chains,
which make a prediction for one class after the other, taking
the prediction of the rst classi ers into account as a feature
for the remaining classi cations [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>
        In our experiments, RDF links have not been exploited
beyond dataset in- and out-degree. For the task of web
page classi cation, link-based classi cation techniques, that
exploit the contents of web pages linking to a particular
page, often yields good results [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and it is possible that
such techniques could also work well for classifying LOD
datasets.
      </p>
    </sec>
    <sec id="sec-11">
      <title>Acknowledgements</title>
      <p>This research has been supported in part by FP7/2013-2015
COMSODE (under contract number FP7-ICT-611358).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>C. C.</given-names>
            <surname>Aggarwal</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhai</surname>
          </string-name>
          .
          <article-title>A survey of text clustering algorithms</article-title>
          .
          <source>In Mining Text Data</source>
          , pages
          <volume>77</volume>
          {
          <fpage>128</fpage>
          .
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>T.</given-names>
            <surname>Basu</surname>
          </string-name>
          and
          <string-name>
            <given-names>C. A.</given-names>
            <surname>Murthy</surname>
          </string-name>
          .
          <article-title>E ective text classi cation by a supervised feature selection approach</article-title>
          .
          <source>In 12th IEEE International Conference on Data Mining Workshops</source>
          , ICDM Workshops, Brussels, Belgium, December
          <volume>10</volume>
          ,
          <year>2012</year>
          , pages
          <fpage>918</fpage>
          {
          <fpage>925</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Heath</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Berners-Lee</surname>
          </string-name>
          .
          <article-title>Linked data - the story so far</article-title>
          .
          <source>Int. J. Semantic Web Inf. Syst.</source>
          ,
          <volume>5</volume>
          (
          <issue>3</issue>
          ):1{
          <fpage>22</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>C.</given-names>
            <surname>Bo</surname>
          </string-name>
          hm, G. Kasneci, and
          <string-name>
            <given-names>F.</given-names>
            <surname>Naumann</surname>
          </string-name>
          .
          <article-title>Latent topics in graph-structured data</article-title>
          .
          <source>In 21st ACM International Conference on Information and Knowledge Management</source>
          ,
          <source>CIKM'12</source>
          ,
          <string-name>
            <surname>Maui</surname>
            ,
            <given-names>HI</given-names>
          </string-name>
          , USA, October 29 - November 02,
          <year>2012</year>
          , pages
          <fpage>2663</fpage>
          {
          <fpage>2666</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M. B.</given-names>
            <surname>Elle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Bellahsene</surname>
          </string-name>
          , F. Schar e, and
          <string-name>
            <given-names>K.</given-names>
            <surname>Todorov</surname>
          </string-name>
          .
          <article-title>Towards semantic dataset pro ling</article-title>
          .
          <source>In Proceedings of the 1st International Workshop on Dataset PROFIling</source>
          &amp;
          <article-title>fEderated Search for Linked Data co-located with the 11th Extended Semantic Web Conference</article-title>
          ,
          <source>PROFILES@ESWC</source>
          <year>2014</year>
          , Anissaras, Crete, Greece, May
          <volume>26</volume>
          ,
          <year>2014</year>
          .,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>B.</given-names>
            <surname>Fetahu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dietze</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. P.</given-names>
            <surname>Nunes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Casanova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Taibi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W.</given-names>
            <surname>Nejdl</surname>
          </string-name>
          .
          <article-title>A scalable approach for e ciently generating structured dataset topic pro les</article-title>
          .
          <source>In The Semantic Web: Trends and Challenges - 11th International Conference, ESWC</source>
          <year>2014</year>
          , Anissaras, Crete, Greece, May
          <volume>25</volume>
          -29,
          <year>2014</year>
          . Proceedings, pages
          <volume>519</volume>
          {
          <fpage>534</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>L.</given-names>
            <surname>Getoor</surname>
          </string-name>
          and
          <string-name>
            <given-names>C. P.</given-names>
            <surname>Diehl</surname>
          </string-name>
          .
          <article-title>Link mining: a survey</article-title>
          .
          <source>ACM SIGKDD Explorations Newsletter</source>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>T.</given-names>
            <surname>Heath</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          .
          <article-title>Linked data: Evolving the web into a global data space</article-title>
          .
          <source>Synthesis lectures on the semantic web: theory and technology</source>
          ,
          <volume>1</volume>
          (
          <issue>1</issue>
          ):1{
          <fpage>136</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>R.</given-names>
            <surname>Isele</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Umbrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Harth. LDSpider</surname>
          </string-name>
          :
          <article-title>An open-source crawling framework for the web of linked data</article-title>
          .
          <source>In Proc. ISWC '10 {Posters and Demos</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>B.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. S.</given-names>
            <surname>Lee</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P. S.</given-names>
            <surname>Yu</surname>
          </string-name>
          .
          <article-title>Text classi cation by labeling words</article-title>
          .
          <source>In Proceedings of the Nineteenth National Conference on Arti cial Intelligence, Sixteenth Conference on Innovative Applications of Arti cial Intelligence, July 25-29</source>
          ,
          <year>2004</year>
          , San Jose, California, USA, pages
          <volume>425</volume>
          {
          <fpage>430</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Nam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <surname>E.</surname>
          </string-name>
          <article-title>Loza Menc a, I. Gurevych, and</article-title>
          <string-name>
            <given-names>J.</given-names>
            <surname>Fu</surname>
          </string-name>
          <article-title>rnkranz. Large-scale multi-label text classi cation - revisiting neural networks</article-title>
          .
          <source>In Machine Learning and Knowledge Discovery in Databases - European Conference, ECML PKDD</source>
          <year>2014</year>
          , Nancy, France,
          <source>September 15-19</source>
          ,
          <year>2014</year>
          . Proceedings,
          <string-name>
            <surname>Part</surname>
            <given-names>II</given-names>
          </string-name>
          , pages
          <volume>437</volume>
          {
          <fpage>452</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Quinlan</surname>
          </string-name>
          .
          <source>C4</source>
          .
          <article-title>5: Programs for Machine Learning</article-title>
          . Morgan Kaufmann Publishers Inc., San Francisco, CA, USA,
          <year>1993</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J.</given-names>
            <surname>Read</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Pfahringer</surname>
          </string-name>
          , G. Holmes, and
          <string-name>
            <given-names>E.</given-names>
            <surname>Frank</surname>
          </string-name>
          .
          <article-title>Classi er chains for multi-label classi cation</article-title>
          .
          <source>In European Conference on Machine Learning and Knowledge Discovery in Databases</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>I. Rish.</surname>
          </string-name>
          <article-title>An empirical study of the naive bayes classi er</article-title>
          .
          <source>In IJCAI 2001 workshop on empirical methods in arti cial intelligence</source>
          , volume
          <volume>3</volume>
          , pages
          <fpage>41</fpage>
          {
          <fpage>46</fpage>
          . IBM New York,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Schmachtenberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Bizer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Paulheim</surname>
          </string-name>
          .
          <article-title>Adoption of the linked data best practices in di erent topical domains</article-title>
          .
          <source>In The Semantic Web{ISWC</source>
          <year>2014</year>
          , pages
          <fpage>245</fpage>
          {
          <fpage>260</fpage>
          . Springer,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>P.</given-names>
            <surname>Shivane</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Rajani</surname>
          </string-name>
          .
          <article-title>A survey on e ective quality enhancement of text clustering &amp; classi cation using metadata.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>G.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Huang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Bie</surname>
          </string-name>
          .
          <article-title>Short text classi cation: A survey</article-title>
          .
          <source>Journal of Multimedia</source>
          ,
          <volume>9</volume>
          (
          <issue>5</issue>
          ):
          <volume>635</volume>
          {
          <fpage>643</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>G.</given-names>
            <surname>Tsoumakas</surname>
          </string-name>
          and
          <string-name>
            <surname>I. Katakis.</surname>
          </string-name>
          <article-title>Multi label classi cation: An overview</article-title>
          .
          <source>International Journal of Data Warehousing and Mining</source>
          ,
          <volume>3</volume>
          (
          <issue>3</issue>
          ):1{
          <fpage>13</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>