<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Ensemble Learning for Multi-type Classi cation in Heterogeneous Networks (Discussion Paper)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Francesco Sera no</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gianvito Pio</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michelangelo Ceci</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Bari, Dept. of Computer Science</institution>
          ,
          <addr-line>Via Orabona, 4, 70125 Bari</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <fpage>24</fpage>
      <lpage>27</lpage>
      <abstract>
        <p>In the literature, several methods have been proposed for the analysis of network data, but they usually focus on homogeneous networks. More recently, the complexity of real scenarios has impelled researchers to design classi cation and clustering methods able to work also on heterogeneous networks, which consist of di erent types of objects and links. However, they often make assumptions on the structure of the network that are too restrictive or do not exploit di erent forms of network correlation and autocorrelation. Moreover, when several nodes of the network have missing values, standard methods can lead to either building incomplete classi cation models or to discarding relevant dependencies (correlation or autocorrelation). In this discussion paper, we describe an ensemble learning approach for multi-type classi cation that we proposed recently, which is able to exploit i) the possible presence of correlation and autocorrelation phenomena, and ii) the classi cation of instances (with missing values) of other node types in the network. Experiments performed on real-world datasets show that the proposed method is able to signi cantly outperform state-of-the-art algorithms.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>In the real world we can easily nd objects which appear to be connected to
each other, thus forming complex networks. Connections among these objects
can represent di erent types of relationships, which can be found in several elds,
including biology, epidemiology, geography, nance, etc. Most of the works in the
literature about mining networked data focus on homogeneous networks, where
all the objects are of the same type (and can, accordingly, be represented by a
prede ned set of features/characteristics) and the links among them describe a
single type of relationship. A common example of a homogeneous network is that
of social networks: objects represent people and links represent the friendship
relationships. However, real scenarios are more complex due to the presence of
multiple types of objects that are connected through di erent types of links,
forming heterogeneous networks. For example, in well-known databases about
movies (e.g., IMDb) we have movies, actors, users, tags, etc. In the bio-medical
domain, databases contain genes, proteins, tissues, pathways and diseases. In all
these cases, objects of di erent types establish di erent types of relationships.</p>
      <p>
        Consequently, recent works have proposed new data mining methods that
work on heterogeneous information networks. For example, in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] and [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] the
authors propose new clustering solutions, whereas in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] the authors
propose classi cation/prediction methods. Other methods that were initially
proposed for relational data can be almost directly applied for the analysis of
heterogeneous information networks [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. However, existing methods su er from one
or more of the following limitations: i) they impose strict restrictions on the
structure of the network, which must be know apriori; ii) they are not able to
take into account (and possibly exploit) one of the main peculiarities of network
data, i.e. the presence of di erent forms of autocorrelation [
        <xref ref-type="bibr" rid="ref1 ref14">1, 14</xref>
        ], according to
which two connected nodes in the network tend to share some properties; iii)
they are not able to consider the possible presence of missing values for some
attributes, which can lead to either learning incomplete classi cation models or
to discarding possibly relevant dependencies; iv) they are not able to classify
objects of di erent types, where each type can have a di erent set of labels.
      </p>
      <p>
        In this discussion paper, we describe the approach proposed in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], which
is able to work on heterogeneous networks with arbitrary structures and is able
to capture both correlation and autocorrelation phenomena which involve the
target objects (i.e., objects which are the main subject of the classi cation task).
Moreover, we exploit the same strategy to predict possibly relevant missing
values belonging to other objects, appearing strongly related to the target objects.
Methodologically, we extend the method Mr-SBC [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], in order to also capture
network autocorrelation phenomena, handle relevant missing values and
perform multi-type classi cation. These last three issues are tackled by resorting to
a combined bagging-boosting ensemble learning solution able to exploit
information conveyed by objects (also of the same type) directly or indirectly connected
to the main subject(s) of the classi cation task. Speci cally, we propose two
extensions of the Mr-SBC algorithm: i) ST-MrSBC (Self-Training MrSBC),
which is able to capture possible autocorrelation phenomena by resorting to a
variant of the self-training method; ii) MT-MrSBC (Multi-Type MrSBC),
which iteratively analyzes objects of multiple types, in order to predict possible
missing values, belonging either to the target type of the main classi cation task
or to other object types that are strongly related to the main classi cation task.
We consider the network classi cation task according to the within-network
setting [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]: objects for which the class is known are linked to objects for which
the class must be estimated [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] (which can be either the subject of the main
classi cation task or other objects related to the main classi cation task). This
semi-supervised setting, which allows the classi cation phase to take advantage
of both labeled and unlabeled examples, leads to smoother predictions [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In
multi-type classi cation, we add an additional mechanism to smooth the
prediction function: capturing relationships among objects of di erent types, i.e,
capturing correlations among labels of objects of di erent types.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Problem Statement and Background</title>
      <p>Before describing the proposed method, in the following we introduce the
notation used. We work on heterogeneous networks, which we formally de ne as
G = (V; E), where V is the set of nodes and E is the set of edges among nodes.
Both nodes and edges can be of di erent types. Moreover:</p>
      <p>Each node type Tp implicitly de nes a subset of nodes Vp V .</p>
      <p>Each node v0 2 V is associated with a node type tv(v0) 2 T , where T is the
nite set fTpg of all the possible types of nodes in the network.</p>
      <p>A node type Tp de nes a set of attributes Xp = fXp;1; Xp;2; : : : ; Xp;mp g.</p>
      <p>An edge type Rj de nes a subset of edges Ej (Vp Vq) E, where Vp and
Vq are not necessarily based on di erent types.</p>
      <p>An edge e between two nodes v0 and v00 is associated with an edge type Rj 2 R,
where R is the nite set fRj g of possible edge types in the network.</p>
      <p>In the considered task, we de ne a role for each node type: Tt (primary
targets), which are considered as the targets of the main classi cation task; Tst
(secondary targets), which are strongly related to the main classi cation task, for
which a prediction of missing values is considered relevant; Ttr (task-relevant),
which are the other node types. An example is reported in Figure 1. Only nodes
of target (primary and secondary) types are actually classi ed, on the basis of
all the nodes. However, we are actually interested in the maximization of the
prediction accuracy only of objects of the primary target types.
3</p>
    </sec>
    <sec id="sec-3">
      <title>The proposed ensemble learning method</title>
      <p>
        In this section we describe the two solutions ST-MrSBC and MT-MrSBC. Both
take as input a partially labeled heterogeneous network and work iteratively. At
each iteration, they build an ensemble of Mr-SBC [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] classi ers from di erent
subsets of labeled nodes (either known or predicted in the previous iterations),
whose combination of the output will possibly lead to a stronger model.
      </p>
      <p>
        Following the idea in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] (for multi-label classi cation tasks), MT-MrSBC
shu es all the target types. This is motivated by the fact that a prede ned
ordering of the analysis of target types can negatively a ect the classi cation
accuracy, since a wrong decision can inhibit the exploitation of relevant dependencies
and, consequently, can possibly enforce the exploitation of irrelevant/wrong
dependencies. In the literature, this phenomenon is also observed in random-scan
Gibbs sampling, as opposed to systematic-scan Gibbs sampling [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>Once the order of analysis has been de ned, MT-MrSBC samples a subset
of nodes of the rst target type and builds a weak (since built from a subset
of instances) predictive model through Mr-SBC. This predictive model is
applied to classify unlabeled nodes which are then added to the labeled network.
The obtained probabilities are also stored for the nal combination of the
outputs. Then, we select a new target type and repeat the process. Note that, at
this stage, predictions performed for the previous target types are available to
Mr-SBC when building a prediction model for the new target. Target types are
shu ed again every time they are all processed. The number of iterations is
limited by a user-de ned threshold and the outputs obtained for each target type
over all the iterations are combined to obtain the nal (strong) predictive model.
ST-MrSBC uses a similar process, but works on a single target type. Therefore, it
exploits predictions obtained on the same target type in the previous iterations.
The main di erence with respect to the standard self-training is in the
construction of the labeled network when a new iteration starts. Speci cally, ST-MrSBC
builds an ensemble of weak classi ers from di erent subsets of labeled nodes,
instead of considering only the output of the last iteration.</p>
      <sec id="sec-3-1">
        <title>Construction of labeled and unlabeled networks</title>
        <p>At each iteration, we select the target attribute of the next target type and build
two separate networks: the rst contains only labeled nodes, and the second
contains only unlabeled nodes. These networks are built by considering the complete
network and by removing unlabeled and labeled nodes, respectively. This means
that all the nodes of other target types, as well as nodes of task-relevant types,
are included in both networks. While the network of labeled nodes, for each
target type, changes at di erent iterations, the network of unlabeled nodes remains
stable and the algorithm classi es the same nodes several times.</p>
        <p>
          The way nodes are considered as training instances at each iteration is
random. In fact, the algorithm randomly selects, according to a uniform distribution,
a given percentage perc of labeled nodes from the labeled network. Note that,
at a given iteration, the labeled network could contain also nodes that were
initially unlabeled but that were classi ed during a previous iteration. In the case
of ST-MrSBC, the last available predictions will regard the same target type,
whereas in the case of MT-MrSBC, the last available predictions will regard all
the target types (primary and secondary). This behavior is coherent with the
assumption that predictions performed on other target types (primary or
secondary) can help in the predictions of the label of nodes of the current target
type. This implies that, at the next iteration, the algorithm will take into account
some labels predicted at the previous iterations. This choice is, apparently, in
contrast with most of the self-training approaches, which usually select the most
reliable predictions for the next iterations. However, 1) this solution does not
necessarily lead to better results with respect to a random selection [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], and 2)
we do not use predictions as \hard" constraints, but in the next iterations we are
able to retract decisions that are not coherent with the new state of the network.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Estimation of probabilities</title>
        <p>
          At the end of each iteration, we build a classi cation model through Mr-SBC for
the current target type from the network of labeled nodes. We remind that
MrSBC is a nave Bayes classi er which relies on a set of rst-order rules induced
from data stored in the tables of a relational database (see [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] for details).
        </p>
        <p>For each unlabeled node v0, the identi ed classi cation model is exploited
to compute the posterior probability (which takes into account both network
correlation and network autocorrelation) and its label t(v0), according to the
following equation based on the Bayes' theorem:
t(v0) = argmaxYc P (YcjRv0 ) = argmaxcP (Yc) P (Rv0 jYc)=P (Rv0 );
(1)
where Yc is a possible class value of the target attribute y, and Rv0 RD is the
subset of rst-order rules, identi ed by Mr-SBC, covering the object v0.</p>
        <p>The labeled nodes are then added to the labeled network for the subsequent
iterations, as described before. The probabilities are exploited at the end of the
entire process. In fact, after the last iteration, the nal classi er combines the
probabilities computed during all the iterations: for each target type Tt and for
each node v0, the method computes the nal probability as the average of the
probabilities computed over the ensemble in the following way:</p>
        <p>z
t(v0) = argmaxc 1 X P (YcjRv0k) = argmaxc
z k=1
k=1</p>
        <p>
          z
1 X P (Yc)P (Rv0kjYc) ;
z P (Rv0k)
(2)
where z is the number of iterations, i.e., the number of classi ers in the ensemble,
for each target type (the total number of iterations is z jLtj) and Rv0k is the set
of rules identi ed from the labeled network at the k-th iteration. The rationale
behind the combination in Equation (2) is twofold: a) the predictions obtained at
each iteration are based on di erent training sets, which may focus on di erent
properties of the concept to be learned; b) the predictions are made more stable.
We performed our experiments on ve heterogeneous networks: MOVIE, NBA,
YELP, IMDB and STACK. Some quantitative information about these datasets
can be found in Table 4, while additional information can be found in [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ].
In order to evaluate the di erent variants of the ensemble learning solutions we
propose, we compared them with the original version of Mr-SBC. This allows us
to evaluate the contribution of each aspect of our method, i.e., the iterative
nature of the ensemble-based self-training approach (introduced in ST-MrSBC)
and the multi-type classi cation of both primary and secondary targets
(implemented in MT-MrSBC). Moreover, we evaluate the performance of
MTMrSBC by considering both a lexicographic (LexicographicMT-MrSBC) and
a random ordering (RandomMT-MrSBC) of target types.
        </p>
        <p>
          In order to perform a comparison with other systems, we also ran the
experiments with four competitor methods. In particular, we considered i) the
relational version of the nearest neighbour algorithm (RelIBK), ii) the SVM-based
algorithm SMO (RelSMO), iii) the algorithm GNetMine [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], which is natively
able to work on heterogeneous networks, and iv) the algorithm HENPC [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ],
recently proposed to solve the multi-type classi cation task in heterogeneous
networks. However, RelIBK, RelSMO and HENPC were not able to nish within 3
days of execution, while GNetMine was not able to compute the results for Yelp
dataset, since the system went out of memory (on a server with 32GB of RAM).
In these cases, we ran the experiments on a reduced set of nodes (about 1; 000
nodes for each target type) of the datasets IMDB, STACK and YELP.
        </p>
        <p>As regards the parameter setting of ST-MrSBC and MT-MrSBC, we
considered a random sampling of 20% of nodes for each classi er in the ensemble and
the number of iterations for each target type z ranging from 1 to 50.</p>
        <p>We collected all the classi cation accuracy values and performed the
Nemenyi post-hoc tests on single datasets, on all the datasets and on the reduced
datasets for the comparative evaluation. The result of the statistical tests are
depicted in Figure 3. In particular, we can observe that for the datasets NBA
and for the target type users of the dataset Yelp, the improvement provided
by LexicographicMT-MrSBC and RandomMT-MrSBC over the competitors is
statistically signi cant at p-value= 0:05. Focusing on LexicographicMT-MrSBC,
we observe two specular and interesting cases: for the dataset IMDB it
provides the best result, while for the target type Business of the dataset Yelp
it appears to be the worst approach. This instability is again caused by the
static ordering on the target types. Indeed, in the rst case, the exploitation
of the dependency movies ! users (which is surely strong, if we observe the
result) led LexicographicMT-MrSBC to obtain a better result with respect to
RandomMT-MrSBC, which alternatively (and randomly) exploited the
dependencies movies ! users and users ! movies (which appears weak). On the
contrary, in the second case, the static (unlucky) choice of the dependency to be
exploited brought LexicographicMT-MrSBC to the bottom of the ranking.</p>
        <p>Finally, the results obtained by our comparative evaluation (Figure 3 - last
chart) shows that the improvement provided by the proposed method is able
to give Mr-SBC the advantage of outperforming the competitors Indeed, the
original Mr-SBC is not able to outperform the considered competitors, while
the ensemble learning (ST-MrSBC) leads to outperform all the competitors
(although not statistically w.r.t. HENPC and RelIBk). The advantage comes from
the combination of capturing label dependencies between multiple types and of
the ensemble learning approach (MT-MrSBC), which improves the accuracy.</p>
        <p>Overall, we can conclude that the application of the proposed method, which
is able to capture both correlation and autocorrelation phenomena, as well as
to predict missing values, by exploiting the same classi cation method adopted
for primary targets, can lead to better, more stable predictions when applied to
real-world data organized in heterogeneous networks.
Fig. 3. Results of the Nemenyi post-hoc test on the average accuracy. Better algorithms
are positioned on the right-hand side, and those that do not signi cantly di er in
performance (at p-value = 0:05) are connected with a line.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusions</title>
      <p>In this paper, we proposed an extension of the system Mr-SBC which works on
heterogeneous networks and that is able to solve multi-type classi cation tasks,
by capturing both correlation and autocorrelation phenomena. Experiments
performed on real datasets show that both the proposed variants are able to
significantly outperform the original Mr-SBC, especially with a random ordering of
the target types, as well as four other well-known competitor algorithms.</p>
      <p>As future work, we plan to study the sensitivity to the sampling size and the
possibility to select only a subset of labeled nodes for the next iteration.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>We would like to acknowledge the European project MAESTRA - Learning from
Massive, Incompletely annotated, and Structured Data (ICT-2013-612944).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>P.</given-names>
            <surname>Angin</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Neville</surname>
          </string-name>
          .
          <article-title>A shrinkage approach for modeling non-stationary relational autocorrelation</article-title>
          .
          <source>In ICDM</source>
          , pages
          <volume>707</volume>
          {
          <fpage>712</fpage>
          . IEEE Computer Society,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>M.</given-names>
            <surname>Ceci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Appice</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Malerba</surname>
          </string-name>
          .
          <article-title>Mr-SBC: a multi-relational naive bayes classi er</article-title>
          .
          <source>In PKDD 2003</source>
          , pages
          <fpage>95</fpage>
          {
          <fpage>106</fpage>
          . Springer,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>O.</given-names>
            <surname>Chapelle</surname>
          </string-name>
          ,
          <string-name>
            <surname>B.</surname>
          </string-name>
          <article-title>Scholkopf, and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Zien</surname>
          </string-name>
          .
          <article-title>Semi-Supervised Learning</article-title>
          .
          <source>The MIT Press, 2nd edition</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>C.</given-names>
            <surname>Desrosiers</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.</given-names>
            <surname>Karypis</surname>
          </string-name>
          .
          <article-title>Within-network classi cation using local structure similarity</article-title>
          .
          <source>In ECML PKDD '09</source>
          , pages
          <fpage>260</fpage>
          {
          <fpage>275</fpage>
          , Berlin,
          <year>2009</year>
          . Springer-Verlag.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Y.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Niu</surname>
          </string-name>
          , and
          <string-name>
            <surname>H. Zhang.</surname>
          </string-name>
          <article-title>An extensive empirical study on semi-supervised learning</article-title>
          .
          <source>In ICDM</source>
          , pages
          <volume>186</volume>
          {
          <fpage>195</fpage>
          . IEEE,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>M.</given-names>
            <surname>Ji</surname>
          </string-name>
          , J. Han, and
          <string-name>
            <given-names>M.</given-names>
            <surname>Danilevsky</surname>
          </string-name>
          .
          <article-title>Ranking-based classi cation of heterogeneous information networks</article-title>
          .
          <source>In SIGKDD '11</source>
          , pages
          <fpage>1298</fpage>
          {
          <fpage>1306</fpage>
          , NY, USA,
          <year>2011</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>M.</given-names>
            <surname>Ji</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Danilevsky</surname>
          </string-name>
          , J. Han, and
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          .
          <article-title>Graph regularized transductive classi cation on heterogeneous information networks</article-title>
          .
          <source>In ECML PKDD</source>
          <year>2010</year>
          , volume
          <volume>6321</volume>
          <source>of LNCS</source>
          , pages
          <volume>570</volume>
          {
          <fpage>586</fpage>
          . Springer Berlin,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Liu</surname>
          </string-name>
          . Monte Carlo Strategies in Scienti c Computing. Springer Publishing Company, Incorporated,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>C.</given-names>
            <surname>Loglisci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ceci</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Malerba</surname>
          </string-name>
          .
          <article-title>Relational mining for discovering changes in evolving networks</article-title>
          .
          <source>Neurocomputing</source>
          ,
          <volume>150</volume>
          :
          <fpage>265</fpage>
          {
          <fpage>288</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Macskassy</surname>
          </string-name>
          and
          <string-name>
            <given-names>F.</given-names>
            <surname>Provost</surname>
          </string-name>
          .
          <article-title>Classi cation in networked data: A toolkit and a univariate case study</article-title>
          .
          <source>J. Mach. Learn. Res.</source>
          ,
          <volume>8</volume>
          :
          <fpage>935</fpage>
          {
          <fpage>983</fpage>
          , May
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>G.</given-names>
            <surname>Pio</surname>
          </string-name>
          , F. Sera no,
          <string-name>
            <given-names>D.</given-names>
            <surname>Malerba</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Ceci</surname>
          </string-name>
          <article-title>. Multi-type clustering and classi - cation from heterogeneous networks</article-title>
          .
          <source>Inf. Sci.</source>
          ,
          <volume>425</volume>
          :
          <fpage>107</fpage>
          {
          <fpage>126</fpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>J. Read</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Pfahringer</surname>
            , G. Holmes, and
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Frank</surname>
          </string-name>
          .
          <article-title>Classi er chains for multi-label classi cation</article-title>
          .
          <source>Machine Learning</source>
          ,
          <volume>85</volume>
          (
          <issue>3</issue>
          ):
          <volume>333</volume>
          {
          <fpage>359</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13. F. Sera no, G. Pio, and
          <string-name>
            <given-names>M.</given-names>
            <surname>Ceci</surname>
          </string-name>
          .
          <article-title>Ensemble learning for multi-type classi cation in heterogeneous networks (doi: 10</article-title>
          .1109/tkde.
          <year>2018</year>
          .
          <volume>2822307</volume>
          ).
          <source>IEEE TKDE</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <given-names>D.</given-names>
            <surname>Stojanova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ceci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Appice</surname>
          </string-name>
          , and S. Dzeroski.
          <article-title>Network regression with predictive clustering trees</article-title>
          .
          <source>Data Min. Knowl. Discov.</source>
          ,
          <volume>25</volume>
          (
          <issue>2</issue>
          ):
          <volume>378</volume>
          {
          <fpage>413</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sun</surname>
          </string-name>
          , J. Han,
          <string-name>
            <given-names>P.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yin</surname>
          </string-name>
          , H. Cheng, and
          <string-name>
            <given-names>T.</given-names>
            <surname>Wu</surname>
          </string-name>
          .
          <article-title>Rankclus: integrating clustering with ranking for heterogeneous information network analysis</article-title>
          .
          <source>In EDBT '09</source>
          , pages
          <fpage>565</fpage>
          {
          <fpage>576</fpage>
          , New York, NY, USA,
          <year>2009</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yu</surname>
          </string-name>
          , and J. Han.
          <article-title>Ranking-based clustering of heterogeneous information networks with star network schema</article-title>
          .
          <source>In ACM SIGKDD</source>
          , pages
          <volume>797</volume>
          {
          <fpage>806</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>