<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Do We Need to Observe Features to Perform Feature Selection?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jan Motl</string-name>
          <email>jan.motl@fit.cvut.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pavel Kordík</string-name>
          <email>pavel.kordik@fit.cvut.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Czech Technical University in Prague</institution>
          ,
          <addr-line>Thákurova 9, 160 00 Praha 6</addr-line>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <volume>2203</volume>
      <fpage>44</fpage>
      <lpage>51</lpage>
      <abstract>
        <p>Many feature selection methods were developed in the past, but in the core, they all work the same way you pass a set of features to the algorithm and get a reduced set of the features. But can we perform a non-trivial feature selection without first observing the features? This is an important question because if we were actually able to predict feature importance before observing the features, we would reduce computation requirements of all stages of machine learning process beginning with feature engineering. In this article, we argue that it is possible to predict feature importance before feature vector observation. The trick is that we use meta-features about the features to perform the feature selection. We evaluate the concept on 15 relational databases. On average, it was enough to generate the top decile of all features to get the same model accuracy as if we generated all features and passed them to the model.</p>
      </abstract>
      <kwd-group>
        <kwd>meta-learning</kwd>
        <kwd>feature engineering</kwd>
        <kwd>feature selection</kwd>
        <kwd>relational database</kwd>
        <kwd>propositionalization</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Data in relational databases are in the form of many tables,
but common classification algorithms require input data
in the form of a single table. Propositionalization solves
this discrepancy by converting data from the form of many
tables into a single table.</p>
      <p>
        But there are two significant problems with the
propositionalization [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. It produces a lot of features. And many of
them are redundant. These two issues result in high
computational requirements during both, propositionalization and
classification.
      </p>
      <p>
        Contrary to the common approach (e.g., [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]),
we deal with these two issues by performing feature
selection before the propositionalization and not after the
propositionalization. The key idea is that we collect
metadata about the attributes in the database (e.g., attribute data
type), meta-data about the feature generative functions (e.g.,
id of the feature function), calculate landmarking features
on a small subset of all features and pass their performance
to a meta-learner, which predicts the optimal order, in which
the remaining features should be calculated.
2.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>Meta-learning for Feature Engineering</title>
      <p>
        Meta-learning was originally concerned with algorithm
selection[21]. Nevertheless, Nargesian [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] trained a neural
network to predict, which feature transformations are going
to improve the accuracy of a classifier based on the feature
histograms.
      </p>
      <p>We extend the idea of using the data-based meta-features
(in Nargesian’s case a histogram) for feature engineering
with landmarking.
2.2</p>
    </sec>
    <sec id="sec-3">
      <title>Meta-learning for Feature Selection</title>
      <p>
        Reif [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] applies meta-learning to accelerate forward
selection. The key concept is that the performance of all
candidate feature subsets in each forward step is first
estimated with a meta-learner. And only the top x percent
of the candidates get evaluated on the data to get the true
subset performance. Based on the reported results, it is
sufficient to evaluate only the top 10% of all candidate subsets
on the data to get results comparable to classical forward
selection.
      </p>
      <p>The difference between our approach and Reif’s
approach is that Reif calculates meta-features from the
features, while we calculate meta-features directly from the
attributes that are used to calculate the features (in Figure
1 we use only the left table, while Reif uses the right
table). Consequently, in Reif’s case, we have to calculate the
features first, to perform feature selection. While in our
case, we can perform the feature selection before feature
calculation.
3</p>
      <sec id="sec-3-1">
        <title>Method</title>
        <p>A high-level schema of our approach is in Figure 2. The
whole process is divided into two phases. During the offline
phase, meta-features and feature performance are collected
on many databases and passed to a meta-learner as training
data. During the online phase, the trained meta-learner is
used to rank candidate features in the descending order
of their estimated utility. Following paragraphs define the
feature utility.</p>
        <p>
          There are many properties that a feature should posses
[
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], but we focus on predicting properties measurable
directly from the data: relevance to the task, redundancy to
other features and runtime of the feature calculation.
function
( 1
i d
1
2
3
3
. .
. .
. .
. .
. .
. .
efatur
+
4
3
space
problem into a single-instance problem solvable with a common attribute value classifier. In this trivial example, the feature
space contains only a single feature vector but it may generally contain thousands of feature vectors.
        </p>
        <p>A numerical feature f1 is redundant to numerical feature f2
if a linear transformation from f1 to f2 and back exists.</p>
        <p>We use this (weaker) definition of redundancy instead
of the identity of the features because it corresponds better
with the notion of redundancy in many models (e.g., in
logistic regression with one shot encoding of categorical
features). To speed up the identification of redundant
features, we use Chi2 as a hash function to identify potential
redundant features [19, Section 2.1].
3.1</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Feature Utility</title>
      <p>
        We calculate features1 in descending order of the estimated
relevance/runtime ratio [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] since we prefer to calculate
highly relevant and fast features first. Furthermore, we
penalize the feature i proportionally to the estimated
probability that the feature is redundant pˆi. Because each dataset
has a different proportion of redundant features (see Table
1) and the tested meta-learning models had difficulties to
model these differences, we employ median thresholding
instead of a fixed threshold:
utilityi = (pˆi &gt; median(pˆ) ? 1 − pˆi : 1)
relevancei
runtimei
, (1)
where redundancy is a vector of estimated redundancy
probabilities for a database.
      </p>
      <p>1In the production, we would calculate only the top n features that
we would use to build a production classifier. But to demonstrate the
meaningfulness of such approach, we calculate all features.</p>
      <sec id="sec-4-1">
        <title>Experiment</title>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Data</title>
      <p>
        We used 15 databases listed in Table 1 from relational
repository [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
4.2
For propositionalization, we used Relaggs [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], which
was modified to work with 31 different feature
(generative) functions, listed in Figure 2. The detail
description of the employed feature functions is at http://
predictorfactory.com.
We employ three sources of meta-features: landmarking
features, database meta-data and feature function meta-data.
Landmarking features Just like the accuracy of a few
classifiers can be used as meta-features for the recommendation
of the best classifier on the data (e.g., [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]), we define a
subset of feature functions as landmarking feature functions
for the recommendation of the best features.
      </p>
      <p>
        Without loss of generality, we used following set of
landmarking features: Direct field (a simple copy of the value),
Aggregate (e.g., min, max,...), WOE (Weight of Evidence),
Count (of tuples), Aggregate WOE, Time aggregate since.
These feature functions were selected for their low
runtime (see Table 11 in the appendix) and good coverage
of different data types (numerical/character/temporal) and
relationships between the label and the data (1:1/1:n). Note
that we do not use multivariate feature functions for
landmarking due to the potential combinatorial explosion.
Database meta-data Basic descriptive and statistical
metafeatures are frequently employed in meta-learning (e.g.,
[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]) and we do not differ in this respect. A noteworthy
difference is that we do not calculate statistics of the
attributes but rather reuse statistics maintained by the
relational database for query plan optimization [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. This slight
deviation allows us to collect estimates of the statistics in
time independent on the count of tuples (records) in the
database.
      </p>
      <p>Feature function meta-data Feature function meta-data
consists of feature function name (e.g., Aggregate) and
feature function parameters (e.g., min).
4.4</p>
    </sec>
    <sec id="sec-6">
      <title>Measures</title>
      <p>Anytime algorithm We formulate feature engineering as
anytime algorithm [25], which aims to deliver the best
subset of calculated features in any time. The quality of
anytime algorithm can be expressed with a performance profile,
where we measure quality of the solution at the given time
(see example in Figure 3). To assign a single number to
the performance profile, we calculate the area between the
archived curve a(t) and the expected random curve r(t)
(which we obtain from averaging the curve from many
random permutations), divided by the area between the perfect
curve p(t) and the expected random curve r(t):</p>
      <p>
        R a(t)dt − R r(t)dt
POP = R p(t)dt − R r(t)dt
(2)
where t is time. The obtained ratio then represents the
“percentage of perfect” solution [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In our case, a(t), r(t)
and p(t) are the Chi2 of the feature calculated at time t. The
only difference between these functions is then the order,
in which the features are calculated. The perfect feature
ordering is based on a complete knowledge of relevance,
redundancy and runtime of all the features. While archived
ordering is based only on the estimates of these feature
properties (the only exception are landmarking features,
which are calculated in a pseudorandom order).
      </p>
      <p>Mutagenesis, POP=0.595
100
80
e
c
an 60
v
e
le 40
R
20
0
0
10
20 30
Runtime [s]</p>
      <p>Perfect
Archived
Random
95% PI
Diagonal
40
Individual models To assess the ability of relevance and
runtime prediction models to rank, we use Spearman
correlation (ρ). The quality of redundancy estimation (a
classification task) is evaluated with area under receiver operating
characteristic curve (AUROC).</p>
      <p>Permutation testing To assess, whether the obtained
performance profiles are significantly better than random, we
generate 1,000 random orderings of the features to estimate
95% prediction intervals.
The obtained accuracies are depicted in Table 3. Since the
difference between the models is not significant, we use
GLM for all following experiments.
5.2</p>
    </sec>
    <sec id="sec-7">
      <title>Feature Importance</title>
      <p>Relevance The most important meta-features for feature
relevance prediction is the average relevance of the
landmarking features on individual attributes and the type of
the employed feature function (see Table 4).</p>
      <p>
        Redundancy There are two main sources of redundant
features [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]: redundancy in the input data and redundancy
introduced by the feature functions. The redundancy in the
input data is covered by landmarking landmark_is_redundant
and data_type. While the introduced redundancy is
explained with feature_function and feature_parameters (see
Table 5).
      </p>
      <p>Runtime The runtime of a feature function calculation is a
function of two factors: the type of the feature function and
data property. Nevertheless, these two factors are dominated
by the landmarking landmark_runtime (see Table 6).
5.3</p>
    </sec>
    <sec id="sec-8">
      <title>Percentage of Perfect</title>
      <p>The quality of anytime learning for all 15 datasets is
reported in Table 7 in the penultimate column.
6
6.1</p>
      <sec id="sec-8-1">
        <title>Discussion</title>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>What is the contribution of the individual models to POP?</title>
      <p>To evaluate the contribution of the individual models to
POP, we performed an experiment with a 2-level full
factorial design for presence/absence of runtime, relevance
and redundancy models (8 combinations in total) on all
databases. To deal with the variability across databases
(some are easier than others), we treat the database name
as a random factor.</p>
      <p>Conclusion: The result of the factor analysis is in Table 8.
As expected, the intercept is not significantly different from
zero, since POP measure should on average be 0 when we
randomly rank the features. The biggest contributions to
the accuracy are from redundancy and relevance prediction.
The interaction between redundancy and relevance has a
negative estimate because we do not reward calculation of
redundant features even if they are highly relevant. Hence,
prediction of the relevance helps only on the subset of
unique features from the set of all candidate features.
6.2</p>
    </sec>
    <sec id="sec-10">
      <title>What is the effect of meta-learning on model accuracy?</title>
      <p>To evaluate the effectivity of the meta-learning, we
iteratively train a classification model on increasing percentage
of the top features, as estimated with meta-learning. As the
classification model, we use a decision tree because it can
model interactions between the features, it is
undemanding on data preprocessing and it is reasonably fast. As the
evaluation measure, we use misclassification error as all
databases have reasonably balanced classes in the label.</p>
      <p>
        An example of the obtained curve is depicted in Figure
4, where we can observe that the decision tree slightly
overfits when we use all the features. Nevertheless, forward
selection still outperforms meta-learning feature selection,
as it can observe all the features (our approach does not)
and it is a wrapper (our approach is a filter [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]).
      </p>
      <p>Financial</p>
      <p>Meta-learning (AUC=0.23)
Random (AUC=0.33)
All features</p>
      <p>Forward selection
20</p>
      <p>40 60
Percentage of used features
80
100</p>
      <p>Conclusion: The result of the factor analysis is in Table
9. Prediction of relevance significantly reduces
misclassification error. Prediction of runtime insignificantly increases
the misclassification error, because this evaluation does not
reward fast features. Redundancy prediction does not
significantly decrease the classification error. Based on our
inspection of the results, this is because this evaluation
rewards early discovery of a few highly relevant features
much higher (since the best possible decision tree may use
just a few features) than it penalizes redundancy (a
redunTo analyze the importance of the three categories of the
meta-features (landmarking, database, feature-function),
we design an experiment, in which we vary the set of the
used meta-features.</p>
      <p>Conclusion: Based on the results reported in Table 10,
only landmarking meta-features help to significantly2
reduce the count of features that have to be engineered to
reach model accuracy obtained on all features. Table 10
also tells us that if all meta-features are used, it is in average
sufficient to engineer only the top 8.25% of the features to
match or surpass the classification accuracy of the model
trained on all features.
features. Based on Wilcoxon signed-rank test, we have to
reject the null hypothesis that the additional features do
not improve accuracy (p-value = 0.00048). The median
improvement is 1.2 percent point in classification accuracy
(average improvement is 2.7 percent point).</p>
      <p>Conclusion: The additional features improve the
accuracy of the model over the accuracy of the model build
only on the landmarking features by a small but significant
amount.
6.5</p>
    </sec>
    <sec id="sec-11">
      <title>Feature Selection vs. Feature Meta-learning</title>
      <p>
        The described feature meta-learning bears similarity with
filter-type feature selection methods likeCorrelated
Feature Selection (CFS)[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and Minimum Redundancy
Maximum Relevance (mRMR)[
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. Both these methods attempt
to quickly select relevant non-redundant features. And so
does our method. But in comparison to these methods, we
perform feature selection before the feature engineering.
Difficulty It can be argued that feature meta-learning is
at least as difficult problem as feature selection since we
can always convert feature selection problem to feature
meta-learning by throwing away the computed features
(and recalculating them on request).
6.6
      </p>
    </sec>
    <sec id="sec-12">
      <title>Limitations</title>
      <p>We performed experiments only on relational data and
features from propositionalization. Propositionalization is
known to produce a lot of duplicate features (38% on
average on the tested databases) and many of the features
are irrelevant to the task (64% on average on the tested
databases based on backward selection). These properties
make it possible to obtain substantial gains from feature
selection. However, the performed experiments do not tell
us how the described approach is going to generalize on
non-relational data.</p>
      <p>
        Another limitation of the reported work is that it ignores
interactions between the feature vectors in the downstream
model. This can reduce the accuracy of the downstream
model because a univariate oraculum meta-learner would
not recommend calculation of features that are useful only
in the combination with other features (a trivial example
where this may happen is XOR problem [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]). Possible
solutions to this problem are briefly mentioned in the future
work Section 7.
6.7
      </p>
    </sec>
    <sec id="sec-13">
      <title>Applications</title>
      <p>Feature meta-learning is desirable in domains, where a
single universal approach to feature extraction does not
exist or is not known ahead. An exemplary domain are
relational data, which may contain highly diverse content
ranging from structured to unstructured data.</p>
      <p>Additionally, feature selection before feature
engineering is applicable to complex or large data, where it is not
feasible or convenient to calculate and evaluate all possible
features due to limited resources.
7</p>
      <sec id="sec-13-1">
        <title>Conclusion</title>
        <p>
          In this article, we evaluated an idea of performing feature
selection before feature engineering. To guide the search, we
exploited meta-learning. Nargesian [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] used meta-features
calculated from the original data. But we found out that
landmarking meta-features work better. When we evaluated
the implementation on 15 databases, we concluded that it
is on average enough to engineer only the top decile of
features to get accuracy comparable to accuracy obtained
on all features. This finding is similar to Reif’s [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] finding,
who applied meta-learning to feature selection. However,
Reif performs feature selection after feature engineering
while we perform feature selection before feature
engineering.
7.1
        </p>
      </sec>
    </sec>
    <sec id="sec-14">
      <title>Future Work</title>
      <p>In this exploratory work, we optimize an ersatz measure
called POP, which is easy to reason about. One possible
extension of this work is to improve individual components
of the meta-learning model. For example, we can detect
redundancy based on a fuzzy comparison of equal-height
histograms estimated with the database engine or quantile
sketch. With this change, we would detect duplicates that
are identical up to a monotonic transformation, leading to
better alignment with models that are invariant to
monotonic transformations of the features (e.g., decision trees in
theory). Or we could replace the redundancy detection with
a precomputed correlation matrix describing similarities
between the feature functions [24, p. 148]. Alternatively,
we could estimate a transition matrix describing the
optimal order in which to apply feature functions (or give up
on the given attribute). To improve non-redundancy and
relevance together, we could train a fast model (e.g., naive
Bayes) on streaming features (e.g., [23] or [22, p. 19]). The
possibilities are vast.</p>
      <p>
        Another possible direction is to directly optimize the
measure we are interested in (e.g., improvement to model’s
AUROC over time). This can be done by training a single
model (e.g., [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]). And this article provides an extended set
of meta-features, on which such model could be trained.
8
      </p>
      <sec id="sec-14-1">
        <title>Acknowledgement</title>
        <p>We would like to thank Adéla Blažková for her help. We
furthermore thank the anonymous reviewers, their comments
helped to improve this paper. The reported research has
been supported by the Grant Agency of the Czech
Technical University in Prague (SGS17/210/OHK3/3T/18) and
the Czech Science Foundation (GACˇ R 18-18080S).</p>
      </sec>
      <sec id="sec-14-2">
        <title>Appendix</title>
      </sec>
    </sec>
    <sec id="sec-15">
      <title>Reproducibility</title>
      <p>The used code is published at:</p>
      <p>https://github.com/janmotl/metalearning.
The used databases are published at:</p>
      <p>https://relational.fit.cvut.cz.
Aggregate frame
Time aggregate diff
Time diff
Time day part
Time since
Null ratio
Time frequency
Existential count
Slope
Time WOE
Time is weekend
Time part
Text length
Intercept
Correlation
Time aggregate
Aggregate text length
Aggregate range
Duplicate ratio
Aggregate distinct
Time range
Coefficient of variation
Time aggregate since event
Direct field
Distinct count
Aggregate
Log product
Time aggregate since
Count
WOE
Aggregate WOE
−2.11
−1.53
−0.92
−1.33
−0.89
−0.99
−0.17
−0.67
−0.49</p>
      <p>0.51
−1.03
−0.45
−0.76
1.42
1.32
0.57
−0.24
0.34
0.32
0.56
0.06
0.36
0.19
0.26
0.21
0.49
0.62
1.25
0.30
1.27
1.04</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Abdulrahman</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Brazdil</surname>
          </string-name>
          .
          <article-title>Measures for combining accuracy and time for meta-learning</article-title>
          .
          <source>CEUR Workshop Proc</source>
          .,
          <volume>1201</volume>
          :
          <fpage>49</fpage>
          -
          <lpage>50</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>T.</given-names>
            <surname>Brandenburger</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Furth. Cumulative Gains Model Quality Metric</surname>
          </string-name>
          .
          <source>J. Appl. Math. Decis. Sci.</source>
          ,
          <year>2009</year>
          :
          <fpage>1</fpage>
          -
          <lpage>14</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>De Raedt</surname>
          </string-name>
          .
          <source>Inductive Logic Programming</source>
          , volume
          <volume>1446</volume>
          of Lecture Notes in Computer Science. Springer, Berlin, Heidelberg,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R. O.</given-names>
            <surname>Duda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. E.</given-names>
            <surname>Hart</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D. G. Stork. Pattern</given-names>
            <surname>Classification</surname>
          </string-name>
          . Wiley Interscience,
          <volume>2</volume>
          <fpage>edition</fpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>I.</given-names>
            <surname>Guyon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Weston</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Barnhill</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V.</given-names>
            <surname>Vapnik</surname>
          </string-name>
          .
          <article-title>An Introduction to Variable and Feature Selection Isabelle</article-title>
          . Mach. Learn.,
          <volume>46</volume>
          (
          <issue>1</issue>
          /3):
          <fpage>389</fpage>
          -
          <lpage>422</lpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Hall</surname>
          </string-name>
          .
          <article-title>Correlation-based Feature Selection for Machine Learning</article-title>
          .
          <source>PhD thesis</source>
          , The University of Waikato,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Kanter</surname>
          </string-name>
          . Deep Feature Synthesis:
          <article-title>Towards Automating Data Science Endeavors</article-title>
          .
          <source>IEEE DSAA</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>C. S.</given-names>
            <surname>Kheau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Alfred</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L. H.</given-names>
            <surname>Keng</surname>
          </string-name>
          .
          <article-title>Dimensionality Reduction in Data Summarization Approach to Learning Relational Data</article-title>
          . ACIIDS,
          <volume>7802</volume>
          :
          <fpage>166</fpage>
          -
          <lpage>175</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.-A.</given-names>
            <surname>Krogel</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Wrobel</surname>
          </string-name>
          .
          <article-title>Transformation-Based Learning Using Multirelational Aggregation</article-title>
          .
          <source>In ILP</source>
          , pages
          <fpage>142</fpage>
          -
          <lpage>155</lpage>
          . Springer, London,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>M.-A. Krogel</surname>
            and
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Wrobel</surname>
          </string-name>
          .
          <article-title>Propositionalization and Redundancy Treatment</article-title>
          . In Databases, Doc. Inf. Fusion, Hannover,
          <year>2002</year>
          . CEUR.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>C.</given-names>
            <surname>Lemke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Budka</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Gabrys</surname>
          </string-name>
          .
          <article-title>Metalearning: a survey of trends and technologies</article-title>
          . Artif. Intell.,
          <volume>44</volume>
          (
          <issue>1</issue>
          ):
          <fpage>117</fpage>
          -
          <lpage>130</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>A.</given-names>
            <surname>McNab</surname>
          </string-name>
          and
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Ladd</surname>
          </string-name>
          .
          <article-title>Information quality: The importance of context and trade-offs</article-title>
          .
          <source>In Proc. Annu. Hawaii Int. Conf. Syst. Sci.</source>
          , pages
          <fpage>3525</fpage>
          -
          <lpage>3532</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>M.</given-names>
            <surname>Minsky</surname>
          </string-name>
          and
          <string-name>
            <surname>S. A. Papert. Perceptrons.</surname>
          </string-name>
          <article-title>An Introduction to Computational Geometry</article-title>
          . MIT, jan
          <year>1969</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>J.</given-names>
            <surname>Motl</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Kordík</surname>
          </string-name>
          .
          <article-title>Foreign Key Constraint Identification in Relational Databases</article-title>
          .
          <source>In ITAT</source>
          , pages
          <fpage>106</fpage>
          -
          <lpage>111</lpage>
          . CEUR,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>J.</given-names>
            <surname>Motl</surname>
          </string-name>
          and
          <string-name>
            <given-names>O.</given-names>
            <surname>Schulte</surname>
          </string-name>
          .
          <source>The CTU Prague Relational Learning Repository. arXiv, page 7</source>
          , nov
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>F.</given-names>
            <surname>Nargesian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Samulowitz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Khurana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. B.</given-names>
            <surname>Khalil</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Turaga</surname>
          </string-name>
          .
          <article-title>Learning Feature Engineering for Classification</article-title>
          . IJCAI, (
          <year>August</year>
          ):
          <fpage>2529</fpage>
          -
          <lpage>2535</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>H.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Long</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Ding</surname>
          </string-name>
          .
          <article-title>Feature selection based on mutual information criteria of max-dependency, maxrelevance, and min-redundancy</article-title>
          .
          <source>IEEE TPAMI</source>
          ,
          <volume>27</volume>
          (
          <issue>8</issue>
          ):
          <fpage>1226</fpage>
          -
          <lpage>1238</lpage>
          , aug
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>B.</given-names>
            <surname>Pfahringer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Bensusan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Giraud-Carrier</surname>
          </string-name>
          .
          <article-title>MetaLearning by Landmarking Various Learning Algorithms</article-title>
          . In ICML, volume
          <volume>951</volume>
          , pages
          <fpage>743</fpage>
          -
          <lpage>750</lpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>A.</given-names>
            <surname>Popescul</surname>
          </string-name>
          and
          <string-name>
            <given-names>L. H.</given-names>
            <surname>Ungar</surname>
          </string-name>
          .
          <article-title>Structural Logistic Regression for Link Analysis</article-title>
          . MRDM, (
          <year>August</year>
          ):
          <fpage>92</fpage>
          -
          <lpage>106</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>M.</given-names>
            <surname>Reif</surname>
          </string-name>
          and
          <string-name>
            <given-names>F.</given-names>
            <surname>Shafait</surname>
          </string-name>
          .
          <article-title>Efficient feature size reduction via predictive forward selection</article-title>
          .
          <source>Pattern Recognit</source>
          .,
          <volume>47</volume>
          (
          <issue>4</issue>
          ):
          <fpage>1664</fpage>
          -
          <lpage>1673</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>