<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Meta-QSAR: learning how to learn QSARs</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ivan Olier</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Crina Grosan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Noureddin Sadawi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Larisa Soldatova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ross D. King</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science University of Brunel</institution>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Manchester Institute of Biotechnology University of Manchester</institution>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Quantitative structure activity relationships (QSARs) are functions that predict bioactivity from compound structure. Although almost every form of statistical and machine learning method has been applied to learning QSARs, there is no single best way of learning QSARs. Therefore, currently the QSAR scientist has little to guide her/him on which QSAR approach to choose for a specific problem. The aim of this work is to introduce Meta-QSAR, a meta-learning approach aimed to learning which QSAR method is most appropriate for a particular problem. For the preliminary results presented here, we used ChEMBL1, a public available chemoinformatic database, to systematically run extensive comparative QSAR experiments. We further apply meta-learning in order to generalise these results.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>The datasets involved in this research have been formed by computing molecular
properties and fingerprints of chemical compounds with associated bioactivity to a
particular target (protein). Learning a QSAR model consists on fitting a regression
method to a dataset which has as input variables the descriptors, as response variable
(output) the associated bioactivities, and as instances, the chemical compounds. We
extracted 2,750 targets from ChEMBL with a very diverse number of chemical
compounds, ranging from 10 to about 6,000. Two sets of properties – one, using 43
constitutional properties, and another, using 1,683 additional properties – and one
fingerprint (FCFP4, 1024bits) were used to form the datasets. Further datasets were
generated by imputing missing values using the median and performing feature selection
based on the chi-squared test. For the QSAR methods, we have selected 20
algorithms typically used in QSAR experiments, which include: linear regression, support
vector machines, artificial neural networks, regression trees, and random forest,
amongst others. Model performance in all experiments has been assessed by taking
1 ChEMBL database is available from: https://www.ebi.ac.uk/chembl/
the average root mean squared error (RMSE) after 10-fold crossvalidation of the
datasets.</p>
      <p>For the meta-learning stage, we conceived a classification problem that indicates
which QSAR method should be used for a particular QSAR problem. The training
and learning dataset is formed by meta-features extracted from the datasets of the base
learning level and are based on target properties (hydrophobicity, molecular weight,
aliphatic index, etc) and on information theory (mean, mutual information, entropy,
etc). We used random forests as meta-learning algorithm.
3</p>
    </sec>
    <sec id="sec-2">
      <title>Results and Discussion</title>
      <p>rforest.fpFCFP4
rforest.all.mp.miss.fs
ksvmfp.fpFCFP4
hod ridge.fpFCFP4
t
em rforest.all.mp
R
A
Sksvm.all.mp.miss.fs
Q
t
seB ksvm.fpFCFP4
glmnet.all.mp.miss.fs
glmnet.fpFCFP4
earth.fpFCFP4.fs
meta.QSAR
rforest.all.mp
ridge.fpFCFP4
ksvmfp.fpFCFP4
ksvm.all.mp.miss.fs
todhem rrttrreeee..cfpoFnCstF.mP4p
AR ridge.const.mp
tseSBQ rfkkossrvevmmst...cffppoFFnCCstFF.mPP44p
glmnet.all.mp.miss.fs
fnn.fpFCFP4
fnn.const.mp
earth.all.mp.miss.fs
0
200 400
Target counts
600
0</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>